mooney-spark repo, built the engine, downloaded and hash-verified all 90 GB, and ran it on one of our own idle Sparks the same afternoon. One of our predictions about his design was wrong — which turned out to be the more interesting result.Qwen3.8-Flash-Next is a big mixture-of-experts model — roughly 180B total parameters, but only a fraction of those experts activate on any given token. In full precision (BF16, 16 bits per weight) it needs about 335 GiB of memory, which puts it firmly out of reach of a single DGX Spark (121 GiB unified memory, GB10 chip).
Mooney quantizes the routed experts down to 2.125 bits each using a ternary format (each weight is roughly −1, 0, or +1, scaled), and the whole transformer averages 2.39 bits per weight once you count the parts that stay higher-precision. That shrinks the download to 92 GB and the resident memory footprint to somewhere between 39 and 57 GiB depending on which runtime and context length you use — small enough that a single Spark has room left over for a long context window.
Here's the part that makes this more than "just quantize harder." Part of the model's data — a per-layer table that the engine calls PLE — is about 54.4 GB on its own. Instead of loading it into GPU memory like everything else, Daniel's engine leaves it on the SSD and reads the pieces it needs on demand. We confirmed this literally in our own boot log: the memory planner printed CUDA left 50.66 GiB SSD-resident before it even finished loading.
That one design choice is why a 92 GB download can run in 39–57 GiB of memory instead of the full 92. The rest of the weights (~86 GiB on our box, including the ternary expert tensors, attention layers, and a small speculative-decoding "MTP" draft head) sit resident on the GPU the whole time; the big table is paged in only as needed.
A quant file is just numbers until something knows how to run it fast. Mooney’s speed comes from ds4, a CUDA inference engine originally built by the cuda.fast project (vendored from Layr-Labs/ds4) and specifically tuned for GB10 — the chip inside a DGX Spark. Daniel didn’t write ds4 from scratch; he forked it (cudafast-qwen38-125b-a6b-engine, branch lbf/pq2-rot, pinned at commit a77755c6 on top of upstream 5707d4f2) and added exactly the pieces Mooney needed on top of it.
Three pieces of that fork are worth understanding on their own, because they explain why Mooney is fast and not just small:
Ggml-style engines identify tensor formats by a type number; Daniel registered a new one, type 142, for his 2.125-bit ternary experts (each weight ≈ −1, 0, or +1, scaled per block). The loader and the matmul kernels both had to learn this format — it didn’t exist in upstream ds4 before this fork.
Naively rounding weights to 3 values loses a lot. His kernels apply a fused Walsh-Hadamard rotation plus a sign flip before the int8 dequant step — a known trick (also used by rotation-based quant methods like QuIP/QuaRot) for spreading outlier weight values more evenly so low-bit rounding hurts less. It’s fail-closed: the engine checks the rotation metadata’s version, segment sizes, and signs before trusting it.
MTP (multi-token prediction) is speculative decoding: a small 2.58 GiB draft head guesses the next couple of tokens, and the full Mooney model verifies them in one batched pass instead of generating one token at a time. When the guesses are right — which is often, for easy continuations — you get multiple tokens per forward pass. That’s the direct cause of the jump from 33.3 tok/s (serial) to 45.8 tok/s (MTP) on Daniel’s numbers.
Shipping Mooney as a standalone fork repo (DJLougen/cudafast-qwen38-125b-a6b-engine) rather than a patch file keeps upstream cuda.fast’s license and commit history intact, and lets setup_spark.sh clone it directly by URL — the same reproducibility habit as the manifest-driven hash verification elsewhere in the project.
All of this is cross-checked against docs/PINS.md and engine/README.md in the mooney-spark repo, which name the exact base commit, the exact fork commit, and list every delta between them — not just prose claims.
The part of Daniel's thread we liked best wasn't the headline number — it was that he didn't stop at one. His 536-task verifier suite scores Mooney at 88.71% against BF16's 92.91%, which is the "95.5% of BF16" figure everywhere, but he broke out where the 4.2 points went instead of hiding it in an aggregate:
Long reasoning chains, arithmetic, and coding. These lean on precise, compounding computation — exactly what you'd expect a 2.1-bit ternary expert format to struggle with.
Tool calls, JSON/schema adherence, long instructions, multi-turn state, and output formatting. Short instructions actually scored better than BF16 in his suite.
Long-context retrieval held up too: needle-in-haystack was 30/30 across 32k–128k (identical to BF16), and 5/5 at every depth he tested at 256k with the fast runtime. Quality loss didn't compound with context length — a real risk with aggressive quantization that didn't materialize here.
We didn't just read the thread. We cloned DJLougen/mooney-spark onto one of our own idle DGX Sparks and ran his one-command setup for real:
git clone https://github.com/DJLougen/mooney-spark
cd mooney-spark
./setup_spark.sh --mtp-source upstream
What that actually does, end to end: preflight-checks the GPU (must be GB10/sm_121) and CUDA toolkit, builds his ds4 engine fork at a pinned commit (a77755c6, CUDA_ARCH=sm_121), downloads all four GGUF shards plus the vision projector and MTP head from Hugging Face, verifies every single file's size and SHA-256 against a manifest fetched from the repo at run time, and only then writes a launcher script. Any hash mismatch aborts before anything runs — a genuinely fail-closed design, which matters when you're pulling 90 GB of someone else's quantized weights onto a box you intend to trust.
| Step | Result on our Spark |
|---|---|
| Engine build (sm_121) | Clean build, ds4-server binary produced, HEAD matched the pinned commit |
| Shard 1 (1.5 GB) | Size + SHA-256 verified |
| Shard 2 — the 54.4 GB PLE table (52 GB) | Size + SHA-256 verified |
| Shards 3–4 + mmproj + MTP head | Size + SHA-256 verified, ~90 GB total on disk |
| Server boot | Listening on 127.0.0.1:8000, OpenAI-compatible, bound to loopback only |
Setup-to-serving took under 20 minutes end to end on our connection, almost all of it the download. The server binds to 127.0.0.1 by default — it never touched our LAN.
Last week we found that DGX OS ships a nvidia-nvme-interrupt-coalescing.service that sets every Spark's NVMe drive to batch small reads — great for throughput, bad for latency on uncached random 4K reads (we measured 8–10µs p50 with it off vs ~200µs with it on, a ~20× hit). A 50 GB table read on demand from SSD looked like the exact shape of workload that would eat that penalty, so we ran the discriminating test:
| Condition | Fixed prompt, 300 tokens, 3 reps each |
|---|---|
Coalescing ON (DGX OS default, feature 8 = 0x107) | 5.90s, 5.97s, 5.90s |
Coalescing OFF (nvme set-feature -f 8 -v 0) | 5.91s, 5.91s, 5.91s |
/proc/diskstats, our fixed prompt triggered only about 7 NVMe reads per request — nowhere near the thousands of small random reads that would actually expose a per-read latency penalty. Either the table is consulted far less often per token than we assumed, or enough of it gets cached in the Spark's unused memory headroom (121 GiB total, ~89 GiB resident model+KV, leaving ~30 GiB free) that a short prompt mostly hits page cache instead of disk.That's a real finding, just not the one we went in looking for. It means our NVMe-coalescing concern doesn't apply to this specific workload shape on the fast ds4 runtime — but it's a one-prompt test, not a clearance. The open question we're leaving in our research queue: does a longer generation, a cold start (page cache dropped), or many concurrent sessions push enough reads through that table to make coalescing visible? We restored coalescing to its DGX OS default before closing out.
We don't have a 744B model we can safely push to 2-bit experts — our GB300 lane runs an NVFP4 incumbent we're not touching, and "IQ2-class quant on a 744B production model" is a rule we set for ourselves after past low-bit experiments went badly. But three ideas here are genuinely portable:
This is the one we're adopting outright, independent of anything Spark-specific. Every quant or serving-config change we publish from now on reports quality broken out by task category — reasoning, arithmetic, coding, tool calls, JSON, multi-turn, formatting — not just one aggregate number. Daniel's suite is the model for the shape of that table.
"Leave the part you read least often on the fastest storage you're not using" is the same idea behind our own Grace-offload / Engram-host work on the GB300 — just applied to a different memory tier (NVMe vs CPU DRAM). Worth re-reading his engine's paging logic for ideas, even though the specific numbers don't transfer.
Our one-prompt A/B came back clean, which is suspicious in a good way — it means the real test (cold cache, long context, concurrent sessions) hasn't been run yet. That's queued as follow-up work, on our hardware, before we conclude anything either direction.
At ~57 GiB resident with 256k context and MTP, Mooney comfortably shares a Spark with other light work. Worth considering as a third model option on our Spark pair for jobs where 95.5%-of-BF16 quality is an acceptable trade for freeing a Spark's remaining ~60 GiB — not a GB300 candidate, but a legitimate option for the auxiliary tier.
Credit where it's due: this release leans on real upstream work — Layr-Labs' cuda.fast ds4 engine, the PrismML Bonsai quantization line and llama.cpp fork, and compute from Lambda and Zach Mueller. Daniel's own contribution is the ternary PQ2_0 expert format, the rotation-based dequant kernels, the SSD-paging memory plan, and the per-category verifier — and he shipped the whole thing as a reproducible, hash-verified one-command setup instead of just a number on a chart. That reproducibility is why we were able to check his claims on our own hardware the same afternoon instead of just trusting the thread.
Provenance note: every number in the "what Mooney is" and "quality" sections above is from Daniel's public thread and model card, cited inline. Every number in the "we rebuilt it" and "A/B" sections is from our own boot log and curl timings on our Spark, October 1, 2026.