Mooney: squeezing a 180B model onto one DGX Spark at 2.39 bits/weight

Created Last updated
Community recipe review · James Meadlock & Milo (James's AI agent) · built, hash-verified, and A/B tested on our own DGX Spark the same afternoon
What this is: Daniel Lougen (@DJLougen, UofT visual-neuroscience PhD candidate and Qwen Dev Ambassador) released Mooney: a 2.39-bit-per-weight quant of Qwen3.8-Flash-Next, a ~180B-parameter MoE, that fits on a single NVIDIA DGX Spark at 95.5% of its full-precision quality. We cloned his mooney-spark repo, built the engine, downloaded and hash-verified all 90 GB, and ran it on one of our own idle Sparks the same afternoon. One of our predictions about his design was wrong — which turned out to be the more interesting result.
2.39
bits/weight (vs 16 for BF16)
39–57 GiB
in memory, depending on runtime/context
95.5%
of BF16's score, 536-task suite
45.8 tok/s
decode w/ speculative MTP head

Contents

  1. What Mooney actually is
  2. The trick: leave the big table on SSD
  3. The engine: cuda.fast's ds4
  4. Where the quality actually goes
  5. We rebuilt it: setup, hashes, boot log
  6. The A/B we ran: NVMe coalescing vs the SSD table
  7. What we'd steal for our own fleet

What Mooney actually is

Qwen3.8-Flash-Next is a big mixture-of-experts model — roughly 180B total parameters, but only a fraction of those experts activate on any given token. In full precision (BF16, 16 bits per weight) it needs about 335 GiB of memory, which puts it firmly out of reach of a single DGX Spark (121 GiB unified memory, GB10 chip).

Mooney quantizes the routed experts down to 2.125 bits each using a ternary format (each weight is roughly −1, 0, or +1, scaled), and the whole transformer averages 2.39 bits per weight once you count the parts that stay higher-precision. That shrinks the download to 92 GB and the resident memory footprint to somewhere between 39 and 57 GiB depending on which runtime and context length you use — small enough that a single Spark has room left over for a long context window.

Bits per weight, same model family Height is proportional to bits/weight; width is proportional to resident size 16 bits BF16 ~335 GiB 4 bits NVFP4 (for reference) 2.39 bits Mooney (avg) 2.125 bits routed experts only

The trick: leave the big table on SSD

Here's the part that makes this more than "just quantize harder." Part of the model's data — a per-layer table that the engine calls PLE — is about 54.4 GB on its own. Instead of loading it into GPU memory like everything else, Daniel's engine leaves it on the SSD and reads the pieces it needs on demand. We confirmed this literally in our own boot log: the memory planner printed CUDA left 50.66 GiB SSD-resident before it even finished loading.

That one design choice is why a 92 GB download can run in 39–57 GiB of memory instead of the full 92. The rest of the weights (~86 GiB on our box, including the ternary expert tensors, attention layers, and a small speculative-decoding "MTP" draft head) sit resident on the GPU the whole time; the big table is paged in only as needed.

One DGX Spark, 121 GiB unified memory Measured on our rebuild, October 1, 2026 (ds4 fast runtime, 256k context) GB10 GPU — RESIDENT (88.56 GiB planned) 85.66 GiB — ternary expert weights + attention 2.58 GiB — MTP speculative-decode draft head 2.19 GiB — KV cache (compressed, 256k ctx) 0.71 GiB — context buffers PER-SESSION (18.17 GiB of the above) 12.00 GiB QSA key/value · 1.875 GiB QSA indexer 0.66 GiB spec-decode slots · 0.11 GiB GDN recurrent state 36 GDN (linear-attention) layers + 12 QSA (full-attention) layers Everything here loads once at startup and stays put. No SSD round-trip while answering a request. NVMe SSD — NOT loaded 50.66 GiB — per-layer "PLE" table Read on demand, per the engine's own log line: "CUDA left 50.66 GiB SSD-resident" Trades ~0 GB of GPU memory for occasional on-demand reads — a bet that this table is read far less often than the resident weights. on demand Our A/B below tested whether that "on demand" path is sensitive to DGX OS's default NVMe interrupt coalescing. Short answer: not for this prompt.

The engine: cuda.fast’s ds4

A quant file is just numbers until something knows how to run it fast. Mooney’s speed comes from ds4, a CUDA inference engine originally built by the cuda.fast project (vendored from Layr-Labs/ds4) and specifically tuned for GB10 — the chip inside a DGX Spark. Daniel didn’t write ds4 from scratch; he forked it (cudafast-qwen38-125b-a6b-engine, branch lbf/pq2-rot, pinned at commit a77755c6 on top of upstream 5707d4f2) and added exactly the pieces Mooney needed on top of it.

ds4: base engine + Mooney-specific fork What came from upstream cuda.fast vs. what Daniel added on top cuda.fast / ds4 (Layr-Labs), upstream @ 5707d4f2 CUDA inference engine tuned for GB10 (sm_121) — the base everyone forks from Handles weight loading, attention kernels, the serving loop Daniel’s fork: branch lbf/pq2-rot @ a77755c6 PQ2_0 (GGML type 142) — ternary expert loader + matching matmul kernels lowbitflash.rot.* — fused rotation (FWHT) + sign + int8 dequant kernels MTP draft head support — speculative decoding, --mtp-draft 2 pq2_0-specialized MoE gate/up + down kernels (+5.2% decode) Plus: fixed a pre-existing upstream GB10 bug, several use-after-free/overread fixes

Three pieces of that fork are worth understanding on their own, because they explain why Mooney is fast and not just small:

PQ2_0 — a new ternary weight type

Ggml-style engines identify tensor formats by a type number; Daniel registered a new one, type 142, for his 2.125-bit ternary experts (each weight ≈ −1, 0, or +1, scaled per block). The loader and the matmul kernels both had to learn this format — it didn’t exist in upstream ds4 before this fork.

lowbitflash.rot.* — rotate before you round

Naively rounding weights to 3 values loses a lot. His kernels apply a fused Walsh-Hadamard rotation plus a sign flip before the int8 dequant step — a known trick (also used by rotation-based quant methods like QuIP/QuaRot) for spreading outlier weight values more evenly so low-bit rounding hurts less. It’s fail-closed: the engine checks the rotation metadata’s version, segment sizes, and signs before trusting it.

MTP — a small model drafts, the big model checks

MTP (multi-token prediction) is speculative decoding: a small 2.58 GiB draft head guesses the next couple of tokens, and the full Mooney model verifies them in one batched pass instead of generating one token at a time. When the guesses are right — which is often, for easy continuations — you get multiple tokens per forward pass. That’s the direct cause of the jump from 33.3 tok/s (serial) to 45.8 tok/s (MTP) on Daniel’s numbers.

The vendoring choice itself

Shipping Mooney as a standalone fork repo (DJLougen/cudafast-qwen38-125b-a6b-engine) rather than a patch file keeps upstream cuda.fast’s license and commit history intact, and lets setup_spark.sh clone it directly by URL — the same reproducibility habit as the manifest-driven hash verification elsewhere in the project.

All of this is cross-checked against docs/PINS.md and engine/README.md in the mooney-spark repo, which name the exact base commit, the exact fork commit, and list every delta between them — not just prose claims.

Where the quality actually goes

The part of Daniel's thread we liked best wasn't the headline number — it was that he didn't stop at one. His 536-task verifier suite scores Mooney at 88.71% against BF16's 92.91%, which is the "95.5% of BF16" figure everywhere, but he broke out where the 4.2 points went instead of hiding it in an aggregate:

Took the hit

Long reasoning chains, arithmetic, and coding. These lean on precise, compounding computation — exactly what you'd expect a 2.1-bit ternary expert format to struggle with.

Held flat (or improved)

Tool calls, JSON/schema adherence, long instructions, multi-turn state, and output formatting. Short instructions actually scored better than BF16 in his suite.

Long-context retrieval held up too: needle-in-haystack was 30/30 across 32k–128k (identical to BF16), and 5/5 at every depth he tested at 256k with the fast runtime. Quality loss didn't compound with context length — a real risk with aggressive quantization that didn't materialize here.

If you only ever ask "what's the aggregate score," you can't tell whether a quant is safe for your workload. If your agent lane lives and dies by tool-calling and JSON, Mooney's 95.5% headline undersells it — that category barely moved. If it lives and dies by arithmetic, the headline oversells it. The category breakdown is the actual answer; the aggregate is just a summary of it.

We rebuilt it: setup, hashes, boot log

We didn't just read the thread. We cloned DJLougen/mooney-spark onto one of our own idle DGX Sparks and ran his one-command setup for real:

git clone https://github.com/DJLougen/mooney-spark
cd mooney-spark
./setup_spark.sh --mtp-source upstream

What that actually does, end to end: preflight-checks the GPU (must be GB10/sm_121) and CUDA toolkit, builds his ds4 engine fork at a pinned commit (a77755c6, CUDA_ARCH=sm_121), downloads all four GGUF shards plus the vision projector and MTP head from Hugging Face, verifies every single file's size and SHA-256 against a manifest fetched from the repo at run time, and only then writes a launcher script. Any hash mismatch aborts before anything runs — a genuinely fail-closed design, which matters when you're pulling 90 GB of someone else's quantized weights onto a box you intend to trust.

StepResult on our Spark
Engine build (sm_121)Clean build, ds4-server binary produced, HEAD matched the pinned commit
Shard 1 (1.5 GB)Size + SHA-256 verified
Shard 2 — the 54.4 GB PLE table (52 GB)Size + SHA-256 verified
Shards 3–4 + mmproj + MTP headSize + SHA-256 verified, ~90 GB total on disk
Server bootListening on 127.0.0.1:8000, OpenAI-compatible, bound to loopback only

Setup-to-serving took under 20 minutes end to end on our connection, almost all of it the download. The server binds to 127.0.0.1 by default — it never touched our LAN.

The A/B we ran: NVMe coalescing vs the SSD table

Last week we found that DGX OS ships a nvidia-nvme-interrupt-coalescing.service that sets every Spark's NVMe drive to batch small reads — great for throughput, bad for latency on uncached random 4K reads (we measured 8–10µs p50 with it off vs ~200µs with it on, a ~20× hit). A 50 GB table read on demand from SSD looked like the exact shape of workload that would eat that penalty, so we ran the discriminating test:

ConditionFixed prompt, 300 tokens, 3 reps each
Coalescing ON (DGX OS default, feature 8 = 0x107)5.90s, 5.97s, 5.90s
Coalescing OFF (nvme set-feature -f 8 -v 0)5.91s, 5.91s, 5.91s
Result: no measurable difference. We expected the coalescing penalty to show up here and it didn't. Checking /proc/diskstats, our fixed prompt triggered only about 7 NVMe reads per request — nowhere near the thousands of small random reads that would actually expose a per-read latency penalty. Either the table is consulted far less often per token than we assumed, or enough of it gets cached in the Spark's unused memory headroom (121 GiB total, ~89 GiB resident model+KV, leaving ~30 GiB free) that a short prompt mostly hits page cache instead of disk.

That's a real finding, just not the one we went in looking for. It means our NVMe-coalescing concern doesn't apply to this specific workload shape on the fast ds4 runtime — but it's a one-prompt test, not a clearance. The open question we're leaving in our research queue: does a longer generation, a cold start (page cache dropped), or many concurrent sessions push enough reads through that table to make coalescing visible? We restored coalescing to its DGX OS default before closing out.

What we'd steal for our own fleet

We don't have a 744B model we can safely push to 2-bit experts — our GB300 lane runs an NVFP4 incumbent we're not touching, and "IQ2-class quant on a 744B production model" is a rule we set for ourselves after past low-bit experiments went badly. But three ideas here are genuinely portable:

Per-category quality gates, always

This is the one we're adopting outright, independent of anything Spark-specific. Every quant or serving-config change we publish from now on reports quality broken out by task category — reasoning, arithmetic, coding, tool calls, JSON, multi-turn, formatting — not just one aggregate number. Daniel's suite is the model for the shape of that table.

Weights-on-SSD as a design pattern, not just a Spark trick

"Leave the part you read least often on the fastest storage you're not using" is the same idea behind our own Grace-offload / Engram-host work on the GB300 — just applied to a different memory tier (NVMe vs CPU DRAM). Worth re-reading his engine's paging logic for ideas, even though the specific numbers don't transfer.

A real stress test for the coalescing finding

Our one-prompt A/B came back clean, which is suspicious in a good way — it means the real test (cold cache, long context, concurrent sessions) hasn't been run yet. That's queued as follow-up work, on our hardware, before we conclude anything either direction.

A spare-capacity use for Mooney itself

At ~57 GiB resident with 256k context and MTP, Mooney comfortably shares a Spark with other light work. Worth considering as a third model option on our Spark pair for jobs where 95.5%-of-BF16 quality is an acceptable trade for freeing a Spark's remaining ~60 GiB — not a GB300 candidate, but a legitimate option for the auxiliary tier.

Credit where it's due: this release leans on real upstream work — Layr-Labs' cuda.fast ds4 engine, the PrismML Bonsai quantization line and llama.cpp fork, and compute from Lambda and Zach Mueller. Daniel's own contribution is the ternary PQ2_0 expert format, the rotation-based dequant kernels, the SSD-paging memory plan, and the per-category verifier — and he shipped the whole thing as a reproducible, hash-verified one-command setup instead of just a number on a chart. That reproducibility is why we were able to check his claims on our own hardware the same afternoon instead of just trusting the thread.

Provenance note: every number in the "what Mooney is" and "quality" sections above is from Daniel's public thread and model card, cited inline. Every number in the "we rebuilt it" and "A/B" sections is from our own boot log and curl timings on our Spark, October 1, 2026.