GLM-5.3 is the big sibling of the GLM-5.3-Flash we usually run. Flash is 320B parameters and fits in GPU memory. The full model is 744B. Even after NVIDIA's 4-bit NVFP4 quantization, the checkpoint is 465 GB (97 files; the README's "433 GB" is the same thing in GiB). The GB300's HBM, the fast memory the GPU actually computes from, is 250.7 GiB visible. The model is almost twice too big.
The DGX Station's trick is that the GPU is welded to a 72-core Grace CPU with 494 GiB of its own memory, and the two are joined by NVLink-C2C, a link fast enough that the GPU can read CPU memory directly. Not as fast as HBM, roughly 330–350 GB/s effective versus HBM's multiple terabytes per second, but fast enough that "the GPU reads some weights from the CPU side" is a serving strategy, not a failure.
If GLM-5.3 were a plain "dense" model, every one of its 744 billion weights would be touched on every token, and reading 465 GB over a 340 GB/s link per token would mean roughly one token per second. Unusable.
It is not dense. It is a mixture of experts. Each of the 75 MoE layers has 256 small "expert" networks plus one shared one, and for each token a tiny router picks only 8 of the 256. So only about 40B of the 744B parameters do work on any given token. Across the whole model that is 75 layers × 8 experts = 600 expert reads per token out of 19,200 possible.
That is what makes caching possible. If the 8 experts a layer wants are already sitting in HBM, the layer runs at full speed. If they are not, the GPU reads them over C2C. The whole recipe is about making the first case common.
Each MoE layer gets its own small set of slots in HBM, between 48 and 96 depending on how "spread out" that layer's routing was in a profiling trace; 7,360 slots in total. A slot holds one expert. When the router asks for an expert that is already in a slot, it is a hit. When it is not, the least-recently-used slot is evicted and the expert is copied in from Grace over C2C. That copy is the miss cost.
Two more pieces finish the recipe. First, MTP(2): the checkpoint ships a small "draft" head that guesses two tokens ahead, and the full model verifies all three in one pass. When the guesses are right (about 2.1–2.8 of 3 accepted, depending on whether it is prose or code), you get up to three tokens for one step's worth of expert misses. Second, the fused hook collapses the per-layer bookkeeping that moves experts into slots from three small GPU launches into one, worth +4.6% on its own.
At 54.7 tok/s with MTP accepting ~2 tokens per step, a decode step is about 40 ms. The recipe's September 15 profile splits it like this:
That pink bar is the whole story of every optimization attempt since. The copy kernel is already at the link's ceiling, so nothing about how we copy can help. We can only reduce how many misses there are (better placement) or start copies before the router asks for them (prediction). Every closed lever in the recipe is one of those two ideas failing a measurement.
Three experiments, all stop-and-keep on the Station, all against the promoted daily with byte-identical arguments except the one axis under test.
The daily runs --max-num-seqs 1, one user at a time. The recipe has listed "try 2–4" as the next untested axis since September 15, and GrokMilo raised it again this week. Today's Card N answered it.
| setting | 1 stream | 2 streams | 4 streams | time to first token (1 / 2 / 4) | output identical? |
|---|---|---|---|---|---|
| seqs 1 (daily) | 51.6 agg | — | — | 9.9 s | reference |
| seqs 2 | 53.7 | 58.2 agg (29/user) | — | 9.5 / 17.1 s | 20/20 greedy, TF Δ 0.0 |
| seqs 4 | 52.5 | 55.9 (28/user) | 26.4 agg (6.6/user) | 9.8 / 17.8 / 76.4 s | 20/20 greedy, TF Δ 0.0 |
conc_bench, 512-token answers, 3 scored reps after a discarded warm rep, same 8 prompts with a nonce so the prefix cache cannot help. Per-user = aggregate ÷ streams; the script's own per-stream column was broken (TTFT detector fired on an empty chunk) and is not used. The slot-cache hit/miss line captured was the single-stream window, so the mechanism below is inferred from the throughput cliff, not from a logged miss count.
Two users together get 58 tok/s total, about 13% more than one user alone, but each of them sees 29 instead of 54 and waits nearly twice as long to start. Four users is a collapse: 26 tok/s total, 76 s to first token. What is happening: four different prompts route to four different sets of experts, so each decode step has to copy roughly four times as many misses over the same saturated link, and the slot cache (sized for one user's locality) thrashes. The 8K-token prefill chunking queues the four prompts behind each other on top. Concurrency on this lane is closed with receipts. The daily stays at one user.
Halving the KV cache's precision would nearly double how much context fits. Both variants vLLM offers were tested with the full quality gate, same argv otherwise:
| KV dtype | KV tokens (vs bf16) | mean |Δ logprob| | tokens shifted > 1.0 nat | verdict |
|---|---|---|---|---|
| bf16 (daily) | 274,368 (1.00×) | 0.000 (repeat) | 0% | instrument proof: greedy 20/20 ×2, needle 18/18 |
| fp8_ds_mla | 470,848 (1.72×) | 0.538 | 16.2% | HARD NO |
| fp8_e4m3 | 532,288 (1.94×) | 0.185 | 5.0% | HARD NO |
The bar was mean ≤ 0.02 and zero tokens shifted by more than a full nat. Both failed by an order of magnitude. Needle retrieval still passed at 8K and 128K, so this is not "the model forgets things"; it is "the model's next-token distribution moved a lot," which is exactly what the teacher-forced gate exists to catch and exactly what eyeballing output would miss. Both runs logged that the checkpoint ships no KV scaling factors (scale 1.0), which is the likely reason the drift is this large. A properly calibrated KV quant would be a different, future candidate.
HelixML published a sharp finding on September 26 about GLM-5.3-Flash: its hybrid attention layers need saved state "snapshots" to reuse cached prompts, and one long prompt can evict every other conversation's snapshots while the token cache looks empty. We reproduced their probe on our Flash daily this morning. It did not happen here: six 20K-token sessions stayed warm (0.21 s) through a 300K prompt and a 12-session loop, on stock settings. Our snapshot pool is bigger relative to traffic (48 slots, 9 running, one box) than theirs (28 slots, 4 running). Their fix (--mamba-max-states-per-path 2) cost us +50% on branch-from-the-middle latency for no gain. Not adopted; worth knowing it exists if the Flash lane ever serves many long sessions at once.
GrokMilo (James's other agent, on the Cursor side) sent six speedup ideas through our new encrypted drop this week. Here is each one against what the ledger already knows. This is the honest part: four of six are already dead, and the two that are alive are alive for specific, checkable reasons.
| idea | what the receipts say | verdict |
|---|---|---|
| Learned expert predictor + async prefetch (SpecPrefetch-style: a small model guesses next step's experts and starts copying them early; the real router still decides) | Not tried in this form. But the bar is high: LRU already hits 0.802 on exact replay, and every wrong prediction costs C2C bytes we do not have. Our September 7 offline screen of simpler predictors was negative (adjacent-layer precision 0.084). The MTP-draft variant was retracted: the draft head has its own router, zero correlation with the main layers. | OPEN, needs offline proof |
| Pack active slots into contiguous HBM | The 18 ms is bytes crossing C2C, not scattered HBM reads. Packing where experts land in HBM does not change how many must be fetched. The kernel is already at the link ceiling. | REJECT, wrong mechanism |
| FP8 KV cache | Measured September 30, both dtypes, HARD NO on the quality gate (table above). | CLOSED |
| 2–4 concurrent streams | Measured today (Card N). +13% aggregate at 2 for half the per-user speed; collapse at 4. | CLOSED |
| Move to vLLM's upstream expert offload (#37190 CachedWeightProvider, RFC #38256) | The right long-term home for the hook, and already in the ledger as the fork-retirement evaluation. Not a speedup; a maintenance bet. | PARKED |
| Dead-ideas list (prev-step prefetch, agent remap, draft correlation, KV offload, launch coalescing) | Matches our receipts exactly. Good carry-forward. | agreed |
One correction sent back: the note said our Flash recipe "already uses FP8 KV." It does not; on GLM-5.3-Flash the MLA KV stays bf16 even when you ask for fp8 (sglang #36830), and only the small DSA indexer pool is fp8.
Everything above is measured. This section is not. It is where I think the next real gain could come from, ranked by how much I would bet on it, with what it would take to find out.
A. The learned predictor is the only idea that attacks the whole 18 ms. GrokMilo is right that SpecPrefetch's design (predict for transfer only; the real router still decides what runs) means wrong guesses cost speed, never correctness. That matters here because our quality gate is strict and this design passes it by construction. What I would do: train a tiny predictor on the frozen 71K-token routing trace that already lives in the recipe's research/ directory, replay it offline against the same trace with LRU as the baseline, and count (misses hidden − bytes wasted on wrong guesses). If that number is not clearly positive on the held-out half, it does not get a Station boot. If it is, the card is one axis: hook with prefetch on versus off, same greedy/TF gates, and the SLOT_CACHE STATS line tells us in real time whether misses per step actually dropped.
B. A calibrated FP8 KV is a slot play, not a context play. The September 30 HARD NO was for an uncalibrated quant (scale 1.0, no k/v scales in the checkpoint). If we generate proper scales on a calibration set, the quality might clear the gate, and then the interesting move is not "more context" but "same 256K context, 12 GiB less KV, ~800 more expert slots." The recipe already measured the exchange rate (24 GiB of KV buys 1,568 slots), and the context-vs-slot curve shows decode tracks misses per layer-step closely (256K/7,360 slots → 51.3 tok/s; 1M/2,672 slots → 34.1). More slots is a direct lever on the pink bar. The risk is that the calibration does not clear the gate either; the fp8_e4m3 drift (0.185 mean) is a lot to recover.
C. More tokens per miss is cheap to check and probably small. MTP(2) is the current drafter. MTP(3) is one config change and one boot: if code acceptance stays near 2.8 of 4 instead of 2.8 of 3, that is more tokens for the same expert copies. DFlash2 as a drafter already failed on this lane (accepted length 1.57 against a 3.0 bar), so I do not expect a big win here, but the test is nearly free.
What I would not do: more cache-policy tuning (LRU is at the ceiling), more kernel work on the copy (at the link ceiling), concurrency (closed today), or anything that trades output fidelity for speed without going through the teacher-forced gate first. The 744B's reason to exist on this box is that it is the smartest thing we can run locally. Fifty-five tokens a second of the real model beats two hundred of a model that has quietly drifted.
:30001 at the daily argv (--max-num-seqs 1, 256K, bf16 KV, MTP(2)); the Flash daily is stopped and kept. That is James's call for the evening, not a recipe change. Receipts for today's cards are on the box under glmf/cardM-2026-10-01 and glm53-big-opt-20260930/cardN-2026-10-01.