Mia 489af95 is the new recipe pin. It is not faster.
DeepSeek-V4-Flash-0731
@ 9e165c30… on two GB10 Sparks, still
ghcr.io/anemll/dspark-vllm-gx10:0.1.1
digest sha256:a8394849….
We canaried
MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark
489af95 (today) against production 018c6bc (August 13).
Tony count 1→300 went 91.6 → 90.7 t/s (−0.9%). That is not a speed win.
We promoted anyway for fail-closed hotfixes, the #109 empty-encoder fix, and the GB10 busy_loop_s 1s→2ms patch.
What we actually tested
This GitHub repo is not a new checkpoint. Funland has been on that recipe since July. The question today was whether nine days of main past our August 13 pin was worth a sole-pair bounce.
Copy-checkout only. We did not git pull into 018c6bc. Funland knobs stayed: util 0.85, DEFAULT_THINKING=off, batch 8192, dynamic K 5/4/3, post-readiness warmup, no VL sidecar. Their README defaults we refused: thinking max, util 0.835, optional 16k/seqs=4 coding profile.
Independent 55-minute systemd deadman pointed at the August 13 tree. Arm → verify timer → stop production was one host command. Cold bind ~12 minutes. Both ranks had ~116 GiB free after stop.
What 489af95 adds
| Change | Why it matters here |
|---|---|
| Issue #109 empty encoder output | Stops phantom token 0 / empty-output loops that corrupt spec accounting. Live log: outputs=applied, scheduler=applied. |
| Fail-closed Python + transactional DSV4 hotfixes | A drifted patch no longer lets vllm serve start on stale code. |
| Issue #79 GB10 spin-wait | Default-on. Live: busy_loop_s 1 → 0.002. This was our parked next-restore item. |
| Issue #105 inflight-prefill parse | Malformed DSPARK_MAX_INFLIGHT_PREFILLS no longer crashes the admission loop. |
| Issue #31 GPU thinking budget | Now opt-in. We stay thinking-off / stock V2. Better than the old always-on cliff. |
Same Anemll 0.1.1 image. No new weights. Hermes route unchanged: spark-ds4 → deepseek-v4-flash-0731.
Same-session board
Loopback on Spark1, JSON thinking: false, Tony board uses stream:false and completion_tokens / wall. Keys is a different fixture — do not mix the two into one winner number.
| Fixture | 018c6bc baseline | 489af95 canary | Δ |
|---|---|---|---|
| Tony count 1→300 best / med | 91.6 / 91.19 | 90.74 / 90.29 | −0.9% |
| Tony count 1→150 best | 89.21 | 89.12 | flat |
| Tony repeat best | 92.96 | 92.39 | −0.6% |
| Tony short code best | 70.43 | 68.64 | −2.5% |
| Tony JSON best | 48.47 | 52.82 | +9% (noisy) |
| Spec accept on Tony board | 0.890 | 0.899 | +0.9 pt |
| Keys C1 best-of-2 | 65.5 | 61.8 | −5.6% |
| Keys C4 best-of-2 | 144.3 | 139.2 | −3.5% |
| Keys C6 best-of-2 | 172.3 | 162.2 | −5.9% |
| KV tokens @ 1M | ~2.24M (Aug 13 pin) | 2,161,287 | still ~2.1× |
stream:false. Peak decode did not move.
Gates
/v1/models→deepseek-v4-flash-0731,max_model_len=1048576- Exact + tools, thinking bool false:
ALL_GATES_OK - Hermes
spark-ds4:ROUTE_OK - Warmup pass-2: C4 mean K 4.01, C6 mean K 3.10, JIT lines 0
- Queues idle after benches
- Live cmdline still has
num_speculative_tokens_per_batch_size:[[1,1,5],[2,4,4],[5,6,3]]
What this is not
- Not a new local model. Official 0731 weights did not change.
- Not a Stage-C / 200K-16 / Keys Power Pack swap.
- Not permission to copy Mia’s
DEFAULT_THINKING=maxor util0.835. Funland 1M still needs 0.85. - Not a reopen of batch 16384. That canary already lost 32% single-stream.
Ops pin
checkout /home/milo/ds4-f-mia-anemll-0731-mia-aug22
sha 489af9589684127ece879dd9e53395c5ff9157c8
rollback /home/milo/ds4-f-mia-anemll-0731-mia-aug13 @ 018c6bc
older /home/milo/ds4-f-mia-anemll-0731 (Jul 29 Funland)
Public neighbors: August 13 apply · 0731 recipe.