August 22, 2026 · James Meadlock & Milo · Funland dual DGX Spark

Mia 489af95 is the new recipe pin. It is not faster.

Created

Created · Updated

LIVE Same model. Same image. Newer launcher. Official DeepSeek-V4-Flash-0731 @ 9e165c30… on two GB10 Sparks, still ghcr.io/anemll/dspark-vllm-gx10:0.1.1 digest sha256:a8394849…. We canaried MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark 489af95 (today) against production 018c6bc (August 13). Tony count 1→300 went 91.6 → 90.7 t/s (−0.9%). That is not a speed win. We promoted anyway for fail-closed hotfixes, the #109 empty-encoder fix, and the GB10 busy_loop_s 1s→2ms patch.
Tony count30090.7was 91.6 · stream:false
Spec accept0.899was 0.890 on the same board
Keys C4 / C6139 / 162was 144 / 172 · best-of-2
KV @ 1M2.16Mutil 0.85 · thinking off

What we actually tested

This GitHub repo is not a new checkpoint. Funland has been on that recipe since July. The question today was whether nine days of main past our August 13 pin was worth a sole-pair bounce.

Copy-checkout only. We did not git pull into 018c6bc. Funland knobs stayed: util 0.85, DEFAULT_THINKING=off, batch 8192, dynamic K 5/4/3, post-readiness warmup, no VL sidecar. Their README defaults we refused: thinking max, util 0.835, optional 16k/seqs=4 coding profile.

Independent 55-minute systemd deadman pointed at the August 13 tree. Arm → verify timer → stop production was one host command. Cold bind ~12 minutes. Both ranks had ~116 GiB free after stop.

What 489af95 adds

ChangeWhy it matters here
Issue #109 empty encoder outputStops phantom token 0 / empty-output loops that corrupt spec accounting. Live log: outputs=applied, scheduler=applied.
Fail-closed Python + transactional DSV4 hotfixesA drifted patch no longer lets vllm serve start on stale code.
Issue #79 GB10 spin-waitDefault-on. Live: busy_loop_s 1 → 0.002. This was our parked next-restore item.
Issue #105 inflight-prefill parseMalformed DSPARK_MAX_INFLIGHT_PREFILLS no longer crashes the admission loop.
Issue #31 GPU thinking budgetNow opt-in. We stay thinking-off / stock V2. Better than the old always-on cliff.

Same Anemll 0.1.1 image. No new weights. Hermes route unchanged: spark-ds4deepseek-v4-flash-0731.

Same-session board

Loopback on Spark1, JSON thinking: false, Tony board uses stream:false and completion_tokens / wall. Keys is a different fixture — do not mix the two into one winner number.

Fixture018c6bc baseline489af95 canaryΔ
Tony count 1→300 best / med91.6 / 91.1990.74 / 90.29−0.9%
Tony count 1→150 best89.2189.12flat
Tony repeat best92.9692.39−0.6%
Tony short code best70.4368.64−2.5%
Tony JSON best48.4752.82+9% (noisy)
Spec accept on Tony board0.8900.899+0.9 pt
Keys C1 best-of-265.561.8−5.6%
Keys C4 best-of-2144.3139.2−3.5%
Keys C6 best-of-2172.3162.2−5.9%
KV tokens @ 1M~2.24M (Aug 13 pin)2,161,287still ~2.1×
Keys C6 first rep had a 5.9 t/s straggler stream (113 agg) then 162. That harness is noisier than Tony. The speed oracle stays count 1→300 stream:false. Peak decode did not move.

Gates

What this is not

Ops pin

checkout  /home/milo/ds4-f-mia-anemll-0731-mia-aug22
sha       489af9589684127ece879dd9e53395c5ff9157c8
rollback  /home/milo/ds4-f-mia-anemll-0731-mia-aug13  @ 018c6bc
older     /home/milo/ds4-f-mia-anemll-0731            (Jul 29 Funland)

Public neighbors: August 13 apply · 0731 recipe.