al-engr.com · lab note · August 23, 2026
Local AI is catching up
Same Hermes turn, four lanes, one session. Local is in the pack. It does not sweep the board. Cloud still wins some fixtures outright. The interesting part is not a headline speed crown. It is that the dual-Spark box no longer feels like a second-class agent once you stop staring at peak tok/s.
What we measured
August 23, 2026. Single session. The M4 Max issued every call. The local lane is DeepSeek V4 Flash 0731 (DS4-F) on 2× NVIDIA DGX Spark, TP=2 over 200G RoCE, Anemll vLLM, 1M context, MTP speculative decode. Felt time is an identical hermes chat -Q -t safe harness turn. Cloud lanes ran at HIGH reasoning effort. Seconds, lower better. n is 1–2 per cell. Smoke-grade, not a bake-off.
Raw local decode · not the agent clock
Count-to-300, stream:false, usage.completion_tokens / wall: 899 tokens / 9.79s = 91.8 tok/s, then 899 / 9.87s = 91.1 tok/s. That is a peak fixture. It is not the sustained agent mix, and it is not what a person waits on.
Felt wall-clock
The number that matters here is time-to-finished-turn through the real harness. CLI spin-up, config and skill load, and a ~10K-token system prompt prefill sit under every lane. That floor is about 15 seconds. It dominates the table.
| fixture | ds4-local | grok-4.6 hi | gpt-5.6-sol hi | claude-fable-5 hi |
|---|---|---|---|---|
| 1-word reply | 18.9 | 16.0 | 16.3 | 18.3 |
| count300 med/2 | 25.1 | 55.7* | 28.0 | 24.9 |
| code task | 29.9 | 22.8 | 26.2 | 24.0 |
* grok count300 runs were 32.5s and 78.9s — large tail variance on the identical prompt.
Reading it without the hype
Local does not win a single fixture on this board. Grok takes the one-word reply (16.0s) and the code task (22.8s, still at high reasoning). Claude takes count300 by two tenths (24.9 vs local 25.1). That is a near four-way tie once you remember every lane is sitting on a ~15s harness floor. The decode race is mostly hidden under CLI and prefill.
Local's real win is variance. The local spread we saw is about ±3s. Grok's two count300 runs on the same prompt were 32.5s and 78.9s. When cloud feels slow in this harness, it is usually a tail, not a worse median. A 55.7s cell that is really 32.5-or-78.9 is a different product than a 25.1s cell that stays near 25s.
Cloud still wins some fixtures outright. Grok's 22.8s code turn is the cleanest example: high reasoning on, still first. Catching up is not overtaking.
Caveats
- High reasoning is a designed handicap on the cloud lanes. This is not local-vs-cloud at matched thinking budget.
- n = 1–2 per cell. Smoke-grade. Do not promote any cell into a ranking.
- 91.8 tok/s is a peak count-to-300 fixture, not a sustained agent mix.
- Felt time includes a ~15s floor shared by every lane. Subtracting it would make the remaining gaps look larger and less honest about what a person waits on.
- One session, one harness, one local stack (Anemll / 1M / MTP / TP=2). Different prompts, timeouts, or reasoning settings would be a different measurement.
What this is not
Not a claim that DS4-F is faster than grok-4.6, GPT-5.6, or Claude Fable 5. Not a quality score. Not a cost model. Not permission to retire the cloud lanes. The honest sentence is narrower: through this Hermes turn, on this day, local is close enough that tail latency is the remaining cloud tax, and peak tok/s is the wrong poster.
Round two: a real agent workload
After the single-turn fixtures, we added the workload that actually matters: a 20-step tool-call chain (each terminal call reveals the next file to read) plus a ~6K-token log to analyze — one Hermes turn, ~21 sequential tool hops, context growing every hop. All four lanes found the hidden token. The clock told a different story than round one:
| Lane | 20-call chain + prefill | vs fastest | Correct |
|---|---|---|---|
| gpt-5.6-sol high | 190.9s (3.2 min) | 1.0× | yes |
| grok-4.6 high | 345.2s (5.8 min) | 1.8× | yes |
| DS4-F local | 1024.6s (17.1 min) | 5.4× | yes |
| claude-fable-5 high | 1254.6s (20.9 min) | 6.6× | yes |
Two honest observations. First: sustained agent work is where the cloud still runs away — per-hop re-prefill of a growing context is exactly what a local 1M-context MoE pays for, and peak decode tok/s predicts none of this. Second: local wasn't last. A frontier model at high reasoning effort thinks hard before every one of 21 hops, and that compounds worse than local prefill. n=1 per lane; treat the ordering as robust, the exact ratios as indicative.
Round three: the bench writes its own blog post
Final fixture, suggested mid-session: give all four lanes the identical brief — write this very post as a self-contained HTML file with a to-scale SVG chart, using only the measured data — then time them, and have the outputs blind-judged. Candidates were anonymized (shuffled labels, model names scrubbed) and scored by two judges from different model families (grok-4.6 and gpt-5.6-sol) on data fidelity, chart, design, prose, and structure. The judging family that authored this assembly (Anthropic) was excluded from judging. Mechanical gates ran first: all four produced valid HTML with balanced SVG.
| Lane | Write time | Blind score (mean /50) |
|---|---|---|
| claude-fable-5 high | 86.8s | 41.5 |
| DS4-F local | 121.7s | 34.5 |
| grok-4.6 high | 288.5s | 45.5 — winner |
| gpt-5.6-sol high | 291.3s | 44.0 |
The article you are reading is the winning candidate — grok-4.6's blind-judged draft — with this section, the agent-workload section, and the acknowledgements added afterward. The judges' one-line on the local candidate: numbers faithful, chart legend wrong. Which is the whole story in miniature: the local model is in the pack on speed and correctness, and the remaining gap is polish.
Scoreboard across all three rounds
| Workload | Winner | Local's position |
|---|---|---|
| Single-turn agent latency | near four-way tie | in the pack, tightest variance (±3s) |
| 20-call agent gauntlet | gpt-5.6-sol (3.2 min) | 3rd of 4 (17.1 min) — beat a frontier lane |
| Long-form writing, blind-judged | grok-4.6 on quality; fable on speed | 2nd fastest, 4th on quality |
Each workload shape crowned a different winner, and the two 20W ARM boxes in the closet were never embarrassed. That's what catching up looks like: not a sweep — a seat at the table, at zero dollars per token.
Thank you
The local lane's speed didn't come from this lab. It came from a small crew of builders and maintainers who spent the summer making dual DGX Spark inference actually fast, and published everything. This post's 91.8 tok/s stands on their work:
- Anemll — the two-node GB10 port of DeepSeek V4 Flash to vLLM: the FlashInfer bridge, the B12X MXFP4 MoE backend, and the reproducible Docker deployment this lab runs in production.
- Mia (MiaAI-Lab) and contributors plotarmordev, de1tydev, sethforprivacy, and 0xSero — the maintained 0731 recipe for 2× Sparks that our production lane is pinned to. Maintainers who keep a recipe current week after week are the quiet engine of all of this.
- tonyd2wild — the 1M NVFP4-KV recipe and relentless public benchmarking that set the parity bars this lab tunes against.
- Weschera — the qualified vLLM TP2 + DSpark community lane for official 0731.
- Daegwon "nacyot" Kim — found the vLLM GB10 spin-wait bug (CPU cores spinning in shm_broadcast while decoding), worth up to −24°C SoC when fixed.
- drowzeys (Keys) — the English guide and one-command patcher for nacyot's fix, plus the anchored-tensor 0731 weights the production lane serves.
Open weights plus open recipes plus people like these: that's the mechanism by which local AI is catching up.