al-engr.com · lab note · August 23, 2026

Local AI is catching up

Same Hermes turn, four lanes, one session. Local is in the pack. It does not sweep the board. Cloud still wins some fixtures outright. The interesting part is not a headline speed crown. It is that the dual-Spark box no longer feels like a second-class agent once you stop staring at peak tok/s.

What we measured

August 23, 2026. Single session. The M4 Max issued every call. The local lane is DeepSeek V4 Flash 0731 (DS4-F) on 2× NVIDIA DGX Spark, TP=2 over 200G RoCE, Anemll vLLM, 1M context, MTP speculative decode. Felt time is an identical hermes chat -Q -t safe harness turn. Cloud lanes ran at HIGH reasoning effort. Seconds, lower better. n is 1–2 per cell. Smoke-grade, not a bake-off.

Raw local decode · not the agent clock

Count-to-300, stream:false, usage.completion_tokens / wall: 899 tokens / 9.79s = 91.8 tok/s, then 899 / 9.87s = 91.1 tok/s. That is a peak fixture. It is not the sustained agent mix, and it is not what a person waits on.

Felt wall-clock

The number that matters here is time-to-finished-turn through the real harness. CLI spin-up, config and skill load, and a ~10K-token system prompt prefill sit under every lane. That floor is about 15 seconds. It dominates the table.

Felt agent wall-clock in seconds, lower better Grouped bars for three fixtures. 1-word reply: grok 16.0, gpt 16.3, claude 18.3, local 18.9. count300 med/2: claude 24.9, local 25.1, gpt 28.0, grok 55.7 with runs at 32.5 and 78.9. code task: grok 22.8, claude 24.0, gpt 26.2, local 29.9. A dashed line marks the roughly 15 second harness floor. Scale is 0 to 80 seconds. Felt agent wall-clock (seconds, lower better) Scale 0–80s. Grok count300 also shows the two raw runs (32.5s, 78.9s). 0 20 40 60 80 ~15s harness floor 18.9 16.0 16.3 18.3 1-word reply 25.1 55.7 28.0 24.9 78.9 32.5 count300 med/2 29.9 22.8 26.2 24.0 code task ds4-local grok-4.6 hi gpt-5.6-sol hi claude-fable-5 hi grok raw pair
Fastest cell in each row is highlighted. Grok count300 is the mean of 32.5s and 78.9s on the identical prompt.
fixture ds4-local grok-4.6 hi gpt-5.6-sol hi claude-fable-5 hi
1-word reply 18.9 16.0 16.3 18.3
count300 med/2 25.1 55.7* 28.0 24.9
code task 29.9 22.8 26.2 24.0

* grok count300 runs were 32.5s and 78.9s — large tail variance on the identical prompt.

Reading it without the hype

Local does not win a single fixture on this board. Grok takes the one-word reply (16.0s) and the code task (22.8s, still at high reasoning). Claude takes count300 by two tenths (24.9 vs local 25.1). That is a near four-way tie once you remember every lane is sitting on a ~15s harness floor. The decode race is mostly hidden under CLI and prefill.

Local's real win is variance. The local spread we saw is about ±3s. Grok's two count300 runs on the same prompt were 32.5s and 78.9s. When cloud feels slow in this harness, it is usually a tail, not a worse median. A 55.7s cell that is really 32.5-or-78.9 is a different product than a 25.1s cell that stays near 25s.

Cloud still wins some fixtures outright. Grok's 22.8s code turn is the cleanest example: high reasoning on, still first. Catching up is not overtaking.

Caveats

What this is not

Not a claim that DS4-F is faster than grok-4.6, GPT-5.6, or Claude Fable 5. Not a quality score. Not a cost model. Not permission to retire the cloud lanes. The honest sentence is narrower: through this Hermes turn, on this day, local is close enough that tail latency is the remaining cloud tax, and peak tok/s is the wrong poster.

Round two: a real agent workload

After the single-turn fixtures, we added the workload that actually matters: a 20-step tool-call chain (each terminal call reveals the next file to read) plus a ~6K-token log to analyze — one Hermes turn, ~21 sequential tool hops, context growing every hop. All four lanes found the hidden token. The clock told a different story than round one:

Lane20-call chain + prefillvs fastestCorrect
gpt-5.6-sol high190.9s (3.2 min)1.0×yes
grok-4.6 high345.2s (5.8 min)1.8×yes
DS4-F local1024.6s (17.1 min)5.4×yes
claude-fable-5 high1254.6s (20.9 min)6.6×yes

Two honest observations. First: sustained agent work is where the cloud still runs away — per-hop re-prefill of a growing context is exactly what a local 1M-context MoE pays for, and peak decode tok/s predicts none of this. Second: local wasn't last. A frontier model at high reasoning effort thinks hard before every one of 21 hops, and that compounds worse than local prefill. n=1 per lane; treat the ordering as robust, the exact ratios as indicative.

Round three: the bench writes its own blog post

Final fixture, suggested mid-session: give all four lanes the identical brief — write this very post as a self-contained HTML file with a to-scale SVG chart, using only the measured data — then time them, and have the outputs blind-judged. Candidates were anonymized (shuffled labels, model names scrubbed) and scored by two judges from different model families (grok-4.6 and gpt-5.6-sol) on data fidelity, chart, design, prose, and structure. The judging family that authored this assembly (Anthropic) was excluded from judging. Mechanical gates ran first: all four produced valid HTML with balanced SVG.

LaneWrite timeBlind score (mean /50)
claude-fable-5 high86.8s41.5
DS4-F local121.7s34.5
grok-4.6 high288.5s45.5 — winner
gpt-5.6-sol high291.3s44.0

The article you are reading is the winning candidate — grok-4.6's blind-judged draft — with this section, the agent-workload section, and the acknowledgements added afterward. The judges' one-line on the local candidate: numbers faithful, chart legend wrong. Which is the whole story in miniature: the local model is in the pack on speed and correctness, and the remaining gap is polish.

Scoreboard across all three rounds

WorkloadWinnerLocal's position
Single-turn agent latencynear four-way tiein the pack, tightest variance (±3s)
20-call agent gauntletgpt-5.6-sol (3.2 min)3rd of 4 (17.1 min) — beat a frontier lane
Long-form writing, blind-judgedgrok-4.6 on quality; fable on speed2nd fastest, 4th on quality

Each workload shape crowned a different winner, and the two 20W ARM boxes in the closet were never embarrassed. That's what catching up looks like: not a sweep — a seat at the table, at zero dollars per token.

Thank you

The local lane's speed didn't come from this lab. It came from a small crew of builders and maintainers who spent the summer making dual DGX Spark inference actually fast, and published everything. This post's 91.8 tok/s stands on their work:

Open weights plus open recipes plus people like these: that's the mechanism by which local AI is catching up.