Benchmarks
Benchmark scorecards, methodology notes, throughput reports, and non-comparable result boundaries.
- DeepSeek-V4-Flash-Vision-Exp on One GB300: Native FP4, SGLang, DSpark, 1M Context September 3, 2026
Single DGX Station TP=1, no requantization: 2,420 tok/s at C32, exact recall to 810K, zero repetition at C64. Recipe + harness on GitHub. - Vision-Exp thinking levels on 2× DGX Spark September 1, 2026
Matched off/low/high/max thinking probes. Effort labels are not a ladder. - NVFP4 vs EXL3: GLM-5.3-Flash on Two DGX Sparks August 29, 2026
Matched cold-prefill / decode / concurrency probes of two 4-bit quants on the same Spark pair. - GLM Flash 5.3 Testing — Dual DGX Spark August 27, 2026
Day-0 SGLang bring-up of GLM-5.3-Flash NVFP4: GB10 memory trap, 20-hop gauntlet, 24.7 tok/s prose. - DeepSeek Flash 0731 Testing: weights vs thinking level August 26, 2026
2×2 A/B on dual Sparks: official vs ablit weights, low vs high thinking. Neither explained the tangents. - Local LLM Testing — Terminal-Bench 2, June 2026 June 22, 2026
A narrowed benchmark page focused on one measurement regime: Terminal-Bench core, terminus-2, native tool calling, and comparable local model results. - GLM-5.2: Terminal-Bench Benchmark — MLX vs GGUF June 22, 2026
GLM-5.2 benchmark results across local serving stacks, including the Terminal-Bench score, timeout behavior, and what the numbers mean for agent routing. - DS4-F Under Three Lights: Tool Discipline, Throughput, and Hermes Fit June 25, 2026
A three-part read on the live DeepSeek V4 Flash route: tool discipline, throughput, and whether the endpoint is a good Hermes fit. - The Tool-Calling Benchmark: 9 Models, Local vs Cloud April 12, 2026
Seven models, same 20 prompts, deterministic scoring. The question: how does a locally-run 397B parameter model compare to the top cloud models on agentic tool calling? The answer was surprising. - MiniMax M2.7 vs Qwen3.5-397B vs Claude Sonnet 4.6: Tool Calling on Apple Silicon April 12, 2026
Three models, same benchmark. Two run locally on a Mac Studio M3 Ultra. One is Claude Sonnet 4.6 via API. How close can local get to cloud on agentic tool calling? - Making an Agentic Benchmark Modeled on Doing Agentic Benchmarks April 12, 2026
Most benchmarks are single-shot snapshots that rot the moment you change hardware or models. Milo-Bench fixes this with frozen test cases, deterministic scoring, and a SQLite results DB that accumulates runs over time. 27 tests across 6 categories, open source. - Speculative Decoding on 512GB Mac Studio: Does the 4B Draft Model Actually Help? April 12, 2026
Long reasoning tasks: +58% speedup. Large-context tool calls: -88%, catastrophic. The answer depends entirely on what you are asking the model to do.