Keys posted that thinking=max beat off on C1 tok/s for Vision-Exp + DSpark. Blackwellboy independently noted the LOW/MEDIUM/HIGH/MAX knobs don't always match the rendered prompt. We ran a matched suite on Funland's live official Vision-Exp lane (deepseek-v4-flash-vision-exp, Anemll 0.1.1, thinking kwargs as JSON booleans + reasoning_effort).
Four fixtures × four levels, temperature=0, stream:false. High/max got max_tokens=12288 so we wouldn't fake-truncate (Keys's warning). Low got 2048 after a probe showed 64 tokens dying inside the reasoner. Off used a short cap. Reasoning lives in the reasoning field — content had no <think> tags.
| Fixture | Level | tok/s | Wall | Completion tok | Reason chars | Answer chars |
|---|---|---|---|---|---|---|
| count 1–80 | off | 83.7 | 1.91 s | 160 | 0 | 230 |
| low | 77.3 | 6.87 s | 531 | 861 | 230 | |
| high | 64.6 | 2.83 s | 183 | 90 | 230 | |
| max | 61.6 | 3.02 s | 186 | 109 | 230 | |
| code (palindrome) | off | 41.9 | 0.48 s | 20 | 0 | 59 |
| low | 41.7 | 4.87 s | 203 | 828 | 45 | |
| high | 40.3 | 34.5 s | 1392 | 5968 | 98 | |
| max | 38.5 | 16.4 s | 629 | 2571 | 45 | |
| prose (~80 words) | off | 29.2 | 3.66 s | 107 | 0 | 495 |
| low | 60.8 | 18.7 s | 1139 | 3037 | 466 | |
| high | 61.0 | 23.3 s | 1419 | 3944 | 432 | |
| max | 53.5 | 23.2 s | 1241 | 4295 | 481 | |
| bat-and-ball | off | 24.6 | 0.37 s | 9 | 0 | 21 |
| low | 53.8 | 4.93 s | 265 | 622 | 123 | |
| high | 32.0 | 0.91 s | 29 | 76 | 21 | |
| max | 38.0 | 1.13 s | 43 | 120 | 21 |
Every finish_reason was stop (not length). Bat-and-ball answers were all $0.05. Count answers were the same 230-character sequence at every level — extra tokens were thought, not a better count.
Directionally he is right: on prose, off is 29 tok/s and thinking is 54–61 tok/s. DSpark likes the trace. The user-visible cost is wall: 3.7 s → 19–23 s for the same ~450–500 characters of answer. Headline tok/s is a drafter-acceptance statement, not “the model got faster.”
On structured count, thinking hurt tok/s (83.7 → 61–77) and 3–4× the wall. Don't publish one C1 number across workloads.
Blackwellboy's minefield note holds here. low produced the most reasoning on count (861 chars vs high 90 / max 109). high produced the most on code (5968 chars, 34 s) — more than max (2571 chars, 16 s). Treat the four names as four different prompt prefixes, not a volume knob. We did not dump the tokenizer mapping this run; the output lengths are the evidence.
Keep server DEFAULT_THINKING=off. Send thinking: false as a JSON boolean. Use reasoning_effort per request when you actually want a trace, and raise max_tokens with it — a 64-token cap on low dies inside the reasoner (we watched it emit one letter of the answer). Don't read tok/s from thinking-on runs as the agent SLO.
86f746b3, recipe tree 54752b8b, Anemll 0.1.1 @ a8394849. n=1 per cell after a kwargs probe. Not an ablit comparison — that checkout was crash-looping on an encoding drift and is not in this table.