By Simon Gonzalez de Cruz. 2026-08-20. Canon: FACTS-PACK; research lane: DSH/ICCO.
This is a methods note, not a benchmark. We are publishing our calibration procedure because the interesting number is not a single tok/s or a single score — it is the knee where paying for more effort stops buying accuracy.
The method. Take a fixed problem set, split by difficulty. Run each arm twice: once with the cheap configuration, once with the expensive one, everything else identical (same model, same quant, same temperature, same grader). Plot accuracy against cost. The point where the expensive arm stops beating the cheap one is the knee. Below it, you are paying for nothing.
Our numbers. On a 40-problem hard/medium set, thinking-on rescued 15/40 hards vs 4/40 with thinking-off at temp 0 — a real gain exactly where the tasks are hard. On HumanEval-30, the same thinking budget bought nothing (28/30 identical). Full-sample: reasoning pays where tasks are hard, nothing where they are easy.
Why publish it. The industry shipped this shape (effort-as-dial in reasoning models, complexity routers in coding assistants), but nobody publishes the local calibration curve — the number that says "on this hardware, with this quant, thinking stops paying past this difficulty." That curve is hardware-specific, and it is what we measure before we ship a default.
The rule we run. Think-off by default. Think-on per request when the task is hard-class. Let the data, not the hype, decide the default. If a model ever auto-thinks, it must keep a manual override — every shipped auto-think without an escape hatch produces the same complaint: burnt tokens, no better output.