Try Hanzo

Benchmarks / GPQA-Diamond

measured

The router on GPQA-Diamond

One set of 198 graduate-level questions, answered by three routing tiers that differ only in what they are willing to spend: 194, 190 and 184 correct, at $25, $20 and $4 per million output tokens.

dataset
GPQA
licence
CC BY 4.0
introduced by
GPQA: A Graduate-Level Google-Proof Q&A BenchmarkRein, Hou, Stickland, Petty, Pang, Dirani, Michael, Bowman, COLM 2024
harness
hanzo.ai/enso

What is measured

GPQA-Diamond is the hardest subset of GPQA: 198 multiple-choice questions in biology, physics and chemistry, written by domain PhDs and filtered to the ones that other PhDs get right and skilled non-experts get wrong even with the web open. It is a knowledge benchmark rather than an agentic one — one question, one answer, no tools — which is what makes it useful here: it isolates the choice of model from everything else a system does.

What the router adds is that the choice is per question. Enso reads each one and dispatches it, so a tier is not a model — it is a pool and a willingness to spend. The three tiers answer the same 198 questions and differ only in that budget, which is why their scores and their prices belong in one table and why a score quoted without its price says nothing about the tier it came from.

198 questions · accuracy and price as hanzo.ai/enso publishes them, read 2026-09-11 · sorted by accuracy

modelGPQA-Diamondcorrect of 198output $/MTokscored by
enso-ultra98.0194$25Hanzo-measured
enso96.0190$20Hanzo-measured
gpt-5.593.6$8.25vendor-reported
gpt-5.2-pro93.2$138.60vendor-reported
enso-flash92.9184$4Hanzo-measured
gpt-5.6-sol92.9$25Hanzo-measured
kimi-k2.689.1$2.71vendor-reported
qwen3.5-397b-a17b88.4$2.04vendor-reported
opus-4.887.4$21Hanzo-measured
glm-5.285.6$3.73vendor-reported
gemma-4-31b84.3$0.44vendor-reported
fable-581.3$42Hanzo-measured

The top row is not the cheap row, and the table is arranged so that shows. enso-ultra is the highest-scoring row at 98.0, 4.4 points above gpt-5.5, and it costs 3.0× as much per million output tokens — so the lead is bought rather than free. The sharper row is lower down: enso-flash scores 92.9 at $4 against gpt-5.6-sol's identical 92.9 at $25, which is 6.25× the price for the same score.

Two kinds of evidence are in one table and the last column says which. A Hanzo-measured row went through our harness; a vendor-reported row is the model's publisher quoting itself, and the two are not the same claim. gpt-5.5 — the row our lead is measured against — is one of the reported ones. Counts appear only for the three tiers, because those are the only rows a count is published for; multiplying a rounded percentage by 198 would manufacture a figure nobody measured.

No public command prints this board. Every other lane on this page is either regenerated from committed runs by one line or says plainly that it is transcribed. This one is transcribed, from hanzo.ai/enso, and the harness behind it is not published. Until it is, read these as figures we stand behind and cannot hand you the means to check — which is a weaker position than the memory tables are in, and the difference is the point of saying so.

What this does not establish

A benchmark score is not a capability. GPQA-Diamond is four-way multiple choice on settled science. It measures what a pool knows and rewards a router that recognises a hard question; it says nothing about whether the answer is reasoned, whether the model holds up over a long autonomous run, or whether it can use a tool. The routing overhead is real work this table does not score either.

The field rows were not all measured the same way. Six of them came through our harness and six are the publishers' own figures, marked apart in the last column because the two are different evidence. A vendor-reported number is not re-run here, and where our lead is measured against one, the lead inherits that weakness.

Accuracy is not the only axis, and the board makes the trade visible rather than hiding it. enso-ultra leads at 98.0 and costs $25 per million output tokens; enso-flash is 5.1 points behind at $4. Which of those is the right row depends entirely on what a wrong answer costs in the work you are doing, and that is not a quantity any benchmark reports.

The other benchmarks

LoCoMo · MemoryAgentBench · LongMemEval · RepoBench-R · LoCoMo · subject scope · LoCoMo-Conv · Fleet residency · Live agent footprint · Sandbox cold start · Inference vs llama.cpp · all of them, and the head-to-head

figures as published at hanzo.ai/enso · the dataset: GPQA under CC BY 4.0