# GPQA-Diamond routing benchmark — Hanzo AI

> One set of 198 graduate-level questions through three Enso routing tiers: 194, 190 and 184 correct at $25, $20 and $4 per million output tokens — against the field each tier dispatches into, with Hanzo-measured and vendor-reported rows marked apart.

[Benchmarks](https://hanzo.ai/benchmarks) / GPQA-Diamond

measured

# The router on GPQA-Diamond

One set of 198 graduate-level questions, answered by three routing tiers that differ only in what they are willing to spend: 194, 190 and 184 correct, at $25, $20 and $4 per million output tokens.

dataset

[GPQA](https://huggingface.co/datasets/Idavidrein/gpqa)

licence

CC BY 4.0

introduced by

[GPQA: A Graduate-Level Google-Proof Q&A Benchmark](https://arxiv.org/abs/2311.12022) — Rein, Hou, Stickland, Petty, Pang, Dirani, Michael, Bowman, COLM 2024

harness

[hanzo.ai/enso](https://hanzo.ai/enso)

## What is measured

GPQA-Diamond is the hardest subset of GPQA: 198 multiple-choice questions in biology, physics and chemistry, written by domain PhDs and filtered to the ones that other PhDs get right and skilled non-experts get wrong even with the web open. It is a knowledge benchmark rather than an agentic one — one question, one answer, no tools — which is what makes it useful here: it isolates the choice of model from everything else a system does.

What the router adds is that the choice is per question. Enso reads each one and dispatches it, so a tier is not a model — it is a pool and a willingness to spend. The three tiers answer the same 198 questions and differ only in that budget, which is why their scores and their prices belong in one table and why a score quoted without its price says nothing about the tier it came from.

198 questions · accuracy and price as hanzo.ai/enso publishes them, read 2026-09-11 · sorted by accuracy

model

GPQA-Diamond

correct of 198

output $/MTok

scored by

enso-ultra

98.0

194

$25

Hanzo-measured

enso

96.0

190

$20

Hanzo-measured

gpt-5.5

93.6

–

$8.25

vendor-reported

gpt-5.2-pro

93.2

–

$138.60

vendor-reported

enso-flash

92.9

184

$4

Hanzo-measured

gpt-5.6-sol

92.9

–

$25

Hanzo-measured

kimi-k2.6

89.1

–

$2.71

vendor-reported

qwen3.5-397b-a17b

88.4

–

$2.04

vendor-reported

opus-4.8

87.4

–

$21

Hanzo-measured

glm-5.2

85.6

–

$3.73

vendor-reported

gemma-4-31b

84.3

–

$0.44

vendor-reported

fable-5

81.3

–

$42

Hanzo-measured

The top row is not the cheap row, and the table is arranged so that shows. enso-ultra is the highest-scoring row at 98.0, 4.4 points above gpt-5.5, and it costs 3.0× as much per million output tokens — so the lead is bought rather than free. The sharper row is lower down: enso-flash scores 92.9 at $4 against gpt-5.6-sol&#x27;s identical 92.9 at $25, which is 6.25× the price for the same score.

Two kinds of evidence are in one table and the last column says which. A Hanzo-measured row went through our harness; a vendor-reported row is the model&#x27;s publisher quoting itself, and the two are not the same claim. gpt-5.5 — the row our lead is measured against — is one of the reported ones. Counts appear only for the three tiers, because those are the only rows a count is published for; multiplying a rounded percentage by 198 would manufacture a figure nobody measured.

No public command prints this board. Every other lane on this page is either regenerated from committed runs by one line or says plainly that it is transcribed. This one is transcribed, from [hanzo.ai/enso](https://hanzo.ai/enso), and the harness behind it is not published. Until it is, read these as figures we stand behind and cannot hand you the means to check — which is a weaker position than the memory tables are in, and the difference is the point of saying so.

## What this does not establish

A benchmark score is not a capability. GPQA-Diamond is four-way multiple choice on settled science. It measures what a pool knows and rewards a router that recognises a hard question; it says nothing about whether the answer is reasoned, whether the model holds up over a long autonomous run, or whether it can use a tool. The routing overhead is real work this table does not score either.

The field rows were not all measured the same way. Six of them came through our harness and six are the publishers&#x27; own figures, marked apart in the last column because the two are different evidence. A vendor-reported number is not re-run here, and where our lead is measured against one, the lead inherits that weakness.

Accuracy is not the only axis, and the board makes the trade visible rather than hiding it. enso-ultra leads at 98.0 and costs $25 per million output tokens; enso-flash is 5.1 points behind at $4. Which of those is the right row depends entirely on what a wrong answer costs in the work you are doing, and that is not a quantity any benchmark reports.

## The other benchmarks

[LoCoMo](https://hanzo.ai/benchmarks/locomo) · [MemoryAgentBench](https://hanzo.ai/benchmarks/memoryagentbench) · [LongMemEval](https://hanzo.ai/benchmarks/longmemeval) · [RepoBench-R](https://hanzo.ai/benchmarks/repobench-r) · [LoCoMo · subject scope](https://hanzo.ai/benchmarks/locomo-subject-scope) · [LoCoMo-Conv](https://hanzo.ai/benchmarks/locomo-conv) · [Fleet residency](https://hanzo.ai/benchmarks/fleet) · [Live agent footprint](https://hanzo.ai/benchmarks/goroutine) · [Sandbox cold start](https://hanzo.ai/benchmarks/sandbox) · [Inference vs llama.cpp](https://hanzo.ai/benchmarks/inference) · [all of them, and the head-to-head](https://hanzo.ai/benchmarks)

figures as published at [hanzo.ai/enso](https://hanzo.ai/enso) · the dataset: [GPQA](https://huggingface.co/datasets/Idavidrein/gpqa) under CC BY 4.0
