Try Hanzo
Hanzo-measured · GPQA-Diamond, one common harness

Enso, measured

Enso is our frontier family, and the router that picks among 500+ models when you ask for auto. A vendor’s published score and your own are rarely the same number, so everything below was run here, on one harness, ours and theirs alike — and every figure still says which it is.

Instant Guest API Keyguest_temp_key
No account or sign up required

Test right now from your terminal. No credentials needed:

cURL · instant guest request
curl https://api.hanzo.ai/v1/chat/completions \
  -H "Authorization: Bearer guest_temp_key" \
  -d '{"model": "zen-free", "messages": [{"role": "user", "content": "hello"}]}'

Three tiers that stay in order

Ultra above Pro above Flash, on quality and on price, and it stays that way when the underlying models change. A tier is a contract about cost and quality, not another name for one model. GPQA is measured here; the price band is what you are billed, input then output, per million tokens.

Enso Ultra

Flagship
enso-ultra
98%GPQA$5 → $25

For work where a wrong answer is expensive: research, security review and long autonomous runs.

Enso Pro

Default
enso
96%GPQA$4 → $20

For coding, code review and agents that answer while someone waits.

Enso Flash

enso-flash
92.9%GPQA$2 → $4

For chat, extraction, classification and simple agent steps at volume.

Accuracy at cost

Accuracy alone picks the most expensive model every time, so both axes are here at once. Top-left is where you want to be: right answers, cheap. Each dot carries its lab’s mark; a solid ring means we measured it, a dashed ring means the vendor reported it. Every dot is labelled, and hovering gives the exact figure.

76828894100$0.5$8$120GPQA %cheap ← output $/MTok → expensivegpt-5.5 — 93.6% GPQA-Diamond · $8.25/MTok (vendor-reported)gpt-5.5 93.6%gpt-5.2-pro — 93.2% GPQA-Diamond · $138.6/MTok (vendor-reported)gpt-5.2-pro 93.2%gpt-5.6-sol — 92.9% GPQA-Diamond · $25/MTokgpt-5.6-sol 92.9%opus-4.8 — 87.4% GPQA-Diamond · $21/MTokopus-4.8 87.4%gemma-4-31b — 84.3% GPQA-Diamond · $0.44/MTok (vendor-reported)gemma-4-31b 84.3%fable-5 — 81.3% GPQA-Diamond · $42/MTokfable-5 81.3%enso-ultra — 98.0% GPQA-Diamond · $25/MTokenso-ultra 98.0%enso — 96.0% GPQA-Diamond · $20/MTokenso 96.0%enso-flash — 92.9% GPQA-Diamond · $4/MTokenso-flash 92.9%
768084889296100$0.5$2$8$30$120GPQA %cheap ← output $/MTok → expensivegpt-5.5 — 93.6% GPQA-Diamond · $8.25/MTok (vendor-reported)gpt-5.5 93.6%gpt-5.2-pro — 93.2% GPQA-Diamond · $138.6/MTok (vendor-reported)gpt-5.2-pro 93.2%gpt-5.6-sol — 92.9% GPQA-Diamond · $25/MTokgpt-5.6-sol 92.9%opus-4.8 — 87.4% GPQA-Diamond · $21/MTokopus-4.8 87.4%gemma-4-31b — 84.3% GPQA-Diamond · $0.44/MTok (vendor-reported)gemma-4-31b 84.3%fable-5 — 81.3% GPQA-Diamond · $42/MTokfable-5 81.3%enso-ultra — 98.0% GPQA-Diamond · $25/MTokenso-ultra 98.0%enso — 96.0% GPQA-Diamond · $20/MTokenso 96.0%enso-flash — 92.9% GPQA-Diamond · $4/MTokenso-flash 92.9%

Reported vs. what we measured

Pick a benchmark, then filter by provenance. Enso numbers are all Hanzo-measured; the rest of the field shows a mix of what we measured and what vendors report. The benchmarks here are the ones an Enso tier tops.

ModelGPQA-DiamondSource$/MTok out
enso-ultraenso98Hanzo$25
ensoenso96Hanzo$20
gemini-3.1-pro94.3Provider-reported—
gpt-5.593.6Provider-reported$8.25
gpt-5.2-pro93.2LLM Stats$139
gpt-5.6-sol92.9Hanzo$25
gpt-5.291.7Vals AI$11.55
gpt-5.491.7Vals AI$12.5
gpt-5.6-terra87.9Hanzo$12.5
opus-4.887.4Hanzo$21
nemotron-3-ultra-550b-a55b86.1Vals AI$1.54
opus-4.585.9Vals AI—
gpt-585.6Vals AI$8.25
sonnet-4.685.6Vals AI—
gemma-4-31b84.3LLM Stats$0.44
o384.1Vals AI$6.8
gpt-5.6-luna82.8Hanzo$5
nemotron-3-super82.7LLM Stats$0.4
minimax-m2.582.1Vals AI$0.76
sonnet-4.581.6Vals AI—
mimo-v2.581.6Vals AI$0.24
fable-581.3Hanzo$42
deepseek-v3.280.3Vals AI—
gpt-5-mini80.3Vals AI$1.65
deepseek-v3.2-exp79.9DeepSeek-V3.2-Exp model …—
gpt-oss-120b78.5Vals AI$0.41
deepseek-4-flash76.9Hanzo$0.2
opus-4.176.3Vals AI$63
o3-mini75.5Vals AI$3.74
deepseek-v4-pro75.3Hanzo$2.5
o173.2Vals AI$51
haiku-4.572.2Vals AI—
deepseek-r171.5DeepSeek-R1 model card (…—

Vendors report on their own harness; Hanzo measures everyone on one. Where both exist the gap is the harness talking — not the model getting better. Hover a source for its exact provenance.

Build on the tier that fits

Start on Flash. Move a request up to Pro or Ultra when it earns the cost, by changing the model id and nothing else.

Build what’s next.