Hanzo-measured · GPQA-Diamond, one common harness

Enso, measured

Enso orchestrates 400+ models behind one API. Here is how it scores when we run it — and the field — on a single harness: three differentiated tiers, accuracy-at-cost, and every number kept with its source.

Three tiers, monotonic in quality

Ultra > Pro > Flash — a cost/quality contract, not a model alias. GPQA is Hanzo-measured; price bands are published input→output $/MTok.

Enso Ultra

Flagship
enso-ultra
92.9%GPQA$12.5 → $75

Adaptive fan-out + conviction-weighted selection on the hardest problems. A confident probe bills one call, so Ultra prices near Pro despite being top-tier.

Enso Pro

Default
enso
87.9%GPQA$20 → $75

Routed to the best-fit model per request across coding, review, and responsive agents. Routes down to a cheap model whenever one suffices. 1M context.

Enso Flash

enso-flash
75.8%GPQA$2 → $6

The high-volume default — a single lean model for chat, extraction, and simple steps, escalating only when a task needs it. Cheapest per request.

Accuracy at cost

enso-ultra lands at 92.9% — the strongest model on our own harness, and 4th once vendor-reported frontier numbers (run on their harness, not ours) are included. Solid dots are Hanzo-measured; hollow dots are vendor-reported.

768084889296$0.5$2$8$30$120GPQA %cheap ← output $/MTok → expensivegpt-5.5 93.6%gpt-5.2-pro 93.2%enso-ultra 92.9%gpt-5.6-sol 90.4%kimi-k2.6 89.1%qwen3.5-397b-a17b 88.4%enso 87.9%opus-4.8 86.9%glm-5.2 85.6%gemma-4-31b 84.3%fable-5 81.3%enso-flash 75.8%

Reported vs. what we measured

Pick a benchmark, then filter by provenance. Enso numbers are all Hanzo-measured; the rest of the field shows a mix of what we measured and what vendors report. 134 models, 12 benchmarks.

ModelGPQA-DiamondSource$/MTok out
gemini-3.1-pro94.3Provider-reported
gpt-5.593.6Provider-reported$8.25
gpt-5.2-pro93.2LLM Stats$139
enso-ultraenso92.9Hanzo$75
gpt-5.291.7Vals AI$11.55
gpt-5.491.7Vals AI$12.5
gpt-5.6-sol90.4Hanzo$25
kimi-k2.689.1Vals AI$2.71
qwen3.5-397b-a17b88.4LLM Stats$2.04
ensoenso87.9Hanzo$75
gpt-5.6-terra87.9Hanzo$12.5
opus-4.886.9Hanzo$21
nemotron-3-ultra-550b-a55b86.1Vals AI$1.54
opus-4.585.9Vals AI
gpt-585.6Vals AI$8.25
glm-5.285.6Vals AI$3.73
sonnet-4.685.6Vals AI
glm-5.184.5Vals AI$3.63
gemma-4-31b84.3LLM Stats$0.44
o384.1Vals AI$6.8
kimi-k2.584.1Vals AI$1.69
glm-583.3Vals AI$2.07
gpt-5.6-luna82.8Hanzo$5
nemotron-3-super82.7LLM Stats$0.4
minimax-m2.582.1Vals AI$0.76
sonnet-4.581.6Vals AI
mimo-v2.581.6Vals AI$0.24
fable-581.3Hanzo$42
deepseek-v3.280.3Vals AI
gpt-5-mini80.3Vals AI$1.65
deepseek-v3.2-exp79.9DeepSeek-V3.2-Exp model …
gpt-oss-120b78.5Vals AI$0.41
opus-4.176.3Vals AI$63
deepseek-v4-pro76.3Hanzo$2.5
enso-flashenso75.8Hanzo$6
o3-mini75.5Vals AI$3.74
o173.2Vals AI$51
haiku-4.572.2Vals AI
deepseek-r171.5DeepSeek-R1 model card (…
deepseek-4-flash70.7Hanzo$0.2

Vendors report on their own harness; Hanzo measures everyone on one. Where both exist the gap is the harness talking — not the model getting better. Hover a source for its exact provenance.

Build on the tier that fits

Flash, Pro, and Ultra behind one OpenAI-compatible API. Switch by changing the model id.