Enso, measured
Enso is our frontier family, and the router that picks among 500+ models when you ask for auto. A vendor’s published score and your own are rarely the same number, so everything below was run here, on one harness, ours and theirs alike — and every figure still says which it is.
Three tiers that stay in order
Ultra above Pro above Flash, on quality and on price, and it stays that way when the underlying models change. A tier is a contract about cost and quality, not another name for one model. GPQA is measured here; the price band is what you are billed, input then output, per million tokens.
Enso Ultra
FlagshipTop-tier accuracy for the hardest, highest-stakes work — the field’s best 98.0% GPQA-Diamond at a fraction of what the priciest frontier models charge to score lower.
Enso Pro
DefaultThe production default: strong 96.0% GPQA-Diamond for coding, review, research, and responsive agents — priced for everyday scale. 1M context.
Enso Flash
Fast answers at production scale — for chat, extraction, classification, and simple steps at a strong 92.9% GPQA-Diamond. Lowest cost per request.
Accuracy at cost
Accuracy alone picks the most expensive model every time, so both axes are here at once. Top-left is where you want to be: right answers, cheap. Each dot carries its lab’s mark; a solid ring means we measured it, a dashed ring means the vendor reported it. Every dot is labelled, and hovering gives the exact figure.
Reported vs. what we measured
Pick a benchmark, then filter by provenance. Enso numbers are all Hanzo-measured; the rest of the field shows a mix of what we measured and what vendors report. 134 models, 12 benchmarks.
| Model | GPQA-Diamond | Source | $/MTok out |
|---|---|---|---|
| enso-ultraenso | 98 | Hanzo | $25 |
| ensoenso | 96 | Hanzo | $20 |
| gemini-3.1-pro | 94.3 | Provider-reported | — |
| gpt-5.5 | 93.6 | Provider-reported | $8.25 |
| gpt-5.2-pro | 93.2 | LLM Stats | $139 |
| enso-flashenso | 92.9 | Hanzo | $4 |
| gpt-5.6-sol | 92.9 | Hanzo | $25 |
| gpt-5.2 | 91.7 | Vals AI | $11.55 |
| gpt-5.4 | 91.7 | Vals AI | $12.5 |
| kimi-k2.6 | 89.1 | Vals AI | $2.71 |
| qwen3.5-397b-a17b | 88.4 | LLM Stats | $2.04 |
| gpt-5.6-terra | 87.9 | Hanzo | $12.5 |
| opus-4.8 | 87.4 | Hanzo | $21 |
| nemotron-3-ultra-550b-a55b | 86.1 | Vals AI | $1.54 |
| opus-4.5 | 85.9 | Vals AI | — |
| gpt-5 | 85.6 | Vals AI | $8.25 |
| glm-5.2 | 85.6 | Vals AI | $3.73 |
| sonnet-4.6 | 85.6 | Vals AI | — |
| glm-5.1 | 84.5 | Vals AI | $3.63 |
| gemma-4-31b | 84.3 | LLM Stats | $0.44 |
| o3 | 84.1 | Vals AI | $6.8 |
| kimi-k2.5 | 84.1 | Vals AI | $1.69 |
| glm-5 | 83.3 | Vals AI | $2.07 |
| gpt-5.6-luna | 82.8 | Hanzo | $5 |
| nemotron-3-super | 82.7 | LLM Stats | $0.4 |
| minimax-m2.5 | 82.1 | Vals AI | $0.76 |
| sonnet-4.5 | 81.6 | Vals AI | — |
| mimo-v2.5 | 81.6 | Vals AI | $0.24 |
| fable-5 | 81.3 | Hanzo | $42 |
| deepseek-v3.2 | 80.3 | Vals AI | — |
| gpt-5-mini | 80.3 | Vals AI | $1.65 |
| deepseek-v3.2-exp | 79.9 | DeepSeek-V3.2-Exp model … | — |
| gpt-oss-120b | 78.5 | Vals AI | $0.41 |
| deepseek-4-flash | 76.9 | Hanzo | $0.2 |
| opus-4.1 | 76.3 | Vals AI | $63 |
| o3-mini | 75.5 | Vals AI | $3.74 |
| deepseek-v4-pro | 75.3 | Hanzo | $2.5 |
| o1 | 73.2 | Vals AI | $51 |
| haiku-4.5 | 72.2 | Vals AI | — |
| deepseek-r1 | 71.5 | DeepSeek-R1 model card (… | — |
Vendors report on their own harness; Hanzo measures everyone on one. Where both exist the gap is the harness talking — not the model getting better. Hover a source for its exact provenance.
Build on the tier that fits
Start on Flash. Move a request up to Pro or Ultra when it earns the cost, by changing the model id and nothing else.