Enso, measured
Enso orchestrates 400+ models behind one API. Here is how it scores when we run it — and the field — on a single harness: three differentiated tiers, accuracy-at-cost, and every number kept with its source.
Three tiers, monotonic in quality
Ultra > Pro > Flash — a cost/quality contract, not a model alias. GPQA is Hanzo-measured; price bands are published input→output $/MTok.
Enso Ultra
FlagshipAdaptive fan-out + conviction-weighted selection on the hardest problems. A confident probe bills one call, so Ultra prices near Pro despite being top-tier.
Enso Pro
DefaultRouted to the best-fit model per request across coding, review, and responsive agents. Routes down to a cheap model whenever one suffices. 1M context.
Enso Flash
The high-volume default — a single lean model for chat, extraction, and simple steps, escalating only when a task needs it. Cheapest per request.
Accuracy at cost
enso-ultra lands at 92.9% — the strongest model on our own harness, and 4th once vendor-reported frontier numbers (run on their harness, not ours) are included. Solid dots are Hanzo-measured; hollow dots are vendor-reported.
Reported vs. what we measured
Pick a benchmark, then filter by provenance. Enso numbers are all Hanzo-measured; the rest of the field shows a mix of what we measured and what vendors report. 134 models, 12 benchmarks.
| Model | GPQA-Diamond | Source | $/MTok out |
|---|---|---|---|
| gemini-3.1-pro | 94.3 | Provider-reported | — |
| gpt-5.5 | 93.6 | Provider-reported | $8.25 |
| gpt-5.2-pro | 93.2 | LLM Stats | $139 |
| enso-ultraenso | 92.9 | Hanzo | $75 |
| gpt-5.2 | 91.7 | Vals AI | $11.55 |
| gpt-5.4 | 91.7 | Vals AI | $12.5 |
| gpt-5.6-sol | 90.4 | Hanzo | $25 |
| kimi-k2.6 | 89.1 | Vals AI | $2.71 |
| qwen3.5-397b-a17b | 88.4 | LLM Stats | $2.04 |
| ensoenso | 87.9 | Hanzo | $75 |
| gpt-5.6-terra | 87.9 | Hanzo | $12.5 |
| opus-4.8 | 86.9 | Hanzo | $21 |
| nemotron-3-ultra-550b-a55b | 86.1 | Vals AI | $1.54 |
| opus-4.5 | 85.9 | Vals AI | — |
| gpt-5 | 85.6 | Vals AI | $8.25 |
| glm-5.2 | 85.6 | Vals AI | $3.73 |
| sonnet-4.6 | 85.6 | Vals AI | — |
| glm-5.1 | 84.5 | Vals AI | $3.63 |
| gemma-4-31b | 84.3 | LLM Stats | $0.44 |
| o3 | 84.1 | Vals AI | $6.8 |
| kimi-k2.5 | 84.1 | Vals AI | $1.69 |
| glm-5 | 83.3 | Vals AI | $2.07 |
| gpt-5.6-luna | 82.8 | Hanzo | $5 |
| nemotron-3-super | 82.7 | LLM Stats | $0.4 |
| minimax-m2.5 | 82.1 | Vals AI | $0.76 |
| sonnet-4.5 | 81.6 | Vals AI | — |
| mimo-v2.5 | 81.6 | Vals AI | $0.24 |
| fable-5 | 81.3 | Hanzo | $42 |
| deepseek-v3.2 | 80.3 | Vals AI | — |
| gpt-5-mini | 80.3 | Vals AI | $1.65 |
| deepseek-v3.2-exp | 79.9 | DeepSeek-V3.2-Exp model … | — |
| gpt-oss-120b | 78.5 | Vals AI | $0.41 |
| opus-4.1 | 76.3 | Vals AI | $63 |
| deepseek-v4-pro | 76.3 | Hanzo | $2.5 |
| enso-flashenso | 75.8 | Hanzo | $6 |
| o3-mini | 75.5 | Vals AI | $3.74 |
| o1 | 73.2 | Vals AI | $51 |
| haiku-4.5 | 72.2 | Vals AI | — |
| deepseek-r1 | 71.5 | DeepSeek-R1 model card (… | — |
| deepseek-4-flash | 70.7 | Hanzo | $0.2 |
Vendors report on their own harness; Hanzo measures everyone on one. Where both exist the gap is the harness talking — not the model getting better. Hover a source for its exact provenance.
Build on the tier that fits
Flash, Pro, and Ultra behind one OpenAI-compatible API. Switch by changing the model id.