Benchmarks / Enso / Humanity's Last Exam
Enso · Model boardHumanity's Last Exam
Humanity's Last Exam: every model with a score, each row naming who measured it. Enso is the router, so its row is the pool it chose from.
Latest
History
2 results, newest first. Superseded runs stay, so a revision's progress can be read.
| Measured | Subject | Metric | Value | Baseline | Result | Status |
|---|---|---|---|---|---|---|
| snapshot of Aug 16, 2026 | Enso Ultra | Score | 34.6% | opus-4.8, Artificial Analysis 45.7% | Loss | Frozen · Hanzo-measured |
| snapshot of Aug 16, 2026 | Enso Pro | Score | 29.4% | opus-4.8, Artificial Analysis 45.7% | Loss | Frozen · Hanzo-measured |
The whole board
Every model with a score, as the leaderboard snapshot holds it. Hanzo-measured and publisher-reported rows are different evidence and are marked apart.
| Model | Vendor | Score | Measured by |
|---|---|---|---|
| opus-4.8 | Anthropic | 45.7% | Artificial Analysis |
| gemini-3.1-pro | 44.7% | Artificial Analysis | |
| gpt-5.5 | OpenAI | 44.3% | Artificial Analysis |
| gpt-5.4 | OpenAI | 41.6% | Artificial Analysis |
| glm-5.2 | Zhipu | 40.1% | Artificial Analysis |
| gpt-5.3-codex | OpenAI | 39.9% | Artificial Analysis |
| deepseek-v4-pro | DeepSeek | 35.9% | Artificial Analysis |
| kimi-k2.6 | Moonshot | 35.9% | Artificial Analysis |
| gpt-5.2 | OpenAI | 35.5% | Artificial Analysis |
| enso-ultra | Hanzo | 34.6% | Hanzo-measured |
| deepseek-v4-flash | DeepSeek | 32.1% | Artificial Analysis |
| sonnet-4.6 | Anthropic | 30.0% | Artificial Analysis |
| enso | Hanzo | 29.4% | Hanzo-measured |
| kimi-k2.5 | Moonshot | 29.4% | Artificial Analysis |
| opus-4.5 | Anthropic | 28.4% | Artificial Analysis |
| glm-5.1 | Zhipu | 28.0% | Artificial Analysis |
| qwen3.5-397b-a17b | Alibaba | 27.3% | Artificial Analysis |
| glm-5 | Zhipu | 27.2% | Artificial Analysis |
| nemotron-3-ultra-550b-a55b | NVIDIA | 26.6% | Artificial Analysis |
| gpt-5 | OpenAI | 25.3% | Humanity's Last Exam off… |
| mimo-v2.5 | Xiaomi | 25.2% | Artificial Analysis |
| gemma-4-31b | 22.7% | Artificial Analysis | |
| deepseek-v3.2 | DeepSeek | 22.2% | Artificial Analysis |
| gemini-2.5-pro | 21.6% | Humanity's Last Exam off… | |
| nvidia-nemotron-3-super-120b-a12b | NVIDIA | 19.2% | Artificial Analysis |
| minimax-m2.5 | MiniMax | 19.1% | Artificial Analysis |
| gpt-oss-120b | OpenAI | 18.5% | Artificial Analysis |
| 4.5-sonnet | Anthropic | 17.3% | Artificial Analysis |
| gpt-5-mini | OpenAI | 14.6% | Artificial Analysis |
| o3-mini | OpenAI | 12.3% | Artificial Analysis |
| 4.1-opus | Anthropic | 11.9% | Artificial Analysis |
| gpt-oss-20b | OpenAI | 9.8% | Artificial Analysis |
| 4.5-haiku | Anthropic | 9.7% | Artificial Analysis |
| deepseek-r1* | DeepSeek | 8.5% | Humanity's Last Exam off… |
| qwen3-32b | Alibaba | 8.3% | Artificial Analysis |
| o1 | OpenAI | 8.0% | Humanity's Last Exam off… |
| gpt-5-nano | OpenAI | 7.7% | Artificial Analysis |
| gpt-4o | OpenAI | 2.7% | Humanity's Last Exam off… |
Conditions
38 models on the board2 Hanzo-measured, the rest as their publishers reportHardware
The run records no device or host.
Source
the leaderboard snapshot (lib/data/leaderboard.json)
Reproduce
No public command regenerates this result. The source above is the record it is read from.
More Enso benchmarks
GPQA-Diamond · LiveCodeBench v6 · CharXiv Reasoning
Every result for Humanity's Last Exam on the shelf · How these are measured