Try Hanzo

Benchmarks / Enso / Humanity's Last Exam

Enso · Model board

Humanity's Last Exam

Humanity's Last Exam: every model with a score, each row naming who measured it. Enso is the router, so its row is the pool it chose from.

Latest

History

2 results, newest first. Superseded runs stay, so a revision's progress can be read.

27%38%48%08-16Enso Ultra: 34.6%, snapshot of Aug 16, 2026Enso Pro: 29.4%, snapshot of Aug 16, 2026
MeasuredSubjectMetricValueBaselineResultStatus
snapshot of Aug 16, 2026Enso UltraScore34.6%opus-4.8, Artificial Analysis 45.7%LossFrozen · Hanzo-measured
snapshot of Aug 16, 2026Enso ProScore29.4%opus-4.8, Artificial Analysis 45.7%LossFrozen · Hanzo-measured

The whole board

Every model with a score, as the leaderboard snapshot holds it. Hanzo-measured and publisher-reported rows are different evidence and are marked apart.

ModelVendorScoreMeasured by
opus-4.8Anthropic45.7%Artificial Analysis
gemini-3.1-proGoogle44.7%Artificial Analysis
gpt-5.5OpenAI44.3%Artificial Analysis
gpt-5.4OpenAI41.6%Artificial Analysis
glm-5.2Zhipu40.1%Artificial Analysis
gpt-5.3-codexOpenAI39.9%Artificial Analysis
deepseek-v4-proDeepSeek35.9%Artificial Analysis
kimi-k2.6Moonshot35.9%Artificial Analysis
gpt-5.2OpenAI35.5%Artificial Analysis
enso-ultraHanzo34.6%Hanzo-measured
deepseek-v4-flashDeepSeek32.1%Artificial Analysis
sonnet-4.6Anthropic30.0%Artificial Analysis
ensoHanzo29.4%Hanzo-measured
kimi-k2.5Moonshot29.4%Artificial Analysis
opus-4.5Anthropic28.4%Artificial Analysis
glm-5.1Zhipu28.0%Artificial Analysis
qwen3.5-397b-a17bAlibaba27.3%Artificial Analysis
glm-5Zhipu27.2%Artificial Analysis
nemotron-3-ultra-550b-a55bNVIDIA26.6%Artificial Analysis
gpt-5OpenAI25.3%Humanity's Last Exam off…
mimo-v2.5Xiaomi25.2%Artificial Analysis
gemma-4-31bGoogle22.7%Artificial Analysis
deepseek-v3.2DeepSeek22.2%Artificial Analysis
gemini-2.5-proGoogle21.6%Humanity's Last Exam off…
nvidia-nemotron-3-super-120b-a12bNVIDIA19.2%Artificial Analysis
minimax-m2.5MiniMax19.1%Artificial Analysis
gpt-oss-120bOpenAI18.5%Artificial Analysis
4.5-sonnetAnthropic17.3%Artificial Analysis
gpt-5-miniOpenAI14.6%Artificial Analysis
o3-miniOpenAI12.3%Artificial Analysis
4.1-opusAnthropic11.9%Artificial Analysis
gpt-oss-20bOpenAI9.8%Artificial Analysis
4.5-haikuAnthropic9.7%Artificial Analysis
deepseek-r1*DeepSeek8.5%Humanity's Last Exam off…
qwen3-32bAlibaba8.3%Artificial Analysis
o1OpenAI8.0%Humanity's Last Exam off…
gpt-5-nanoOpenAI7.7%Artificial Analysis
gpt-4oOpenAI2.7%Humanity's Last Exam off…

Conditions

38 models on the board2 Hanzo-measured, the rest as their publishers report

Hardware

The run records no device or host.

Source

the leaderboard snapshot (lib/data/leaderboard.json)

Reproduce

No public command regenerates this result. The source above is the record it is read from.

More Enso benchmarks

GPQA-Diamond · LiveCodeBench v6 · CharXiv Reasoning

Every result for Humanity's Last Exam on the shelf · How these are measured

Build what’s next.