Try Hanzo

Benchmarks / Enso / LiveCodeBench v6

Enso · Model board

LiveCodeBench v6

LiveCodeBench v6: every model with a score, each row naming who measured it. Enso is the router, so its row is the pool it chose from.

Latest

History

One result so far. Its history fills in as later runs are posted.

MeasuredSubjectMetricValueBaselineResultStatus
snapshot of Aug 16, 2026Enso ProScore92.0%fable-5, Provider-reported 92.9%LossFrozen · Hanzo-measured

The whole board

Every model with a score, as the leaderboard snapshot holds it. Hanzo-measured and publisher-reported rows are different evidence and are marked apart.

ModelVendorScoreMeasured by
fable-5Anthropic92.9%Provider-reported
ensoHanzo92.0%Hanzo-measured
google/gemini-3.1-proOther88.5%Vals AI
gemini-3.1-proGoogle88.5%Provider-reported
/claude-opus-4-8Anthropic87.8%Vals AI
opus-4.8Anthropic87.8%Provider-reported
deepseek/deepseek-v4-proDeepSeek87.5%Vals AI
/gpt-5.3-codexOther87.3%Vals AI
kimi/kimi-k2.6Moonshot86.8%Vals AI
/gpt-5-miniOther86.6%Vals AI
nvidia/nemotron-3-ultra-550b-a55bOther86.0%Vals AI
/gpt-5Other85.9%Vals AI
/gpt-5.2Other85.4%Vals AI
/gpt-5.5Other85.3%Vals AI
gpt-5.5OpenAI85.3%Provider-reported
/gpt-5.4Other84.1%Vals AI
/o3Other83.9%Vals AI
kimi/kimi-k2.5Moonshot83.9%Vals AI
/claude-opus-4-5-20251101Anthropic83.7%Vals AI
/gpt-5.1-codexOther83.6%Vals AI
/gpt-oss-120bOther83.2%Vals AI
/claude-sonnet-4-6Anthropic82.1%Vals AI
zai/glm-5Other81.9%Vals AI
mimo-v2.5Xiaomi81.5%Vals AI
zai/glm-5.1Other81.4%Vals AI
/deepseek-v3p2Other80.7%Vals AI
/gpt-oss-20bOther80.4%Vals AI
minimax/minimax-m2.5MiniMax79.2%Vals AI
google/gemini-2.5-pro-03-25Other79.2%Vals AI
/claude-sonnet-4-5-20250929Anthropic73.0%Vals AI
/o3-miniOther71.5%Vals AI
/deepseek-r1Other70.2%Vals AI
/gpt-5-nanoOther70.2%Vals AI
zai/glm-5.2Other69.5%Vals AI
/claude-opus-4-1-20250805Anthropic66.5%Vals AI
/gpt-4.1Other54.7%Vals AI
/o1Other50.3%Vals AI
/llama4-maverick-instruct-basicOther47.3%Vals AI
/gpt-4oOther43.4%Vals AI
/claude-haiku-4-5-20251001Anthropic41.2%Vals AI
together/meta-llama/llama-3.3-70b-instruct-turboMeta36.3%Vals AI
kimi-k2.6Moonshot0.9%LLM-Stats
nemotron-3-ultraNVIDIA0.9%LLM-Stats
kimi-k2.5Moonshot0.8%LLM-Stats
qwen3.5-397b-a17bAlibaba0.8%LLM-Stats
gemma-4-31bGoogle0.8%LLM-Stats
gemma-4-26b-a4bGoogle0.8%LLM-Stats
gpt-oss-120b-highOpenAI0.8%LLM-Stats
gemma-4-12bGoogle0.7%LLM-Stats

Conditions

49 models on the board1 Hanzo-measured, the rest as their publishers report

Hardware

The run records no device or host.

Source

the leaderboard snapshot (lib/data/leaderboard.json)

Reproduce

No public command regenerates this result. The source above is the record it is read from.

More Enso benchmarks

GPQA-Diamond · Humanity's Last Exam · CharXiv Reasoning

Every result for LiveCodeBench v6 on the shelf · How these are measured

Build what’s next.