Benchmarks / Enso / LiveCodeBench v6
Enso · Model boardLiveCodeBench v6
LiveCodeBench v6: every model with a score, each row naming who measured it. Enso is the router, so its row is the pool it chose from.
Latest
History
One result so far. Its history fills in as later runs are posted.
| Measured | Subject | Metric | Value | Baseline | Result | Status |
|---|---|---|---|---|---|---|
| snapshot of Aug 16, 2026 | Enso Pro | Score | 92.0% | fable-5, Provider-reported 92.9% | Loss | Frozen · Hanzo-measured |
The whole board
Every model with a score, as the leaderboard snapshot holds it. Hanzo-measured and publisher-reported rows are different evidence and are marked apart.
| Model | Vendor | Score | Measured by |
|---|---|---|---|
| fable-5 | Anthropic | 92.9% | Provider-reported |
| enso | Hanzo | 92.0% | Hanzo-measured |
| google/gemini-3.1-pro | Other | 88.5% | Vals AI |
| gemini-3.1-pro | 88.5% | Provider-reported | |
| /claude-opus-4-8 | Anthropic | 87.8% | Vals AI |
| opus-4.8 | Anthropic | 87.8% | Provider-reported |
| deepseek/deepseek-v4-pro | DeepSeek | 87.5% | Vals AI |
| /gpt-5.3-codex | Other | 87.3% | Vals AI |
| kimi/kimi-k2.6 | Moonshot | 86.8% | Vals AI |
| /gpt-5-mini | Other | 86.6% | Vals AI |
| nvidia/nemotron-3-ultra-550b-a55b | Other | 86.0% | Vals AI |
| /gpt-5 | Other | 85.9% | Vals AI |
| /gpt-5.2 | Other | 85.4% | Vals AI |
| /gpt-5.5 | Other | 85.3% | Vals AI |
| gpt-5.5 | OpenAI | 85.3% | Provider-reported |
| /gpt-5.4 | Other | 84.1% | Vals AI |
| /o3 | Other | 83.9% | Vals AI |
| kimi/kimi-k2.5 | Moonshot | 83.9% | Vals AI |
| /claude-opus-4-5-20251101 | Anthropic | 83.7% | Vals AI |
| /gpt-5.1-codex | Other | 83.6% | Vals AI |
| /gpt-oss-120b | Other | 83.2% | Vals AI |
| /claude-sonnet-4-6 | Anthropic | 82.1% | Vals AI |
| zai/glm-5 | Other | 81.9% | Vals AI |
| mimo-v2.5 | Xiaomi | 81.5% | Vals AI |
| zai/glm-5.1 | Other | 81.4% | Vals AI |
| /deepseek-v3p2 | Other | 80.7% | Vals AI |
| /gpt-oss-20b | Other | 80.4% | Vals AI |
| minimax/minimax-m2.5 | MiniMax | 79.2% | Vals AI |
| google/gemini-2.5-pro-03-25 | Other | 79.2% | Vals AI |
| /claude-sonnet-4-5-20250929 | Anthropic | 73.0% | Vals AI |
| /o3-mini | Other | 71.5% | Vals AI |
| /deepseek-r1 | Other | 70.2% | Vals AI |
| /gpt-5-nano | Other | 70.2% | Vals AI |
| zai/glm-5.2 | Other | 69.5% | Vals AI |
| /claude-opus-4-1-20250805 | Anthropic | 66.5% | Vals AI |
| /gpt-4.1 | Other | 54.7% | Vals AI |
| /o1 | Other | 50.3% | Vals AI |
| /llama4-maverick-instruct-basic | Other | 47.3% | Vals AI |
| /gpt-4o | Other | 43.4% | Vals AI |
| /claude-haiku-4-5-20251001 | Anthropic | 41.2% | Vals AI |
| together/meta-llama/llama-3.3-70b-instruct-turbo | Meta | 36.3% | Vals AI |
| kimi-k2.6 | Moonshot | 0.9% | LLM-Stats |
| nemotron-3-ultra | NVIDIA | 0.9% | LLM-Stats |
| kimi-k2.5 | Moonshot | 0.8% | LLM-Stats |
| qwen3.5-397b-a17b | Alibaba | 0.8% | LLM-Stats |
| gemma-4-31b | 0.8% | LLM-Stats | |
| gemma-4-26b-a4b | 0.8% | LLM-Stats | |
| gpt-oss-120b-high | OpenAI | 0.8% | LLM-Stats |
| gemma-4-12b | 0.7% | LLM-Stats |
Conditions
49 models on the board1 Hanzo-measured, the rest as their publishers reportHardware
The run records no device or host.
Source
the leaderboard snapshot (lib/data/leaderboard.json)
Reproduce
No public command regenerates this result. The source above is the record it is read from.
More Enso benchmarks
GPQA-Diamond · Humanity's Last Exam · CharXiv Reasoning
Every result for LiveCodeBench v6 on the shelf · How these are measured