Benchmarks / Enso / CharXiv Reasoning
Enso · Model boardCharXiv Reasoning
CharXiv Reasoning: every model with a score, each row naming who measured it. Enso is the router, so its row is the pool it chose from.
Latest
History
2 results, newest first. Superseded runs stay, so a revision's progress can be read.
| Measured | Subject | Metric | Value | Baseline | Result | Status |
|---|---|---|---|---|---|---|
| snapshot of Aug 16, 2026 | Enso Ultra | Score | 85.5% | o3, CharXiv official leaderb… 78.6% | Win | Frozen · Hanzo-measured |
| snapshot of Aug 16, 2026 | Enso Pro | Score | 78.0% | o3, CharXiv official leaderb… 78.6% | Loss | Frozen · Hanzo-measured |
The whole board
Every model with a score, as the leaderboard snapshot holds it. Hanzo-measured and publisher-reported rows are different evidence and are marked apart.
| Model | Vendor | Score | Measured by |
|---|---|---|---|
| enso-ultra | Hanzo | 85.5% | Hanzo-measured |
| o3 | OpenAI | 78.6% | CharXiv official leaderb… |
| enso | Hanzo | 78.0% | Hanzo-measured |
| gpt-4.1 | OpenAI | 56.7% | CharXiv official leaderb… |
| o1 | OpenAI | 55.1% | CharXiv official leaderb… |
| gpt-4o-241120 | OpenAI | 50.5% | CharXiv official leaderb… |
| gpt-4o-240513 | OpenAI | 47.1% | CharXiv official leaderb… |
Conditions
7 models on the board2 Hanzo-measured, the rest as their publishers reportHardware
The run records no device or host.
Source
the leaderboard snapshot (lib/data/leaderboard.json)
Reproduce
No public command regenerates this result. The source above is the record it is read from.
More Enso benchmarks
GPQA-Diamond · LiveCodeBench v6 · Humanity's Last Exam
Every result for CharXiv Reasoning on the shelf · How these are measured