Try Hanzo

Benchmarks / Enso / CharXiv Reasoning

Enso · Model board

CharXiv Reasoning

CharXiv Reasoning: every model with a score, each row naming who measured it. Enso is the router, so its row is the pool it chose from.

Latest

History

2 results, newest first. Superseded runs stay, so a revision's progress can be read.

77%82%87%08-16Enso Ultra: 85.5%, snapshot of Aug 16, 2026Enso Pro: 78.0%, snapshot of Aug 16, 2026
MeasuredSubjectMetricValueBaselineResultStatus
snapshot of Aug 16, 2026Enso UltraScore85.5%o3, CharXiv official leaderb… 78.6%WinFrozen · Hanzo-measured
snapshot of Aug 16, 2026Enso ProScore78.0%o3, CharXiv official leaderb… 78.6%LossFrozen · Hanzo-measured

The whole board

Every model with a score, as the leaderboard snapshot holds it. Hanzo-measured and publisher-reported rows are different evidence and are marked apart.

ModelVendorScoreMeasured by
enso-ultraHanzo85.5%Hanzo-measured
o3OpenAI78.6%CharXiv official leaderb…
ensoHanzo78.0%Hanzo-measured
gpt-4.1OpenAI56.7%CharXiv official leaderb…
o1OpenAI55.1%CharXiv official leaderb…
gpt-4o-241120OpenAI50.5%CharXiv official leaderb…
gpt-4o-240513OpenAI47.1%CharXiv official leaderb…

Conditions

7 models on the board2 Hanzo-measured, the rest as their publishers report

Hardware

The run records no device or host.

Source

the leaderboard snapshot (lib/data/leaderboard.json)

Reproduce

No public command regenerates this result. The source above is the record it is read from.

More Enso benchmarks

GPQA-Diamond · LiveCodeBench v6 · Humanity's Last Exam

Every result for CharXiv Reasoning on the shelf · How these are measured

Build what’s next.