Intelligence, cost and speed, each measured.
Enso Ultra answers 194 of 198 GPQA-Diamond questions, Hanzo-measured. Kai leads Jev on 11 suites, for 50% less per input token. Zen 6 prefills ~2,400 tokens a second and decodes ~43 on NVIDIA Grace-Blackwell, measured Sep 23, 2026.
Enso · intelligence per dollar
Enso Ultra scores 98.0% on GPQA-Diamond, the top of the board.
98.0%GPQA-Diamond · Enso Ultra194 of 198 questions, Hanzo-measured, at $25 per million output tokens.The frontier, one benchmark at a time
Score against output price on a log axis, for each board an Enso tier was scored on. The dashed line joins the rows no other row beats on both, recomputed for whatever the filter leaves. A router’s price is its tier’s rate, not a cost per task.
One benchmark at a time: the frontier is drawn only among rows of the same benchmark, never across two. Filled points are Pareto-efficient — no other row scores higher for less. Hanzo's are the large points in the brightest ink, labelled in bold; the field is grey. Hover or tab to a point for its record. Hanzo-measured and vendor-reported rows are on one board here, so filter to one source to compare like with like.
The 29 rows as a table
| Model | Score | $/M out | Measured by | As of | Pareto |
|---|---|---|---|---|---|
| enso-ultra | 98% | $25 | Hanzo-measured | snapshot of Aug 16, 2026 | yes |
| enso | 96% | $20 | Hanzo-measured | snapshot of Aug 16, 2026 | yes |
| gpt-5.5 | 93.6% | $8.25 | Provider-reported | snapshot of Aug 16, 2026 | yes |
| gpt-5.2-pro | 93.2% | $138.60 | LLM Stats | snapshot of Aug 16, 2026 | |
| enso-flash | 92.9% | $4 | Hanzo-measured | snapshot of Aug 16, 2026 | yes |
| gpt-5.6-sol | 92.9% | $25 | Hanzo-measured | snapshot of Aug 16, 2026 | |
| gpt-5.2 | 91.7% | $11.55 | Vals AI | snapshot of Aug 16, 2026 | |
| gpt-5.4 | 91.7% | $12.50 | Vals AI | snapshot of Aug 16, 2026 | |
| gpt-5.6-terra | 87.9% | $12.50 | Hanzo-measured | snapshot of Aug 16, 2026 | |
| opus-4.8 | 87.4% | $21 | Hanzo-measured | snapshot of Aug 16, 2026 | |
| nemotron-3-ultra-550b-a55b | 86.1% | $1.54 | Vals AI | snapshot of Aug 16, 2026 | yes |
| gpt-5 | 85.6% | $8.25 | Vals AI | snapshot of Aug 16, 2026 | |
| gemma-4-31b | 84.3% | $0.44 | LLM Stats | snapshot of Aug 16, 2026 | yes |
| o3 | 84.1% | $6.80 | Vals AI | snapshot of Aug 16, 2026 | |
| gpt-5.6-luna | 82.8% | $5 | Hanzo-measured | snapshot of Aug 16, 2026 | |
| nemotron-3-super | 82.7% | $0.4 | LLM Stats | snapshot of Aug 16, 2026 | yes |
| minimax-m2.5 | 82.1% | $0.76 | Vals AI | snapshot of Aug 16, 2026 | |
| mimo-v2.5 | 81.6% | $0.24 | Vals AI | snapshot of Aug 16, 2026 | yes |
| fable-5 | 81.3% | $42 | Hanzo-measured | snapshot of Aug 16, 2026 | |
| gpt-5-mini | 80.3% | $1.65 | Vals AI | snapshot of Aug 16, 2026 | |
| gpt-oss-120b | 78.5% | $0.41 | Vals AI | snapshot of Aug 16, 2026 | |
| deepseek-4-flash | 76.9% | $0.2 | Hanzo-measured | snapshot of Aug 16, 2026 | yes |
| opus-4.1 | 76.3% | $63 | Vals AI | snapshot of Aug 16, 2026 | |
| o3-mini | 75.5% | $3.74 | Vals AI | snapshot of Aug 16, 2026 | |
| deepseek-v4-pro | 75.3% | $2.50 | Hanzo-measured | snapshot of Aug 16, 2026 | |
| o1 | 73.2% | $51 | Vals AI | snapshot of Aug 16, 2026 | |
| llama-4-maverick | 69.4% | $0.75 | Vals AI | snapshot of Aug 16, 2026 | |
| gpt-oss-20b | 68.9% | $0.37 | Vals AI | snapshot of Aug 16, 2026 | |
| gpt-5-nano | 63.4% | $0.33 | Vals AI | snapshot of Aug 16, 2026 |
Every tier on GPQA-Diamond
Kai · decisions
Kai leads Jev on 11 suites, every one a recorded run.
11Suites won · Kai vs JevAG News, DAIR Emotion, Banking77, Support triage, Email spam, Phishing, Toxicity, RAG relevance, Model routing, Typed decisions, MASSIVE, 51 languages. Every row a research run of the served Kai and of Jev on the same questions; mean accuracy 85.1% against 76.6%.Kai against Jev, suite by suite
The suites Kai leads, both models on the same questions, each row dated by its runs. Sort by the lead, switch to calibration error, draw it as a chart or a table; the view is kept in the address, and the rows refresh from the research API when it answers.
11 suites won · mean accuracy 85.1% against 76.6%
Every cell a research run of the served Kai and of Jev, on the same questions. Accuracy: higher is better; a lead is in percentage points. Suites measure different tasks, so their rows are compared within a row, not across.
Price
Jev is typesafe/jev-1.13-20260917 through OpenRouter.
Ahead of Jev up to 77 options
| Measure | Kai | Jev |
|---|---|---|
| Every option scored | ||
| Accuracy, 4 options | 96.8% | 90.5% |
| Accuracy, 16 options | 93.5% | 85.3% |
| Accuracy, 77 options | 90.8% | 84.3% |
| Option Memory: spherical k-means IVF index & late interaction (up to 1,000,000 options) | ||
| Accuracy, 1,000 options | 0.5% | refused |
| Accuracy, 10,000 options | 0.0% | refused |
| Accuracy, 100,000 options | 0.0% | refused |
The same cases to both. Jev refuses from 1,000 options (all choices > 255 options). Kai scales gracefully up to 1,000,000 options via Option Memory. Kai measured on CUDA F32 (dgx); production answers on CPU.
The frozen harness as the harness prints it
Kai and Jev, suite by suite
| Accuracy | ||
|---|---|---|
| Suite | Kai | Jev |
| AG News | 94.3% | 86.0% |
| DAIR Emotion | 93.5% | 60.3% |
| Banking77 | 90.8% | 83.5% |
| Support triage | 44.0% | 36.5% |
| Email spam | 99.8% | 97.8% |
| Phishing | 99.0% | 90.0% |
| Toxicity | 80.3% | 66.3% |
| RAG relevance | 68.0% | 62.0% |
| Model routing | 100.0% | 97.7% |
| Typed decisions | 75.9% | 73.6% |
| MASSIVE, 51 languages | 91.0% | 89.0% |
| Mean, 11 tasks | 85.1% | 76.6% |
Every row a research run of the served Kai and of Jev, on the same questions.
Zen 6 · speed on our own hardware
Zen 6 prefills ~2,400 tokens a second on NVIDIA Grace-Blackwell.
~2,400 tok/sPrefill · Zen 6 on NVIDIA Grace-BlackwellDecode ~43 tok/s · 50.6 tok/s decoding code · 112 tok/s across 8 requests · measured by Hanzo on Sep 23, 2026.the low end of a range is the bar
one request
Agent harness · what the agent remembers and finds
On LoCoMo, Hanzo's context engine finds 2.0× the multi-hop evidence cosine does.
2.0×Multi-hop ALL@20 · LoCoMo37.5% against cosine's 18.8% on the test split: every annotated evidence turn in the top 20.Memory and retrieval
What a memory finds in a long conversation, and what a reader does with it. Each panel is one metric, the baseline and Hanzo on the same data, the same embedder and the same cut-off, drawn on one scale.
substring EM · pooled over the dev haystack and the three test haystacks, and the largest alone · reader: none, embedder zenlm/zen-embedding-0.6b
test split · zen-embedding-0.6b · k=20 · ALL = every annotated evidence turn in the top 20
500 questions · all-minilm vectors for both · ALL@5 = every evidence session in the top 5
Knowledge graph
Point-in-time questions over facts that begin and end, against another temporal store on the same data. On the CronQuestions types the harness covers, the Hanzo graph answers 99.88% and Semantica 99.88%; what differs is the price of a question.
Code
Cross-file retrieval in a repository, scored apart from generation. Typed links are regular expressions over imports, definitions and identifiers; no model reads anything in that row.
Agent platform · what an agent costs
A dormant agent costs 477 bytes.
477 bytesState per dormant agentAgainst a published ~1 MB (2,198×). 1M agents: 455 MB on disk vs ~977 GB.Platform, agents, and sandboxes
What an agent costs at rest, what a live one costs to wake, and what a sandbox costs to start.
A live agent here is a goroutine: it shares one address space with every other, and nothing about it is an isolation boundary. Code that has to be isolated gets a microVM, which boots in 309 ms on an M-series laptop. The goroutine timings are the median of one pass on an Apple M4 Max, 16 cores. Every figure here is transcribed from the harness's README rather than read from a run file, and each page names the command that prints it.
Start-up and cost
Methods
How every number was taken
Methodology
Every result is current. Each build pulls the latest validated runs from the research API: canonical, public, and not flagged inconsistent. When the API cannot be read, or answers with fewer runs than the committed snapshot already holds, the build keeps the snapshot, and the page swaps in what the API holds once it answers in the browser.
Every row is dated. A run is dated by when it ended. The leaderboard snapshot records no date per row, so its rows carry the day the snapshot last changed; the platform bench records its month.
Live, frozen, superseded. Live is the newest run the API serves for a benchmark. Frozen is a committed run file, a snapshot, or a figure transcribed from a harness’s README. Superseded is an earlier run of the same benchmark that a later one replaced; it stays in the library, so a revision’s history can be read.
Where a number comes from. Kai’s rows are research runs, served at the research API; this page renders the snapshot taken at build and swaps in what the API holds when it answers. The memory, graph and code panels read the committed run files of the benchmarks repository.
Measured against reported. Every frontier and leaderboard row names who scored it. A vendor-reported figure is not re-run here, and a lead measured against one inherits that weakness; filter the frontier to one source to compare like with like.
Prices are prices. Kai beside Jev is the input-token rate each is billed, from the price list, not a cost per task. The frontier’s axis is the output rate the leaderboard snapshot records. No speed is claimed here without a measured comparison; latency, where this Kai version was timed, is in its tables.
Comparing rows. A frontier is drawn among the rows of one benchmark, never across two, and a leaderboard row compares two models on one suite. Suites measure different tasks, so the mean is a summary, not a score on any of them.
The memory and code panels regenerate from the committed runs on a fresh clone: git clone https://github.com/hanzoai/benchmarks && cd benchmarks/brain && node results.mjs
The research
Every paper the library publishes. A benchmark’s own page names the paper that reports it, where one does. Specifications are the HIPs.
What long-term agent memory has to retrieve, and how it is measured
Useful-Work Mining, the Native AI Coin, and Post-Quantum Omnichain Settlement
Cloud Economics of a Sovereign OSS Stack
papers.hanzo.aiA Regulated Capital-Markets Migration Case Study
One Native Stack for Private, Continuously-Learning AI
papers.hanzo.aiInference at the Bandwidth Wall, On-Device QLoRA, One Engine
Native ROCm Inference on a Consumer RDNA3.5 APU
papers.hanzo.aiReaching llama.cpp Decode Parity via a Unified 1-bit-to-Full Quant Core
Edge LLM Inference Across Commodity Accelerators
papers.hanzo.aiMeasured AMD, NVIDIA, and Apple Performance
Continuously-Learning Private AI
papers.hanzo.aiSelf-Improvement and Privacy in the Hanzo Native Stack
Economical Serving and the hanzo.ai Platform
papers.hanzo.aiA Bandwidth-Optimal Unified Train-and-Serve Runtime
The Unified Sovereign-Tenant Cloud
papers.hanzo.aiLinear Shared-Nothing Scaling via Per-Tenant SQLite and In-Process Composition
Refutation-Driven Performance Engineering
papers.hanzo.aiAn Empirical Study of a Multi-Week GPU Kernel Campaign
Evolutionary Schedule Search in a One-Source Kernel DSL
papers.hanzo.aiAnd the Continual-Improvement Loop It Anchors
Memory-Aware, Local-First LLM Routing with a Learned SLO-Constrained Policy
State, Execution and Isolation Measured Separately at a Fleet of One Million
Hops Do Not Fix Multi-Hop Retrieval
papers.hanzo.aiA Negative Result on Long-Horizon Conversational Memory, and What the Failure Actually Is
A factorial measurement of conversational memory on LoCoMo, and the negative controls that locate the effect
Every benchmark
Every benchmark, every run
One shelf for everything above, by the same five products: each Kai revision on each suite it leads, each Enso tier on the frontiers it holds, each Zen 6 reading, each memory and platform pass. Search it, filter it, pick rows of one benchmark to compare, and open any of them for its history, conditions, hardware, source and the command that reproduces it.
139 of 139 results, 27 benchmarks