The Pareto frontier of AI has moved.
Enso sits on the price–quality frontier of GPQA-Diamond (every tier scored) and LiveCodeBench v6, and off it on Humanity's Last Exam. Kai is measured against Jev on the same questions, and every loss is on this page.
Proof
Kai against Jev, suite by suite
Every suite of the decision harness, both models on the same questions, each row dated by its runs. Sort by the lead, switch to calibration error or to the second index, draw it as a chart or a table; the view is kept in the address, and the rows refresh from the research API when it answers.
11 of 12 suites won, behind on Jailbreak · mean accuracy 85.7% against 78.0%
The frozen harness: twelve suites, every cell a research run of the same Kai revision and of Jev, on the same questions. Accuracy: higher is better; a lead is in percentage points. Suites measure different tasks, so their rows are compared within a row, not across.
The index as the paper prints it
| Accuracy | ||
|---|---|---|
| Suite | Kai | Jev |
| AG News | 94.3% | 86.0% |
| DAIR Emotion | 93.5% | 60.3% |
| Banking77 | 90.8% | 83.5% |
| Support triage | 44.0% | 36.5% |
| Email spam | 99.8% | 97.8% |
| Phishing | 99.0% | 90.0% |
| HarmBench (benign prompts) | 98.6% | 83.6% |
| Toxicity | 80.3% | 66.3% |
| RAG relevance | 68.0% | 62.0% |
| Model routing | 100.0% | 97.7% |
| Typed decisions | 75.9% | 73.6% |
| MASSIVE, 51 languages | 91.0% | 89.0% |
| Mean, 12 tasks | 86.2% | 77.2% |
Decision Index v2. The jailbreak task is HarmBench's 73 benign prompts, scored as the share each model leaves unflagged.
Frontier
The price–quality frontier
Score against output price, one benchmark at a time, on a log price axis. The dashed line joins the rows no other row beats on both, recomputed for whatever the filter leaves. A router’s price is its tier’s rate, not a cost per task.
One benchmark at a time: the frontier is drawn only among rows of the same benchmark, never across two. Filled points are Pareto-efficient — no other row scores higher for less. Coral is Hanzo. Hover or tab to a point for its record. Hanzo-measured and vendor-reported rows are on one board here, so filter to one source to compare like with like.
The 29 rows as a table
| Model | Score | $/M out | Measured by | As of | Pareto |
|---|---|---|---|---|---|
| enso-ultra | 98% | $25 | Hanzo-measured | snapshot of Aug 16, 2026 | yes |
| enso | 96% | $20 | Hanzo-measured | snapshot of Aug 16, 2026 | yes |
| gpt-5.5 | 93.6% | $8.25 | Provider-reported | snapshot of Aug 16, 2026 | yes |
| gpt-5.2-pro | 93.2% | $138.60 | LLM Stats | snapshot of Aug 16, 2026 | |
| enso-flash | 92.9% | $4 | Hanzo-measured | snapshot of Aug 16, 2026 | yes |
| gpt-5.6-sol | 92.9% | $25 | Hanzo-measured | snapshot of Aug 16, 2026 | |
| gpt-5.2 | 91.7% | $11.55 | Vals AI | snapshot of Aug 16, 2026 | |
| gpt-5.4 | 91.7% | $12.50 | Vals AI | snapshot of Aug 16, 2026 | |
| gpt-5.6-terra | 87.9% | $12.50 | Hanzo-measured | snapshot of Aug 16, 2026 | |
| opus-4.8 | 87.4% | $21 | Hanzo-measured | snapshot of Aug 16, 2026 | |
| nemotron-3-ultra-550b-a55b | 86.1% | $1.54 | Vals AI | snapshot of Aug 16, 2026 | yes |
| gpt-5 | 85.6% | $8.25 | Vals AI | snapshot of Aug 16, 2026 | |
| gemma-4-31b | 84.3% | $0.44 | LLM Stats | snapshot of Aug 16, 2026 | yes |
| o3 | 84.1% | $6.80 | Vals AI | snapshot of Aug 16, 2026 | |
| gpt-5.6-luna | 82.8% | $5 | Hanzo-measured | snapshot of Aug 16, 2026 | |
| nemotron-3-super | 82.7% | $0.4 | LLM Stats | snapshot of Aug 16, 2026 | yes |
| minimax-m2.5 | 82.1% | $0.76 | Vals AI | snapshot of Aug 16, 2026 | |
| mimo-v2.5 | 81.6% | $0.24 | Vals AI | snapshot of Aug 16, 2026 | yes |
| fable-5 | 81.3% | $42 | Hanzo-measured | snapshot of Aug 16, 2026 | |
| gpt-5-mini | 80.3% | $1.65 | Vals AI | snapshot of Aug 16, 2026 | |
| gpt-oss-120b | 78.5% | $0.41 | Vals AI | snapshot of Aug 16, 2026 | |
| deepseek-4-flash | 76.9% | $0.2 | Hanzo-measured | snapshot of Aug 16, 2026 | yes |
| opus-4.1 | 76.3% | $63 | Vals AI | snapshot of Aug 16, 2026 | |
| o3-mini | 75.5% | $3.74 | Vals AI | snapshot of Aug 16, 2026 | |
| deepseek-v4-pro | 75.3% | $2.50 | Hanzo-measured | snapshot of Aug 16, 2026 | |
| o1 | 73.2% | $51 | Vals AI | snapshot of Aug 16, 2026 | |
| llama-4-maverick | 69.4% | $0.75 | Vals AI | snapshot of Aug 16, 2026 | |
| gpt-oss-20b | 68.9% | $0.37 | Vals AI | snapshot of Aug 16, 2026 | |
| gpt-5-nano | 63.4% | $0.33 | Vals AI | snapshot of Aug 16, 2026 |
Products
Enso routes. Kai decides. Zen reasons.
Three models on one call path. Enso takes every call, in every modality, and hands it to the model that should answer it: Kai to decide, Zen to reason, or another frontier model.
The router · its board: GPQA-Diamond
Enso on GPQA-Diamond
Each tier is the router over a different pool, scored by Hanzo on the same questions. The board with every model Enso dispatches to, and what it does not establish, is its own page.
Kai leads Jev on 11 of 12 suites.
On the frozen harness, every row a research run, Kai leads on 11 of 12 and trails on Jailbreak, 91.8% against 94.0%. Decision Index v2 scores the jailbreak task on HarmBench's benign prompts instead, a count from the sealed board rather than a run, and reads 12 of 12. The leaderboard holds both.
Jev is typesafe/jev-1.13-20260917 through OpenRouter.
Ahead of Jev up to 77 options
| Measure | Kai | Jev |
|---|---|---|
| Every option scored (160 or fewer) | ||
| Accuracy, 4 options | 96.8% | 90.5% |
| Accuracy, 16 options | 93.5% | 85.3% |
| Accuracy, 77 options | 90.8% | 84.3% |
| Accuracy, 150 options | 90.8% | 94.0% |
| Shortlisted to 160 by retrieval first | ||
| Accuracy, 1,000 options | 0.5% | refused |
| Accuracy, 10,000 options | 0.0% | refused |
| Accuracy, 100,000 options | 0.0% | refused |
The same cases to both. Jev refuses from 1,000 options. Kai measured on CUDA F32 (dgx); production answers on CPU.
Zen, on the engine
Zen’s open weights are served by the Hanzo engine, so its board here is the engine against llama.cpp on the same weights and prompts. A cell is won only where its whole confidence interval clears parity, and a cell the harness flags as noisy is left out of the count.
Infrastructure
The platform under the models
What an agent costs at rest and awake, what a sandbox costs to start, and what memory finds in a long conversation, each beside the baseline it was measured against.
Platform, agents, and sandboxes
What an agent costs at rest, what a live one costs to wake, and what a sandbox costs to start.
A live agent here is a goroutine: it shares one address space with every other, and nothing about it is an isolation boundary. Code that has to be isolated gets a microVM, which boots in 309 ms on an M-series laptop. The goroutine timings are the median of one pass on an Apple M4 Max, 16 cores. Every figure here is transcribed from the harness's README rather than read from a run file, and each page names the command that prints it.
Memory and retrieval
What a memory finds in a long conversation, and what a reader does with it. Each panel is one metric, the baseline and Hanzo on the same data, the same embedder and the same cut-off, drawn on one scale.
substring EM · pooled over the dev haystack and the three test haystacks, and the largest alone · reader: none, embedder zenlm/zen-embedding-0.6b
test split · zen-embedding-0.6b · k=20 · ALL = every annotated evidence turn in the top 20
500 questions · all-minilm vectors for both · ALL@5 = every evidence session in the top 5
Knowledge graph
Point-in-time questions over facts that begin and end, against another temporal store on the same data. On the CronQuestions types the harness covers, the Hanzo graph answers 99.88% and Semantica 99.88%; what differs is the price of a question.
Code
Cross-file retrieval in a repository, scored apart from generation. Typed links are regular expressions over imports, definitions and identifiers; no model reads anything in that row.
Boundaries
Where each result stops
Read from the same rows as the wins: every suite Kai trails, every size it does not reach, every panel where the baseline holds.
Kai trails Jev on Jailbreak. 91.8% against 94.0% on the frozen harness. Decision Index v2 scores that task on HarmBench's benign prompts instead; both are in the leaderboard. The row
At 150 options, Kai trails. Kai 90.8% against Jev's 94.0%. The row
From 1,000 options, Kai does not scale. 0.5% at 1,000 options, 0.0% at 10,000 options, 0.0% at 100,000 options; Jev refused at those sizes. Large choice sets need retrieval first. The row
Finding the evidence is not answering from it. On LoCoMo, multi-hop ALL@20 goes from 18.8 to 37.5, but the answers one reader writes from it move from 36.3 to 39.0 token F1 over 208 questions, inside the baseline's interval of [31.5, 40.9] — and the annotated gold turns reach only 50.8. LoCoMo
MemoryAgentBench stops at 262k multi-hop. 67.0 is where the typed search stands there. Most of the misses want an earlier version of a fact that a later line overwrites, and nothing in the haystack says which edit the question follows. MemoryAgentBench
The engine is slower than cosine. On LongMemEval its retrieval takes 1.96 ms at the median against 0.20 ms, and session recall is not answer accuracy. LongMemEval
The code rows compare generators; they are not a held-out result. No configuration was frozen on dev before the test split ran, and on cross-file-random the full combination scores 30.2 R@1 against 31.2 for dense retrieval with typed links alone. RepoBench-R
Decode trails llama.cpp on every backend measured, and prefill loses on Metal, ROCm, and Vulkan at one prompt length or more. The inference page has every cell with its interval. Inference
Methods
How every number was taken
Methodology
Every result is current. Each build pulls the latest validated runs from the research API: canonical, public, and not flagged inconsistent. When the API cannot be read, or answers with fewer runs than the committed snapshot already holds, the build keeps the snapshot, and the page swaps in what the API holds once it answers in the browser.
Every row is dated. A run is dated by when it ended. The leaderboard snapshot records no date per row, so its rows carry the day the snapshot last changed. A count from the sealed jailbreak board records none at all, and says so.
Live, frozen, superseded. Live is the newest run the API serves for a benchmark. Frozen is a committed run file, a snapshot, or a figure transcribed from a harness’s README. Superseded is an earlier run of the same benchmark that a later one replaced; it stays in the library, so a revision’s history can be read.
Where a number comes from. Kai’s rows are research runs, served at the research API; this page renders the snapshot taken at build and swaps in what the API holds when it answers. The second index’s HarmBench row is a count from the sealed jailbreak board, not a run, and the leaderboard marks it. The memory, graph and code panels read the committed run files of the benchmarks repository.
Measured against reported. Every frontier and leaderboard row names who scored it. A vendor-reported figure is not re-run here, and a lead measured against one inherits that weakness; filter the frontier to one source to compare like with like.
GPQA-Diamond. Enso’s tiers are Hanzo-measured and not independently reproduced. The Enso paper’s own run over the same questions reads lower, on one tier by a wide margin; the GPQA-Diamond page prints both, and the spread between them is the honest uncertainty.
Prices are prices. Kai beside Jev is the input-token rate each is billed, from the price list, not a cost per task. The frontier’s axis is the output rate the leaderboard snapshot records. No speed is claimed here without a measured comparison; latency, where this Kai version was timed, is in its tables.
Comparing rows. A frontier is drawn among the rows of one benchmark, never across two, and a leaderboard row compares two models on one suite. Suites measure different tasks, so the mean is a summary, not a score on any of them.
The memory and code panels regenerate from the committed runs on a fresh clone: git clone https://github.com/hanzoai/benchmarks && cd benchmarks/brain && node results.mjs
The research
Every paper the library publishes. A benchmark’s own page names the paper that reports it, where one does. Specifications are the HIPs.
What long-term agent memory has to retrieve, and how it is measured
Useful-Work Mining, the Native AI Coin, and Post-Quantum Omnichain Settlement
Cloud Economics of a Sovereign OSS Stack →
papers.hanzo.aiA Regulated Capital-Markets Migration Case Study
One Native Stack for Private, Continuously-Learning AI →
papers.hanzo.aiInference at the Bandwidth Wall, On-Device QLoRA, One Engine
Native ROCm Inference on a Consumer RDNA3.5 APU →
papers.hanzo.aiReaching llama.cpp Decode Parity via a Unified 1-bit-to-Full Quant Core
Edge LLM Inference Across Commodity Accelerators →
papers.hanzo.aiMeasured AMD, NVIDIA, and Apple Performance
Continuously-Learning Private AI →
papers.hanzo.aiSelf-Improvement and Privacy in the Hanzo Native Stack
Economical Serving and the hanzo.ai Platform →
papers.hanzo.aiA Bandwidth-Optimal Unified Train-and-Serve Runtime
The Unified Sovereign-Tenant Cloud →
papers.hanzo.aiLinear Shared-Nothing Scaling via Per-Tenant SQLite and In-Process Composition
Refutation-Driven Performance Engineering →
papers.hanzo.aiAn Empirical Study of a Multi-Week GPU Kernel Campaign
Evolutionary Schedule Search in a One-Source Kernel DSL →
papers.hanzo.aiAnd the Continual-Improvement Loop It Anchors
Memory-Aware, Local-First LLM Routing with a Learned SLO-Constrained Policy
State, Execution and Isolation Measured Separately at a Fleet of One Million
Hops Do Not Fix Multi-Hop Retrieval →
papers.hanzo.aiA Negative Result on Long-Horizon Conversational Memory, and What the Failure Actually Is
A factorial measurement of conversational memory on LoCoMo, and the negative controls that locate the effect
Every benchmark
Every benchmark, every run
One shelf for everything above and everything behind it: each Kai revision on each suite, each Enso tier on each board, each engine cell, each memory and runtime pass. Search it, filter it, pick rows of one benchmark to compare, and open any of them for its history, conditions, hardware, source and the command that reproduces it.
216 of 216 results, 33 benchmarks
Every benchmark page
filter the shelf to it