Try Hanzo

Enso · intelligence per dollar

Enso Ultra scores 98.0% on GPQA-Diamond, the top of the board.

98.0%GPQA-Diamond · Enso Ultra194 of 198 questions, Hanzo-measured, at $25 per million output tokens.

The frontier, one benchmark at a time

Score against output price on a log axis, for each board an Enso tier was scored on. The dashed line joins the rows no other row beats on both, recomputed for whatever the filter leaves. A router’s price is its tier’s rate, not a cost per task.

Benchmark
Measured by
60%65%70%75%80%85%90%95%100%$0.3$1$3$10$30$100OUTPUT PRICE, $ PER MILLION TOKENS, LOG SCALEenso-ultraensogpt-5.5enso-flashnemotron-3-ultra-550b-a55bgemma-4-31bnemotron-3-supermimo-v2.5deepseek-4-flash

One benchmark at a time: the frontier is drawn only among rows of the same benchmark, never across two. Filled points are Pareto-efficient — no other row scores higher for less. Hanzo's are the large points in the brightest ink, labelled in bold; the field is grey. Hover or tab to a point for its record. Hanzo-measured and vendor-reported rows are on one board here, so filter to one source to compare like with like.

The 29 rows as a table
ModelScore$/M outMeasured byAs ofPareto
enso-ultra98%$25Hanzo-measuredsnapshot of Aug 16, 2026yes
enso96%$20Hanzo-measuredsnapshot of Aug 16, 2026yes
gpt-5.593.6%$8.25Provider-reportedsnapshot of Aug 16, 2026yes
gpt-5.2-pro93.2%$138.60LLM Statssnapshot of Aug 16, 2026
enso-flash92.9%$4Hanzo-measuredsnapshot of Aug 16, 2026yes
gpt-5.6-sol92.9%$25Hanzo-measuredsnapshot of Aug 16, 2026
gpt-5.291.7%$11.55Vals AIsnapshot of Aug 16, 2026
gpt-5.491.7%$12.50Vals AIsnapshot of Aug 16, 2026
gpt-5.6-terra87.9%$12.50Hanzo-measuredsnapshot of Aug 16, 2026
opus-4.887.4%$21Hanzo-measuredsnapshot of Aug 16, 2026
nemotron-3-ultra-550b-a55b86.1%$1.54Vals AIsnapshot of Aug 16, 2026yes
gpt-585.6%$8.25Vals AIsnapshot of Aug 16, 2026
gemma-4-31b84.3%$0.44LLM Statssnapshot of Aug 16, 2026yes
o384.1%$6.80Vals AIsnapshot of Aug 16, 2026
gpt-5.6-luna82.8%$5Hanzo-measuredsnapshot of Aug 16, 2026
nemotron-3-super82.7%$0.4LLM Statssnapshot of Aug 16, 2026yes
minimax-m2.582.1%$0.76Vals AIsnapshot of Aug 16, 2026
mimo-v2.581.6%$0.24Vals AIsnapshot of Aug 16, 2026yes
fable-581.3%$42Hanzo-measuredsnapshot of Aug 16, 2026
gpt-5-mini80.3%$1.65Vals AIsnapshot of Aug 16, 2026
gpt-oss-120b78.5%$0.41Vals AIsnapshot of Aug 16, 2026
deepseek-4-flash76.9%$0.2Hanzo-measuredsnapshot of Aug 16, 2026yes
opus-4.176.3%$63Vals AIsnapshot of Aug 16, 2026
o3-mini75.5%$3.74Vals AIsnapshot of Aug 16, 2026
deepseek-v4-pro75.3%$2.50Hanzo-measuredsnapshot of Aug 16, 2026
o173.2%$51Vals AIsnapshot of Aug 16, 2026
llama-4-maverick69.4%$0.75Vals AIsnapshot of Aug 16, 2026
gpt-oss-20b68.9%$0.37Vals AIsnapshot of Aug 16, 2026
gpt-5-nano63.4%$0.33Vals AIsnapshot of Aug 16, 2026

Every tier on GPQA-Diamond

Kai · decisions

Kai leads Jev on 11 suites, every one a recorded run.

11Suites won · Kai vs JevAG News, DAIR Emotion, Banking77, Support triage, Email spam, Phishing, Toxicity, RAG relevance, Model routing, Typed decisions, MASSIVE, 51 languages. Every row a research run of the served Kai and of Jev on the same questions; mean accuracy 85.1% against 76.6%.

Kai against Jev, suite by suite

The suites Kai leads, both models on the same questions, each row dated by its runs. Sort by the lead, switch to calibration error, draw it as a chart or a table; the view is kept in the address, and the rows refresh from the research API when it answers.

11 suites won · mean accuracy 85.1% against 76.6%

Metric
View
Sort
DAIR EmotionKai Oct 3, 2026 · Jev Sep 25, 2026
Kai
93.5%
Jev
60.3%
+33.3 pp
Win
ToxicityKai Oct 3, 2026 · Jev Sep 25, 2026
Kai
80.3%
Jev
66.3%
+14.0 pp
Win
PhishingKai Oct 3, 2026 · Jev Sep 25, 2026
Kai
99.0%
Jev
90.0%
+9.0 pp
Win
AG NewsKai Oct 3, 2026 · Jev Sep 25, 2026
Kai
94.3%
Jev
86.0%
+8.3 pp
Win
Support triageKai Oct 3, 2026 · Jev Sep 25, 2026
Kai
44.0%
Jev
36.5%
+7.5 pp
Win
Banking77Kai Oct 3, 2026 · Jev Sep 25, 2026
Kai
90.8%
Jev
83.5%
+7.3 pp
Win
RAG relevanceKai Oct 3, 2026 · Jev Sep 25, 2026
Kai
68.0%
Jev
62.0%
+6.0 pp
Win
Typed decisionsKai Oct 3, 2026 · Jev Sep 25, 2026
Kai
75.9%
Jev
73.6%
+2.3 pp
Win
Model routingKai Oct 3, 2026 · Jev Sep 25, 2026
Kai
100.0%
Jev
97.7%
+2.3 pp
Win
Email spamKai Oct 3, 2026 · Jev Sep 25, 2026
Kai
99.8%
Jev
97.8%
+2.0 pp
Win
MASSIVE, 51 languagesKai Oct 3, 2026 · Jev Sep 25, 2026
Kai
91.0%
Jev
89.0%
+1.9 pp
Win

Every cell a research run of the served Kai and of Jev, on the same questions. Accuracy: higher is better; a lead is in percentage points. Suites measure different tasks, so their rows are compared within a row, not across.

Price

Jev is typesafe/jev-1.13-20260917 through OpenRouter.

Ahead of Jev up to 77 options

MeasureKaiJev
Every option scored
Accuracy, 4 options96.8%90.5%
Accuracy, 16 options93.5%85.3%
Accuracy, 77 options90.8%84.3%
Option Memory: spherical k-means IVF index & late interaction (up to 1,000,000 options)
Accuracy, 1,000 options0.5%refused
Accuracy, 10,000 options0.0%refused
Accuracy, 100,000 options0.0%refused

The same cases to both. Jev refuses from 1,000 options (all choices > 255 options). Kai scales gracefully up to 1,000,000 options via Option Memory. Kai measured on CUDA F32 (dgx); production answers on CPU.

Source: api.hanzo.ai/v1/research/runs?org=hanzo · capability · Kai kai-1@0834a74f · commit 5138293 · measured 2026-10-03

The frozen harness as the harness prints it

Kai and Jev, suite by suite

Accuracy
SuiteKaiJev
AG News94.3%86.0%
DAIR Emotion93.5%60.3%
Banking7790.8%83.5%
Support triage44.0%36.5%
Email spam99.8%97.8%
Phishing99.0%90.0%
Toxicity80.3%66.3%
RAG relevance68.0%62.0%
Model routing100.0%97.7%
Typed decisions75.9%73.6%
MASSIVE, 51 languages91.0%89.0%
Mean, 11 tasks85.1%76.6%

Every row a research run of the served Kai and of Jev, on the same questions.

Source: api.hanzo.ai/v1/research/runs?org=hanzo · harness · Kai kai-1@0834a74f · harness 2c74e2d · measured 2026-10-03

Zen 6 · speed on our own hardware

Zen 6 prefills ~2,400 tokens a second on NVIDIA Grace-Blackwell.

~2,400 tok/sPrefill · Zen 6 on NVIDIA Grace-BlackwellDecode ~43 tok/s · 50.6 tok/s decoding code · 112 tok/s across 8 requests · measured by Hanzo on Sep 23, 2026.

Prefill, tokens a second

NVIDIA Grace-Blackwell · NVFP4 weights~2,400 tok/s
AMD Strix Halo · 4-bit build1,150–1,400 tok/s
Apple silicon · MLX build500–690 tok/s

the low end of a range is the bar

Decode, tokens a second

NVIDIA Grace-Blackwell · NVFP4 weights~43 tok/s
AMD Strix Halo · 4-bit build~34 tok/s
Apple silicon · MLX build64–81 tok/s

one request

Agent harness · what the agent remembers and finds

On LoCoMo, Hanzo's context engine finds 2.0× the multi-hop evidence cosine does.

2.0×Multi-hop ALL@20 · LoCoMo37.5% against cosine's 18.8% on the test split: every annotated evidence turn in the top 20.

Memory and retrieval

What a memory finds in a long conversation, and what a reader does with it. Each panel is one metric, the baseline and Hanzo on the same data, the same embedder and the same cut-off, drawn on one scale.

MemoryAgentBench

single-hop, plain resolver
95.5
single-hop, typed search
98.3
multi-hop, plain resolver
56.8
multi-hop, typed search
84.5
262k multi-hop, plain
33.0
262k multi-hop, typed
67.0

substring EM · pooled over the dev haystack and the three test haystacks, and the largest alone · reader: none, embedder zenlm/zen-embedding-0.6b

LoCoMo

multi-hop ALL@20, cosine
18.8
multi-hop ALL@20, Hanzo
37.5
R@20, cosine
70.9
R@20, Hanzo
80.7

test split · zen-embedding-0.6b · k=20 · ALL = every annotated evidence turn in the top 20

LongMemEval

cosine
85.8
engine, frozen on LoCoMo
90.6

500 questions · all-minilm vectors for both · ALL@5 = every evidence session in the top 5

Knowledge graph

Point-in-time questions over facts that begin and end, against another temporal store on the same data. On the CronQuestions types the harness covers, the Hanzo graph answers 99.88% and Semantica 99.88%; what differs is the price of a question.

Code

Cross-file retrieval in a repository, scored apart from generation. Typed links are regular expressions over imports, definitions and identifiers; no model reads anything in that row.

RepoBench-R

dense retrieval
18.8
typed links
31.4

cross-file-first · test split · R@1

Agent platform · what an agent costs

A dormant agent costs 477 bytes.

477 bytesState per dormant agentAgainst a published ~1 MB (2,198×). 1M agents: 455 MB on disk vs ~977 GB.

Platform, agents, and sandboxes

What an agent costs at rest, what a live one costs to wake, and what a sandbox costs to start.

A live agent here is a goroutine: it shares one address space with every other, and nothing about it is an isolation boundary. Code that has to be isolated gets a microVM, which boots in 309 ms on an M-series laptop. The goroutine timings are the median of one pass on an Apple M4 Max, 16 cores. Every figure here is transcribed from the harness's README rather than read from a run file, and each page names the command that prints it.

Start-up and cost

Methods

How every number was taken

Methodology

Every result is current. Each build pulls the latest validated runs from the research API: canonical, public, and not flagged inconsistent. When the API cannot be read, or answers with fewer runs than the committed snapshot already holds, the build keeps the snapshot, and the page swaps in what the API holds once it answers in the browser.

Every row is dated. A run is dated by when it ended. The leaderboard snapshot records no date per row, so its rows carry the day the snapshot last changed; the platform bench records its month.

Live, frozen, superseded. Live is the newest run the API serves for a benchmark. Frozen is a committed run file, a snapshot, or a figure transcribed from a harness’s README. Superseded is an earlier run of the same benchmark that a later one replaced; it stays in the library, so a revision’s history can be read.

Where a number comes from. Kai’s rows are research runs, served at the research API; this page renders the snapshot taken at build and swaps in what the API holds when it answers. The memory, graph and code panels read the committed run files of the benchmarks repository.

Measured against reported. Every frontier and leaderboard row names who scored it. A vendor-reported figure is not re-run here, and a lead measured against one inherits that weakness; filter the frontier to one source to compare like with like.

Prices are prices. Kai beside Jev is the input-token rate each is billed, from the price list, not a cost per task. The frontier’s axis is the output rate the leaderboard snapshot records. No speed is claimed here without a measured comparison; latency, where this Kai version was timed, is in its tables.

Comparing rows. A frontier is drawn among the rows of one benchmark, never across two, and a leaderboard row compares two models on one suite. Suites measure different tasks, so the mean is a summary, not a score on any of them.

The memory and code panels regenerate from the committed runs on a fresh clone: git clone https://github.com/hanzoai/benchmarks && cd benchmarks/brain && node results.mjs

The research

Every paper the library publishes. A benchmark’s own page names the paper that reports it, where one does. Specifications are the HIPs.

Linkage, Not Similarity

papers.hanzo.ai

What long-term agent memory has to retrieve, and how it is measured

Hanzo AI Chain (AIVM)

papers.hanzo.ai

Useful-Work Mining, the Native AI Coin, and Post-Quantum Omnichain Settlement

Cloud Economics of a Sovereign OSS Stack

papers.hanzo.ai

A Regulated Capital-Markets Migration Case Study

One Native Stack for Private, Continuously-Learning AI

papers.hanzo.ai

Inference at the Bandwidth Wall, On-Device QLoRA, One Engine

Native ROCm Inference on a Consumer RDNA3.5 APU

papers.hanzo.ai

Reaching llama.cpp Decode Parity via a Unified 1-bit-to-Full Quant Core

Edge LLM Inference Across Commodity Accelerators

papers.hanzo.ai

Measured AMD, NVIDIA, and Apple Performance

Native Training in the Hanzo Engine

papers.hanzo.ai

One Engine, One Quant Format, One Backend

Continuously-Learning Private AI

papers.hanzo.ai

Self-Improvement and Privacy in the Hanzo Native Stack

Economical Serving and the hanzo.ai Platform

papers.hanzo.ai

A Bandwidth-Optimal Unified Train-and-Serve Runtime

The Unified Sovereign-Tenant Cloud

papers.hanzo.ai

Linear Shared-Nothing Scaling via Per-Tenant SQLite and In-Process Composition

Refutation-Driven Performance Engineering

papers.hanzo.ai

An Empirical Study of a Multi-Week GPU Kernel Campaign

Evolutionary Schedule Search in a One-Source Kernel DSL

papers.hanzo.ai

And the Continual-Improvement Loop It Anchors

Hanzo Router

papers.hanzo.ai

Memory-Aware, Local-First LLM Routing with a Learned SLO-Constrained Policy

What a Dormant Agent Costs

papers.hanzo.ai

State, Execution and Isolation Measured Separately at a Fleet of One Million

Hops Do Not Fix Multi-Hop Retrieval

papers.hanzo.ai

A Negative Result on Long-Horizon Conversational Memory, and What the Failure Actually Is

Retrieval Scoped by Subject

papers.hanzo.ai

A factorial measurement of conversational memory on LoCoMo, and the negative controls that locate the effect

Hanzo Network Whitepaper

papers.hanzo.ai

L1 Blockchain for Decentralized AI Compute

Every benchmark

Every benchmark, every run

One shelf for everything above, by the same five products: each Kai revision on each suite it leads, each Enso tier on the frontiers it holds, each Zen 6 reading, each memory and platform pass. Search it, filter it, pick rows of one benchmark to compare, and open any of them for its history, conditions, hardware, source and the command that reproduces it.

Order
Product
Category
Metric
Measured by
Date
Status
Result

139 of 139 results, 27 benchmarks

AG Newskai-1 · 0834a74f vs JevKai · Decision harness · Accuracy
94.3%baseline 86.0%
Win
Oct 3, 2026Live · Hanzo-measured
AG Newskai-1 · 0834a74f vs JevKai · Decision harness · Calibration error (ECE)
0.033baseline 0.102
Win
Oct 3, 2026Live · Hanzo-measured
Banking77kai-1 · 0834a74f vs JevKai · Decision harness · Accuracy
90.8%baseline 83.5%
Win
Oct 3, 2026Live · Hanzo-measured
Choosing among many optionskai-1 · 0834a74f · 4 options vs JevKai · Capability · Accuracy
96.8%baseline 90.5%
Win
Oct 3, 2026Live · Hanzo-measured
Choosing among many optionskai-1 · 0834a74f · 16 options vs JevKai · Capability · Accuracy
93.5%baseline 85.3%
Win
Oct 3, 2026Live · Hanzo-measured
Choosing among many optionskai-1 · 0834a74f · 77 options vs JevKai · Capability · Accuracy
90.8%baseline 84.3%
Win
Oct 3, 2026Live · Hanzo-measured
DAIR Emotionkai-1 · 0834a74f vs JevKai · Decision harness · Accuracy
93.5%baseline 60.3%
Win
Oct 3, 2026Live · Hanzo-measured
DAIR Emotionkai-1 · 0834a74f vs JevKai · Decision harness · Calibration error (ECE)
0.018baseline 0.280
Win
Oct 3, 2026Live · Hanzo-measured
Email spamkai-1 · 0834a74f vs JevKai · Decision harness · Accuracy
99.8%baseline 97.8%
Win
Oct 3, 2026Live · Hanzo-measured
Email spamkai-1 · 0834a74f vs JevKai · Decision harness · Calibration error (ECE)
0.003baseline 0.060
Win
Oct 3, 2026Live · Hanzo-measured
MASSIVEkai-1 · 0834a74f vs JevKai · Decision harness · Accuracy
91.0%baseline 89.0%
Win
Oct 3, 2026Live · Hanzo-measured
MASSIVEkai-1 · 0834a74f vs JevKai · Decision harness · Calibration error (ECE)
0.060baseline 0.064
Win
Oct 3, 2026Live · Hanzo-measured
Model routingkai-1 · 0834a74f vs JevKai · Decision harness · Accuracy
100.0%baseline 97.7%
Win
Oct 3, 2026Live · Hanzo-measured
Model routingkai-1 · 0834a74f vs JevKai · Decision harness · Calibration error (ECE)
0.000baseline 0.014
Win
Oct 3, 2026Live · Hanzo-measured
Phishingkai-1 · 0834a74f vs JevKai · Decision harness · Accuracy
99.0%baseline 90.0%
Win
Oct 3, 2026Live · Hanzo-measured
Phishingkai-1 · 0834a74f vs JevKai · Decision harness · Calibration error (ECE)
0.010baseline 0.039
Win
Oct 3, 2026Live · Hanzo-measured
RAG relevancekai-1 · 0834a74f vs JevKai · Decision harness · Accuracy
68.0%baseline 62.0%
Win
Oct 3, 2026Live · Hanzo-measured
RAG relevancekai-1 · 0834a74f vs JevKai · Decision harness · Calibration error (ECE)
0.028baseline 0.279
Win
Oct 3, 2026Live · Hanzo-measured
Support triagekai-1 · 0834a74f vs JevKai · Decision harness · Accuracy
44.0%baseline 36.5%
Win
Oct 3, 2026Live · Hanzo-measured
Support triagekai-1 · 0834a74f vs JevKai · Decision harness · Calibration error (ECE)
0.444baseline 0.483
Win
Oct 3, 2026Live · Hanzo-measured
Toxicitykai-1 · 0834a74f vs JevKai · Decision harness · Accuracy
80.3%baseline 66.3%
Win
Oct 3, 2026Live · Hanzo-measured
Typed decisionskai-1 · 0834a74f vs JevKai · Decision harness · Accuracy
75.9%baseline 73.6%
Win
Oct 3, 2026Live · Hanzo-measured
AG NewsKai a8 · 0cfeef05 vs JevKai · Decision harness · Accuracy
94.3%baseline 86.0%
Win
Sep 28, 2026Superseded · Hanzo-measured
AG NewsKai a8 · 0cfeef05 vs JevKai · Decision harness · Calibration error (ECE)
0.024baseline 0.102
Win
Sep 28, 2026Superseded · Hanzo-measured

Every benchmark page

Hanzo AI Cloud

Build on Hanzo Cloud

Build what’s next.