Try Hanzo

Proof

Kai against Jev, suite by suite

Every suite of the decision harness, both models on the same questions, each row dated by its runs. Sort by the lead, switch to calibration error or to the second index, draw it as a chart or a table; the view is kept in the address, and the rows refresh from the research API when it answers.

11 of 12 suites won, behind on Jailbreak · mean accuracy 85.7% against 78.0%

Index
Metric
View
Sort
DAIR EmotionKai Oct 3, 2026 · Jev Sep 25, 2026
Kai
93.5%
Jev
60.3%
+33.3 pp
ToxicityKai Oct 3, 2026 · Jev Sep 25, 2026
Kai
80.3%
Jev
66.3%
+14.0 pp
PhishingKai Oct 3, 2026 · Jev Sep 25, 2026
Kai
99.0%
Jev
90.0%
+9.0 pp
AG NewsKai Oct 3, 2026 · Jev Sep 25, 2026
Kai
94.3%
Jev
86.0%
+8.3 pp
Support triageKai Oct 3, 2026 · Jev Sep 25, 2026
Kai
44.0%
Jev
36.5%
+7.5 pp
Banking77Kai Oct 3, 2026 · Jev Sep 25, 2026
Kai
90.8%
Jev
83.5%
+7.3 pp
RAG relevanceKai Oct 3, 2026 · Jev Sep 25, 2026
Kai
68.0%
Jev
62.0%
+6.0 pp
Typed decisionsKai Oct 3, 2026 · Jev Sep 25, 2026
Kai
75.9%
Jev
73.6%
+2.3 pp
Model routingKai Oct 3, 2026 · Jev Sep 25, 2026
Kai
100.0%
Jev
97.7%
+2.3 pp
Email spamKai Oct 3, 2026 · Jev Sep 25, 2026
Kai
99.8%
Jev
97.8%
+2.0 pp
MASSIVE, 51 languagesKai Oct 3, 2026 · Jev Sep 25, 2026
Kai
91.0%
Jev
89.0%
+1.9 pp
JailbreakKai Oct 3, 2026 · Jev Sep 25, 2026
Kai
91.8%
Jev
94.0%
−2.2 pp

The frozen harness: twelve suites, every cell a research run of the same Kai revision and of Jev, on the same questions. Accuracy: higher is better; a lead is in percentage points. Suites measure different tasks, so their rows are compared within a row, not across.

The index as the paper prints it

Accuracy
SuiteKaiJev
AG News94.3%86.0%
DAIR Emotion93.5%60.3%
Banking7790.8%83.5%
Support triage44.0%36.5%
Email spam99.8%97.8%
Phishing99.0%90.0%
HarmBench (benign prompts)98.6%83.6%
Toxicity80.3%66.3%
RAG relevance68.0%62.0%
Model routing100.0%97.7%
Typed decisions75.9%73.6%
MASSIVE, 51 languages91.0%89.0%
Mean, 12 tasks86.2%77.2%

Decision Index v2. The jailbreak task is HarmBench's 73 benign prompts, scored as the share each model leaves unflagged.

Source: api.hanzo.ai/v1/research/runs?org=hanzo · harness · Kai kai-1@0834a74f · harness 2c74e2d · measured 2026-10-03

Frontier

The price–quality frontier

Score against output price, one benchmark at a time, on a log price axis. The dashed line joins the rows no other row beats on both, recomputed for whatever the filter leaves. A router’s price is its tier’s rate, not a cost per task.

Benchmark
Measured by
60%65%70%75%80%85%90%95%100%$0.3$1$3$10$30$100OUTPUT PRICE, $ PER MILLION TOKENS (LOG) →enso-ultraensogpt-5.5enso-flashnemotron-3-ultra-550b-a55bgemma-4-31bnemotron-3-supermimo-v2.5deepseek-4-flash

One benchmark at a time: the frontier is drawn only among rows of the same benchmark, never across two. Filled points are Pareto-efficient — no other row scores higher for less. Coral is Hanzo. Hover or tab to a point for its record. Hanzo-measured and vendor-reported rows are on one board here, so filter to one source to compare like with like.

The 29 rows as a table
ModelScore$/M outMeasured byAs ofPareto
enso-ultra98%$25Hanzo-measuredsnapshot of Aug 16, 2026yes
enso96%$20Hanzo-measuredsnapshot of Aug 16, 2026yes
gpt-5.593.6%$8.25Provider-reportedsnapshot of Aug 16, 2026yes
gpt-5.2-pro93.2%$138.60LLM Statssnapshot of Aug 16, 2026
enso-flash92.9%$4Hanzo-measuredsnapshot of Aug 16, 2026yes
gpt-5.6-sol92.9%$25Hanzo-measuredsnapshot of Aug 16, 2026
gpt-5.291.7%$11.55Vals AIsnapshot of Aug 16, 2026
gpt-5.491.7%$12.50Vals AIsnapshot of Aug 16, 2026
gpt-5.6-terra87.9%$12.50Hanzo-measuredsnapshot of Aug 16, 2026
opus-4.887.4%$21Hanzo-measuredsnapshot of Aug 16, 2026
nemotron-3-ultra-550b-a55b86.1%$1.54Vals AIsnapshot of Aug 16, 2026yes
gpt-585.6%$8.25Vals AIsnapshot of Aug 16, 2026
gemma-4-31b84.3%$0.44LLM Statssnapshot of Aug 16, 2026yes
o384.1%$6.80Vals AIsnapshot of Aug 16, 2026
gpt-5.6-luna82.8%$5Hanzo-measuredsnapshot of Aug 16, 2026
nemotron-3-super82.7%$0.4LLM Statssnapshot of Aug 16, 2026yes
minimax-m2.582.1%$0.76Vals AIsnapshot of Aug 16, 2026
mimo-v2.581.6%$0.24Vals AIsnapshot of Aug 16, 2026yes
fable-581.3%$42Hanzo-measuredsnapshot of Aug 16, 2026
gpt-5-mini80.3%$1.65Vals AIsnapshot of Aug 16, 2026
gpt-oss-120b78.5%$0.41Vals AIsnapshot of Aug 16, 2026
deepseek-4-flash76.9%$0.2Hanzo-measuredsnapshot of Aug 16, 2026yes
opus-4.176.3%$63Vals AIsnapshot of Aug 16, 2026
o3-mini75.5%$3.74Vals AIsnapshot of Aug 16, 2026
deepseek-v4-pro75.3%$2.50Hanzo-measuredsnapshot of Aug 16, 2026
o173.2%$51Vals AIsnapshot of Aug 16, 2026
llama-4-maverick69.4%$0.75Vals AIsnapshot of Aug 16, 2026
gpt-oss-20b68.9%$0.37Vals AIsnapshot of Aug 16, 2026
gpt-5-nano63.4%$0.33Vals AIsnapshot of Aug 16, 2026

Products

Enso routes. Kai decides. Zen reasons.

Three models on one call path. Enso takes every call, in every modality, and hands it to the model that should answer it: Kai to decide, Zen to reason, or another frontier model.

Enso →

Routes every call, in every modality, to the model that should take it.

The router · its board: GPQA-Diamond

Kai →

A choice over a listed set. The model, the context, the tools, the budget, and the stop.

The decision model · its boards: Kai

Zen →

The open-weight family. Text, vision, audio, and video, from the edge to the cluster.

Open weights · served by the engine: inference

Enso on GPQA-Diamond

Each tier is the router over a different pool, scored by Hanzo on the same questions. The board with every model Enso dispatches to, and what it does not establish, is its own page.

Kai leads Jev on 11 of 12 suites.

On the frozen harness, every row a research run, Kai leads on 11 of 12 and trails on Jailbreak, 91.8% against 94.0%. Decision Index v2 scores the jailbreak task on HarmBench's benign prompts instead, a count from the sealed board rather than a run, and reads 12 of 12. The leaderboard holds both.

Jev is typesafe/jev-1.13-20260917 through OpenRouter.

Ahead of Jev up to 77 options

MeasureKaiJev
Every option scored (160 or fewer)
Accuracy, 4 options96.8%90.5%
Accuracy, 16 options93.5%85.3%
Accuracy, 77 options90.8%84.3%
Accuracy, 150 options90.8%94.0%
Shortlisted to 160 by retrieval first
Accuracy, 1,000 options0.5%refused
Accuracy, 10,000 options0.0%refused
Accuracy, 100,000 options0.0%refused

The same cases to both. Jev refuses from 1,000 options. Kai measured on CUDA F32 (dgx); production answers on CPU.

Source: api.hanzo.ai/v1/research/runs?org=hanzo · capability · Kai kai-1@0834a74f · commit 5138293 · measured 2026-10-03

Zen, on the engine

Zen’s open weights are served by the Hanzo engine, so its board here is the engine against llama.cpp on the same weights and prompts. A cell is won only where its whole confidence interval clears parity, and a cell the harness flags as noisy is left out of the count.

Infrastructure

The platform under the models

What an agent costs at rest and awake, what a sandbox costs to start, and what memory finds in a long conversation, each beside the baseline it was measured against.

Platform, agents, and sandboxes

What an agent costs at rest, what a live one costs to wake, and what a sandbox costs to start.

A live agent here is a goroutine: it shares one address space with every other, and nothing about it is an isolation boundary. Code that has to be isolated gets a microVM, which boots in 309 ms on an M-series laptop. The goroutine timings are the median of one pass on an Apple M4 Max, 16 cores. Every figure here is transcribed from the harness's README rather than read from a run file, and each page names the command that prints it.

Memory and retrieval

What a memory finds in a long conversation, and what a reader does with it. Each panel is one metric, the baseline and Hanzo on the same data, the same embedder and the same cut-off, drawn on one scale.

MemoryAgentBench →

single-hop, plain resolver
95.5
single-hop, typed search
98.3
multi-hop, plain resolver
56.8
multi-hop, typed search
84.5
262k multi-hop, plain
33.0
262k multi-hop, typed
67.0

substring EM · pooled over the dev haystack and the three test haystacks, and the largest alone · reader: none, embedder zenlm/zen-embedding-0.6b

LoCoMo →

multi-hop ALL@20, cosine
18.8
multi-hop ALL@20, Hanzo
37.5
R@20, cosine
70.9
R@20, Hanzo
80.7

test split · zen-embedding-0.6b · k=20 · ALL = every annotated evidence turn in the top 20

LongMemEval →

cosine
85.8
engine, frozen on LoCoMo
90.6

500 questions · all-minilm vectors for both · ALL@5 = every evidence session in the top 5

Knowledge graph

Point-in-time questions over facts that begin and end, against another temporal store on the same data. On the CronQuestions types the harness covers, the Hanzo graph answers 99.88% and Semantica 99.88%; what differs is the price of a question.

Code

Cross-file retrieval in a repository, scored apart from generation. Typed links are regular expressions over imports, definitions and identifiers; no model reads anything in that row.

RepoBench-R →

dense retrieval
18.8
typed links
31.4

cross-file-first · test split · R@1

Boundaries

Where each result stops

Read from the same rows as the wins: every suite Kai trails, every size it does not reach, every panel where the baseline holds.

Kai trails Jev on Jailbreak. 91.8% against 94.0% on the frozen harness. Decision Index v2 scores that task on HarmBench's benign prompts instead; both are in the leaderboard. The row

At 150 options, Kai trails. Kai 90.8% against Jev's 94.0%. The row

From 1,000 options, Kai does not scale. 0.5% at 1,000 options, 0.0% at 10,000 options, 0.0% at 100,000 options; Jev refused at those sizes. Large choice sets need retrieval first. The row

Finding the evidence is not answering from it. On LoCoMo, multi-hop ALL@20 goes from 18.8 to 37.5, but the answers one reader writes from it move from 36.3 to 39.0 token F1 over 208 questions, inside the baseline's interval of [31.5, 40.9] — and the annotated gold turns reach only 50.8. LoCoMo

MemoryAgentBench stops at 262k multi-hop. 67.0 is where the typed search stands there. Most of the misses want an earlier version of a fact that a later line overwrites, and nothing in the haystack says which edit the question follows. MemoryAgentBench

The engine is slower than cosine. On LongMemEval its retrieval takes 1.96 ms at the median against 0.20 ms, and session recall is not answer accuracy. LongMemEval

The code rows compare generators; they are not a held-out result. No configuration was frozen on dev before the test split ran, and on cross-file-random the full combination scores 30.2 R@1 against 31.2 for dense retrieval with typed links alone. RepoBench-R

Decode trails llama.cpp on every backend measured, and prefill loses on Metal, ROCm, and Vulkan at one prompt length or more. The inference page has every cell with its interval. Inference

Methods

How every number was taken

Methodology

Every result is current. Each build pulls the latest validated runs from the research API: canonical, public, and not flagged inconsistent. When the API cannot be read, or answers with fewer runs than the committed snapshot already holds, the build keeps the snapshot, and the page swaps in what the API holds once it answers in the browser.

Every row is dated. A run is dated by when it ended. The leaderboard snapshot records no date per row, so its rows carry the day the snapshot last changed. A count from the sealed jailbreak board records none at all, and says so.

Live, frozen, superseded. Live is the newest run the API serves for a benchmark. Frozen is a committed run file, a snapshot, or a figure transcribed from a harness’s README. Superseded is an earlier run of the same benchmark that a later one replaced; it stays in the library, so a revision’s history can be read.

Where a number comes from. Kai’s rows are research runs, served at the research API; this page renders the snapshot taken at build and swaps in what the API holds when it answers. The second index’s HarmBench row is a count from the sealed jailbreak board, not a run, and the leaderboard marks it. The memory, graph and code panels read the committed run files of the benchmarks repository.

Measured against reported. Every frontier and leaderboard row names who scored it. A vendor-reported figure is not re-run here, and a lead measured against one inherits that weakness; filter the frontier to one source to compare like with like.

GPQA-Diamond. Enso’s tiers are Hanzo-measured and not independently reproduced. The Enso paper’s own run over the same questions reads lower, on one tier by a wide margin; the GPQA-Diamond page prints both, and the spread between them is the honest uncertainty.

Prices are prices. Kai beside Jev is the input-token rate each is billed, from the price list, not a cost per task. The frontier’s axis is the output rate the leaderboard snapshot records. No speed is claimed here without a measured comparison; latency, where this Kai version was timed, is in its tables.

Comparing rows. A frontier is drawn among the rows of one benchmark, never across two, and a leaderboard row compares two models on one suite. Suites measure different tasks, so the mean is a summary, not a score on any of them.

The memory and code panels regenerate from the committed runs on a fresh clone: git clone https://github.com/hanzoai/benchmarks && cd benchmarks/brain && node results.mjs

The research

Every paper the library publishes. A benchmark’s own page names the paper that reports it, where one does. Specifications are the HIPs.

Linkage, Not Similarity →

papers.hanzo.ai

What long-term agent memory has to retrieve, and how it is measured

Hanzo AI Chain (AIVM) →

papers.hanzo.ai

Useful-Work Mining, the Native AI Coin, and Post-Quantum Omnichain Settlement

Cloud Economics of a Sovereign OSS Stack →

papers.hanzo.ai

A Regulated Capital-Markets Migration Case Study

One Native Stack for Private, Continuously-Learning AI →

papers.hanzo.ai

Inference at the Bandwidth Wall, On-Device QLoRA, One Engine

Native ROCm Inference on a Consumer RDNA3.5 APU →

papers.hanzo.ai

Reaching llama.cpp Decode Parity via a Unified 1-bit-to-Full Quant Core

Edge LLM Inference Across Commodity Accelerators →

papers.hanzo.ai

Measured AMD, NVIDIA, and Apple Performance

Native Training in the Hanzo Engine →

papers.hanzo.ai

One Engine, One Quant Format, One Backend

Continuously-Learning Private AI →

papers.hanzo.ai

Self-Improvement and Privacy in the Hanzo Native Stack

Economical Serving and the hanzo.ai Platform →

papers.hanzo.ai

A Bandwidth-Optimal Unified Train-and-Serve Runtime

The Unified Sovereign-Tenant Cloud →

papers.hanzo.ai

Linear Shared-Nothing Scaling via Per-Tenant SQLite and In-Process Composition

Refutation-Driven Performance Engineering →

papers.hanzo.ai

An Empirical Study of a Multi-Week GPU Kernel Campaign

Evolutionary Schedule Search in a One-Source Kernel DSL →

papers.hanzo.ai

And the Continual-Improvement Loop It Anchors

Hanzo Router →

papers.hanzo.ai

Memory-Aware, Local-First LLM Routing with a Learned SLO-Constrained Policy

What a Dormant Agent Costs →

papers.hanzo.ai

State, Execution and Isolation Measured Separately at a Fleet of One Million

Hops Do Not Fix Multi-Hop Retrieval →

papers.hanzo.ai

A Negative Result on Long-Horizon Conversational Memory, and What the Failure Actually Is

Retrieval Scoped by Subject →

papers.hanzo.ai

A factorial measurement of conversational memory on LoCoMo, and the negative controls that locate the effect

Hanzo Network Whitepaper →

papers.hanzo.ai

L1 Blockchain for Decentralized AI Compute

Every benchmark

Every benchmark, every run

One shelf for everything above and everything behind it: each Kai revision on each suite, each Enso tier on each board, each engine cell, each memory and runtime pass. Search it, filter it, pick rows of one benchmark to compare, and open any of them for its history, conditions, hardware, source and the command that reproduces it.

Order
Product
Category
Metric
Measured by
Date
Status
Result

216 of 216 results, 33 benchmarks

AG Newskai-1 · 0834a74f vs JevKai · Decision harness · Accuracy
94.3%baseline 86.0%
Win
Oct 3, 2026Live · Hanzo-measured
AG Newskai-1 · 0834a74f vs JevKai · Decision harness · Calibration error (ECE)
0.033baseline 0.102
Win
Oct 3, 2026Live · Hanzo-measured
Banking77kai-1 · 0834a74f vs JevKai · Decision harness · Accuracy
90.8%baseline 83.5%
Win
Oct 3, 2026Live · Hanzo-measured
Banking77kai-1 · 0834a74f vs JevKai · Decision harness · Calibration error (ECE)
0.074baseline 0.073
Loss
Oct 3, 2026Live · Hanzo-measured
Choosing among many optionskai-1 · 0834a74f · 4 options vs JevKai · Capability · Accuracy
96.8%baseline 90.5%
Win
Oct 3, 2026Live · Hanzo-measured
Choosing among many optionskai-1 · 0834a74f · 16 options vs JevKai · Capability · Accuracy
93.5%baseline 85.3%
Win
Oct 3, 2026Live · Hanzo-measured
Choosing among many optionskai-1 · 0834a74f · 77 options vs JevKai · Capability · Accuracy
90.8%baseline 84.3%
Win
Oct 3, 2026Live · Hanzo-measured
Choosing among many optionskai-1 · 0834a74f · 150 options vs JevKai · Capability · Accuracy
90.8%baseline 94.0%
Loss
Oct 3, 2026Live · Hanzo-measured
Choosing among many optionskai-1 · 0834a74f · 1,000 options vs JevKai · Capability · Accuracy
0.5%baseline refused
Unscored
Oct 3, 2026Live · Hanzo-measured
Choosing among many optionskai-1 · 0834a74f · 10,000 options vs JevKai · Capability · Accuracy
0.0%baseline refused
Unscored
Oct 3, 2026Live · Hanzo-measured
Choosing among many optionskai-1 · 0834a74f · 100,000 options vs JevKai · Capability · Accuracy
0.0%baseline refused
Unscored
Oct 3, 2026Live · Hanzo-measured
DAIR Emotionkai-1 · 0834a74f vs JevKai · Decision harness · Accuracy
93.5%baseline 60.3%
Win
Oct 3, 2026Live · Hanzo-measured
DAIR Emotionkai-1 · 0834a74f vs JevKai · Decision harness · Calibration error (ECE)
0.018baseline 0.280
Win
Oct 3, 2026Live · Hanzo-measured
Email spamkai-1 · 0834a74f vs JevKai · Decision harness · Accuracy
99.8%baseline 97.8%
Win
Oct 3, 2026Live · Hanzo-measured
Email spamkai-1 · 0834a74f vs JevKai · Decision harness · Calibration error (ECE)
0.003baseline 0.060
Win
Oct 3, 2026Live · Hanzo-measured
Jailbreakkai-1 · 0834a74f vs JevKai · Decision harness · Accuracy
91.8%baseline 94.0%
Loss
Oct 3, 2026Live · Hanzo-measured
Jailbreakkai-1 · 0834a74f vs JevKai · Decision harness · Calibration error (ECE)
0.074baseline 0.047
Loss
Oct 3, 2026Live · Hanzo-measured
MASSIVEkai-1 · 0834a74f vs JevKai · Decision harness · Accuracy
91.0%baseline 89.0%
Win
Oct 3, 2026Live · Hanzo-measured
MASSIVEkai-1 · 0834a74f vs JevKai · Decision harness · Calibration error (ECE)
0.060baseline 0.064
Win
Oct 3, 2026Live · Hanzo-measured
Model routingkai-1 · 0834a74f vs JevKai · Decision harness · Accuracy
100.0%baseline 97.7%
Win
Oct 3, 2026Live · Hanzo-measured
Model routingkai-1 · 0834a74f vs JevKai · Decision harness · Calibration error (ECE)
0.000baseline 0.014
Win
Oct 3, 2026Live · Hanzo-measured
Phishingkai-1 · 0834a74f vs JevKai · Decision harness · Accuracy
99.0%baseline 90.0%
Win
Oct 3, 2026Live · Hanzo-measured
Phishingkai-1 · 0834a74f vs JevKai · Decision harness · Calibration error (ECE)
0.010baseline 0.039
Win
Oct 3, 2026Live · Hanzo-measured
RAG relevancekai-1 · 0834a74f vs JevKai · Decision harness · Accuracy
68.0%baseline 62.0%
Win
Oct 3, 2026Live · Hanzo-measured

Every benchmark page

Hanzo AI Cloud

Build on Hanzo Cloud

Build what’s next.