Try Hanzo
Hanzo AI · Benchmarks

Enso routes. Zen writes. Kai decides.

Kai first: typed decisions against Laya and Jev, on the same questions. Then Enso and Zen, the platform they run on, memory, and the rest of the suite, each result beside the baseline it was measured against.

Enso, Zen, and Kai

Three models on one call path. Enso takes every call, in every modality, and hands it to the model that should answer it: Zen to write, Kai to decide, or another frontier model.

Enso →

Routes every call, in every modality, to the model that should take it.

The router · its board: GPQA-Diamond

Zen →

The generative family. Text, vision, audio, and video, from the edge to the cluster.

Open weights · served by the engine: inference

Kai →

A choice over a listed set. The model, the context, the tools, the budget, and the stop.

The decision model · its boards: below

Kai, against Laya and Jev.

The top mean accuracy across 11 application suites.

Laya is the upstream kai-1 weights, served natively. Jev is typesafe/jev-1.13-20260917 through OpenRouter.

The frozen decision harness

AccuracyCalibration error (ECE)
SuiteKaiLayaJevKaiLayaJev
AG News94.3%95.0%86.0%0.0340.0320.102
DAIR Emotion93.5%59.5%60.3%0.0190.3060.280
Banking7790.8%42.5%83.5%0.0760.3810.073
Support triage44.3%50.2%36.5%0.4410.0910.483
Email spam99.8%99.8%97.8%0.0030.0110.060
Phishing99.0%98.3%90.0%0.0100.0100.039
Jailbreak92.0%70.5%94.0%0.0720.2620.047
Toxicity80.3%53.0%66.3%0.1910.2970.177
RAG relevance67.5%62.5%62.0%0.0290.1170.279
Model routing100.0%63.9%97.7%0.0000.0890.014
Typed decisions75.9%76.6%73.6%0.1720.2130.047
MASSIVE, 51 languages90.9%38.2%89.0%0.0600.4180.064
Mean, 11 suites85.2%70.2%77.0%0.0950.1640.145

11 application suites and MASSIVE intent in 51 languages, the same questions to all three. Laya answers typed decisions on the checkpoint it ships for them.

Source: api.hanzo.ai/v1/research/runs?org=hanzo · harness · Kai a7@0834a74f · harness a1a4a73

Every language

MASSIVE intent, by language

Accuracy
LanguageKaiLayaJev
Afrikaans91.0%34.0%87.0%
Amharic85.0%15.0%85.0%
Arabic92.0%46.0%87.0%
Azerbaijani94.0%36.0%88.0%
Bangla92.0%45.0%86.0%
Welsh86.0%16.0%76.0%
Danish92.0%46.0%88.0%
German91.0%53.0%93.9%
Greek94.0%44.0%94.0%
English94.0%82.0%92.0%
Spanish92.0%53.0%93.0%
Persian94.0%51.0%95.0%
Finnish92.0%30.0%89.0%
French91.0%60.0%90.0%
Hebrew91.0%37.0%94.0%
Hindi91.0%46.0%95.0%
Hungarian90.0%34.0%92.0%
Armenian89.0%25.0%84.0%
Indonesian96.0%36.0%93.0%
Icelandic85.0%33.0%84.0%
Italian91.0%47.0%96.0%
Japanese97.0%64.0%97.0%
Javanese89.0%16.0%75.0%
Georgian83.0%15.0%88.0%
Khmer84.0%20.0%81.0%
Kannada91.0%30.0%87.0%
Korean91.0%47.0%94.0%
Latvian89.0%30.0%85.0%
Malayalam87.0%28.0%91.0%
Mongolian88.0%16.0%84.0%
Malay93.0%27.0%92.0%
Burmese92.0%16.0%89.0%
Norwegian Bokmål94.0%50.0%86.0%
Dutch92.0%40.0%90.0%
Polish94.0%47.0%92.0%
Portuguese93.0%44.0%90.0%
Romanian91.0%37.0%94.0%
Russian97.0%57.0%94.0%
Slovenian91.0%33.0%89.0%
Albanian93.0%29.0%76.0%
Swedish93.0%48.0%94.0%
Swahili84.0%13.0%70.0%
Tamil88.0%31.0%92.0%
Telugu87.0%22.0%91.0%
Thai93.0%48.0%94.0%
Filipino91.0%30.0%85.0%
Turkish90.0%41.0%92.0%
Urdu90.0%42.0%90.0%
Vietnamese91.0%33.0%89.0%
Chinese (China)94.0%65.0%94.0%
Chinese (Taiwan)95.0%61.0%93.0%

Twenty intents to choose from, a hundred utterances per language.

Source: api.hanzo.ai/v1/research/runs?org=hanzo · harness · Kai a7@0834a74f · harness a1a4a73

What a decision model can do

MeasureKaiLayaJev
Choosing among many options
Accuracy, 4 options–77.8%90.5%
Accuracy, 16 options–67.0%85.3%
Accuracy, 77 options–42.5%84.3%
Accuracy, 150 options–61.0%94.0%
Accuracy, 1,000 options–0.0%0.0%
Accuracy, 10,000 options–0.0%0.0%
Accuracy, 100,000 options–0.0%0.0%
Refused, 1,000 options–100.0%100.0%
Refused, 10,000 options–100.0%100.0%
Refused, 100,000 options–100.0%100.0%
Five dependent questions on one state
Accuracy–33.6%83.7%
All five right–0.5%43.8%
Answers that break a rule–63.7%24.8%
Accuracy, projected onto the rules–32.8%87.9%
All five right, projected–4.0%56.0%
Invariance
Answer changes when the options are reordered–25.5%6.0%
Largest probability shift under reordering–0.9960.470
A yes/no with its sides swapped: answer mirrors–20.5%6.3%
A yes/no negated: answer mirrors–30.0%88.3%
Sensor evidence, accuracy
Human activity, UCI HAR–20.3%15.3%
Turbofan remaining life, C-MAPSS–17.0%30.0%
Valve sound, MIMII–50.0%84.0%
Conformal prediction sets at a 90% target
Coverage–90.1%90.1%
Mean set size–6.8011.304
Programs and deployment
Questions re-asked after one evidence change–100 of 100100 of 100
Runs offline–yesno
Self-hosted–yesno
Evidence leaves the machine–noyes
Pinned by–weights sha256model name

The same cases to all three. Laya and Jev answer each question on its own; Kai refines dependent questions over three passes. Kai a7@0834a74f has no capability run yet, so no Kai figure shows here.

Source: api.hanzo.ai/v1/research/runs?org=hanzo · capability · Kai a7@0834a74f · harness a1a4a73

Latency of one decision

Serving modeDeviceDtypep50 msp95 msp99 ms
Cold encodeROCmBF1619.829.0–
Incremental encode–––––
Cached decision–––––

Kai a7@0834a74f, one decision a call, timed on a production accelerator. A mode with no CUDA or ROCm measurement shows as absent; nothing here is projected.

Source: api.hanzo.ai/v1/research/runs?org=hanzo · latency · Kai a7@0834a74f · harness a1a4a73

Platform, agents, and sandboxes

What an agent costs at rest, what a live one costs to wake, and what a sandbox costs to start.

A live agent here is a goroutine: it shares one address space with every other, and nothing about it is an isolation boundary. Code that has to be isolated gets a microVM, which boots in 309 ms on an M-series laptop. The goroutine timings are the median of one pass on an Apple M4 Max, 16 cores. Every figure here is transcribed from the harness's README rather than read from a run file, and each page names the command that prints it.

Memory and retrieval

What a memory finds in a long conversation, and what a reader does with it. Each panel is one metric, the baseline and Hanzo on the same data, the same embedder and the same cut-off, drawn on one scale.

MemoryAgentBench →

single-hop, plain resolver
95.5
single-hop, typed search
98.3
multi-hop, plain resolver
56.8
multi-hop, typed search
84.5
262k multi-hop, plain
33.0
262k multi-hop, typed
67.0

substring EM · pooled over the dev haystack and the three test haystacks, and the largest alone · reader: none, embedder zenlm/zen-embedding-0.6b

LoCoMo →

multi-hop ALL@20, cosine
18.8
multi-hop ALL@20, Hanzo
37.5
R@20, cosine
70.9
R@20, Hanzo
80.7

test split · zen-embedding-0.6b · k=20 · ALL = every annotated evidence turn in the top 20

LongMemEval →

cosine
85.8
engine, frozen on LoCoMo
90.6

500 questions · all-minilm vectors for both · ALL@5 = every evidence session in the top 5

Knowledge graph

Point-in-time questions over facts that begin and end, against another temporal store on the same data. On the CronQuestions types the harness covers, the Hanzo graph answers 99.88% and Semantica 99.88%; what differs is the price of a question.

Code

Cross-file retrieval in a repository, scored apart from generation. Typed links are regular expressions over imports, definitions and identifiers; no model reads anything in that row.

RepoBench-R →

dense retrieval
18.8
typed links
31.4

cross-file-first · test split · R@1

Every benchmark

Grouped in the order the board runs: our models, the platform, memory, then the rest. Each page carries its protocol, its tables, and — where a public command regenerates the tables — that command.

Enso, Zen, and Kai

Enso routes the call. The published board is the graduate-question set the router is scored on.

GPQA-Diamond →

Accuracy against price, for each routing tier and the models it dispatches to.

measured · CC BY 4.0

Platform, agents, and sandboxes

What an agent costs at rest, what a live one costs to wake, and what a sandbox costs to start.

Fleet residency →

Bytes, write rate and resume latency for a dormant agent, at a million of them.

measured

Live agent footprint →

Heap and spawn cost of a running agent, and of the runtimes it can host.

measured

Sandbox cold start →

Time to a running sandbox, for a context, an isolate and a microVM, cold and from a checkpoint.

measured

Cloud inference

The engine against llama.cpp, on the same box and the same weights.

Inference vs llama.cpp →

Prefill and decode throughput against llama.cpp, by backend and prompt length, with intervals.

measured

Memory and retrieval

What a memory finds in a long conversation, and what a reader does with it.

LoCoMo →

Long-conversation retrieval and answering, with the evidence turns annotated.

measured · CC BY-NC 4.0

MemoryAgentBench →

Conflict resolution — which version of a fact is current, over four haystack sizes.

measured · MIT

LongMemEval →

Session recall over long chat histories, engine against cosine on the same embedder.

measured · MIT

LoCoMo · subject scope →

What subject scope and a graph walk each change, and which arms survive an emptied graph.

measured · CC BY-NC 4.0

LoCoMo-Conv →

Conversational rewrites of LoCoMo questions. Dataset unreleased; proxy only.

baseline only · unreleased

Knowledge graph

Point-in-time questions over facts that begin and end, against another temporal store on the same data.

CronQuestions →

Point-in-time and history questions over a temporal knowledge graph, and the cost of each question.

measured · see the dataset repository

Code

Cross-file retrieval in a repository, scored apart from generation.

RepoBench-R →

Cross-file retrieval for code completion, scored apart from generation.

measured · CC BY-NC-ND 4.0

The rest of the record

Finding the evidence is not answering from it. On LoCoMo, multi-hop ALL@20 goes from 18.8 to 37.5, but the answers one reader writes from it move from 36.3 to 39.0 token F1 over 208 questions, inside the baseline's interval of [31.5, 40.9] — and the annotated gold turns reach only 50.8. LoCoMo

MemoryAgentBench stops at 262k multi-hop. 67.0 is where the typed search stands there. Most of the misses want an earlier version of a fact that a later line overwrites, and nothing in the haystack says which edit the question follows. MemoryAgentBench

The engine is slower than cosine. On LongMemEval its retrieval takes 1.96 ms at the median against 0.20 ms, and session recall is not answer accuracy. LongMemEval

The code rows compare generators; they are not a held-out result. No configuration was frozen on dev before the test split ran, and on cross-file-random the full combination scores 30.2 R@1 against 31.2 for dense retrieval with typed links alone. RepoBench-R

Decode trails llama.cpp on every backend measured, and prefill loses on Metal, ROCm, and Vulkan at one prompt length or more. The inference page has every cell with its interval. Inference

GPQA-Diamond is the router’s published board. It lives on its own page, with every tier and the models Enso dispatches to. GPQA-Diamond

Hanzo AI Cloud

Build on Hanzo Cloud

Everything in the memory and code panels above regenerates from the committed runs on a fresh clone: git clone https://github.com/hanzoai/benchmarks && cd benchmarks/brain && node results.mjs