Enso routes. Zen writes. Kai decides.
Kai first: typed decisions against Laya and Jev, on the same questions. Then Enso and Zen, the platform they run on, memory, and the rest of the suite, each result beside the baseline it was measured against.
Enso, Zen, and Kai
Three models on one call path. Enso takes every call, in every modality, and hands it to the model that should answer it: Zen to write, Kai to decide, or another frontier model.
The router · its board: GPQA-Diamond
Open weights · served by the engine: inference
Kai, against Laya and Jev.
The top mean accuracy across 11 application suites.
Laya is the upstream kai-1 weights, served natively. Jev is typesafe/jev-1.13-20260917 through OpenRouter.
The frozen decision harness
| Accuracy | Calibration error (ECE) | |||||
|---|---|---|---|---|---|---|
| Suite | Kai | Laya | Jev | Kai | Laya | Jev |
| AG News | 94.3% | 95.0% | 86.0% | 0.034 | 0.032 | 0.102 |
| DAIR Emotion | 93.5% | 59.5% | 60.3% | 0.019 | 0.306 | 0.280 |
| Banking77 | 90.8% | 42.5% | 83.5% | 0.076 | 0.381 | 0.073 |
| Support triage | 44.3% | 50.2% | 36.5% | 0.441 | 0.091 | 0.483 |
| Email spam | 99.8% | 99.8% | 97.8% | 0.003 | 0.011 | 0.060 |
| Phishing | 99.0% | 98.3% | 90.0% | 0.010 | 0.010 | 0.039 |
| Jailbreak | 92.0% | 70.5% | 94.0% | 0.072 | 0.262 | 0.047 |
| Toxicity | 80.3% | 53.0% | 66.3% | 0.191 | 0.297 | 0.177 |
| RAG relevance | 67.5% | 62.5% | 62.0% | 0.029 | 0.117 | 0.279 |
| Model routing | 100.0% | 63.9% | 97.7% | 0.000 | 0.089 | 0.014 |
| Typed decisions | 75.9% | 76.6% | 73.6% | 0.172 | 0.213 | 0.047 |
| MASSIVE, 51 languages | 90.9% | 38.2% | 89.0% | 0.060 | 0.418 | 0.064 |
| Mean, 11 suites | 85.2% | 70.2% | 77.0% | 0.095 | 0.164 | 0.145 |
11 application suites and MASSIVE intent in 51 languages, the same questions to all three. Laya answers typed decisions on the checkpoint it ships for them.
Source: api.hanzo.ai/v1/research/runs?org=hanzo · harness · Kai a7@0834a74f · harness a1a4a73
Every language
MASSIVE intent, by language
| Accuracy | |||
|---|---|---|---|
| Language | Kai | Laya | Jev |
| Afrikaans | 91.0% | 34.0% | 87.0% |
| Amharic | 85.0% | 15.0% | 85.0% |
| Arabic | 92.0% | 46.0% | 87.0% |
| Azerbaijani | 94.0% | 36.0% | 88.0% |
| Bangla | 92.0% | 45.0% | 86.0% |
| Welsh | 86.0% | 16.0% | 76.0% |
| Danish | 92.0% | 46.0% | 88.0% |
| German | 91.0% | 53.0% | 93.9% |
| Greek | 94.0% | 44.0% | 94.0% |
| English | 94.0% | 82.0% | 92.0% |
| Spanish | 92.0% | 53.0% | 93.0% |
| Persian | 94.0% | 51.0% | 95.0% |
| Finnish | 92.0% | 30.0% | 89.0% |
| French | 91.0% | 60.0% | 90.0% |
| Hebrew | 91.0% | 37.0% | 94.0% |
| Hindi | 91.0% | 46.0% | 95.0% |
| Hungarian | 90.0% | 34.0% | 92.0% |
| Armenian | 89.0% | 25.0% | 84.0% |
| Indonesian | 96.0% | 36.0% | 93.0% |
| Icelandic | 85.0% | 33.0% | 84.0% |
| Italian | 91.0% | 47.0% | 96.0% |
| Japanese | 97.0% | 64.0% | 97.0% |
| Javanese | 89.0% | 16.0% | 75.0% |
| Georgian | 83.0% | 15.0% | 88.0% |
| Khmer | 84.0% | 20.0% | 81.0% |
| Kannada | 91.0% | 30.0% | 87.0% |
| Korean | 91.0% | 47.0% | 94.0% |
| Latvian | 89.0% | 30.0% | 85.0% |
| Malayalam | 87.0% | 28.0% | 91.0% |
| Mongolian | 88.0% | 16.0% | 84.0% |
| Malay | 93.0% | 27.0% | 92.0% |
| Burmese | 92.0% | 16.0% | 89.0% |
| Norwegian Bokmål | 94.0% | 50.0% | 86.0% |
| Dutch | 92.0% | 40.0% | 90.0% |
| Polish | 94.0% | 47.0% | 92.0% |
| Portuguese | 93.0% | 44.0% | 90.0% |
| Romanian | 91.0% | 37.0% | 94.0% |
| Russian | 97.0% | 57.0% | 94.0% |
| Slovenian | 91.0% | 33.0% | 89.0% |
| Albanian | 93.0% | 29.0% | 76.0% |
| Swedish | 93.0% | 48.0% | 94.0% |
| Swahili | 84.0% | 13.0% | 70.0% |
| Tamil | 88.0% | 31.0% | 92.0% |
| Telugu | 87.0% | 22.0% | 91.0% |
| Thai | 93.0% | 48.0% | 94.0% |
| Filipino | 91.0% | 30.0% | 85.0% |
| Turkish | 90.0% | 41.0% | 92.0% |
| Urdu | 90.0% | 42.0% | 90.0% |
| Vietnamese | 91.0% | 33.0% | 89.0% |
| Chinese (China) | 94.0% | 65.0% | 94.0% |
| Chinese (Taiwan) | 95.0% | 61.0% | 93.0% |
Twenty intents to choose from, a hundred utterances per language.
Source: api.hanzo.ai/v1/research/runs?org=hanzo · harness · Kai a7@0834a74f · harness a1a4a73
What a decision model can do
| Measure | Kai | Laya | Jev |
|---|---|---|---|
| Choosing among many options | |||
| Accuracy, 4 options | – | 77.8% | 90.5% |
| Accuracy, 16 options | – | 67.0% | 85.3% |
| Accuracy, 77 options | – | 42.5% | 84.3% |
| Accuracy, 150 options | – | 61.0% | 94.0% |
| Accuracy, 1,000 options | – | 0.0% | 0.0% |
| Accuracy, 10,000 options | – | 0.0% | 0.0% |
| Accuracy, 100,000 options | – | 0.0% | 0.0% |
| Refused, 1,000 options | – | 100.0% | 100.0% |
| Refused, 10,000 options | – | 100.0% | 100.0% |
| Refused, 100,000 options | – | 100.0% | 100.0% |
| Five dependent questions on one state | |||
| Accuracy | – | 33.6% | 83.7% |
| All five right | – | 0.5% | 43.8% |
| Answers that break a rule | – | 63.7% | 24.8% |
| Accuracy, projected onto the rules | – | 32.8% | 87.9% |
| All five right, projected | – | 4.0% | 56.0% |
| Invariance | |||
| Answer changes when the options are reordered | – | 25.5% | 6.0% |
| Largest probability shift under reordering | – | 0.996 | 0.470 |
| A yes/no with its sides swapped: answer mirrors | – | 20.5% | 6.3% |
| A yes/no negated: answer mirrors | – | 30.0% | 88.3% |
| Sensor evidence, accuracy | |||
| Human activity, UCI HAR | – | 20.3% | 15.3% |
| Turbofan remaining life, C-MAPSS | – | 17.0% | 30.0% |
| Valve sound, MIMII | – | 50.0% | 84.0% |
| Conformal prediction sets at a 90% target | |||
| Coverage | – | 90.1% | 90.1% |
| Mean set size | – | 6.801 | 1.304 |
| Programs and deployment | |||
| Questions re-asked after one evidence change | – | 100 of 100 | 100 of 100 |
| Runs offline | – | yes | no |
| Self-hosted | – | yes | no |
| Evidence leaves the machine | – | no | yes |
| Pinned by | – | weights sha256 | model name |
The same cases to all three. Laya and Jev answer each question on its own; Kai refines dependent questions over three passes. Kai a7@0834a74f has no capability run yet, so no Kai figure shows here.
Source: api.hanzo.ai/v1/research/runs?org=hanzo · capability · Kai a7@0834a74f · harness a1a4a73
Latency of one decision
| Serving mode | Device | Dtype | p50 ms | p95 ms | p99 ms |
|---|---|---|---|---|---|
| Cold encode | ROCm | BF16 | 19.8 | 29.0 | – |
| Incremental encode | – | – | – | – | – |
| Cached decision | – | – | – | – | – |
Kai a7@0834a74f, one decision a call, timed on a production accelerator. A mode with no CUDA or ROCm measurement shows as absent; nothing here is projected.
Source: api.hanzo.ai/v1/research/runs?org=hanzo · latency · Kai a7@0834a74f · harness a1a4a73
Platform, agents, and sandboxes
What an agent costs at rest, what a live one costs to wake, and what a sandbox costs to start.
A live agent here is a goroutine: it shares one address space with every other, and nothing about it is an isolation boundary. Code that has to be isolated gets a microVM, which boots in 309 ms on an M-series laptop. The goroutine timings are the median of one pass on an Apple M4 Max, 16 cores. Every figure here is transcribed from the harness's README rather than read from a run file, and each page names the command that prints it.
Memory and retrieval
What a memory finds in a long conversation, and what a reader does with it. Each panel is one metric, the baseline and Hanzo on the same data, the same embedder and the same cut-off, drawn on one scale.
substring EM · pooled over the dev haystack and the three test haystacks, and the largest alone · reader: none, embedder zenlm/zen-embedding-0.6b
test split · zen-embedding-0.6b · k=20 · ALL = every annotated evidence turn in the top 20
500 questions · all-minilm vectors for both · ALL@5 = every evidence session in the top 5
Knowledge graph
Point-in-time questions over facts that begin and end, against another temporal store on the same data. On the CronQuestions types the harness covers, the Hanzo graph answers 99.88% and Semantica 99.88%; what differs is the price of a question.
Code
Cross-file retrieval in a repository, scored apart from generation. Typed links are regular expressions over imports, definitions and identifiers; no model reads anything in that row.
Every benchmark
Grouped in the order the board runs: our models, the platform, memory, then the rest. Each page carries its protocol, its tables, and — where a public command regenerates the tables — that command.
Enso, Zen, and Kai
Enso routes the call. The published board is the graduate-question set the router is scored on.
measured · CC BY 4.0
Platform, agents, and sandboxes
What an agent costs at rest, what a live one costs to wake, and what a sandbox costs to start.
measured
measured
measured
Cloud inference
The engine against llama.cpp, on the same box and the same weights.
measured
Memory and retrieval
What a memory finds in a long conversation, and what a reader does with it.
measured · CC BY-NC 4.0
measured · MIT
measured · MIT
measured · CC BY-NC 4.0
baseline only · unreleased
Knowledge graph
Point-in-time questions over facts that begin and end, against another temporal store on the same data.
measured · see the dataset repository
Code
Cross-file retrieval in a repository, scored apart from generation.
measured · CC BY-NC-ND 4.0
The rest of the record
Finding the evidence is not answering from it. On LoCoMo, multi-hop ALL@20 goes from 18.8 to 37.5, but the answers one reader writes from it move from 36.3 to 39.0 token F1 over 208 questions, inside the baseline's interval of [31.5, 40.9] — and the annotated gold turns reach only 50.8. LoCoMo
MemoryAgentBench stops at 262k multi-hop. 67.0 is where the typed search stands there. Most of the misses want an earlier version of a fact that a later line overwrites, and nothing in the haystack says which edit the question follows. MemoryAgentBench
The engine is slower than cosine. On LongMemEval its retrieval takes 1.96 ms at the median against 0.20 ms, and session recall is not answer accuracy. LongMemEval
The code rows compare generators; they are not a held-out result. No configuration was frozen on dev before the test split ran, and on cross-file-random the full combination scores 30.2 R@1 against 31.2 for dense retrieval with typed links alone. RepoBench-R
Decode trails llama.cpp on every backend measured, and prefill loses on Metal, ROCm, and Vulkan at one prompt length or more. The inference page has every cell with its interval. Inference
GPQA-Diamond is the router’s published board. It lives on its own page, with every tier and the models Enso dispatches to. GPQA-Diamond
Everything in the memory and code panels above regenerates from the committed runs on a fresh clone: git clone https://github.com/hanzoai/benchmarks && cd benchmarks/brain && node results.mjs