Try Hanzo

Benchmarks / Kai / Model routing

Kai · Decision harness

Model routing

Kai and Jev on the same 399 questions of Model routing, one run per Kai revision.

Latest

History

10 results across 2 metrics, newest first. Superseded runs stay, so a revision's progress can be read.

97.4%98.7%100.0%09-2709-2810-03Kai a5 · df1c16a2: 100.0%, Sep 27, 2026Kai a8 · 0cfeef05: 100.0%, Sep 28, 2026Kai a7 · 0834a74f: 100.0%, Sep 28, 2026Kai a6 · a211cc70: 100.0%, Sep 28, 2026kai-1 · 0834a74f: 100.0%, Oct 3, 2026
MeasuredSubjectMetricValueBaselineResultStatus
Oct 3, 2026kai-1 · 0834a74fAccuracy100.0%Jev 97.7%WinLive · Hanzo-measured
Oct 3, 2026kai-1 · 0834a74fCalibration error (ECE)0.000Jev 0.014WinLive · Hanzo-measured
Sep 28, 2026Kai a8 · 0cfeef05Accuracy100.0%Jev 97.7%WinSuperseded · Hanzo-measured
Sep 28, 2026Kai a7 · 0834a74fAccuracy100.0%Jev 97.7%WinSuperseded · Hanzo-measured
Sep 28, 2026Kai a6 · a211cc70Accuracy100.0%Jev 97.7%WinSuperseded · Hanzo-measured
Sep 28, 2026Kai a8 · 0cfeef05Calibration error (ECE)0.001Jev 0.014WinSuperseded · Hanzo-measured
Sep 28, 2026Kai a7 · 0834a74fCalibration error (ECE)0.000Jev 0.014WinSuperseded · Hanzo-measured
Sep 28, 2026Kai a6 · a211cc70Calibration error (ECE)0.000Jev 0.014WinSuperseded · Hanzo-measured
Sep 27, 2026Kai a5 · df1c16a2Accuracy100.0%Jev 97.7%WinSuperseded · Hanzo-measured
Sep 27, 2026Kai a5 · df1c16a2Calibration error (ECE)0.002Jev 0.014WinSuperseded · Hanzo-measured

Conditions

split: harness399 questions, 399 answereddataset dde140a63fc55 Kai revisions measuredJev: typesafe/jev-1.13-20260917, measured Sep 25, 2026

Hardware

CUDA BF16

Source

research run app.model_routing_domain-harness-kai-1.0834a74f-kai

Reproduce

bench score --harness . --kai results/kai/preds.json.gz --out results/scores.json # in hanzoai/benchmarks/decision @ 2c74e2d

More Kai benchmarks

AG News · DAIR Emotion · Banking77 · Support triage · Email spam · Phishing · Jailbreak · Toxicity · RAG relevance · Typed decisions · MASSIVE · AG News, validation split · Support triage, validation split · Typed decisions, validation split · HarmBench (benign prompts) · Choosing among many options · Joint decisions · Release gate against Jev · Latency of one decision

Every result for Model routing on the shelf · How these are measured

Build what’s next.