Benchmarks / Kai / MASSIVE
Kai · Decision harnessMASSIVE
Kai and Jev on the same 5,100 questions of MASSIVE, one run per Kai revision.
Latest
History
10 results across 2 metrics, newest first. Superseded runs stay, so a revision's progress can be read.
| Measured | Subject | Metric | Value | Baseline | Result | Status |
|---|---|---|---|---|---|---|
| Oct 3, 2026 | kai-1 · 0834a74f | Accuracy | 91.0% | Jev 89.0% | Win | Live · Hanzo-measured |
| Oct 3, 2026 | kai-1 · 0834a74f | Calibration error (ECE) | 0.060 | Jev 0.064 | Win | Live · Hanzo-measured |
| Sep 28, 2026 | Kai a8 · 0cfeef05 | Accuracy | 90.6% | Jev 89.0% | Win | Superseded · Hanzo-measured |
| Sep 28, 2026 | Kai a7 · 0834a74f | Accuracy | 90.9% | Jev 89.0% | Win | Superseded · Hanzo-measured |
| Sep 28, 2026 | Kai a6 · a211cc70 | Accuracy | 91.0% | Jev 89.0% | Win | Superseded · Hanzo-measured |
| Sep 28, 2026 | Kai a8 · 0cfeef05 | Calibration error (ECE) | 0.068 | Jev 0.064 | Loss | Superseded · Hanzo-measured |
| Sep 28, 2026 | Kai a7 · 0834a74f | Calibration error (ECE) | 0.060 | Jev 0.064 | Win | Superseded · Hanzo-measured |
| Sep 28, 2026 | Kai a6 · a211cc70 | Calibration error (ECE) | 0.059 | Jev 0.064 | Win | Superseded · Hanzo-measured |
| Sep 27, 2026 | Kai a5 · df1c16a2 | Accuracy | 88.8% | Jev 89.0% | Level | Superseded · Hanzo-measured |
| Sep 27, 2026 | Kai a5 · df1c16a2 | Calibration error (ECE) | 0.066 | Jev 0.064 | Loss | Superseded · Hanzo-measured |
Conditions
split: harness5,100 questions, 5,100 answereddataset dde140a63fc55 Kai revisions measuredJev: typesafe/jev-1.13-20260917, measured Sep 25, 2026Hardware
CUDA BF16
Source
research run massive-harness-kai-1.0834a74f-kai
Reproduce
bench score --harness . --kai results/kai/preds.json.gz --out results/scores.json # in hanzoai/benchmarks/decision @ 2c74e2d
More Kai benchmarks
AG News · DAIR Emotion · Banking77 · Support triage · Email spam · Phishing · Jailbreak · Toxicity · RAG relevance · Model routing · Typed decisions · AG News, validation split · Support triage, validation split · Typed decisions, validation split · HarmBench (benign prompts) · Choosing among many options · Joint decisions · Release gate against Jev · Latency of one decision
Every result for MASSIVE on the shelf · How these are measured