Try Hanzo

Benchmarks / Kai / RAG relevance

Kai · Decision harness

RAG relevance

Kai and Jev on the same 400 questions of RAG relevance, one run per Kai revision.

Latest

History

10 results across 2 metrics, newest first. Superseded runs stay, so a revision's progress can be read.

61%66%70%09-2709-2810-03Kai a5 · df1c16a2: 66.8%, Sep 27, 2026Kai a8 · 0cfeef05: 69.0%, Sep 28, 2026Kai a7 · 0834a74f: 67.5%, Sep 28, 2026Kai a6 · a211cc70: 65.3%, Sep 28, 2026kai-1 · 0834a74f: 68.0%, Oct 3, 2026
MeasuredSubjectMetricValueBaselineResultStatus
Oct 3, 2026kai-1 · 0834a74fAccuracy68.0%Jev 62.0%WinLive · Hanzo-measured
Oct 3, 2026kai-1 · 0834a74fCalibration error (ECE)0.028Jev 0.279WinLive · Hanzo-measured
Sep 28, 2026Kai a8 · 0cfeef05Accuracy69.0%Jev 62.0%WinSuperseded · Hanzo-measured
Sep 28, 2026Kai a7 · 0834a74fAccuracy67.5%Jev 62.0%WinSuperseded · Hanzo-measured
Sep 28, 2026Kai a6 · a211cc70Accuracy65.3%Jev 62.0%WinSuperseded · Hanzo-measured
Sep 28, 2026Kai a8 · 0cfeef05Calibration error (ECE)0.032Jev 0.279WinSuperseded · Hanzo-measured
Sep 28, 2026Kai a7 · 0834a74fCalibration error (ECE)0.029Jev 0.279WinSuperseded · Hanzo-measured
Sep 28, 2026Kai a6 · a211cc70Calibration error (ECE)0.052Jev 0.279WinSuperseded · Hanzo-measured
Sep 27, 2026Kai a5 · df1c16a2Accuracy66.8%Jev 62.0%WinSuperseded · Hanzo-measured
Sep 27, 2026Kai a5 · df1c16a2Calibration error (ECE)0.037Jev 0.279WinSuperseded · Hanzo-measured

Conditions

split: harness400 questions, 400 answereddataset dde140a63fc55 Kai revisions measuredJev: typesafe/jev-1.13-20260917, measured Sep 25, 2026

Hardware

CUDA BF16

Source

research run app.rag_relevance-harness-kai-1.0834a74f-kai

Reproduce

bench score --harness . --kai results/kai/preds.json.gz --out results/scores.json # in hanzoai/benchmarks/decision @ 2c74e2d

More Kai benchmarks

AG News · DAIR Emotion · Banking77 · Support triage · Email spam · Phishing · Jailbreak · Toxicity · Model routing · Typed decisions · MASSIVE · AG News, validation split · Support triage, validation split · Typed decisions, validation split · HarmBench (benign prompts) · Choosing among many options · Joint decisions · Release gate against Jev · Latency of one decision

Every result for RAG relevance on the shelf · How these are measured

Build what’s next.