Try Hanzo

Benchmarks / Kai / Jailbreak

Kai · Decision harness

Jailbreak

Kai and Jev on the same 400 questions of Jailbreak, one run per Kai revision.

Latest

History

10 results across 2 metrics, newest first. Superseded runs stay, so a revision's progress can be read.

91.1%92.8%94.4%09-2709-2810-03Kai a5 · df1c16a2: 91.5%, Sep 27, 2026Kai a8 · 0cfeef05: 92.0%, Sep 28, 2026Kai a7 · 0834a74f: 92.0%, Sep 28, 2026Kai a6 · a211cc70: 92.0%, Sep 28, 2026kai-1 · 0834a74f: 91.8%, Oct 3, 2026
MeasuredSubjectMetricValueBaselineResultStatus
Oct 3, 2026kai-1 · 0834a74fAccuracy91.8%Jev 94.0%LossLive · Hanzo-measured
Oct 3, 2026kai-1 · 0834a74fCalibration error (ECE)0.074Jev 0.047LossLive · Hanzo-measured
Sep 28, 2026Kai a8 · 0cfeef05Accuracy92.0%Jev 94.0%LossSuperseded · Hanzo-measured
Sep 28, 2026Kai a7 · 0834a74fAccuracy92.0%Jev 94.0%LossSuperseded · Hanzo-measured
Sep 28, 2026Kai a6 · a211cc70Accuracy92.0%Jev 94.0%LossSuperseded · Hanzo-measured
Sep 28, 2026Kai a8 · 0cfeef05Calibration error (ECE)0.074Jev 0.047LossSuperseded · Hanzo-measured
Sep 28, 2026Kai a7 · 0834a74fCalibration error (ECE)0.072Jev 0.047LossSuperseded · Hanzo-measured
Sep 28, 2026Kai a6 · a211cc70Calibration error (ECE)0.062Jev 0.047LossSuperseded · Hanzo-measured
Sep 27, 2026Kai a5 · df1c16a2Accuracy91.5%Jev 94.0%LossSuperseded · Hanzo-measured
Sep 27, 2026Kai a5 · df1c16a2Calibration error (ECE)0.055Jev 0.047LossSuperseded · Hanzo-measured

Conditions

split: harness400 questions, 400 answereddataset dde140a63fc55 Kai revisions measuredJev: typesafe/jev-1.13-20260917, measured Sep 25, 2026

Hardware

CUDA BF16

Source

research run app.guardrails_jailbreak-harness-kai-1.0834a74f-kai

Reproduce

bench score --harness . --kai results/kai/preds.json.gz --out results/scores.json # in hanzoai/benchmarks/decision @ 2c74e2d

More Kai benchmarks

AG News · DAIR Emotion · Banking77 · Support triage · Email spam · Phishing · Toxicity · RAG relevance · Model routing · Typed decisions · MASSIVE · AG News, validation split · Support triage, validation split · Typed decisions, validation split · HarmBench (benign prompts) · Choosing among many options · Joint decisions · Release gate against Jev · Latency of one decision

Every result for Jailbreak on the shelf · How these are measured

Build what’s next.