# RAG relevance benchmark · Kai — Hanzo AI

> Kai and Jev on the same 400 questions of RAG relevance, one run per Kai revision.

[Benchmarks](https://hanzo.ai/benchmarks) / [Kai](https://hanzo.ai/benchmarks?product=Kai#all) / RAG relevance

Kai · Decision harness

# RAG relevance

Kai and Jev on the same 400 questions of RAG relevance, one run per Kai revision.

## Latest

[68.0%Accuracy · kai-1 · 0834a74fJev 62.0% · Win · Oct 3, 2026 · Hanzo-measured](#history)

[0.028Calibration error (ECE) · kai-1 · 0834a74fJev 0.279 · Win · Oct 3, 2026 · Hanzo-measured](#history)

## History

10 results across 2 metrics, newest first. Superseded runs stay, so a revision&#x27;s progress can be read.

Measured

Subject

Metric

Value

Baseline

Result

Status

Oct 3, 2026

kai-1 · 0834a74f

Accuracy

68.0%

Jev 62.0%

Win

Live · Hanzo-measured

Oct 3, 2026

kai-1 · 0834a74f

Calibration error (ECE)

0.028

Jev 0.279

Win

Live · Hanzo-measured

Sep 28, 2026

Kai a8 · 0cfeef05

Accuracy

69.0%

Jev 62.0%

Win

Superseded · Hanzo-measured

Sep 28, 2026

Kai a7 · 0834a74f

Accuracy

67.5%

Jev 62.0%

Win

Superseded · Hanzo-measured

Sep 28, 2026

Kai a6 · a211cc70

Accuracy

65.3%

Jev 62.0%

Win

Superseded · Hanzo-measured

Sep 28, 2026

Kai a8 · 0cfeef05

Calibration error (ECE)

0.032

Jev 0.279

Win

Superseded · Hanzo-measured

Sep 28, 2026

Kai a7 · 0834a74f

Calibration error (ECE)

0.029

Jev 0.279

Win

Superseded · Hanzo-measured

Sep 28, 2026

Kai a6 · a211cc70

Calibration error (ECE)

0.052

Jev 0.279

Win

Superseded · Hanzo-measured

Sep 27, 2026

Kai a5 · df1c16a2

Accuracy

66.8%

Jev 62.0%

Win

Superseded · Hanzo-measured

Sep 27, 2026

Kai a5 · df1c16a2

Calibration error (ECE)

0.037

Jev 0.279

Win

Superseded · Hanzo-measured

## Conditions

split: harness400 questions, 400 answereddataset dde140a63fc55 Kai revisions measuredJev: typesafe/jev-1.13-20260917, measured Sep 25, 2026

## Hardware

CUDA BF16

## Source

[research run app.rag_relevance-harness-kai-1.0834a74f-kai](https://api.hanzo.ai/v1/research/runs?org=hanzo&id=app.rag_relevance-harness-kai-1.0834a74f-kai)

## Reproduce

`bench score --harness . --kai results/kai/preds.json.gz --out results/scores.json # in hanzoai/benchmarks/decision @ 2c74e2d`

## More Kai benchmarks

[AG News](https://hanzo.ai/benchmarks/kai-ag-news) · [DAIR Emotion](https://hanzo.ai/benchmarks/kai-dair-emotion) · [Banking77](https://hanzo.ai/benchmarks/kai-banking77) · [Support triage](https://hanzo.ai/benchmarks/kai-support-triage) · [Email spam](https://hanzo.ai/benchmarks/kai-email-spam) · [Phishing](https://hanzo.ai/benchmarks/kai-phishing) · [Jailbreak](https://hanzo.ai/benchmarks/kai-jailbreak) · [Toxicity](https://hanzo.ai/benchmarks/kai-toxicity) · [Model routing](https://hanzo.ai/benchmarks/kai-model-routing) · [Typed decisions](https://hanzo.ai/benchmarks/kai-typed-decisions) · [MASSIVE](https://hanzo.ai/benchmarks/kai-massive) · [AG News, validation split](https://hanzo.ai/benchmarks/kai-ag-news-validation-split) · [Support triage, validation split](https://hanzo.ai/benchmarks/kai-support-triage-validation-split) · [Typed decisions, validation split](https://hanzo.ai/benchmarks/kai-typed-decisions-validation-split) · [HarmBench (benign prompts)](https://hanzo.ai/benchmarks/kai-harmbench) · [Choosing among many options](https://hanzo.ai/benchmarks/kai-choice-size) · [Joint decisions](https://hanzo.ai/benchmarks/kai-joint) · [Release gate against Jev](https://hanzo.ai/benchmarks/kai-release-gate) · [Latency of one decision](https://hanzo.ai/benchmarks/kai-latency)

[Every result for RAG relevance on the shelf](https://hanzo.ai/benchmarks?q=RAG%20relevance#all) · [How these are measured](https://hanzo.ai/benchmarks#methods)
