# Release gate against Jev benchmark · Kai — Hanzo AI

> Whether a Kai revision regresses past tolerance against Jev on any gated suite. Passing it is the condition for calling a revision state of the art.

[Benchmarks](https://hanzo.ai/benchmarks) / [Kai](https://hanzo.ai/benchmarks?product=Kai#all) / Release gate against Jev

Kai · Release gate

# Release gate against Jev

Whether a Kai revision regresses past tolerance against Jev on any gated suite. Passing it is the condition for calling a revision state of the art.

## Latest

[fails · 14 regressed of 62 suitesGate result · Kai a8 · 0cfeef05Jev – · Loss · Sep 28, 2026 · Hanzo-measured](#history)

## History

4 results, newest first. Superseded runs stay, so a revision&#x27;s progress can be read.

Measured

Subject

Metric

Value

Baseline

Result

Status

Sep 28, 2026

Kai a8 · 0cfeef05

Gate result

fails · 14 regressed of 62 suites

–

Loss

Live · Hanzo-measured

Sep 28, 2026

Kai a7 · 0834a74f

Gate result

fails · 13 regressed of 62 suites

–

Loss

Superseded · Hanzo-measured

Sep 28, 2026

Kai a6 · a211cc70

Gate result

fails · 11 regressed of 62 suites

–

Loss

Superseded · Hanzo-measured

Sep 27, 2026

Kai a5 · df1c16a2

Gate result

fails · 16 regressed of 62 suites

–

Loss

Superseded · Hanzo-measured

## Conditions

tolerance accuracy 0.02, ece 0.03

## Hardware

The run records no device or host.

## Source

[research run gate.jev-drawn-a8.0cfeef05-kai](https://api.hanzo.ai/v1/research/runs?org=hanzo&id=gate.jev-drawn-a8.0cfeef05-kai)

## Reproduce

No public command regenerates this result. The source above is the record it is read from.

## More Kai benchmarks

[AG News](https://hanzo.ai/benchmarks/kai-ag-news) · [DAIR Emotion](https://hanzo.ai/benchmarks/kai-dair-emotion) · [Banking77](https://hanzo.ai/benchmarks/kai-banking77) · [Support triage](https://hanzo.ai/benchmarks/kai-support-triage) · [Email spam](https://hanzo.ai/benchmarks/kai-email-spam) · [Phishing](https://hanzo.ai/benchmarks/kai-phishing) · [Jailbreak](https://hanzo.ai/benchmarks/kai-jailbreak) · [Toxicity](https://hanzo.ai/benchmarks/kai-toxicity) · [RAG relevance](https://hanzo.ai/benchmarks/kai-rag-relevance) · [Model routing](https://hanzo.ai/benchmarks/kai-model-routing) · [Typed decisions](https://hanzo.ai/benchmarks/kai-typed-decisions) · [MASSIVE](https://hanzo.ai/benchmarks/kai-massive) · [AG News, validation split](https://hanzo.ai/benchmarks/kai-ag-news-validation-split) · [Support triage, validation split](https://hanzo.ai/benchmarks/kai-support-triage-validation-split) · [Typed decisions, validation split](https://hanzo.ai/benchmarks/kai-typed-decisions-validation-split) · [HarmBench (benign prompts)](https://hanzo.ai/benchmarks/kai-harmbench) · [Choosing among many options](https://hanzo.ai/benchmarks/kai-choice-size) · [Joint decisions](https://hanzo.ai/benchmarks/kai-joint) · [Latency of one decision](https://hanzo.ai/benchmarks/kai-latency)

[Every result for Release gate against Jev on the shelf](https://hanzo.ai/benchmarks?q=Release%20gate%20against%20Jev#all) · [How these are measured](https://hanzo.ai/benchmarks#methods)
