# Support triage, validation split benchmark · Kai — Hanzo AI

> Kai on Support triage, validation split, one run per revision. No baseline was run on this split.

[Benchmarks](https://hanzo.ai/benchmarks) / [Kai](https://hanzo.ai/benchmarks?product=Kai#all) / Support triage, validation split

Kai · Validation split

# Support triage, validation split

Kai on Support triage, validation split, one run per revision. No baseline was run on this split.

## Latest

[64.9%Accuracy · Kai triage+job_30bfc6e5af09a9f7 · 0cac9e85no baseline · Unscored · Sep 29, 2026 · Hanzo-measured](#history)

[0.282Calibration error (ECE) · Kai triage+job_30bfc6e5af09a9f7 · 0cac9e85no baseline · Unscored · Sep 29, 2026 · Hanzo-measured](#history)

## History

6 results across 2 metrics, newest first. Superseded runs stay, so a revision&#x27;s progress can be read.

Measured

Subject

Metric

Value

Baseline

Result

Status

Sep 29, 2026

Kai triage+job_30bfc6e5af09a9f7 · 0cac9e85

Accuracy

64.9%

–

Unscored

Live · Hanzo-measured

Sep 29, 2026

Kai a7 · 0834a74f

Accuracy

66.4%

–

Unscored

Superseded · Hanzo-measured

Sep 29, 2026

Kai triage+job_cccd2fcfeae9ecb9 · 57535ef4

Accuracy

65.8%

–

Unscored

Superseded · Hanzo-measured

Sep 29, 2026

Kai triage+job_30bfc6e5af09a9f7 · 0cac9e85

Calibration error (ECE)

0.282

–

Unscored

Live · Hanzo-measured

Sep 29, 2026

Kai a7 · 0834a74f

Calibration error (ECE)

0.262

–

Unscored

Superseded · Hanzo-measured

Sep 29, 2026

Kai triage+job_cccd2fcfeae9ecb9 · 57535ef4

Calibration error (ECE)

0.278

–

Unscored

Superseded · Hanzo-measured

## Conditions

split: val:0b8b6476822b343c518 questions, 518 answereddataset sha256:0b8b63 Kai revisions measured

## Hardware

ROCm BF16 on evo

## Source

[research run kai.support_triage-val:0b8b6476822b343c-triage+job_30bfc6e5af09a9f7.0cac9e85-kai](https://api.hanzo.ai/v1/research/runs?org=hanzo&id=kai.support_triage-val%3A0b8b6476822b343c-triage%2Bjob_30bfc6e5af09a9f7.0cac9e85-kai)

## Reproduce

No public command regenerates this result. The source above is the record it is read from.

## More Kai benchmarks

[AG News](https://hanzo.ai/benchmarks/kai-ag-news) · [DAIR Emotion](https://hanzo.ai/benchmarks/kai-dair-emotion) · [Banking77](https://hanzo.ai/benchmarks/kai-banking77) · [Support triage](https://hanzo.ai/benchmarks/kai-support-triage) · [Email spam](https://hanzo.ai/benchmarks/kai-email-spam) · [Phishing](https://hanzo.ai/benchmarks/kai-phishing) · [Jailbreak](https://hanzo.ai/benchmarks/kai-jailbreak) · [Toxicity](https://hanzo.ai/benchmarks/kai-toxicity) · [RAG relevance](https://hanzo.ai/benchmarks/kai-rag-relevance) · [Model routing](https://hanzo.ai/benchmarks/kai-model-routing) · [Typed decisions](https://hanzo.ai/benchmarks/kai-typed-decisions) · [MASSIVE](https://hanzo.ai/benchmarks/kai-massive) · [AG News, validation split](https://hanzo.ai/benchmarks/kai-ag-news-validation-split) · [Typed decisions, validation split](https://hanzo.ai/benchmarks/kai-typed-decisions-validation-split) · [HarmBench (benign prompts)](https://hanzo.ai/benchmarks/kai-harmbench) · [Choosing among many options](https://hanzo.ai/benchmarks/kai-choice-size) · [Joint decisions](https://hanzo.ai/benchmarks/kai-joint) · [Release gate against Jev](https://hanzo.ai/benchmarks/kai-release-gate) · [Latency of one decision](https://hanzo.ai/benchmarks/kai-latency)

[Every result for Support triage, validation split on the shelf](https://hanzo.ai/benchmarks?q=Support%20triage%2C%20validation%20split#all) · [How these are measured](https://hanzo.ai/benchmarks#methods)
