# Choosing among many options benchmark · Kai — Hanzo AI

> Accuracy as the number of listed options grows, the same cases put to both models. Past 160 options Kai shortlists by retrieval first.

[Benchmarks](https://hanzo.ai/benchmarks) / [Kai](https://hanzo.ai/benchmarks?product=Kai#all) / Choosing among many options

Kai · Capability

# Choosing among many options

Accuracy as the number of listed options grows, the same cases put to both models. Past 160 options Kai shortlists by retrieval first.

## Latest

[96.8%Accuracy · kai-1 · 0834a74f · 4 optionsJev 90.5% · Win · Oct 3, 2026 · Hanzo-measured](#history)

[93.5%Accuracy · kai-1 · 0834a74f · 16 optionsJev 85.3% · Win · Oct 3, 2026 · Hanzo-measured](#history)

[90.8%Accuracy · kai-1 · 0834a74f · 77 optionsJev 84.3% · Win · Oct 3, 2026 · Hanzo-measured](#history)

[90.8%Accuracy · kai-1 · 0834a74f · 150 optionsJev 94.0% · Loss · Oct 3, 2026 · Hanzo-measured](#history)

## History

7 results, newest first. Superseded runs stay, so a revision&#x27;s progress can be read.

Measured

Subject

Metric

Value

Baseline

Result

Status

Oct 3, 2026

kai-1 · 0834a74f · 4 options

Accuracy

96.8%

Jev 90.5%

Win

Live · Hanzo-measured

Oct 3, 2026

kai-1 · 0834a74f · 16 options

Accuracy

93.5%

Jev 85.3%

Win

Live · Hanzo-measured

Oct 3, 2026

kai-1 · 0834a74f · 77 options

Accuracy

90.8%

Jev 84.3%

Win

Live · Hanzo-measured

Oct 3, 2026

kai-1 · 0834a74f · 150 options

Accuracy

90.8%

Jev 94.0%

Loss

Live · Hanzo-measured

Oct 3, 2026

kai-1 · 0834a74f · 1,000 options

Accuracy

0.5%

Jev refused

Unscored

Live · Hanzo-measured

Oct 3, 2026

kai-1 · 0834a74f · 10,000 options

Accuracy

0.0%

Jev refused

Unscored

Live · Hanzo-measured

Oct 3, 2026

kai-1 · 0834a74f · 100,000 options

Accuracy

0.0%

Jev refused

Unscored

Live · Hanzo-measured

## Conditions

choice sizes: 4, 16, 77, 150, 1,000, 10,000, 100,000every option scored up to 160, shortlisted past itJev: typesafe/jev-1.13, measured Sep 27, 2026

## Hardware

CUDA F32 on dgx

## Source

[research run cap.cardinality-cases-kai-1-kai](https://api.hanzo.ai/v1/research/runs?org=hanzo&id=cap.cardinality-cases-kai-1-kai)

## Reproduce

`bench cap --harness . --who kai,jev --kai # in hanzoai/benchmarks/decision`

## More Kai benchmarks

[AG News](https://hanzo.ai/benchmarks/kai-ag-news) · [DAIR Emotion](https://hanzo.ai/benchmarks/kai-dair-emotion) · [Banking77](https://hanzo.ai/benchmarks/kai-banking77) · [Support triage](https://hanzo.ai/benchmarks/kai-support-triage) · [Email spam](https://hanzo.ai/benchmarks/kai-email-spam) · [Phishing](https://hanzo.ai/benchmarks/kai-phishing) · [Jailbreak](https://hanzo.ai/benchmarks/kai-jailbreak) · [Toxicity](https://hanzo.ai/benchmarks/kai-toxicity) · [RAG relevance](https://hanzo.ai/benchmarks/kai-rag-relevance) · [Model routing](https://hanzo.ai/benchmarks/kai-model-routing) · [Typed decisions](https://hanzo.ai/benchmarks/kai-typed-decisions) · [MASSIVE](https://hanzo.ai/benchmarks/kai-massive) · [AG News, validation split](https://hanzo.ai/benchmarks/kai-ag-news-validation-split) · [Support triage, validation split](https://hanzo.ai/benchmarks/kai-support-triage-validation-split) · [Typed decisions, validation split](https://hanzo.ai/benchmarks/kai-typed-decisions-validation-split) · [HarmBench (benign prompts)](https://hanzo.ai/benchmarks/kai-harmbench) · [Joint decisions](https://hanzo.ai/benchmarks/kai-joint) · [Release gate against Jev](https://hanzo.ai/benchmarks/kai-release-gate) · [Latency of one decision](https://hanzo.ai/benchmarks/kai-latency)

[Every result for Choosing among many options on the shelf](https://hanzo.ai/benchmarks?q=Choosing%20among%20many%20options#all) · [How these are measured](https://hanzo.ai/benchmarks#methods)
