# HarmBench (benign prompts) benchmark · Kai — Hanzo AI

> The share of HarmBench’s benign prompts each model leaves unflagged: the row Decision Index v2 scores in place of the jailbreak suite.

[Benchmarks](https://hanzo.ai/benchmarks) / [Kai](https://hanzo.ai/benchmarks?product=Kai#all) / HarmBench (benign prompts)

Kai · Decision harness

# HarmBench (benign prompts)

The share of HarmBench’s benign prompts each model leaves unflagged: the row Decision Index v2 scores in place of the jailbreak suite.

## Latest

[98.6%Accuracy · kai-1 · 0834a74fJev 83.6% · Win · sealed board, undated · sealed board](#history)

## History

One result so far. Its history fills in as later runs are posted.

Measured

Subject

Metric

Value

Baseline

Result

Status

sealed board, undated

kai-1 · 0834a74f

Accuracy

98.6%

Jev 83.6%

Win

Frozen · sealed board

## Conditions

benign prompts only, none an attackone threshold, the sealed jailbreak board’s answersa count from the board, not a research run

## Hardware

The run records no device or host.

## Source

[Decision Index v2 on Kai’s page, with the board’s note](https://hanzo.ai/kai#benchmarks)

## Reproduce

No public command regenerates this result. The source above is the record it is read from.

## More Kai benchmarks

[AG News](https://hanzo.ai/benchmarks/kai-ag-news) · [DAIR Emotion](https://hanzo.ai/benchmarks/kai-dair-emotion) · [Banking77](https://hanzo.ai/benchmarks/kai-banking77) · [Support triage](https://hanzo.ai/benchmarks/kai-support-triage) · [Email spam](https://hanzo.ai/benchmarks/kai-email-spam) · [Phishing](https://hanzo.ai/benchmarks/kai-phishing) · [Jailbreak](https://hanzo.ai/benchmarks/kai-jailbreak) · [Toxicity](https://hanzo.ai/benchmarks/kai-toxicity) · [RAG relevance](https://hanzo.ai/benchmarks/kai-rag-relevance) · [Model routing](https://hanzo.ai/benchmarks/kai-model-routing) · [Typed decisions](https://hanzo.ai/benchmarks/kai-typed-decisions) · [MASSIVE](https://hanzo.ai/benchmarks/kai-massive) · [AG News, validation split](https://hanzo.ai/benchmarks/kai-ag-news-validation-split) · [Support triage, validation split](https://hanzo.ai/benchmarks/kai-support-triage-validation-split) · [Typed decisions, validation split](https://hanzo.ai/benchmarks/kai-typed-decisions-validation-split) · [Choosing among many options](https://hanzo.ai/benchmarks/kai-choice-size) · [Joint decisions](https://hanzo.ai/benchmarks/kai-joint) · [Release gate against Jev](https://hanzo.ai/benchmarks/kai-release-gate) · [Latency of one decision](https://hanzo.ai/benchmarks/kai-latency)

[Every result for HarmBench (benign prompts) on the shelf](https://hanzo.ai/benchmarks?q=HarmBench%20(benign%20prompts)#all) · [How these are measured](https://hanzo.ai/benchmarks#methods)
