# Zen 6 throughput benchmark · Zen 6 — Hanzo AI

> Zen 6, the latest open-weight generation, on three machines Hanzo owns: prompt prefill and single-request decode, in tokens per second, one machine each.

[Benchmarks](https://hanzo.ai/benchmarks) / [Zen 6](https://hanzo.ai/benchmarks?product=Zen%206#all) / Zen 6 throughput

Zen 6 · Throughput on Hanzo hardware

# Zen 6 throughput

Zen 6, the latest open-weight generation, on three machines Hanzo owns: prompt prefill and single-request decode, in tokens per second, one machine each.

## Latest

[50.6 tok/s decoding code · 112 tok/s across 8 requestsAlso measured · Zen 6 · NVIDIA Grace-Blackwellno baseline · Unscored · Sep 23, 2026 · Hanzo-measured](#history)

[262K context, in two slotsAlso measured · Zen 6 · AMD Strix Halono baseline · Unscored · Sep 23, 2026 · Hanzo-measured](#history)

[~43 tok/sDecode throughput · Zen 6 · NVIDIA Grace-Blackwellno baseline · Unscored · Sep 23, 2026 · Hanzo-measured](#history)

[~2,400 tok/sPrefill throughput · Zen 6 · NVIDIA Grace-Blackwellno baseline · Unscored · Sep 23, 2026 · Hanzo-measured](#history)

## History

8 results across 3 metrics, newest first. Superseded runs stay, so a revision&#x27;s progress can be read.

Measured

Subject

Metric

Value

Baseline

Result

Status

Sep 23, 2026

Zen 6 · NVIDIA Grace-Blackwell

Also measured

50.6 tok/s decoding code · 112 tok/s across 8 requests

–

Unscored

Frozen · Hanzo-measured

Sep 23, 2026

Zen 6 · AMD Strix Halo

Also measured

262K context, in two slots

–

Unscored

Frozen · Hanzo-measured

Sep 23, 2026

Zen 6 · NVIDIA Grace-Blackwell

Decode throughput

~43 tok/s

–

Unscored

Frozen · Hanzo-measured

Sep 23, 2026

Zen 6 · AMD Strix Halo

Decode throughput

~34 tok/s

–

Unscored

Frozen · Hanzo-measured

Sep 23, 2026

Zen 6 · Apple silicon

Decode throughput

64–81 tok/s

–

Unscored

Frozen · Hanzo-measured

Sep 23, 2026

Zen 6 · NVIDIA Grace-Blackwell

Prefill throughput

~2,400 tok/s

–

Unscored

Frozen · Hanzo-measured

Sep 23, 2026

Zen 6 · AMD Strix Halo

Prefill throughput

1,150–1,400 tok/s

–

Unscored

Frozen · Hanzo-measured

Sep 23, 2026

Zen 6 · Apple silicon

Prefill throughput

500–690 tok/s

–

Unscored

Frozen · Hanzo-measured

## Conditions

NVIDIA Grace-Blackwell: 121 GB unified, NVFP4 weightsAMD Strix Halo: 128 GB unified, 4-bit buildApple silicon: unified memory, MLX builddecode is one request unless a row says otherwisea range is what the runs spread over; the low end is the figure ordered and drawn

## Hardware

NVIDIA Grace-Blackwell · AMD Strix Halo · Apple silicon

## Source

[Zen 6, measured by Hanzo on Sep 23, 2026](https://hanzo.ai/zen)

## Reproduce

No public command regenerates this result. The source above is the record it is read from.

## More Zen 6 benchmarks

[Inference vs llama.cpp](https://hanzo.ai/benchmarks/inference)

[Every result for Zen 6 throughput on the shelf](https://hanzo.ai/benchmarks?q=Zen%206%20throughput#all) · [How these are measured](https://hanzo.ai/benchmarks#methods)
