Try Hanzo

Benchmarks / Zen 6 / Zen 6 throughput

Zen 6 · Throughput on Hanzo hardware

Zen 6 throughput

Zen 6, the latest open-weight generation, on three machines Hanzo owns: prompt prefill and single-request decode, in tokens per second, one machine each.

Latest

History

8 results across 3 metrics, newest first. Superseded runs stay, so a revision's progress can be read.

MeasuredSubjectMetricValueBaselineResultStatus
Sep 23, 2026Zen 6 · NVIDIA Grace-BlackwellAlso measured50.6 tok/s decoding code · 112 tok/s across 8 requests–
Unscored
Frozen · Hanzo-measured
Sep 23, 2026Zen 6 · AMD Strix HaloAlso measured262K context, in two slots–
Unscored
Frozen · Hanzo-measured
Sep 23, 2026Zen 6 · NVIDIA Grace-BlackwellDecode throughput~43 tok/s–
Unscored
Frozen · Hanzo-measured
Sep 23, 2026Zen 6 · AMD Strix HaloDecode throughput~34 tok/s–
Unscored
Frozen · Hanzo-measured
Sep 23, 2026Zen 6 · Apple siliconDecode throughput64–81 tok/s–
Unscored
Frozen · Hanzo-measured
Sep 23, 2026Zen 6 · NVIDIA Grace-BlackwellPrefill throughput~2,400 tok/s–
Unscored
Frozen · Hanzo-measured
Sep 23, 2026Zen 6 · AMD Strix HaloPrefill throughput1,150–1,400 tok/s–
Unscored
Frozen · Hanzo-measured
Sep 23, 2026Zen 6 · Apple siliconPrefill throughput500–690 tok/s–
Unscored
Frozen · Hanzo-measured

Conditions

NVIDIA Grace-Blackwell: 121 GB unified, NVFP4 weightsAMD Strix Halo: 128 GB unified, 4-bit buildApple silicon: unified memory, MLX builddecode is one request unless a row says otherwisea range is what the runs spread over; the low end is the figure ordered and drawn

Hardware

NVIDIA Grace-Blackwell · AMD Strix Halo · Apple silicon

Source

Zen 6, measured by Hanzo on Sep 23, 2026

Reproduce

No public command regenerates this result. The source above is the record it is read from.

More Zen 6 benchmarks

Inference vs llama.cpp

Every result for Zen 6 throughput on the shelf · How these are measured

Build what’s next.