# The Zen family — Hanzo AI

> The Zen family: open-weight frontier models you can self-host anywhere, built by Zoo Labs Foundation and served on the Hanzo API. Benchmarks shown are UPSTREAM-reported for the open ecosystem — only Enso is Hanzo-measured end-to-end.

[Models](https://hanzo.ai/models)/Zen

Weights published · run them anywhere

# The Zen family

45 open-weight models across language, code, vision, image, audio, and retrieval — built by [Zoo Labs Foundation](https://zoo.industries). Free to self-host, or managed on Hanzo Cloud. Benchmarks here are UPSTREAM-reported for the open ecosystem Zen builds on; only Enso is Hanzo-measured end-to-end.

[Full Zen catalog](https://hanzo.ai/zen/models)[Get API key](https://console.hanzo.ai)

## Generations, and what each one is for

A generation is a training run, not a marketing tier: newer does not mean the older ones stop working, and a small model from an older generation is often the right call. The large ones are Zen MoDE — Mixture of Diverse Experts. Open any generation for specs and prices.

[Zen5The current frontier run6models](https://hanzo.ai/zen/models)[Zen3Vision, audio, and the specialists8models](https://hanzo.ai/zen/models)[FoundationCheckpoints to fine-tune from31models](https://hanzo.ai/zen/models)

## Where open weights stand

44 open-weight models across the field, by benchmark, so you can see what running your own hardware costs you in capability before you commit to it. A figure is what its vendor published unless it is tagged Hanzo, which means we ran it ourselves. Toggle the provenance to see which is which.

GPQA-Diamond

Humanity&#x27;s Last Exam

MMLU-Pro

LiveCodeBench v6

LiveCodeBench Pro

SWE-Bench Pro

Terminal-Bench 2.1

SciCode

MRCR v2

All

Hanzo-measured

Vendor-reported

Model

GPQA-Diamond

Source

$/MTok out

kimi-k2.6

89.1

Vals AI

$2.71

qwen3.5-397b-a17b

88.4

LLM Stats

$2.04

nemotron-3-ultra-550b-a55b

86.1

Vals AI

$1.54

glm-5.2

85.6

Vals AI

$3.73

glm-5.1

84.5

Vals AI

$3.63

kimi-k2.5

84.1

Vals AI

$1.69

glm-5

83.3

Vals AI

$2.07

nemotron-3-super

82.7

LLM Stats

$0.4

minimax-m2.5

82.1

Vals AI

$0.76

mimo-v2.5

81.6

Vals AI

$0.24

deepseek-v3.2

80.3

Vals AI

—

deepseek-v3.2-exp

79.9

DeepSeek-V3.2-Exp model …

—

deepseek-4-flash

76.9

Hanzo

$0.2

deepseek-v4-pro

75.3

Hanzo

$2.5

deepseek-r1

71.5

DeepSeek-R1 model card (…

—

llama-4-maverick

69.4

Vals AI

$0.75

llama-3.3-70b-instruct

50.5

LLM Stats

—

Vendors report on their own harness; Hanzo measures everyone on one. Where both exist the gap is the harness talking — not the model getting better. Hover a source for its exact provenance.

Why the two labels. A score is only comparable to another score run the same way, and most published figures were not. So the table keeps the distinction rather than averaging it away: a figure we ran on our own harness says Hanzo, and a figure a vendor published says so too. Enso is the family we measure end to end on one common harness — see it on the Enso page. Everything else on this page is cited, not claimed.

[Full Zen catalog](https://hanzo.ai/zen/models)[See Enso (measured)](https://hanzo.ai/models/enso)
