Try Hanzo
Zen · Open models

Open models.
Yours to run.

Zen is the open-weight family behind Hanzo: seven generations on one line, from a laptop to the datacenter. Download the weights, run them privately, or call the same family through Hanzo Cloud.

Explore the lineage
01
02
03
04
05
5.8
06
Now
07
Next
Available nowZen 627.3B dense. 262K context. Text, images and video.Run Zen 6
Research previewZen 7The next generation, open to researchers by request.Request access

Available now

Zen 6.
Open now.

A 27.3 billion parameter dense model. Hybrid attention, a long context, images and video, and a speculative drafter bundled with the weights. The interesting part is not that it is the biggest. It is that you can run it.

Agentic coding

Long repositories, tool use, long context, and fast speculative decoding.

Multimodal work

Zen 6 reads text, images, and video in one context.
27.3BParameters
262KNative context
1MWith YaRN
On one DGX Spark, the weight card records 62.4 tokens per second on code alone and 141.2 with the drafter. That figure is published with the weights. It is not a Hanzo harness.

Zen 6 Flash

27B.
Laptop-sized.

The same generation, packed ternary. 26,895,998,464 language-model weights in a 5.95 GB file, 1.77 bits per weight, 9.0 times smaller than the 53.79 GB FP16 representation. Text and images. Not video.

FP16 representation53.79 GB
Ternary1.77 bits / weight
5.95 GBDense packed
5.95 GBDense packed
1.77Bits per weight
262KNative context
Apple Silicon or one GPU, through llama.cpp and Metal. The card records 47.2 tokens per second alone and 92.4 with the drafter on M4 and M5 Max, in 7.84 GB with the projector and a 32K cache. Published with the weights. Not a Hanzo harness.

Measured by Hanzo

A frontier model.
On your desk.

Hanzo ran Zen 6 on three machines it owns. These are the readings, in tokens a second.

NVIDIA Grace-Blackwell

121 GB unified · NVFP4 weights
~2,400tok/s prefill
~43tok/s decode
50.6 tok/s decoding code · 112 tok/s across 8 requests

AMD Strix Halo

128 GB unified · 4-bit build
1,150–1,400tok/s prefill
~34tok/s decode
262K context, in two slots

Apple silicon

unified memory · MLX build
500–690tok/s prefill
64–81tok/s decode
Hanzo's runs, 23 September 2026. Decode is one request unless a line says otherwise, and each machine runs its own build of Zen 6. Where the Hanzo engine's prefill on CUDA beats llama.cpp is on the inference benchmark.

Open weights

APIs can disappear.
Weights don't.

Download them, read them, fine-tune them, and run them where you want. Zen 6 is Apache-2.0, published by Zoo Labs Foundation, a 501(c)(3) non-profit.

Zen weightsYour laptopYour GPUHanzo Cloud

2023 to now

Seven generations. One line.

Foundation, then context, then a family, then scale, then the models that actually shipped. Each generation inherits the last, and the current one is the practical one.

01

Zen 1

Q4 2023
HistoricalFoundationDense multilingual experiments at 7B, 14B and 72B, to prove an open foundation could scale. These are not models the gateway serves.One dense block.
02

Zen 2

Q2 2024
HistoricalGo longer128K context, the first production deployments, and code and math specialists. From research model to production model. The context got longer. The architecture did not get more decorative.The same block, wider.
03

Zen 3

2024
Open weightsMore than languageThe line branched: sparse experts, vision, audio, image, retrieval and safety. This is where Zen becomes a family. None of these ids are in the live catalog. Open-weight specialist artifacts are still published.One line, six branches.
LanguageVisionAudioImageRetrievalSafety
04

Zen 4

2025
RetiredScale upFrontier-scale experts and a longer context. Retired. Not a current production model, and not a public download.Large, and sparse.
05

Zen 5

ShippedAgentic generationCode, long context, and smaller and host-specific variants. Older pages described a much larger routed checkpoint. It was not published. These are the models that were.Paths for tools and hosts.
zen5 zen5-coder zen5-evo zen5-flash zen5-mini zen5-pro zen5-spark
5.8

Zen 5.8

Served onlyServing bridgeA served-only bridge between Zen 5 and Zen 6, with no public weights. The gateway does not list it now.A thin serving step.
06

Zen 6

CurrentDense againZen 6 is a practical dense model: hybrid attention, a long context, images and video, speculative decoding, and Zen 6 Flash, a ternary build that fits a laptop. Available now.Smaller, and solid.
zen6 zen6-flashThe current generation
07

Zen 7

Research previewResearch previewZen 7 has no public weights and no callable id. Its architecture, size, modalities and date are not stated. Research preview — not currently available.The release path, unresolved past the preview.

What comes next is still research.

Train
Attest
EvaluateNot published
ReleaseNot published
The path to a release is training, then attestation, then evaluation, then publication. Nothing past the preview is published.

Today

Every sense.
One family.

Zen MoDE, a Mixture of Diverse Experts: a model for each job, from generating and coding to seeing, speaking, retrieving and guarding. Grouped by job, from the models the gateway lists. Hosted means the id is in the catalog. Open weights means a public repository for that exact id.

Generate

zen6
CurrentOpen weightsHosted
API 1MWeights
zen6-flash
CurrentOpen weightsHosted
API 1MWeights
zen-free
Hosted
API 1M
zen5
Hosted
API 1M
zen5-evo
Hosted
API 1M
API 1M
zen5-mini
Hosted
API 1M
zen5-pro
Hosted
API 1M
API 1M

Code

API 1M

See

zen-vl
Hosted
API 1M

Speak and listen

Retrieve

API 33K
API 8K
API 33K

Guard

zen-guard
Hosted
API 128K

Two artifacts

Open weights, and a hosted id.

The same name can be two things. A hosted id is not documented as the same bytes as a public repository with a similar name.

Open weights

Download. Self-host. Modify.

Hosted

An API call. Managed. Metered.

Call it

One call.
Any SDK.

The gateway speaks the OpenAI API, so the SDK you already use works. Point it at api.hanzo.ai and ask for "zen6". The weights are the download.

python
from openai import OpenAI

client = OpenAI(
    base_url="https://api.hanzo.ai/v1",
    api_key="hk-...",  # your Hanzo API key
)

resp = client.chat.completions.create(
    model="zen6",
    messages=[{"role": "user", "content": "Hello"}],
)
print(resp.choices[0].message.content)
bash
curl https://api.hanzo.ai/v1/chat/completions \
  -H "Authorization: Bearer $HANZO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "zen6", "messages": [{"role": "user", "content": "Hello"}]}'

Run it yourself

Three ways.

The commands are the ones on the current weight cards.

Download

hf download zenlm/zen6
hf download zenlm/zen6-flash

Local

Zen 6 Flash. llama.cpp and Metal. Apple Silicon or one GPU.
hf download zenlm/zen6-flash
The file names are the loader's.
bash
llama-server \
  -m Ternary-Bonsai-2-27B-PQ2_0.gguf \
  --mmproj Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf \
  -md Bonsai-2-27B-DFlash2-Q8_0.gguf \
  --spec-draft-n-max 3 \
  -c 32768 \
  --port 8080

Production

Zen 6. SGLang, with the bundled drafter.
bash
hf download zenlm/zen6 --local-dir zen6

python3 -m sglang.launch_server \
  --model-path ./zen6 \
  --speculative-draft-model-path ./zen6/dflash2 \
  --speculative-num-steps 3 \
  --speculative-algorithm DFLASH \
  --kv-cache-dtype fp8_e5m2 \
  --context-length 1048576 \
  --port 30000 \
  --host 0.0.0.0

Specifications

Zen 6 beside Zen 6 Flash.

Card facts only. A cell with no published value is left blank.

Zen 6Zen 6 Flash
Parameters
Zen 627.3 billion, dense, plus a vision encoder
Zen 6 Flash26,895,998,464 language-model weights
Weight format
Zen 6NVFP4 on the MLP and LM head, FP8 on attention
Zen 6 FlashTernary, 1.77 bits per weight, 5.95 GB packed
Native context
Zen 6262,144
Zen 6 Flash262,144
Extended context
Zen 61,048,576 with YaRN
Zen 6 Flash—
Text
Zen 6Yes
Zen 6 FlashYes
Images
Zen 6Yes
Zen 6 FlashYes
Video
Zen 6Yes
Zen 6 Flash—
Runs on
Zen 6A GPU, with SGLang. Measured on one DGX Spark.
Zen 6 FlashA laptop or one GPU, with llama.cpp and Metal
The API catalog lists 1M for zen6 and 1M for zen6-flash. The weight cards state 262,144 tokens native, and YaRN to 1,048,576 for Zen 6 only.
Layers, heads, and the files
Zen 6 has 64 layers, 48 of them linear attention and 16 full attention, one full layer every fourth. Hidden size 5,120. 24 query heads and 4 key and value heads, head size 256. Vocabulary 248,320. The speculative drafter ships in the weights.Zen 6 Flash also publishes a 7.21 GB file at 2.14 bits per weight, which is the one the kernels read, a 629 MB vision projector, and a 2.06 GB drafter. With the projector and a 32K cache it stays under 8 GB.

Scored

Scored where
it counts.

Hanzo scores the family on GPQA-Diamond, 198 graduate-level science questions. The publisher scores the open weights. The two are never mixed.

91.4%Zen 5.8 · GPQA-Diamond
90.8%Zen 5.8 Coder · GPQA-Diamond
98.2%Kept by Zen 6 Flash's ternary build
Zen 5.8 and Zen 5.8 Coder are Hanzo-scored, at $5.40 per million output tokens. The weights report 86.33 at full precision and 84.78 for the ternary build across 14 reasoning suites: the publisher's numbers, not a Hanzo harness. Zen 6 is not on the board yet, so this page does not borrow one.

Your model should
survive your vendor.

Weights you can hold, runtimes you can choose, infrastructure you control.

DownloadInspectFine-tuneDeployFork

Zen reasons.

Open weights, yours to run. Routing and decisions are the other two families. They are compared on the models page.

Run Zen where you want.

Download the weights, run Zen privately, or use the same family through Hanzo.