Open models.
Yours to run.
Zen is the open-weight family behind Hanzo: seven generations on one line, from a laptop to the datacenter. Download the weights, run them privately, or call the same family through Hanzo Cloud.
Explore the lineageAvailable now
Zen 6.
Open now.
A 27.3 billion parameter dense model. Hybrid attention, a long context, images and video, and a speculative drafter bundled with the weights. The interesting part is not that it is the biggest. It is that you can run it.
Agentic coding
Long repositories, tool use, long context, and fast speculative decoding.Multimodal work
Zen 6 reads text, images, and video in one context.Zen 6 Flash
27B.
Laptop-sized.
The same generation, packed ternary. 26,895,998,464 language-model weights in a 5.95 GB file, 1.77 bits per weight, 9.0 times smaller than the 53.79 GB FP16 representation. Text and images. Not video.
Measured by Hanzo
A frontier model.
On your desk.
Hanzo ran Zen 6 on three machines it owns. These are the readings, in tokens a second.
NVIDIA Grace-Blackwell
121 GB unified · NVFP4 weightsAMD Strix Halo
128 GB unified · 4-bit buildApple silicon
unified memory · MLX buildOpen weights
APIs can disappear.
Weights don't.
Download them, read them, fine-tune them, and run them where you want. Zen 6 is Apache-2.0, published by Zoo Labs Foundation, a 501(c)(3) non-profit.
2023 to now
Seven generations. One line.
Foundation, then context, then a family, then scale, then the models that actually shipped. Each generation inherits the last, and the current one is the practical one.
Zen 1
Q4 2023Zen 2
Q2 2024Zen 3
2024Zen 4
2025Zen 5
Zen 5.8
Zen 6
Zen 7
What comes next is still research.
Today
Every sense.
One family.
Zen MoDE, a Mixture of Diverse Experts: a model for each job, from generating and coding to seeing, speaking, retrieving and guarding. Grouped by job, from the models the gateway lists. Hosted means the id is in the catalog. Open weights means a public repository for that exact id.
Generate
Code
See
Guard
Two artifacts
Open weights, and a hosted id.
The same name can be two things. A hosted id is not documented as the same bytes as a public repository with a similar name.
Open weights
Download. Self-host. Modify.Hosted
An API call. Managed. Metered.Call it
One call.
Any SDK.
The gateway speaks the OpenAI API, so the SDK you already use works. Point it at api.hanzo.ai and ask for "zen6". The weights are the download.
from openai import OpenAI
client = OpenAI(
base_url="https://api.hanzo.ai/v1",
api_key="hk-...", # your Hanzo API key
)
resp = client.chat.completions.create(
model="zen6",
messages=[{"role": "user", "content": "Hello"}],
)
print(resp.choices[0].message.content)curl https://api.hanzo.ai/v1/chat/completions \
-H "Authorization: Bearer $HANZO_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "zen6", "messages": [{"role": "user", "content": "Hello"}]}'Run it yourself
Three ways.
The commands are the ones on the current weight cards.
Download
Local
Zen 6 Flash. llama.cpp and Metal. Apple Silicon or one GPU.llama-server \
-m Ternary-Bonsai-2-27B-PQ2_0.gguf \
--mmproj Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf \
-md Bonsai-2-27B-DFlash2-Q8_0.gguf \
--spec-draft-n-max 3 \
-c 32768 \
--port 8080Production
Zen 6. SGLang, with the bundled drafter.hf download zenlm/zen6 --local-dir zen6
python3 -m sglang.launch_server \
--model-path ./zen6 \
--speculative-draft-model-path ./zen6/dflash2 \
--speculative-num-steps 3 \
--speculative-algorithm DFLASH \
--kv-cache-dtype fp8_e5m2 \
--context-length 1048576 \
--port 30000 \
--host 0.0.0.0Specifications
Zen 6 beside Zen 6 Flash.
Card facts only. A cell with no published value is left blank.
Layers, heads, and the files
Scored
Scored where
it counts.
Hanzo scores the family on GPQA-Diamond, 198 graduate-level science questions. The publisher scores the open weights. The two are never mixed.
Your model should
survive your vendor.
Weights you can hold, runtimes you can choose, infrastructure you control.
Zen reasons.
Open weights, yours to run. Routing and decisions are the other two families. They are compared on the models page.
Read
Models, benchmarks, research.
The weights you can run, where they are scored, and the writing.
Run Zen where you want.
Download the weights, run Zen privately, or use the same family through Hanzo.