Try Hanzo
Simple & Transparent

Pricing

Pay for what the agents use. One plan per organization, then metered per call for what the agents consume. No per-seat fee. Nothing at all for a machine that is sitting still.

Most Popular

Plan

One plan per organization, then billed for what the agents consume.

$19.00per month, per organization
  • No per-seat fee, however many people use the organization.
  • Usage metered per tier, booked in integer micro-USD
  • MCP, SDK and CLI — $0
Start — $19.00 / month

Budgeted

Every agent carries a budget. Model calls and renders are quoted before they run, so anything that would breach it is refused rather than discovered on an invoice.

You set the ceiling
  • A cap per period and a cap per task
  • A render with no headroom is refused, not run
  • A failed job is not billed and its hold comes back
  • Compute and web calls meter after the fact, on the ledger either way
How budgets work

Custom

Committed volume, procurement, or anything the meter does not cover.

Talk to us
  • Committed-volume pricing
  • Invoicing and PO terms
  • A monthly allowance arranged against your account
  • Security review and DPA
Contact sales

What a bill is made of

Four components, each on your spend breakdown under its own name. Two are quoted before they run; two are metered from what they used.

Model inference

per token, five tiers
Quoted before it runs
Input, cached reads, cache writes, output and reasoning are metered as five disjoint tiers, at the tariff for the completion window the request asked for.
inputcache_readcache_writeoutputreasoning

Computer

per second, while running
Quoted before it runs
vCPU and memory on what the machine was given. Nothing to create one, nothing while it sleeps, and the boot disk is inside the allowance — so a computer waiting between turns meters nothing.

Web tools

per call
Metered from what it used
Priced from what the provider charged, so a web call is booked after it runs. A fetch the domain policy refuses is stopped before the request and costs nothing.

Media generation

what the job cost
Metered from what it used
Billed once per finished job at its real cost; a render cannot be quoted before it runs, so it is admitted against the headroom under the agent's per-task cap instead. A job that fails is not billed.

The rates

Every rate here is the price you pay — the number the ledger books and the number your spend breakdown reports.

Rate
Computer · vCPU
$0.0504/hr$0.000014/second
Charged by the second, while running.
Computer · memory
$0.0162/GiB-hr$0.0000045/GiB-second
Charged by the second, on the memory the machine was given.
Computer · paused$0No vCPU, no memory, no storage.
Computer · created$0There is no creation fee.
web_search$0.002158/callOne query, up to ten results.
web_fetch$0.001079/callOne page read as text.
MCP, SDK and CLI$0The interfaces are not metered.
Seats$0No per-seat fee, however many people use the organization.

The completion window is the biggest lever

The same model at a different window is a different price, chosen per request. Pay the interactive tariff only when a person is waiting.

immediateAnswers now, at the highest tariff.
priorityAnswers soon, at a lower one.
looseAnswers eventually, at the lowest.

Every model, at the rate the ledger books

Model inference is metered per token in five disjoint tiers, at the tariff for the completion window the request asked for.

14Zen Models
397Third-Party
432Total Models
57Providers
11Featured
74Free Models

Hanzo Enso

Flagship frontier models — Flash, Pro, and Ultra. Live per-token retail, one API key.

Explore Enso

Enso Flash

pro

Fast, economical Enso tier for high-volume, low-latency everyday work, with 1M-context overflow.

1M context window
Low latency
Input$2 / MTok
Output$4 / MTok
Cache WriteN/A
Cache ReadN/A

Enso

ultra max

Hanzo's proprietary frontier model — Opus-class reasoning by default with 1M-context overflow.

1M context window
Frontier reasoning
Input$4 / MTok
Output$20 / MTok
Cache WriteN/A
Cache ReadN/A

Enso Ultra

ultra max

Adaptive fan-out — probes a task-appropriate model, escalates to a top-K panel only when needed, then verifies-then-selects the best answer.

200K context window
Adaptive fan-out
Input$5 / MTok
Output$25 / MTok
Cache WriteN/A
Cache ReadN/A

Zen Models

34 models across LLM, embedding, reranker, image, audio, and videoUpdated 8/5/2026

Full Model Catalog
ultra max
ultra
pro max
pro

Zen Guard — Content Safety

starter

Safety classifier for moderation and guardrails across a broad category and language set.

128K context
Safety classifier
Input$0.3 / MTok
Output$0.96 / MTok
Cache WriteN/A
Cache ReadFree

Zen VL — Vision-Language

pro

Reads images and reasons over them — visual Q&A, document and chart understanding, grounded captioning.

128K context
Vision + Language
Input$0.09 / MTok
Output$0.39 / MTok
Cache WriteN/A
Cache ReadFree

Zen 5.8

pro

Zen 5.8 flagship frontier general reasoning and agentic foundation model. Advanced chain-of-thought, 1M context, MoDE architecture.

Parameters: 397B (17B active)Architecture: Zen frontier MoE
1M context window
Frontier MoDE architecture
Native chain-of-thought
Autonomous agent loops
Input$1.8 / MTok
Output$5.4 / MTok
Cache WriteN/A
Cache Read$0.35 / MTok

Zen 5.8 Coder

pro

Zen 5.8 Coder flagship autonomous coding and repository understanding model. 1M context, multi-file code synthesis, agentic refactoring.

Parameters: 397B (MoE)Architecture: Zen Coder MoE
1M context window
Code-specialized MoDE
Repository-scale synthesis
Agentic refactoring loops
Input$1.8 / MTok
Output$5.4 / MTok
Cache WriteN/A
Cache Read$0.35 / MTok

Zen5 (Default)

pro

Canonical Zen5 default. a 35B frontier MoE (3B active) (35B total / 3B active per token, released Apr 2026, Apache-2.0). 256K context, agentic-trained, OpenAI + Anthropic API. The everyday Zen5 chat model.

Parameters: 35B (3B active)Architecture: Zen frontier MoE
256K context window
35B total / 3B active (MoE)
Zen base
Apr 2026 release
OpenAI + Anthropic API
Input$2.28 / MTok
Output$7.26 / MTok
Cache WriteN/A
Cache Read$0.42 / MTok

Zen5 Coder

pro

Code-specialized Zen5 tier. 80B sparse MoE tuned for repo-scale code understanding, agentic refactoring, and tool-use coding loops.

Parameters: 80B (MoE)Architecture: Zen Coder MoE
256K context
80B sparse MoE
Code-specialized
Agentic / tool-use
Input$2.28 / MTok
Output$7.26 / MTok
Cache WriteN/A
Cache Read$0.42 / MTok

Zen5 Flash

starter

Smallest and cheapest text-only Zen5 chat tier. 4B-class dense, sub-100ms TTFT, 32K context. For high-volume routing and simple agent loops.

Parameters: 4BArchitecture: Zen dense
32K context window
4B parameters (dense)
Sub-100ms TTFT
Highest throughput
Input$0.2646 / MTok
Output$0.5292 / MTok
Cache WriteN/A
Cache Read$0.05292 / MTok

Zen5 Mini

pro

Frontier agentic at the lowest cost in the family. Built on a 230B agentic MoE (10B active) (230B MoE / 10B active, released Feb 2026). 80.2% SWE-Bench Verified, 76.3% BrowseComp; trained on 200K+ real-world environments via large-scale RL.

Parameters: 230B (10B active)Architecture: Zen MoE
192K context window
230B total / 10B active (MoE)
Zen agentic base
Frontier agentic / coding
Lowest $/token in family
Input$0.09 / MTok
Output$0.39 / MTok
Cache WriteN/A
Cache Read$0.09 / MTok

Zen5 Pro

ultra

Zen Flash IQ2_XXS-imatrix weights (81 GB GGUF on zenlm/zen-5-pro-gguf). 284B total / 37B active per token, 1M context, asymmetric routed-MoE quant. Fits a single 128 GB Apple Silicon / DGX Spark / H100 80 GB.

Parameters: 284B (37B active)Architecture: Zen Flash MoE
1M context window
284B total / 37B active (MoE)
Zen Flash base
IQ2_XXS-imatrix quant (81 GB)
Runs on 128 GB hardware
Input$1.305 / MTok
Output$2.61 / MTok
Cache WriteN/A
Cache Read$0.010875 / MTok

Router General

pro

Automatic model routing — hosted by Hanzo.

Architecture: router
Automatic model routing
InputN/A
OutputN/A
Cache WriteN/A
Cache ReadN/A

Router Knowledge Base Document

pro

Automatic model routing — hosted by Hanzo.

Architecture: router
Automatic model routing
InputN/A
OutputN/A
Cache WriteN/A
Cache ReadN/A

Router Software Engineering

pro

Automatic model routing — hosted by Hanzo.

Architecture: router
Automatic model routing
InputN/A
OutputN/A
Cache WriteN/A
Cache ReadN/A

Router Software Engineering 01

pro

Automatic model routing — hosted by Hanzo.

Architecture: router
Automatic model routing
InputN/A
OutputN/A
Cache WriteN/A
Cache ReadN/A

Router Writing

pro

Automatic model routing — hosted by Hanzo.

Architecture: router
Automatic model routing
InputN/A
OutputN/A
Cache WriteN/A
Cache ReadN/A

All Zen LLMs available via OpenAI-compatible API at api.hanzo.ai/v1/chat/completions

Embedding Models

High-quality text embeddings via /v1/embeddings

Zen Embedding

starter

Dense text embeddings for semantic search, retrieval, and clustering.

8K context
Embeddings
Input$0.06 / MTok
Output$0.06 / MTok
Cache WriteN/A
Cache ReadFree

All Mini Lm L6 V2

pro

Text embeddings — hosted by Hanzo.

Architecture: embedding
Text embeddings
Input$0.02 / MTok
OutputFree
Cache WriteN/A
Cache ReadN/A

Bge M3

pro

Text embeddings — hosted by Hanzo.

Architecture: embedding
Text embeddings
Input$0.02 / MTok
OutputFree
Cache WriteN/A
Cache ReadN/A

E5 Large V2

pro

Text embeddings — hosted by Hanzo.

Architecture: embedding
Text embeddings
Input$0.02 / MTok
OutputFree
Cache WriteN/A
Cache ReadN/A

Gte Large En V1.5

pro

Text embeddings — hosted by Hanzo.

Architecture: embedding
Text embeddings
Input$0.02 / MTok
OutputFree
Cache WriteN/A
Cache ReadN/A

Multi Qa Mpnet Base Dot V1

pro

Text embeddings — hosted by Hanzo.

Architecture: embedding
Text embeddings
Input$0.02 / MTok
OutputFree
Cache WriteN/A
Cache ReadN/A

Qwen3 Embedding 0.6b

pro

Text embeddings — hosted by Hanzo.

Architecture: embedding
Text embeddings
Input$0.02 / MTok
OutputFree
Cache WriteN/A
Cache ReadN/A

Reranker Models

Improve retrieval quality via /v1/rerank

Zen Rerank

starter

Cross-encoder reranker — scores (query, document) pairs to reorder retrieval results.

Reranking
Retrieval
Price$0.03

Bge Reranker V2 M3

pro

Reranking — hosted by Hanzo.

Architecture: rerank
Reranking
Input$0.02 / MTok
OutputFree
Cache WriteN/A
Cache ReadN/A

Image Generation

FLUX, Stable Diffusion, and more via /v1/images/generations

Zen Image

pro

Text-to-image generation.

Text → Image
Price$0.24 / image

Openai Gpt Image 1

pro

Text-to-image — hosted by Hanzo.

Architecture: image
Text-to-image
Price$0.04 / image
/v1/images/generations

Openai Gpt Image 1.5

pro

Text-to-image — hosted by Hanzo.

Architecture: image
Text-to-image
Price$0.04 / image
/v1/images/generations

Openai Gpt Image 2

pro

Text-to-image — hosted by Hanzo.

Architecture: image
Text-to-image
Price$0.04 / image
/v1/images/generations

Stable Diffusion 3.5 Large

pro

Text-to-image — hosted by Hanzo.

Architecture: image
Text-to-image
256 context window
Price$0.04 / image
/v1/images/generations

Audio & Speech

Speech-to-text, text-to-speech, and audio generation via /v1/audio/*

Zen Foley

pro

Text-to-sound-effects — generates Foley and ambient audio from a prompt.

Text → Sound FX
Price$0.15

Zen Music

pro

Text-to-music generation.

Text → Music
Price$0.3

Zen Voice

pro

Text-to-speech — natural voice synthesis.

Text → Speech
Price$0.045

Qwen3 Tts Voicedesign

pro

Text-to-speech — hosted by Hanzo.

Architecture: speech
Text-to-speech
33k context window
Price$5
/v1/audio/speech

Video Generation

Text-to-video via /v1/videos/generations

Zen Video

pro max

Text-to-video — generates short clips from a prompt (async).

Text → Video
Async
Price$1.8

Wan2 2 T2v A14b

pro

Text-to-video — hosted by Hanzo.

Architecture: video
Text-to-video
100 context window
Price$0.5
/v1/videos/generations

Customers can purchase prioritized API capacity with Priority Tier

Prompt caching pricing is for the standard 5-minute TTL. A longer one costs more.

Explore pricing for tools

Web Search

per query$0.005

Code Interpreter

per session minute$0.03

File Storage

per GB/month$0.2

Image Generation

per image$0.04

Speech-to-Text

per minute$0.006

Text-to-Speech

per 1M characters$15

*Does not include input and output tokens required to process requests

Featured Third-Party Models

Top models from leading AI providers — all accessible through one Hanzo API key

Anthropic: Claude Opus 4.6

Anthropic
Featured
1M context windowtext+image+file->text
Input$5 / MTok
Output$25 / MTok

Anthropic: Claude Sonnet 4.6

Anthropic
Featured
1M context windowtext+image+file->text
Input$3 / MTok
Output$15 / MTok

Anthropic: Claude Haiku 4.5

Anthropic
Featured
200k context windowtext+image+file->text
Input$1 / MTok
Output$5 / MTok

OpenAI: GPT-5

OpenAI
Featured
400k context windowtext+image+file->text
Input$1.25 / MTok
Output$10 / MTok

OpenAI: GPT-5 Mini

OpenAI
Featured
400k context windowtext+image+file->text
Input$0.25 / MTok
Output$2 / MTok

Google: Gemini 2.5 Pro

Google
Featured
1M context windowtext+image+file+audio+video->text
Input$1.25 / MTok
Output$10 / MTok

DeepSeek: R1

DeepSeek
Featured
164k context windowtext->text
Input$0.7 / MTok
Output$2.5 / MTok

DeepSeek: DeepSeek V3

DeepSeek
Featured
164k context windowtext->text
Input$0.257 / MTok
Output$1.03 / MTok

Meta: Llama 4 Maverick

Meta
Featured
1M context windowtext+image->text
Input$0.2 / MTok
Output$0.8 / MTok

Mistral: Mistral Large 3 2512

Mistral
Featured
262k context windowtext+image+file->text
Input$0.5 / MTok
Output$1.5 / MTok

Cohere: Command A

Cohere
Featured
256k context windowtext->text
Input$2.5 / MTok
Output$10 / MTok

All Third-Party Models

397 models from 57 providers — dynamically detected and priced

397 models
ModelProviderContextInputOutput
Anthropic: Claude Opus 4.6
Featured
Anthropic1000K$5 / MTok$25 / MTok
Anthropic: Claude Sonnet 4.6
Featured
Anthropic1000K$3 / MTok$15 / MTok
Anthropic: Claude Haiku 4.5
Featured
Anthropic200K$1 / MTok$5 / MTok
OpenAI: GPT-5
Featured
OpenAI400K$1.25 / MTok$10 / MTok
OpenAI: GPT-5 Mini
Featured
OpenAI400K$0.25 / MTok$2 / MTok
Google: Gemini 2.5 Pro
Featured
Google1049K$1.25 / MTok$10 / MTok
DeepSeek: R1
Featured
DeepSeek164K$0.7 / MTok$2.5 / MTok
DeepSeek: DeepSeek V3
Featured
DeepSeek164K$0.257 / MTok$1.03 / MTok
Meta: Llama 4 Maverick
Featured
Meta1049K$0.2 / MTok$0.8 / MTok
Mistral: Mistral Large 3 2512
Featured
Mistral262K$0.5 / MTok$1.5 / MTok
Cohere: Command A
Featured
Cohere256K$2.5 / MTok$10 / MTok
Gemma SEA LION v4 27B IT
Free
AI SingaporeN/AFreeFree
Qwen SEA LION v4 32B IT
Free
AI SingaporeN/AFreeFree
AI21: Jamba Large 1.7
AI21256K$2 / MTok$8 / MTok
AionLabs: Aion-2.0
Aion Labs131K$0.8 / MTok$1.6 / MTok
AionLabs: Aion-3.0
Aion Labs131K$3 / MTok$6 / MTok
AionLabs: Aion-3.0-Mini
Aion Labs131K$0.7 / MTok$1.4 / MTok
AionLabs: Aion-RP 1.0 (8B)
Aion Labs33K$0.8 / MTok$1.6 / MTok
AllenAI: Olmo 3 32B Think
Allen AI66K$0.15 / MTok$0.5 / MTok
Olmo 3 7B Instruct
Free
Allen AIN/AFreeFree
Amazon: Nova 2 Lite
Amazon1000K$0.3 / MTok$2.5 / MTok
Amazon: Nova Lite 1.0
Amazon300K$0.06 / MTok$0.24 / MTok
Amazon: Nova Micro 1.0
Amazon128K$0.035 / MTok$0.14 / MTok
Amazon: Nova Premier 1.0
Amazon1000K$2.5 / MTok$12.5 / MTok
Amazon: Nova Pro 1.0
Amazon300K$0.8 / MTok$3.2 / MTok
Magnum v4 72B
Anthracite16K$3 / MTok$5 / MTok
Anthropic Claude Haiku Latest
Anthropic200K$1 / MTok$5 / MTok
Anthropic Claude Sonnet Latest
Anthropic1000K$2 / MTok$10 / MTok
Anthropic: Claude 3 Haiku
Anthropic200K$0.25 / MTok$1.25 / MTok
Anthropic: Claude Fable 5
Anthropic1000K$10 / MTok$50 / MTok
Anthropic: Claude Fable Latest
Anthropic1000K$10 / MTok$50 / MTok
Anthropic: Claude Opus 4
Anthropic200K$15 / MTok$75 / MTok
Anthropic: Claude Opus 4.1
Anthropic200K$15 / MTok$75 / MTok
Anthropic: Claude Opus 4.5
Anthropic200K$5 / MTok$25 / MTok
Anthropic: Claude Opus 4.7
Anthropic1000K$5 / MTok$25 / MTok
Anthropic: Claude Opus 4.7 (Fast)
Anthropic1000K$30 / MTok$150 / MTok
Anthropic: Claude Opus 4.8
Anthropic1000K$5 / MTok$25 / MTok
Anthropic: Claude Opus 4.8 (Fast)
Anthropic1000K$10 / MTok$50 / MTok
Anthropic: Claude Opus Latest
Anthropic1000K$5 / MTok$25 / MTok
Anthropic: Claude Sonnet 4
Anthropic1000K$3 / MTok$15 / MTok
Anthropic: Claude Sonnet 4.5
Anthropic1000K$3 / MTok$15 / MTok
Anthropic: Claude Sonnet 5
Anthropic1000K$2 / MTok$10 / MTok
Claude Opus 5
Anthropic1000K$5 / MTok$25 / MTok
Claude Opus 5 (Fast)
Anthropic1000K$10 / MTok$50 / MTok
Arcee AI: Trinity Large Thinking
Arcee262K$0.22 / MTok$0.85 / MTok
Arcee AI: Virtuoso Large
Arcee131K$0.75 / MTok$1.2 / MTok
Baidu: ERNIE 4.5 VL 424B A47B
Baidu123K$0.42 / MTok$1.25 / MTok
ERNIE 4.5 VL 424B A47B Base PT
Free
BaiduN/AFreeFree
ByteDance Seed: Seed 1.6
ByteDance262K$0.25 / MTok$2 / MTok
ByteDance Seed: Seed 1.6 Flash
ByteDance262K$0.075 / MTok$0.3 / MTok

Third-party model pricing includes a passthrough markup. Prices update automatically every 6 hours. All models accessible via api.hanzo.ai/v1/chat/completions using the provider-prefixed model ID (e.g. anthropic/claude-sonnet-4.6).

Frequently Asked Questions

Measured, not modelled.

Every figure is corrected against vendor invoice brackets — list-price rate cards overstate real cost by up to 4.2x. The method, and the cost per completed task, are on the bench.