Pricing
Pay for what the agents use. One plan per organization, then metered per call for what the agents consume. No per-seat fee. Nothing at all for a machine that is sitting still.
Plan
One plan per organization, then billed for what the agents consume.
- No per-seat fee, however many people use the organization.
- Usage metered per tier, booked in integer micro-USD
- MCP, SDK and CLI — $0
Budgeted
Every agent carries a budget. Model calls and renders are quoted before they run, so anything that would breach it is refused rather than discovered on an invoice.
- A cap per period and a cap per task
- A render with no headroom is refused, not run
- A failed job is not billed and its hold comes back
- Compute and web calls meter after the fact, on the ledger either way
Custom
Committed volume, procurement, or anything the meter does not cover.
- Committed-volume pricing
- Invoicing and PO terms
- A monthly allowance arranged against your account
- Security review and DPA
What a bill is made of
Four components, each on your spend breakdown under its own name. Two are quoted before they run; two are metered from what they used.
Model inference
per token, five tiersComputer
per second, while runningWeb tools
per callMedia generation
what the job costThe rates
Every rate here is the price you pay — the number the ledger books and the number your spend breakdown reports.
| Rate | ||
|---|---|---|
| Computer · vCPU | $0.0504/hr$0.000014/second | Charged by the second, while running. |
| Computer · memory | $0.0162/GiB-hr$0.0000045/GiB-second | Charged by the second, on the memory the machine was given. |
| Computer · paused | $0 | No vCPU, no memory, no storage. |
| Computer · created | $0 | There is no creation fee. |
| web_search | $0.002158/call | One query, up to ten results. |
| web_fetch | $0.001079/call | One page read as text. |
| MCP, SDK and CLI | $0 | The interfaces are not metered. |
| Seats | $0 | No per-seat fee, however many people use the organization. |
The completion window is the biggest lever
The same model at a different window is a different price, chosen per request. Pay the interactive tariff only when a person is waiting.
Every model, at the rate the ledger books
Model inference is metered per token in five disjoint tiers, at the tariff for the completion window the request asked for.
Hanzo Enso
Flagship frontier models — Flash, Pro, and Ultra. Live per-token retail, one API key.
Enso Flash
Fast, economical Enso tier for high-volume, low-latency everyday work, with 1M-context overflow.
Enso
Hanzo's proprietary frontier model — Opus-class reasoning by default with 1M-context overflow.
Enso Ultra
Adaptive fan-out — probes a task-appropriate model, escalates to a top-K panel only when needed, then verifies-then-selects the best answer.
Zen Models
34 models across LLM, embedding, reranker, image, audio, and videoUpdated 8/5/2026
Zen Guard — Content Safety
Safety classifier for moderation and guardrails across a broad category and language set.
Zen VL — Vision-Language
Reads images and reasons over them — visual Q&A, document and chart understanding, grounded captioning.
Zen 5.8
Zen 5.8 flagship frontier general reasoning and agentic foundation model. Advanced chain-of-thought, 1M context, MoDE architecture.
Zen 5.8 Coder
Zen 5.8 Coder flagship autonomous coding and repository understanding model. 1M context, multi-file code synthesis, agentic refactoring.
Zen5 (Default)
Canonical Zen5 default. a 35B frontier MoE (3B active) (35B total / 3B active per token, released Apr 2026, Apache-2.0). 256K context, agentic-trained, OpenAI + Anthropic API. The everyday Zen5 chat model.
Zen5 Coder
Code-specialized Zen5 tier. 80B sparse MoE tuned for repo-scale code understanding, agentic refactoring, and tool-use coding loops.
Zen5 Flash
Smallest and cheapest text-only Zen5 chat tier. 4B-class dense, sub-100ms TTFT, 32K context. For high-volume routing and simple agent loops.
Zen5 Mini
Frontier agentic at the lowest cost in the family. Built on a 230B agentic MoE (10B active) (230B MoE / 10B active, released Feb 2026). 80.2% SWE-Bench Verified, 76.3% BrowseComp; trained on 200K+ real-world environments via large-scale RL.
Zen5 Pro
Zen Flash IQ2_XXS-imatrix weights (81 GB GGUF on zenlm/zen-5-pro-gguf). 284B total / 37B active per token, 1M context, asymmetric routed-MoE quant. Fits a single 128 GB Apple Silicon / DGX Spark / H100 80 GB.
Router General
Automatic model routing — hosted by Hanzo.
Router Knowledge Base Document
Automatic model routing — hosted by Hanzo.
Router Software Engineering
Automatic model routing — hosted by Hanzo.
Router Software Engineering 01
Automatic model routing — hosted by Hanzo.
Router Writing
Automatic model routing — hosted by Hanzo.
All Zen LLMs available via OpenAI-compatible API at api.hanzo.ai/v1/chat/completions
Embedding Models
High-quality text embeddings via /v1/embeddings
Zen Embedding
Dense text embeddings for semantic search, retrieval, and clustering.
All Mini Lm L6 V2
Text embeddings — hosted by Hanzo.
Bge M3
Text embeddings — hosted by Hanzo.
E5 Large V2
Text embeddings — hosted by Hanzo.
Gte Large En V1.5
Text embeddings — hosted by Hanzo.
Multi Qa Mpnet Base Dot V1
Text embeddings — hosted by Hanzo.
Qwen3 Embedding 0.6b
Text embeddings — hosted by Hanzo.
Reranker Models
Improve retrieval quality via /v1/rerank
Zen Rerank
Cross-encoder reranker — scores (query, document) pairs to reorder retrieval results.
Bge Reranker V2 M3
Reranking — hosted by Hanzo.
Image Generation
FLUX, Stable Diffusion, and more via /v1/images/generations
Zen Image
Text-to-image generation.
Openai Gpt Image 1
Text-to-image — hosted by Hanzo.
/v1/images/generationsOpenai Gpt Image 1.5
Text-to-image — hosted by Hanzo.
/v1/images/generationsOpenai Gpt Image 2
Text-to-image — hosted by Hanzo.
/v1/images/generationsStable Diffusion 3.5 Large
Text-to-image — hosted by Hanzo.
/v1/images/generationsAudio & Speech
Speech-to-text, text-to-speech, and audio generation via /v1/audio/*
Zen Foley
Text-to-sound-effects — generates Foley and ambient audio from a prompt.
Zen Music
Text-to-music generation.
Zen Voice
Text-to-speech — natural voice synthesis.
Qwen3 Tts Voicedesign
Text-to-speech — hosted by Hanzo.
/v1/audio/speechVideo Generation
Text-to-video via /v1/videos/generations
Zen Video
Text-to-video — generates short clips from a prompt (async).
Wan2 2 T2v A14b
Text-to-video — hosted by Hanzo.
/v1/videos/generationsCustomers can purchase prioritized API capacity with Priority Tier
Prompt caching pricing is for the standard 5-minute TTL. A longer one costs more.
Explore pricing for tools
Web Search
Code Interpreter
File Storage
Image Generation
Speech-to-Text
Text-to-Speech
*Does not include input and output tokens required to process requests
Featured Third-Party Models
Top models from leading AI providers — all accessible through one Hanzo API key
Anthropic: Claude Opus 4.6
Anthropic: Claude Sonnet 4.6
Anthropic: Claude Haiku 4.5
OpenAI: GPT-5
OpenAI: GPT-5 Mini
Google: Gemini 2.5 Pro
DeepSeek: R1
DeepSeek: DeepSeek V3
Meta: Llama 4 Maverick
Mistral: Mistral Large 3 2512
Cohere: Command A
All Third-Party Models
397 models from 57 providers — dynamically detected and priced
| Model | Provider | Context | Input | Output |
|---|---|---|---|---|
Anthropic: Claude Opus 4.6 Featured | Anthropic | 1000K | $5 / MTok | $25 / MTok |
Anthropic: Claude Sonnet 4.6 Featured | Anthropic | 1000K | $3 / MTok | $15 / MTok |
Anthropic: Claude Haiku 4.5 Featured | Anthropic | 200K | $1 / MTok | $5 / MTok |
OpenAI: GPT-5 Featured | OpenAI | 400K | $1.25 / MTok | $10 / MTok |
OpenAI: GPT-5 Mini Featured | OpenAI | 400K | $0.25 / MTok | $2 / MTok |
Google: Gemini 2.5 Pro Featured | 1049K | $1.25 / MTok | $10 / MTok | |
DeepSeek: R1 Featured | DeepSeek | 164K | $0.7 / MTok | $2.5 / MTok |
DeepSeek: DeepSeek V3 Featured | DeepSeek | 164K | $0.257 / MTok | $1.03 / MTok |
Meta: Llama 4 Maverick Featured | Meta | 1049K | $0.2 / MTok | $0.8 / MTok |
Mistral: Mistral Large 3 2512 Featured | Mistral | 262K | $0.5 / MTok | $1.5 / MTok |
Cohere: Command A Featured | Cohere | 256K | $2.5 / MTok | $10 / MTok |
Gemma SEA LION v4 27B IT Free | AI Singapore | N/A | Free | Free |
Qwen SEA LION v4 32B IT Free | AI Singapore | N/A | Free | Free |
AI21: Jamba Large 1.7 | AI21 | 256K | $2 / MTok | $8 / MTok |
AionLabs: Aion-2.0 | Aion Labs | 131K | $0.8 / MTok | $1.6 / MTok |
AionLabs: Aion-3.0 | Aion Labs | 131K | $3 / MTok | $6 / MTok |
AionLabs: Aion-3.0-Mini | Aion Labs | 131K | $0.7 / MTok | $1.4 / MTok |
AionLabs: Aion-RP 1.0 (8B) | Aion Labs | 33K | $0.8 / MTok | $1.6 / MTok |
AllenAI: Olmo 3 32B Think | Allen AI | 66K | $0.15 / MTok | $0.5 / MTok |
Olmo 3 7B Instruct Free | Allen AI | N/A | Free | Free |
Amazon: Nova 2 Lite | Amazon | 1000K | $0.3 / MTok | $2.5 / MTok |
Amazon: Nova Lite 1.0 | Amazon | 300K | $0.06 / MTok | $0.24 / MTok |
Amazon: Nova Micro 1.0 | Amazon | 128K | $0.035 / MTok | $0.14 / MTok |
Amazon: Nova Premier 1.0 | Amazon | 1000K | $2.5 / MTok | $12.5 / MTok |
Amazon: Nova Pro 1.0 | Amazon | 300K | $0.8 / MTok | $3.2 / MTok |
Magnum v4 72B | Anthracite | 16K | $3 / MTok | $5 / MTok |
Anthropic Claude Haiku Latest | Anthropic | 200K | $1 / MTok | $5 / MTok |
Anthropic Claude Sonnet Latest | Anthropic | 1000K | $2 / MTok | $10 / MTok |
Anthropic: Claude 3 Haiku | Anthropic | 200K | $0.25 / MTok | $1.25 / MTok |
Anthropic: Claude Fable 5 | Anthropic | 1000K | $10 / MTok | $50 / MTok |
Anthropic: Claude Fable Latest | Anthropic | 1000K | $10 / MTok | $50 / MTok |
Anthropic: Claude Opus 4 | Anthropic | 200K | $15 / MTok | $75 / MTok |
Anthropic: Claude Opus 4.1 | Anthropic | 200K | $15 / MTok | $75 / MTok |
Anthropic: Claude Opus 4.5 | Anthropic | 200K | $5 / MTok | $25 / MTok |
Anthropic: Claude Opus 4.7 | Anthropic | 1000K | $5 / MTok | $25 / MTok |
Anthropic: Claude Opus 4.7 (Fast) | Anthropic | 1000K | $30 / MTok | $150 / MTok |
Anthropic: Claude Opus 4.8 | Anthropic | 1000K | $5 / MTok | $25 / MTok |
Anthropic: Claude Opus 4.8 (Fast) | Anthropic | 1000K | $10 / MTok | $50 / MTok |
Anthropic: Claude Opus Latest | Anthropic | 1000K | $5 / MTok | $25 / MTok |
Anthropic: Claude Sonnet 4 | Anthropic | 1000K | $3 / MTok | $15 / MTok |
Anthropic: Claude Sonnet 4.5 | Anthropic | 1000K | $3 / MTok | $15 / MTok |
Anthropic: Claude Sonnet 5 | Anthropic | 1000K | $2 / MTok | $10 / MTok |
Claude Opus 5 | Anthropic | 1000K | $5 / MTok | $25 / MTok |
Claude Opus 5 (Fast) | Anthropic | 1000K | $10 / MTok | $50 / MTok |
Arcee AI: Trinity Large Thinking | Arcee | 262K | $0.22 / MTok | $0.85 / MTok |
Arcee AI: Virtuoso Large | Arcee | 131K | $0.75 / MTok | $1.2 / MTok |
Baidu: ERNIE 4.5 VL 424B A47B | Baidu | 123K | $0.42 / MTok | $1.25 / MTok |
ERNIE 4.5 VL 424B A47B Base PT Free | Baidu | N/A | Free | Free |
ByteDance Seed: Seed 1.6 | ByteDance | 262K | $0.25 / MTok | $2 / MTok |
ByteDance Seed: Seed 1.6 Flash | ByteDance | 262K | $0.075 / MTok | $0.3 / MTok |
Third-party model pricing includes a passthrough markup. Prices update automatically every 6 hours. All models accessible via api.hanzo.ai/v1/chat/completions using the provider-prefixed model ID (e.g. anthropic/claude-sonnet-4.6).
Frequently Asked Questions
Measured, not modelled.
Every figure is corrected against vendor invoice brackets — list-price rate cards overstate real cost by up to 4.2x. The method, and the cost per completed task, are on the bench.