# Hanzo Engine — Hanzo AI

> Rust-native LLM and embedding inference engine

hanzoai/engine

# Hanzo Engine

The inference engine, in one binary

Name a model on Hugging Face and it runs. Engine quantizes it for the hardware it finds, then serves an HTTP API and a web chat from the same process — text, vision, audio, speech, image and embeddings, without a second stack underneath.

[Get started](https://docs.hanzo.ai/docs/services/engine)[View on GitHub](https://github.com/hanzoai)

## One process, laptop to GPU cluster

Chat, an API server, and a benchmark are three commands on the same binary. There is no Python interpreter under any of them.

### run, serve, bench

One binary, three verbs. Chat with a model, put it behind an HTTP API, or measure it — the thing you benchmarked is the thing you shipped.

### Built to serve, not just to run

Paged attention and continuous batching keep the card busy across concurrent requests. A disk-first KV cache carries a session across a restart, so an agent that reuses a long prefix does not pay for it twice.

### More than text

Vision, audio, speech, image and embeddings go through the same modality pipeline as generation, in the same process. A multimodal app is one deployment, not four.

### It quantizes on the way in

In-situ quantization of any Hugging Face model, plus GGUF from 2 to 8 bits, GPTQ, AWQ, HQQ, FP8 and bitsandbytes. Ask for a level and it takes a published prebuilt when one exists, or makes one when it doesn&#x27;t.

### The silicon you have

CUDA with FlashAttention and multi-GPU tensor parallelism, Metal on Apple Silicon, or plain CPU. Same engine and same commands on each, and a doctor command that reports what it found.

### Tools, and a UI to watch them

It can call MCP servers over a process, HTTP or a WebSocket. The built-in UI at /ui shows reasoning, code execution and plots inline, and edits branch a conversation with its own state. Pass --no-ui and it is a plain API server.

## Get started with Hanzo Engine

[Read the docs](https://docs.hanzo.ai/docs/services/engine)[View on GitHub](https://github.com/hanzoai)

## Open source

License: Apache-2.0[hanzoai](https://github.com/hanzoai)

## Get Engine

AI inference engine

[Deploy to Cloud](https://console.hanzo.ai/deploy)[Self-host](https://docs.hanzo.ai/docs/services/engine)
