Multimodal Models
Zen VL and Zen Omni for image understanding, OCR, document parsing, and visual Q&A.
Process images and video at production scale
Extract structured data from any visual input. Hanzo's vision pipeline handles ingestion, processing, and structured output — from single images to high-volume video streams.
Every feature you need to ship fast and scale confidently.
Zen VL and Zen Omni for image understanding, OCR, document parsing, and visual Q&A.
Frame extraction, scene detection, and temporal analysis for long-form video content.
Convert visual content to JSON, tables, or any schema. Works on receipts, forms, diagrams, charts.
Object detection, face recognition, and anomaly detection on live camera feeds.
Zen Artist for photorealistic image generation and editing. Zen Artist Edit for precise inpainting.
Embed images in the same space as text for multimodal search and clustering.
Real workloads, real teams, real impact.
Get up and running in minutes. Our documentation covers everything from quick start to production deployment.
Also available on
Enterprise ready
Continual internal audits, a full audit trail, and your own tenancy. Custom SLA and dedicated support engineers on Enterprise.