Try Hanzo
Beta

Evals

Measure model and agent quality.

Open sourceOn-chain settlement
API
/v1/eval
Google Cloud
Vertex AI Eval

Paste and ship

Paste this into any agent. It reads the skill manifest and calls Evals from there.

Agent promptapi /v1/eval
Read https://hanzo.ai/skill.md and use Hanzo Evals in my project. Start with: hanzo eval runs create
Evals — Hanzo

Every request enters here, so the limiter does too.

api.hanzo.ai/v1/evalOpen in the console
Example project. Every name in it is one you would give your own.

Benchmark, grade, and regression-test models and agents. LLM-as-judge, golden sets, and drift detection with reproducible scorecards anchored on-chain.

LLM-as-judge & golden sets
Regression gates in CI
Drift detection
Anchored scorecards

A human signs a company into existence. After genesis, agents can run it.

Start building with Evals

One API key. One credit balance. Every primitive.

Related capabilities

BenchmarkResearchDatasets

Build what’s next.