Try Hanzo
Hanzo Context

Benchmarks

Similarity finds what sounds related. Memory needs what is connected. These tables measure that difference on public benchmarks, with the split, the embedder, the reader and the metric named on every table, and every number regenerated from its raw run.

Protocol

Configurations are chosen on a declared development split and frozen at a commit before the test split runs once. LoCoMo: dev conversations 0–2, test 3–9; adversarial questions excluded, as Mem0 and Zep exclude them. MemoryAgentBench: dev 6k, test 32k / 64k / 262k. Readers: enso-flash (house),gemma4:31b (open, LoCoMo-Conv's own reader, run locally), and the strongest router model that answers. Temperature 0, one prompt file per benchmark shared by every row. Intervals are bootstrap 95% over questions. Retrieval recall is reported as ALL and ANY because a multi-hop question is only answerable when every turn it needs is present — and as a diagnostic, not the headline: a system can ground a correct answer on equivalent evidence the annotator did not mark, which the supported column counts.

Raw runs, traces and scoring: hanzoai/cloud · bench/brain. Papers: papers.hanzo.ai. Product: POST api.hanzo.ai/v1/memory/search.

LoCoMo · retrieval

Ten conversations, 5,882 turns, 1,540 questions with annotated evidence turns. Each row adds one candidate generator to the engine — BM25, the atomic fact index, canonical entities, the timeline, adjacency, one hop of typed expansion, a second retrieval hop asked with what the first resolved — and the rows below the working ones are the shapes that lost and shaped the design. The fact layer here is LoCoMo's own observation annotations, each citing its source turns.

LoCoMo · all ten conversations · zen-embedding-0.6b · k=20

facts LoCoMo's own observations · weights frozen on dev at 33d0f8749bd1 · ALL = every annotated turn in the top 20; ANY = at least one; R@20 = |retrieved ∩ gold| / |gold|

configurationmulti-hop ALLmulti-hop ANYsingle-hop ALLtemporal ALLopen-domain ALLR@20nDCG@20supportedcandidatestokens/qp50 ms
semantic only22.779.878.277.333.771.949.681.0406600.55
+lexical22.778.082.578.837.074.653.283.3526650.61
+facts36.988.382.382.642.478.259.686.4717870.83
+adjacency39.088.784.283.544.679.761.487.8747820.83
+iterative hops39.087.986.884.445.780.962.688.0767881.65
failed: global RRF33.789.086.883.840.280.556.289.01377621.73

LoCoMo · all ten conversations · all-MiniLM-L6-v2 · k=10

facts LoCoMo's own observations · weights frozen on dev at 33d0f8749bd1 · ALL = every annotated turn in the top 10; ANY = at least one; R@10 = |retrieved ∩ gold| / |gold|

configurationmulti-hop ALLmulti-hop ANYsingle-hop ALLtemporal ALLopen-domain ALLR@10nDCG@10supportedcandidatestokens/qp50 ms
semantic only9.649.651.446.116.345.331.153.1402830.25
+lexical13.858.967.262.023.959.243.767.6543010.31
+facts20.973.473.172.025.066.752.075.6733640.42
+adjacency21.674.574.672.627.267.753.376.3773610.42
+iterative hops24.178.077.573.527.270.155.878.5783770.76
failed: global RRF20.274.873.474.128.367.248.676.11433820.43

LoCoMo · test split, conversations 3–9 · zen-embedding-0.6b · k=20

facts LoCoMo's own observations · weights frozen on dev at 33d0f8749bd1 · ALL = every annotated turn in the top 20; ANY = at least one; R@20 = |retrieved ∩ gold| / |gold|

configurationmulti-hop ALLmulti-hop ANYsingle-hop ALLtemporal ALLopen-domain ALLR@20nDCG@20supportedcandidatestokens/qp50 ms
semantic only18.880.378.372.335.670.948.880.5406510.55
+lexical21.278.482.873.638.473.952.983.2536560.62
+facts35.188.583.078.843.877.958.786.4717711.15
+entities35.188.583.078.843.877.958.786.4717711.10
+timeline34.688.583.379.745.278.259.986.8717691.04
+adjacency37.588.584.680.145.279.160.587.5757671.03
+typed graph37.588.584.680.145.279.160.587.5757670.98
+iterative hops37.588.087.781.445.280.761.687.8777711.67
failed: PRF19.775.572.469.335.666.446.375.86346651.10
failed: chain search37.588.083.878.845.278.556.786.3797641.64
failed: surface entities35.688.083.078.843.877.958.286.2827730.98
failed: global RRF32.288.586.780.138.479.655.588.41397511.20

LoCoMo-Conv

The same memories asked the way users ask: dialog, implicit, counterfactual, composed. The data is not yet released (the paper's one address answers 404), so no LoCoMo-Conv number appears here. What does: the paper's retrieval protocol (all-MiniLM-L6-v2, k=10, verbatim containment) reproduced on the original LoCoMo questions the rewrites came from, by memory unit.

conv-proxy-naive-rag-all-minilm-k10

k=10 · run conv-proxy-naive-rag-all-minilm-k10

rownrecallALLANYtokens/q
original197742.939.047.7262
original·cat232049.446.352.5253
original·no-adversarial153145.540.751.6265
original·cat38926.316.939.3250
original·cat128126.79.649.8270
original·cat484152.351.553.2269
original·cat544633.933.234.5250

conv-proxy-naive-rag-st-minilm-k10

k=10 · run conv-proxy-naive-rag-st-minilm-k10

rownrecallALLANYtokens/q
original197742.939.047.9262
original·cat232049.446.352.5253
original·no-adversarial153145.540.751.7265
original·cat38926.316.939.3251
original·cat128127.09.650.5270
original·cat484152.351.553.2269
original·cat544633.933.234.5250

conv-proxy-naive-rag-zenlm_zen-embedding-0.6b-k10

k=10 · run conv-proxy-naive-rag-zenlm_zen-embedding-0.6b-k10

rownrecallALLANYtokens/q
original197759.754.965.6315
original·cat232074.070.976.9332
original·no-adversarial153163.557.670.9319
original·cat38940.433.749.4286
original·cat128138.114.968.7311
original·cat484170.569.371.7320
original·cat544646.445.547.3302

conv-proxy-units-st-minilm-k10

run conv-proxy-units-st-minilm-k10

rownrecalltokens/q
turn·all198242.8279
turn·no-adversarial153645.4283
w2·all198256.3571
w2·no-adversarial153654.3567
w3·all198262.5829
w3·no-adversarial153660.7823
w5·all198260.31258
w5·no-adversarial153658.21240
session·all198262.16855
session·no-adversarial153661.86848

LoCoMo · answers

The same questions answered by a reader from the retrieved context. Token F1 and exact match with LoCoMo's own normalisation. Rows differ only in what was retrieved: cosine, the shipped re-rank, the engine, or the annotated gold turns themselves (the ceiling that separates retrieval's share of the error from the reader's).

locomo-all-cer-k20-gemma4-31b

k=20 · reader gemma4:31b · commit ddc076f5dd0506ce50428416c336e45c4a7ada73 · run locomo-all-cer-k20-gemma4-31b

rownsupportedF1EMtokens/q
all28211.040.7[36.5, 44.9]11.0[7.4, 14.9]968
dev7410.844.5[36.3, 53.2]10.8[4.1, 18.9]964
test20811.139.3[34.8, 43.9]11.1[7.2, 15.4]970

locomo-all-context-k20-gemma4-31b

k=20 · reader gemma4:31b · commit ddc076f5dd0506ce50428416c336e45c4a7ada73 · run locomo-all-context-k20-gemma4-31b

rownsupportedF1EMtokens/q
all28213.540.5[36.8, 44.3]9.9[6.4, 13.5]971
dev7413.544.7[36.5, 52.6]9.5[2.7, 16.2]984
test20813.539.0[34.8, 43.4]10.1[6.3, 14.4]967

locomo-all-oracle-k20-gemma4-31b

k=20 · reader gemma4:31b · commit ddc076f5dd0506ce50428416c336e45c4a7ada73 · run locomo-all-oracle-k20-gemma4-31b

rownsupportedF1EMtokens/q
all28212.453.0[49.2, 57.0]16.0[12.1, 19.9]169
dev7413.559.2[50.7, 66.9]13.5[6.8, 21.6]135
test20812.050.8[46.7, 55.1]16.8[12.5, 22.1]180

locomo-all-single-k20-gemma4-31b

k=20 · reader gemma4:31b · commit ddc076f5dd0506ce50428416c336e45c4a7ada73 · run locomo-all-single-k20-gemma4-31b

rownsupportedF1EMtokens/q
all28211.737.6[33.8, 42.0]9.9[6.7, 13.8]835
dev7412.241.1[33.1, 48.9]6.8[1.4, 12.2]831
test20811.536.3[31.9, 40.9]11.1[6.7, 15.4]836

MemoryAgentBench · Conflict Resolution

FactConsolidation: a haystack of numbered facts with later updates, single-hop and multi-hop questions, 6k to 262k tokens. Metric: substring exact match, the benchmark's own. This is the lane CAR reports on and the one vendor claims cite. The 6k haystacks are the development split; 32k, 64k and 262k are run once.

noreader

split dev · k=10 · reader none · commit ddc076f5dd · run mab-dev-noreader-none

rownsubstring EM
mh_6k9877.5
sh_6k9898.0

noreader

split test · k=10 · reader none · commit ced158ddb3 · run mab-test-noreader-none-v0

rownsubstring EM
mh_32k10032.0
mh_64k9931.3
mh_262k9825.5
sh_32k9898.0
sh_64k9992.9
sh_262k9597.9

Reference points

Numbers other systems publish, quoted with their protocol because the protocols differ. CAR (arXiv 2606.01435): MemoryAgentBench FactConsolidation 78.0 single-hop / 30.2 multi-hop pooled over 6k–262k with gpt-4o-mini, 82 / 27 at 262k; 94.8 / 51.5 pooled with gpt-4o, 93 / 41 at 262k. LoCoMo-Conv (arXiv 2609.03467): best reported retrieval R@10 with query facets, AnchorMem, dialog 0.754 · implicit 0.524 · counterfactual 0.732 · composed 0.432. Vendor pages: Mem0 92.5 and Zep 94.7 on LoCoMo under their own answer-accuracy protocols with LLM judges; Pith 68.0 substring EM on MemoryAgentBench multi-hop 262k. None of these are re-run here; a row of ours is comparable to one of theirs only where the split, the embedder, the reader and the metric match, and the page says where they do.

generated 2026-09-09T12:35:28.042Z