Try Hanzo

Benchmarks / LongMemEval

measured

LongMemEval

Five hundred questions, each with its own haystack of chat sessions around 115k tokens. The engine, frozen on LoCoMo and run here untuned, against the cosine floor on the same vectors across 4,846 sessions.

dataset
LongMemEval
licence
MIT
introduced by
LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory — Wu et al., 2024
harness
hanzoai/benchmarks · brain
node longmemeval/engine.mjs --embed=minilm && node results.mjs

regenerates every table on this page from the raw runs

What is measured

Five hundred questions, each with its own haystack of chat sessions about 115k tokens long. One or more sessions hold the evidence, and retrieval is scored as recall of those sessions against the benchmark’s own answer_session_ids. Every session is embedded turn by turn and ranked by its best turn — 4,846 sessions and 199,641 turns. ALL@5 is the share of questions whose every evidence session is in the top five; ANY@5 is the share with at least one.

Both columns use the same all-minilm vectors. The cosine column ranks a session by its best turn’s cosine. The engine column ranks it by its best turn under the engine’s score, with the configuration frozen on LoCoMo dev at 33d0f8749bd1 — nothing was tuned on this benchmark, so all 500 questions are the report.

Results

embedder all-minilm · engine b467901f3950 · retrieval p50 cosine 0.20 ms, engine 1.96 ms · run longmemeval-all-context-minilm

question typenALL@5 cosineALL@5 engineANY@5 cosineANY@5 engineALL@10 cosineALL@10 engineMRR cosineMRR engine
all questions50085.890.697.498.293.497.00.9030.919
single-session-user7095.798.695.798.697.1100.00.8520.872
multi-session13375.985.798.599.289.594.70.9120.917
single-session-preference3096.796.796.796.796.796.70.8030.832
temporal-reasoning13375.981.294.795.588.794.70.8650.894
knowledge-update7896.298.7100.0100.098.7100.00.9640.981
single-session-assistant56100.0100.0100.0100.0100.0100.01.0001.000

Where it gains

Every evidence session is in the top five for 90.6% of questions against 85.8% for cosine, and in the top ten for 97.0% against 93.4%. No question type goes down. The gain sits where the floor left room: multi-session ALL@5 moves from 75.9 to 85.7, and temporal-reasoning from 75.9 to 81.2. Single-session-assistant was saturated and stays at 100.0.

Where it loses, and why

The engine runs here with one hand tied. LongMemEval has no fact layer, so its fact, entity and typed-graph generators propose nothing; what does the work is lexical match, the session timeline, adjacency and a second hop. Three questions lose ALL@5 to cosine. The questions whose first correct session ranks lower are mostly phrased in relative time — “four weeks ago”, “last Saturday” — and the engine’s timeline fires only on a named month or year and does not read the question’s date. That is the next thing to fix, and it is a fix to the engine, not to this table.

Retrieval costs more: 1.96 ms at the median against 0.20 ms for cosine, once the turns are embedded.

What this table is not

Session recall is not answer accuracy. Systems reporting LongMemEval scores in the high eighties usually report an end-to-end answer metric with a language-model judge, which is a different measurement and not comparable to these columns. The run uses the release the dataset card now marks deprecated in favour of a cleaned one.

raw runs and scoring: hanzoai/benchmarks · brain · regenerated by node longmemeval/engine.mjs --embed=minilm && node results.mjs

Build what’s next.