Benchmarks / LongMemEval
LongMemEval
Five hundred questions, each with its own haystack of chat sessions around 115k tokens. The engine, frozen on LoCoMo and run here untuned, against the cosine floor on the same vectors across 4,846 sessions.
- dataset
- LongMemEval
- licence
- MIT
- introduced by
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory — Wu et al., 2024
- harness
- hanzoai/benchmarks · brain
node longmemeval/engine.mjs --embed=minilm && node results.mjsregenerates every table on this page from the raw runs
What is measured
Five hundred questions, each with its own haystack of chat sessions about 115k tokens long. One or more sessions hold the evidence, and retrieval is scored as recall of those sessions against the benchmark’s own answer_session_ids. Every session is embedded turn by turn and ranked by its best turn — 4,846 sessions and 199,641 turns. ALL@5 is the share of questions whose every evidence session is in the top five; ANY@5 is the share with at least one.
Both columns use the same all-minilm vectors. The cosine column ranks a session by its best turn’s cosine. The engine column ranks it by its best turn under the engine’s score, with the configuration frozen on LoCoMo dev at 33d0f8749bd1 — nothing was tuned on this benchmark, so all 500 questions are the report.
Results
embedder all-minilm · engine b467901f3950 · retrieval p50 cosine 0.20 ms, engine 1.96 ms · run longmemeval-all-context-minilm
| question type | n | ALL@5 cosine | ALL@5 engine | ANY@5 cosine | ANY@5 engine | ALL@10 cosine | ALL@10 engine | MRR cosine | MRR engine |
|---|---|---|---|---|---|---|---|---|---|
| all questions | 500 | 85.8 | 90.6 | 97.4 | 98.2 | 93.4 | 97.0 | 0.903 | 0.919 |
| single-session-user | 70 | 95.7 | 98.6 | 95.7 | 98.6 | 97.1 | 100.0 | 0.852 | 0.872 |
| multi-session | 133 | 75.9 | 85.7 | 98.5 | 99.2 | 89.5 | 94.7 | 0.912 | 0.917 |
| single-session-preference | 30 | 96.7 | 96.7 | 96.7 | 96.7 | 96.7 | 96.7 | 0.803 | 0.832 |
| temporal-reasoning | 133 | 75.9 | 81.2 | 94.7 | 95.5 | 88.7 | 94.7 | 0.865 | 0.894 |
| knowledge-update | 78 | 96.2 | 98.7 | 100.0 | 100.0 | 98.7 | 100.0 | 0.964 | 0.981 |
| single-session-assistant | 56 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 1.000 | 1.000 |
Where it gains
Every evidence session is in the top five for 90.6% of questions against 85.8% for cosine, and in the top ten for 97.0% against 93.4%. No question type goes down. The gain sits where the floor left room: multi-session ALL@5 moves from 75.9 to 85.7, and temporal-reasoning from 75.9 to 81.2. Single-session-assistant was saturated and stays at 100.0.
Where it loses, and why
The engine runs here with one hand tied. LongMemEval has no fact layer, so its fact, entity and typed-graph generators propose nothing; what does the work is lexical match, the session timeline, adjacency and a second hop. Three questions lose ALL@5 to cosine. The questions whose first correct session ranks lower are mostly phrased in relative time — “four weeks ago”, “last Saturday” — and the engine’s timeline fires only on a named month or year and does not read the question’s date. That is the next thing to fix, and it is a fix to the engine, not to this table.
Retrieval costs more: 1.96 ms at the median against 0.20 ms for cosine, once the turns are embedded.
What this table is not
Session recall is not answer accuracy. Systems reporting LongMemEval scores in the high eighties usually report an end-to-end answer metric with a language-model judge, which is a different measurement and not comparable to these columns. The run uses the release the dataset card now marks deprecated in favour of a cleaned one.
raw runs and scoring: hanzoai/benchmarks · brain · regenerated by node longmemeval/engine.mjs --embed=minilm && node results.mjs