©2026 adebench · adecubedbuilt 2026-09-29 from the reports in the repo

ADEBENCH — agent memory benchmark and leaderboard

A benchmark for agent memory. It scores the text a memory actually delivers to the model, on one golden set, with no LLM judge.

Leaderboard · synthetic golden set
MemoryCore · 55Full Chars per answerWrite → visibleReached through
Dakera
REST memory with supersession edges and session-scoped recall
94%
51.9 / 55
97%
96.9 / 100
304 1.1 s REST
ADE Brain
episodes, cards, a fact layer that retires a value when a newer one contradicts it
89%
48.8 / 55
94%
93.8 / 100
1,272 101 ms REST
Aionforge
bi-temporal graph in Rust, hybrid recall, explicit forgetting, MCP only
76%
41.8 / 55
80%
51.8 / 65
419 483 ms MCP
gbrain
entity pages, chronicle, remember/recall/forget, hybrid search
73%
40.4 / 55
84%
75.4 / 90
1,220 2.6 s CLI / MCP
Hindsight
world facts, experiences and observations; LLM extraction on retain; per-bank MCP
66%
36.3 / 55
66%
36.2 / 55
3,225 — MCP
Memoose
local knowledge graph in SQLite, triples plus lexical chunks, 26 MCP tools
63%
34.6 / 55
76%
68.3 / 90
1,760 2.3 s CLI / SQLite
Jev-Mem
graph memory whose decisions are taken by a small System-One model
58%
32.1 / 55
65%
42.1 / 65
943 3.4 s Python library
Nemp
Claude Code plugin: a JSON file in the project, and the model itself as the engine
58%
32.1 / 55
56%
36.4 / 65
29,530 70.5 s inside Claude Code

Core is what every memory can be measured on, a write and a read through the door: door 25, updates 10, time 10, abstention 10. The ranking is on it. Full adds what a memory has behind the door (cards, live state, file search, graph), over the points it could be measured on. Chars per answer is what the model receives for one question: the same score at 300 characters and at 20,000 is not the same memory. Write → visible is how long a value just written takes to reach the door.

On their own data
MemoryScoreProbesRun
ADE Brain
the Brain on its owner's real memory: 165 probes written on that memory, private
98%
98.0 / 100
165
set 9a276b567d
2026-09-28
adebench 0.2.15

The golden set is small on purpose, so it runs anywhere. The bench was built for something else: a memory measured on its owner's data, with probes written on the facts it really holds, rerun after every change. Not comparable across memories, and the number that tells an owner whether a change helped. Totals only: the probes are the owner's facts.

Method

Eight sections: door, cards, updates, time, live state, abstention, file search, graph. A point is earned when the expected words are in the text the model receives; a retired value next to the current one is a failure.

Rules in the README →

Run it

No dependencies. Every adapter and importer is in the repository, and each memory's page names the configuration that produced its numbers.

python -m adebench --help

Your memory

Write your probes, point the harness at your memory, send the totals with a pull request. The set's hash ties a number to the probes that produced it.

GitHub

adebench is MIT. Adapters, golden set, reports: adecubed/adebench. Results are reviewed with each memory's author before they appear here.