RxAi AMP · Benchmarks ·
We test it the way you would test a new study trick. Two identical AI helpers get the same homework. One of them also gets a sticky note from the shared memory. Then we count every word each of them reads and writes. This page explains how it works, simply enough for a ten-year-old and precisely enough to run it again yourself.
Each test is a pair of brand-new sessions with the same AI model. They are identical in every way except one: whether the AMP hooks may hand the helper the memories that match its task.
AMP_DISABLE=1: the same hooks stay silentPrompt, model, tool allowlist and project are identical. The only difference between the two helpers is the memory switch.
AMP_DISABLE=1In one pair the memory helper runs first, in the next pair the other one does. Neither side always gets the warm start of a filled prompt cache.
(pair + model) % 2One try could just be luck, so tasks rotate across pairs and every question is asked more than once.
--cases 16 → 8 pairsThe control task is one no memory covers. It checks that memory stays quiet, and costs nothing, when it has nothing useful to say.
"expectedIssue": nullMemories keep changing weight as the store compiles. The helpers read a pinned copy with its remote removed, so every run sees the same records.
--memory-repo <snapshot>Sessions may read files and run read-only commands, nothing else. Personal connectors are switched off, so AMP is the only memory under test.
--strict-mcp-configA token is a small piece of a word. The helper is charged for every token it reads and every token it writes.
It doesn't do the homework in one go. It works in steps: think, open a file, think again, open another file. Each step is one call to the model, and at every call it re-reads the whole conversation from the beginning, like reading your whole notebook again before you write the next sentence. That is why the reading count grows so fast.
An illustration, not measured data. Call 1 holds the system prompt, the question and (for helper A) the sticky note. Each later call carries everything before it plus the file it just opened. Add up every bar and you get the context tokens.
The runner starts Claude Code with --output-format stream-json, which prints one line per model call with a usage block. The benchmark adds the fields up across the whole session.
| On the receipt | What it means |
|---|---|
| input_tokens | Brand-new pages, read for the first time |
| cache_creation_input_tokens | Pages read and bookmarked so they can be skipped cheaply next time |
| cache_read_input_tokens | Pages already bookmarked; about a tenth of the normal price |
| output_tokens | Words it wrote, including its private thinking |
Read, Grep, Glob).
result event (total_cost_usd).
These are exact numbers from the model's own receipts, not guesses. The monitor's npm run scan estimates token counts with a tokenizer and a calibration factor; the benchmark never estimates.
Being cheaper means nothing if the answer is wrong. Every answer gets three checks: one on its content, one on what the model says it used, and one on what the hooks actually delivered.
Each task lists signals: file and function names a correct answer should mention. The runner checks which ones appear. This is a deterministic check, not a quality grade.
"signals": [ … ]Every answer must end with a line naming the memories it used. Right means the expected memory is named, or none on the trick question.
The hooks keep a diary of what they injected and how much of it: the title only, or a short summary. The runner copies it next to each result.
ledger · pointer | summaryChecks 2 and 3 are kept apart on purpose. In one September 2026 run, two sessions had the right summary in front of them and still wrote Recall used: none. A single "memory hits" number would have called that a retrieval miss; the ledger shows the memory was delivered and simply not credited.
Each measurement is summed over every successful session in each arm. A negative number means the memory helper used less. "Cheaper pairs" counts how often the memory helper beat its own twin.
| Measurement | Change with memory |
|---|---|
| Total readingcontext tokens | −13.6% |
| Full-price readinguncached input | +21.2% |
| Stuff dug up from filestool output | −14.9% |
| Tool calls | −11.1% |
| Time | −9.0% |
| Costcheaper in 4 of 8 pairs | +4.9% |
| Keyword checklistmemory on vs off | 34/34 · 34/34 |
In kid terms: with the sticky note, the helper read less overall, dug up less, and finished a bit faster. But the note itself is new, full-price text, so the total bill came out about the same, slightly higher on average. Both helpers got the homework equally right.
| Same test, same day | Reading | Dug up | Time | Cost |
|---|---|---|---|---|
| Before: the note was a long excerpt | +27.9% | +19.2% | +7.1% | +27.1% |
| After: a short summary, only when the task matches | −13.6% | −14.9% | −9.0% | +4.9% |
| After, repeated with stricter matching (B2) | −6.5% | −4.6% | −3.9% | +14.1% |
The two "after" rows should agree and don't quite: with only two pairs per task, one control session that dug up 8 KB more and one cold-cache first session are enough to move cost by ten points. Read the direction (memory reads less and costs a little more), not the decimals.
The first sticky note listed a person's decisions one by one. The helper treated the list as homework of its own and went to check every item, so memory made it read more. The fix was to hand over only a short goal-and-status summary, and only for tasks the memory matches; the full record stays one fetch away. Same model, same questions: reading flipped from +28% to −14%.
An earlier run with Opus 5 produced the numbers on the overview page: 18% less dug up from files, 5% cheaper, 14% faster. A model that reads little to begin with gained nothing there. Memory saves blind exploration, so the more a model explores, the more memory pays.
Codex runs the same experiment with the same task table and schedule. Its receipts come from the turn.completed event at the end of each turn:
| On the receipt | What it means |
|---|---|
| input_tokens | Everything read, bookmarked pages included |
| cached_input_tokens | The bookmarked part of that; uncached = input − cached |
| output_tokens | Everything written, reasoning included |
| reasoning_output_tokens | The thinking part of that; visible output = output − reasoning |
config/prices.json): uncached input × input price + cached input × cached price + output × output price. On a ChatGPT plan nothing is billed per token; the number is what the same traffic would cost on the API.--ignore-rules and --disable memories switch off Codex's own saved approvals and its built-in memory, so AMP is the only memory being tested.--rollouts the runner also reads each session's rollout file for per-call numbers: calls per case, the first call's size, the biggest call, and cold starts.Every row is memory on against memory off on the same pinned memory snapshot and the same project commit, prompt wording v2. "Named the right note" is the model's own Recall used: line on the three memory-backed tasks. Cost is the API-equivalent from the price list; astra has no price row yet.
| Run | Readinginput tokens | Dug up | Time | Cost | Named the right note |
|---|---|---|---|---|---|
| GPT-6 Sol · xhigh8 pairs · 2026-09-23 | +0.7% | −18.4% | −4.2% | +4.2% | 4/6 |
| GPT-6 Luna · xhigh8 pairs · 2026-09-23 | −12.3% | −8.1% | +8.5% | 0.0% | 4/6 |
| GPT-6 Sol · xhigh4 pairs · 2026-09-24 · Codex memory off | −17.1% | +8.5% | −25.5% | −10.6% | 3/3 |
| GPT-6 Luna · xhigh4 pairs · 2026-09-24 · Codex memory off | +26.1% | +49.1% | +41.9% | +11.4% | 2/3 |
| gpt-6-astra · medium4 pairs · 2026-09-24 · Codex memory off | +9.9% | −4.5% | −1.7% | — | 2/3 |
| gpt-6-astra · xhigh4 pairs · 2026-09-24 · Codex memory off | −3.4% | −13.6% | −10.7% | — | 3/3 |
Four pairs per model is half a normal run. Sol's 4-pair −17% sits next to its 8-pair +0.7% from the day before, and Luna flips the other way, so neither is a headline yet; an 8-pair repeat with the isolation flags decides it. What holds across Codex is what held on Claude Code: the harder the model thinks, the more it reads, and the more there is for memory to replace. At xhigh, astra dug up 13.6% less and finished 10.7% sooner; at medium it broke even. Every session's ledger shows the same three notes delivered.
Prompt rules: vN.id, a question, the signals a good answer names, and the expectedIssue it should recall (null for the trick question).
cd tokenMonitor
cp config/tasks.example.json config/tasks.json # then edit
git clone <memory-repo> ../memory-snapshot git -C ../memory-snapshot remote remove origin
--cases counts sessions, not pairs, and every case is a full, paid session.
npm run benchmark:claude -- --cases 16 \ --models claude-opus-5-5 --effort xhigh \ --project /path/to/target-project \ --memory-repo ../memory-snapshot npm run benchmark:codex -- --cases 20 --hooks \ --models <model-a>,<model-b> --reasoning medium \ --project /path/to/target-project \ --memory-repo ../memory-snapshot
report.md and summary.json, and keeps every session's raw output next to them. An interrupted run resumes from results.ndjson.
Full options and caveats: tokenMonitor/README.md.