RxAi AMP · Benchmarks ·

Does memory pay for itself?

We test it the way you would test a new study trick. Two identical AI helpers get the same homework. One of them also gets a sticky note from the shared memory. Then we count every word each of them reads and writes. This page explains how it works, simply enough for a ten-year-old and precisely enough to run it again yourself.

tokenMonitor ↗ PROTOCOL.md ↗ paired A/B exact token counts read-only sessions
THE TEST

Twin helpers, one sticky note

Each test is a pair of brand-new sessions with the same AI model. They are identical in every way except one: whether the AMP hooks may hand the helper the memories that match its task.

AMP ON

Helper A

  • same question
  • same model and effort level
  • same tools, read-only
  • + a sticky note: the hooks inject the matching memories
AMP OFF

Helper B

  • same question
  • same model and effort level
  • same tools, read-only
  • AMP_DISABLE=1: the same hooks stay silent

Keeping it fair

ONE SWITCH

Only one thing differs

Prompt, model, tool allowlist and project are identical. The only difference between the two helpers is the memory switch.

AMP_DISABLE=1
TAKE TURNS

They take turns going first

In one pair the memory helper runs first, in the next pair the other one does. Neither side always gets the warm start of a filled prompt cache.

(pair + model) % 2
REPEAT

Do it many times

One try could just be luck, so tasks rotate across pairs and every question is asked more than once.

--cases 16 → 8 pairs
TRICK QUESTION

One question has no answer in memory

The control task is one no memory covers. It checks that memory stays quiet, and costs nothing, when it has nothing useful to say.

"expectedIssue": null
FROZEN MEMORY

The memory box is frozen

Memories keep changing weight as the store compiles. The helpers read a pinned copy with its remote removed, so every run sees the same records.

--memory-repo <snapshot>
HANDS OFF

Look, don't touch

Sessions may read files and run read-only commands, nothing else. Personal connectors are switched off, so AMP is the only memory under test.

--strict-mcp-config
RECEIPTS

How the tokens are counted

A token is a small piece of a word. The helper is charged for every token it reads and every token it writes.

It doesn't do the homework in one go. It works in steps: think, open a file, think again, open another file. Each step is one call to the model, and at every call it re-reads the whole conversation from the beginning, like reading your whole notebook again before you write the next sentence. That is why the reading count grows so fast.

An illustration, not measured data. Call 1 holds the system prompt, the question and (for helper A) the sticky note. Each later call carries everything before it plus the file it just opened. Add up every bar and you get the context tokens.

The receipt printed after every call

The runner starts Claude Code with --output-format stream-json, which prints one line per model call with a usage block. The benchmark adds the fields up across the whole session.

On the receipt What it means
input_tokensBrand-new pages, read for the first time
cache_creation_input_tokensPages read and bookmarked so they can be skipped cheaply next time
cache_read_input_tokensPages already bookmarked; about a tenth of the normal price
output_tokensWords it wrote, including its private thinking

What gets added up

Context tokens All the reading, bookmarked or not: Σ (input + cache creation + cache read) over every call.
Uncached input Only the full-price reading: Σ (input + cache creation).
Start context How heavy the backpack is before any work: the first call's context, which includes the setup text and the sticky note.
Tool output How much it dug up: the bytes of every file and command result it received before writing its answer.
Calls and look-ups Model calls, tool calls, and reads (Read, Grep, Glob).
Time and cost The runner's own stopwatch, and the list price Claude Code reports in its final result event (total_cost_usd).

These are exact numbers from the model's own receipts, not guesses. The monitor's npm run scan estimates token counts with a tokenizer and a calibration factor; the benchmark never estimates.

GRADING

Checking the homework

Being cheaper means nothing if the answer is wrong. Every answer gets three checks: one on its content, one on what the model says it used, and one on what the hooks actually delivered.

CHECK 1 · KEYWORDS

Did it name the right things?

Each task lists signals: file and function names a correct answer should mention. The runner checks which ones appear. This is a deterministic check, not a quality grade.

"signals": [ … ]
CHECK 2 · RECALL LINE

Did it say which note helped?

Every answer must end with a line naming the memories it used. Right means the expected memory is named, or none on the trick question.

Recall used: #N | none
CHECK 3 · LEDGER

Was the note really handed over?

The hooks keep a diary of what they injected and how much of it: the title only, or a short summary. The runner copies it next to each result.

ledger · pointer | summary

Checks 2 and 3 are kept apart on purpose. In one September 2026 run, two sessions had the right summary in front of them and still wrote Recall used: none. A single "memory hits" number would have called that a retrieval miss; the ledger shows the memory was delivered and simply not credited.

SCORE

Keeping score

change = (AMP on − AMP off) ÷ AMP off × 100%

Each measurement is summed over every successful session in each arm. A negative number means the memory helper used less. "Cheaper pairs" counts how often the memory helper beat its own twin.

A real result: Opus 5.5 at xhigh effort, 8 pairs, 2026-09-23

Measurement Change with memory
Total readingcontext tokens−13.6%
Full-price readinguncached input+21.2%
Stuff dug up from filestool output−14.9%
Tool calls−11.1%
Time−9.0%
Costcheaper in 4 of 8 pairs+4.9%
Keyword checklistmemory on vs off34/34 · 34/34

In kid terms: with the sticky note, the helper read less overall, dug up less, and finished a bit faster. But the note itself is new, full-price text, so the total bill came out about the same, slightly higher on average. Both helpers got the homework equally right.

What the benchmark taught us

Same test, same day Reading Dug up Time Cost
Before: the note was a long excerpt +27.9%+19.2%+7.1%+27.1%
After: a short summary, only when the task matches −13.6%−14.9%−9.0%+4.9%
After, repeated with stricter matching (B2) −6.5%−4.6%−3.9%+14.1%

The two "after" rows should agree and don't quite: with only two pairs per task, one control session that dug up 8 KB more and one cold-cache first session are enough to move cost by ten points. Read the direction (memory reads less and costs a little more), not the decimals.

The first sticky note listed a person's decisions one by one. The helper treated the list as homework of its own and went to check every item, so memory made it read more. The fix was to hand over only a short goal-and-status summary, and only for tasks the memory matches; the full record stays one fetch away. Same model, same questions: reading flipped from +28% to −14%.

An earlier run with Opus 5 produced the numbers on the overview page: 18% less dug up from files, 5% cheaper, 14% faster. A model that reads little to begin with gained nothing there. Memory saves blind exploration, so the more a model explores, the more memory pays.

CODEX

Same test, different receipts

Codex runs the same experiment with the same task table and schedule. Its receipts come from the turn.completed event at the end of each turn:

On the receipt What it means
input_tokensEverything read, bookmarked pages included
cached_input_tokensThe bookmarked part of that; uncached = input − cached
output_tokensEverything written, reasoning included
reasoning_output_tokensThe thinking part of that; visible output = output − reasoning

Same snapshot, same questions: the GPT-6 family on Codex, hooks delivering

Every row is memory on against memory off on the same pinned memory snapshot and the same project commit, prompt wording v2. "Named the right note" is the model's own Recall used: line on the three memory-backed tasks. Cost is the API-equivalent from the price list; astra has no price row yet.

Run Readinginput tokens Dug up Time Cost Named the right note
GPT-6 Sol · xhigh8 pairs · 2026-09-23+0.7%−18.4%−4.2%+4.2%4/6
GPT-6 Luna · xhigh8 pairs · 2026-09-23−12.3%−8.1%+8.5%0.0%4/6
GPT-6 Sol · xhigh4 pairs · 2026-09-24 · Codex memory off−17.1%+8.5%−25.5%−10.6%3/3
GPT-6 Luna · xhigh4 pairs · 2026-09-24 · Codex memory off+26.1%+49.1%+41.9%+11.4%2/3
gpt-6-astra · medium4 pairs · 2026-09-24 · Codex memory off+9.9%−4.5%−1.7%—2/3
gpt-6-astra · xhigh4 pairs · 2026-09-24 · Codex memory off−3.4%−13.6%−10.7%—3/3

Four pairs per model is half a normal run. Sol's 4-pair −17% sits next to its 8-pair +0.7% from the day before, and Luna flips the other way, so neither is a headline yet; an 8-pair repeat with the isolation flags decides it. What holds across Codex is what held on Claude Code: the harder the model thinks, the more it reads, and the more there is for memory to replace. At xhigh, astra dug up 13.6% less and finished 10.7% sooner; at medium it broke even. Every session's ledger shows the same three notes delivered.

LIMITS

What it can't tell you

RUN IT

Run it yourself

  1. Write the questions. They are specific to the project you test, so they are not committed. Each task has an id, a question, the signals a good answer names, and the expectedIssue it should recall (null for the trick question).
    cd tokenMonitor
    cp config/tasks.example.json config/tasks.json   # then edit
  2. Freeze the memory. Clone the memory repo to a snapshot and remove its remote so no hook can pull newer records into it.
    git clone <memory-repo> ../memory-snapshot
    git -C ../memory-snapshot remote remove origin
  3. Run the pairs. --cases counts sessions, not pairs, and every case is a full, paid session.
    npm run benchmark:claude -- --cases 16 \
      --models claude-opus-5-5 --effort xhigh \
      --project /path/to/target-project \
      --memory-repo ../memory-snapshot
    
    npm run benchmark:codex -- --cases 20 --hooks \
      --models <model-a>,<model-b> --reasoning medium \
      --project /path/to/target-project \
      --memory-repo ../memory-snapshot
  4. Read the report. Each run writes report.md and summary.json, and keeps every session's raw output next to them. An interrupted run resumes from results.ndjson.

Full options and caveats: tokenMonitor/README.md.