Benchmarks · BEAM-100K
Memory, measured.
BEAM asks 400 questions about 20 long conversations, across ten abilities. We run it the way a customer's company runs: every message goes into the company's memory, and the agent answers each question in a fresh session.
reader DeepSeek V4 Flash · judge GPT-4.1 mini · 2026-09-05
BEAM-100K, as posted
Each row is the number its vendor posted, with the reader and judge they stated. No row was re-run by us.
Our score by ability
Forty questions per ability. Temporal reasoning and summarization are where the next points are.
How the number moved
- Sep 1
- 73.0
- first full run
- Sep 4
- 71.8 · 73.9
- the same memory, answered twice — the noise band
- Sep 5
- 77.8
- search renderers, read path, chunking and timeline fixed
Every point came from a defect fixed in the product, not in the benchmark harness. Between the first run and this one, injected context per question fell from 23K characters to 1.2K and tool calls from 15.9 to 6.2.
What it cost
- Tool calls per question
- 6.2
- Searches per question
- 3.1
- Prompt tokens per question
- 41K
- Cost of 400 answers
- $1.22
- Cost of ingesting 20 conversations
- $5.87
DeepSeek pricing via OpenRouter. The judge added about $4.
Notes
- Readers and judges differ from row to row, so a gap under five points is a tie, not a rank.
- Our per-question standard deviation is 38 points; at 400 questions the smallest real difference is about 5.
- Two gold answers in one conversation are wrong in the dataset. We scored them as misses.
- The lowest-scoring conversation was used to find and fix defects with its questions visible; the other 19 were not read. Gains showed in 13 of those 19.
- Per-question answers, judge verdicts, and run configuration are kept on disk and available on request.
Questions: 0x@aeqi.io