Skip to content

Benchmarks · BEAM-100K

Memory, measured.

BEAM asks 400 questions about 20 long conversations, across ten abilities. We run it the way a customer's company runs: every message goes into the company's memory, and the agent answers each question in a fresh session.

77.8/ 100

reader DeepSeek V4 Flash · judge GPT-4.1 mini · 2026-09-05

BEAM-100K, as posted

SystemScoreReaderJudgePosted
AEQI77.8DeepSeek V4 FlashGPT-4.1 miniSep 2026
Exabase M-176.9Gemini 3 FlashGemini 3 Flash2026
Hindsight73.4Apr 2026
AutoMem67.52026
Mnemosyne65.2Llama 3.3 70BDeepSeek V4 Flash2026
Honcho63.0Gemini 2.5 Flash LiteGPT-4o2026
Ogham55.42026
BEAM paper, best baseline35.8Qwen2.5-32BICLR 2026

Each row is the number its vendor posted, with the reader and judge they stated. No row was re-run by us.

Our score by ability

Information extraction93.0
Preference following91.5
Abstention90.0
Instruction following85.4
Multi-session reasoning79.5
Knowledge update75.0
Event ordering73.7
Contradiction resolution72.5
Summarization62.7
Temporal reasoning55.0

Forty questions per ability. Temporal reasoning and summarization are where the next points are.

How the number moved

Sep 1
73.0
first full run
Sep 4
71.8 · 73.9
the same memory, answered twice — the noise band
Sep 5
77.8
search renderers, read path, chunking and timeline fixed

Every point came from a defect fixed in the product, not in the benchmark harness. Between the first run and this one, injected context per question fell from 23K characters to 1.2K and tool calls from 15.9 to 6.2.

What it cost

Tool calls per question
6.2
Searches per question
3.1
Prompt tokens per question
41K
Cost of 400 answers
$1.22
Cost of ingesting 20 conversations
$5.87

DeepSeek pricing via OpenRouter. The judge added about $4.

Notes

  • Readers and judges differ from row to row, so a gap under five points is a tie, not a rank.
  • Our per-question standard deviation is 38 points; at 400 questions the smallest real difference is about 5.
  • Two gold answers in one conversation are wrong in the dataset. We scored them as misses.
  • The lowest-scoring conversation was used to find and fix defects with its questions visible; the other 19 were not read. Gains showed in 13 of those 19.
  • Per-question answers, judge verdicts, and run configuration are kept on disk and available on request.

Questions: 0x@aeqi.io

Cookies for sign-in and analytics. No third-party tracking.