Measured, not assumed

Evidence for the
evidence layer.

We publish the completed result, the measurement boundary and the limitation together. AletheionAGI is evaluated separately for retrieval, reader outcomes and fail-closed delivery.

Inspect the benchmark

Completed retrieval result

The right evidence,
without the history dump.

On 979 frozen, support-valid MultiWOZ questions with the same Qwen3 14B reader, ASM-CM + Bridge 8.1 produced the strongest completed Recall@5 and diagnostic answer score in this compact comparison.

93.6%Recall@5ASM-CM + Bridge 8.1
66.5%Diagnostic answer scoreSame 979-question protocol
1.09KReader input tokens / questionMeasured context volume
979/979Questions completedFrozen retrieval comparison
Frozen support benchmarkMultiWOZ
ReaderQwen3 14B · 979 questions
System
Recall@5
Diagnostic answer score
Reader input / question
ASM-CM + Bridge
93.6%
66.5%
1.09K
Vector RAG
70%
49.7%
2.04K
BM25
75.9%
56.8%
2.20K
Protocol-scoped results. Quality depends on workload, reader and retrieval profile; the Bridge is measured separately as the delivery boundary.

A budget is a safety control

More context is not automatically a better answer.

In a separate LongMemEval-S control—500 questions, GPT-4o reader and judge—raising the evidence budget from 2K to 28K increased reader latency and reduced measured answer accuracy for four of five retrieval systems.

The result is intentionally narrow: it does not prove that more context is always harmful. It shows why evidence selection, token budgets and a delivery gate must be measured rather than guessed.

4 / 5systems lost answer accuracy at the larger measured evidence budget

How to read these numbers

A benchmark is a boundary,
not a slogan.

01

Completed artifacts only

Marketing claims use completed question sets. Incomplete or diagnostic paths stay labelled as such and are excluded from promoted comparisons.

02

Separate the layers

Retrieval quality, reader quality, context volume, latency and Bridge containment answer different questions. We do not blend them into one vanity score.

03

Choose profiles by workload

ASM-CM, lexical retrieval and vector retrieval are profiles, not a universal winner. The production invariant is authorized evidence and fail-closed delivery.

See the security suite