Measured, not assumed
Evidence for the
evidence layer.
We publish the completed result, the measurement boundary and the limitation together. AletheionAGI is evaluated separately for retrieval, reader outcomes and fail-closed delivery.
Inspect the benchmark ↗Completed retrieval result
The right evidence,
without the history dump.
On 979 frozen, support-valid MultiWOZ questions with the same Qwen3 14B reader, ASM-CM + Bridge 8.1 produced the strongest completed Recall@5 and diagnostic answer score in this compact comparison.
1.09K2.04K2.20KA budget is a safety control
More context is not automatically a better answer.
In a separate LongMemEval-S control—500 questions, GPT-4o reader and judge—raising the evidence budget from 2K to 28K increased reader latency and reduced measured answer accuracy for four of five retrieval systems.
The result is intentionally narrow: it does not prove that more context is always harmful. It shows why evidence selection, token budgets and a delivery gate must be measured rather than guessed.
How to read these numbers
A benchmark is a boundary,
not a slogan.
Completed artifacts only
Marketing claims use completed question sets. Incomplete or diagnostic paths stay labelled as such and are excluded from promoted comparisons.
Separate the layers
Retrieval quality, reader quality, context volume, latency and Bridge containment answer different questions. We do not blend them into one vanity score.
Choose profiles by workload
ASM-CM, lexical retrieval and vector retrieval are profiles, not a universal winner. The production invariant is authorized evidence and fail-closed delivery.