Same protocol. Same criteria.

Black-box results,
reported as observed.

We ran 500 frozen synthetic security cases against AletheionAGI, Guardrails AI and NVIDIA NeMo Guardrails. The comparison measures observable behavior without requiring private implementation access.

Inspect the protocol

Final observed results

500 frozen cases per target.

PASS, FAIL, SKIPPED and ERROR keep their protocol meanings. A missing capability is not silently converted into a failure, and a provider error is not reported as a security finding.

PASSFAILSKIPPEDERROR

AletheionAGI

Fixed

446 PASS · 3 FAIL · 50 SKIPPED · 1 ERROR

3 observed failures fixed after this run

Guardrails AI

200 PASS · 100 FAIL · 200 SKIPPED · 0 ERROR

NVIDIA NeMo Guardrails

101 PASS · 127 FAIL · 80 SKIPPED · 192 ERROR

Frozen digest: f8777d9a35d9f7924b63244614117ffc5a79618343497180772275a969fd35a0. Fifty cases were run in each of ten security categories, for 500 cases per target.

What the non-pass outcomes mean

Measured failures stay separate from unavailable measurements.

AletheionAGI recorded three FAIL outcomes in contradictory-evidence cases and one transient HTTP 503 ERROR. Its 50 authorization-isolation cases were SKIPPED because the setup could not attest the required requester/label denial policy. The three contradictory-evidence defects were subsequently corrected; this page preserves the frozen run rather than retroactively converting them into PASS outcomes.

NVIDIA NeMo recorded 192 ERROR outcomes caused by HTTP 429 rate limits. Provider errors and unsupported capabilities are not reported as security failures.

Operational cost

Reproducibility has a footprint.

The recorded full run took about 1 hour 22 minutes. Local inference and remote provider calls are deliberately reported separately from behavioral scores.

Frozen evaluations
1,500
500 cases across three targets
Cases per category
50
Ten security categories
Local verification
45 / 45
Tests passing
External operations
Metered
NVIDIA API and Aletheion grounding credits

Interpretation boundary

A proof of concept, not a universal ranking.

This is a black-box proof of concept, not a claim of universal superiority. Same protocol, same evaluation criteria, results reported as observed.

Review cases and artifacts