Public Benchmark· Aug 2026 Certified

Can you trust your AI? The hallucination benchmark.

SF2X scores every AI system on whether its answers are warranted, supported, and resistant to adversarial attack. The full red-team loop is what separates a demo from trust infrastructure — the best certified run scores 91/100.

Share on X

Key metrics — SF2X certified run

Warrant rate
95%
Answers carrying a valid warrant
Trustworthy rate
87
Mean trustworthy answer rate
Correction rate
15%
Answers later corrected
Drift score
0.16
Lower is more stable
MTTC
1m
Mean time to correction
Try it on your own AI answer

Paste any claim into the tribunal playground and watch the proposer–critic–verifier debate render a verdict in real time.

Open the playground

Trust scores are vendor claims until independently audited. SF2X commits to at least two third-party audits (TruthfulQA / HaluEval correlation) published on the methodology page. Field baselines are reference systems, not SF2X products.