Public Benchmark· Aug 2026 Certified
Can you trust your AI? The hallucination benchmark.
SF2X scores every AI system on whether its answers are warranted, supported, and resistant to adversarial attack. The full red-team loop is what separates a demo from trust infrastructure — the best certified run scores 91/100.
Key metrics — SF2X certified run
Warrant rate
95%
Answers carrying a valid warrant
Trustworthy rate
87
Mean trustworthy answer rate
Correction rate
15%
Answers later corrected
Drift score
0.16
Lower is more stable
MTTC
1m
Mean time to correction
Try it on your own AI answer
Paste any claim into the tribunal playground and watch the proposer–critic–verifier debate render a verdict in real time.
Trust scores are vendor claims until independently audited. SF2X commits to at least two third-party audits (TruthfulQA / HaluEval correlation) published on the methodology page. Field baselines are reference systems, not SF2X products.