Beyond Benchmarks: Disagreement Among Frontier LLMs on Real-World Fact-Checks
Published August 2026 · Acuityio AI Research Team · 1,000 Claims Empirical Benchmark v1.0
Executive Abstract
While frontier large language models (LLMs) achieve comparable leaderboard scores on standardized evaluation benchmarks, this study examines whether frontier LLMs are functionally interchangeable when adjudicating real-world factual claims. Across 1,000 non-benchmark, user-submitted factual claims evaluated by 5 leading frontier models on a 5-point truth scale, models failed to reach consensus on 63% of claims. Furthermore, high confidence self-ratings (76% of responses rated 9 or 10 on a 10-point scale) failed to predict agreement (Krippendorff's alpha = 0.44 for confidence vs. 0.77 for verdicts). Our findings demonstrate that single-model verification introduces significant random variance, demonstrating the necessity of structured, multi-model adversarial evaluation pipelines.
Model Comparison Matrix
Frontier LLM Agreement & Confidence Variance
Models evaluated on 1,000 user-submitted claims. 63% resulted in split verdicts.