Beyond Standard Leaderboards
Standard benchmarks measure multiple-choice performance. Our empirical evaluations test how frontier models behave on real-world, open-ended factual adjudications.
Beyond Benchmarks: Disagreement Among Frontier LLMs on Real-World Fact-Checks
While frontier large language models (LLMs) achieve comparable leaderboard scores on standardized evaluation benchmarks, this study examines whether frontier LLMs are functionally interchangeable when adjudicating real-world factual claims. Across 1,000 non-benchmark, user-submitted factual claims evaluated by 5 leading frontier models on a 5-point truth scale, models failed to reach consensus on 63% of claims. Furthermore, high confidence self-ratings (76% of responses rated 9 or 10 on a 10-point scale) failed to predict agreement (Krippendorff's alpha = 0.44 for confidence vs. 0.77 for verdicts). Our findings demonstrate that single-model verification introduces significant random variance, demonstrating the necessity of structured, multi-model adversarial evaluation pipelines.
Frontier LLM Agreement & Confidence Variance
Models evaluated on 1,000 user-submitted claims. 63% resulted in split verdicts.