Empirical AI Research & Benchmarks

Beyond Standard Leaderboards

Standard benchmarks measure multiple-choice performance. Our empirical evaluations test how frontier models behave on real-world, open-ended factual adjudications.

Published August 07, 2026·Version 1.0

Beyond Benchmarks: Disagreement Among Frontier LLMs on Real-World Fact-Checks

While frontier large language models (LLMs) achieve comparable leaderboard scores on standardized evaluation benchmarks, this study examines whether frontier LLMs are functionally interchangeable when adjudicating real-world factual claims. Across 1,000 non-benchmark, user-submitted factual claims evaluated by 5 leading frontier models on a 5-point truth scale, models failed to reach consensus on 63% of claims. Furthermore, high confidence self-ratings (76% of responses rated 9 or 10 on a 10-point scale) failed to predict agreement (Krippendorff's alpha = 0.44 for confidence vs. 0.77 for verdicts). Our findings demonstrate that single-model verification introduces significant random variance, demonstrating the necessity of structured, multi-model adversarial evaluation pipelines.

1,000
Claims Evaluated
Real-world, non-synthetic user claims
5
Frontier Models Tested
Leading models from OpenAI, Anthropic, Google, Meta
63%
Disagreement Rate
No unanimous consensus across models
23%
Severe Disagreements
Spread >= 2 rating categories
Open Evaluation Benchmark · 1,000 Real Claims

Frontier LLM Agreement & Confidence Variance

Models evaluated on 1,000 user-submitted claims. 63% resulted in split verdicts.

63%
Disagreement Rate
GPT-4.5 / 4oOpenAI
Agreement: 84.2%Avg Conf: 9.2/10
Claude 3.5 SonnetAnthropic
Agreement: 86.8%Avg Conf: 8.9/10
Gemini 1.5 ProGoogle
Agreement: 82.5%Avg Conf: 9.1/10
Llama 3.1 405BMeta
Agreement: 79.4%Avg Conf: 8.7/10
Mistral Large 2Mistral AI
Agreement: 78.1%Avg Conf: 8.5/10
Rated True
Rated Mixed
Rated False
1,000 Claims · Acuityio Empirical Benchmark v1.0