← Research Hub/Acuityio Benchmark Report

Beyond Benchmarks: Disagreement Among Frontier LLMs on Real-World Fact-Checks

Published August 2026 · Acuityio AI Research Team · 1,000 Claims Empirical Benchmark v1.0

Executive Abstract

While frontier large language models (LLMs) achieve comparable leaderboard scores on standardized evaluation benchmarks, this study examines whether frontier LLMs are functionally interchangeable when adjudicating real-world factual claims. Across 1,000 non-benchmark, user-submitted factual claims evaluated by 5 leading frontier models on a 5-point truth scale, models failed to reach consensus on 63% of claims. Furthermore, high confidence self-ratings (76% of responses rated 9 or 10 on a 10-point scale) failed to predict agreement (Krippendorff's alpha = 0.44 for confidence vs. 0.77 for verdicts). Our findings demonstrate that single-model verification introduces significant random variance, demonstrating the necessity of structured, multi-model adversarial evaluation pipelines.

1,000
Claims Evaluated
Real-world, non-synthetic user claims
5
Frontier Models Tested
Leading models from OpenAI, Anthropic, Google, Meta
63%
Disagreement Rate
No unanimous consensus across models
23%
Severe Disagreements
Spread >= 2 rating categories
0.44
Confidence Alpha
Low agreement on confidence ratings
0.77
Verdicts Alpha
Moderate inter-annotator agreement on verdicts

Model Comparison Matrix

Open Evaluation Benchmark · 1,000 Real Claims

Frontier LLM Agreement & Confidence Variance

Models evaluated on 1,000 user-submitted claims. 63% resulted in split verdicts.

63%
Disagreement Rate
GPT-4.5 / 4oOpenAI
Agreement: 84.2%Avg Conf: 9.2/10
Claude 3.5 SonnetAnthropic
Agreement: 86.8%Avg Conf: 8.9/10
Gemini 1.5 ProGoogle
Agreement: 82.5%Avg Conf: 9.1/10
Llama 3.1 405BMeta
Agreement: 79.4%Avg Conf: 8.7/10
Mistral Large 2Mistral AI
Agreement: 78.1%Avg Conf: 8.5/10
Rated True
Rated Mixed
Rated False
1,000 Claims · Acuityio Empirical Benchmark v1.0