Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking
Summary
Multimodal automated fact-checking (MAFC) benchmarks often inflate performance due to "contamination," where claims are verifiable using an LLM's internal knowledge rather than external evidence. While emerging dynamic benchmarks collect claims published after LLMs' knowledge cut-off dates to assume uncontaminated status, this work empirically challenges that assumption. The study found that dynamic evaluation reduces but does not eliminate contamination risks, with 17.09%--29.30% of post-cut-off claims still potentially contaminated. Many newly published claims can be verified using pre-cut-off public knowledge. Critically, contamination can induce statistically significant inflation in MAFC performance, increasing Macro-F1 by up to 11.34 points and distorting system rankings. The research re-evaluates state-of-the-art LLMs under strictly controlled settings and provides practical guidelines for trustworthy MAFC evaluation.
Key takeaway
For AI Scientists and Machine Learning Engineers developing or evaluating multimodal automated fact-checking systems, you must critically reassess "contamination-free" dynamic benchmarks. Your current performance metrics may be inflated by up to 11.34 Macro-F1 points, distorting system rankings. Implement stricter contamination controls, ensuring claims genuinely require external, post-cut-off evidence to accurately gauge true MAFC capabilities.
Key insights
Dynamic evaluation for multimodal automated fact-checking still faces significant contamination risks from pre-cut-off knowledge.
Principles
- Static MAFC benchmarks often inflate performance due to outdated claims.
- Post-LLM knowledge cut-off claims are not inherently contamination-free.
- Contamination significantly distorts MAFC system rankings and Macro-F1 scores.
Method
The study empirically investigates contamination risks in static (AVeriTeC) and dynamic (ClaimReview2025Q4) MAFC benchmarks, re-evaluating SOTA LLMs under strictly controlled settings.
In practice
- Scrutinize "contamination-free" claims in dynamic benchmarks.
- Verify if post-cut-off claims require truly novel external evidence.
- Implement strict contamination controls for MAFC evaluation.
Topics
- Multimodal Fact-Checking
- Benchmark Contamination
- Dynamic Evaluation
- LLM Knowledge Cut-off
- AVeriTeC
- ClaimReview2025Q4
Best for: Research Scientist, AI Scientist, Machine Learning Engineer, AI Ethicist
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Computation and Language.