ARTFEED — Contemporary Art Intelligence

LLM-Based Trace Ranking for Numerical Claim Verification at CheckThat! 2026

ai-technology · 2026-07-29

For the CLEF 2026 CheckThat! Task 2, a system was created to automate the verification of numerical claims, concentrating on ranking reasoning traces derived from large language models (LLMs) and forecasting final decisions for numerical assertions in both English and Arabic. The initial method involves fine-tuning an LLM-based verifier via LoRA for binary classification of each trace, employing Best-of-N selection for the concluding verdict, and testing adaptive sub-claim decomposition. The alternative method utilizes a streamlined TF-IDF reward model that incorporates custom numeric and temporal overlap features to evaluate traces, with scores aggregated by verdict category. Comparisons for Arabic verification are conducted using generative techniques.

Key facts

  • System developed for CLEF 2026 CheckThat! Task 2
  • Focuses on numerical claim verification in English and Arabic
  • First approach: fine-tunes LLM verifier with LoRA and Best-of-N selection
  • Experiments with adaptive sub-claim decomposition
  • Second approach: TF-IDF reward model with numeric and temporal overlap features
  • Aggregates scores by verdict group for final prediction
  • Arabic verification uses generative comparison
  • Described in arXiv paper 2607.25069

Entities

Institutions

  • CLEF
  • CheckThat!

Sources