Evidence-Ledger Adjudication Boosts Claim-Evidence Traceability in AI
A recent publication on arXiv presents a novel approach called evidence-ledger adjudication, which connects each AI-generated assertion with a corresponding evidence packet, establishes a support relationship, and sends unsupported, contradictory, or mixed-evidence claims back to the original author. The study's empirical foundation consists of a blind benchmark with 2,335 rows, derived from independent external labels in AVeriTeC, CLIMATE-FEVER, and SciFact. During prediction, gold relations and source evidence labels remain concealed and are only revealed for scoring purposes. In this benchmark, the agent evidence-ledger condition achieves relation accuracy of 0.676 and a macro-F1 score of 0.601, outperforming the best non-agent baseline, which has 0.383 accuracy and 0.303 macro-F1. Additionally, it successfully routes 1270 out of 1435 claims indicating contradiction, missing, or mixed evidence, while handling 295 out of 900 supported claims. These findings demonstrate that evidence-ledger adjudication enhances both traceability and accuracy in verifying claims against evidence.
Key facts
- Study introduces evidence-ledger adjudication for claim-evidence traceability
- Workflow pairs claims with evidence packets and assigns support relations
- Built a 2,335-row blind benchmark from AVeriTeC, CLIMATE-FEVER, and SciFact
- Agent condition achieves 0.676 relation accuracy and 0.601 macro-F1
- Non-agent baseline achieves 0.383 accuracy and 0.303 macro-F1
- Routes 1270/1435 unsupported or contradicted claims back to author
- Routes 295/900 supported claims
- Published on arXiv with ID 2607.26512
Entities
Institutions
- arXiv