TRACES: New Benchmark Probes LLM Reliability in Scientific Reasoning
A new benchmark called TRACES has been introduced to evaluate how well large language models (LLMs) can assess scientific reasoning. Detailed in a paper on arXiv (2608.11415), TRACES seeks to tackle a major issue: LLMs often struggle to tell apart credible scientific research from unreliable materials, something existing benchmarks overlook. It includes a set of 42 documents that are either retracted, fraudulent, or pseudoscientific, along with a method to measure the model's interaction with these papers in one go. The benchmark focuses on five types of misleading claims and provides two scores to measure how effectively the model rejects these flawed ideas. This tool is vital for ensuring LLMs can accurately evaluate the reliability of scientific literature.
Key facts
- TRACES is a benchmark for epistemic reliability in scientific reasoning by LLMs.
- It uses a probe corpus of 42 retracted, fraudulent, and pseudoscientific papers.
- The benchmark measures whether models can distinguish reliable from unreliable scientific literature.
- Probes pair near-verbatim preambles from target papers with study-design requests.
- Five claim types are covered: fabricated observation, pseudophysical mechanism, magical premise, legitimization bridge, and cargo-cult experiment.
- Two complementary scores evaluate model rejection of flawed premises.
- The paper is available on arXiv with ID 2608.11415.
- The benchmark addresses a gap in existing evaluations that focus on factuality with known answers.
Entities
Institutions
- arXiv