ARTFEED — Contemporary Art Intelligence

New Framework Detects Data Contamination in Tabular LLM Evaluations

ai-technology · 2026-08-06

A new framework for assessing data contamination in tabular datasets used to evaluate large language models (LLMs) has been proposed in a preprint on arXiv (arXiv:2510.20351). The authors argue that existing contamination detection methods, which rely on memorization tests, are too coarse to identify performance gains from prior exposure to test data. Their approach involves generating controlled multiple-choice queries that preserve task structure while systematically transforming the underlying data. These transformations selectively disrupt dataset information while retaining partial knowledge, allowing the isolation of performance attributable to contamination. The framework also incorporates non-neural baselines to provide reference performance. The paper is an update (replace-cross) and addresses a largely unexplored area, as contamination in tabular data has received little attention compared to text and image domains. The work is significant for ensuring the reliability of LLM benchmarks, particularly in structured data applications.

Key facts

  • The framework is proposed in arXiv:2510.20351.
  • It targets data contamination in tabular datasets.
  • Existing methods rely on memorization tests, which are too coarse.
  • The framework generates controlled multiple-choice queries.
  • Transformations disrupt dataset information while preserving partial knowledge.
  • Non-neural baselines provide reference performance.
  • The paper is an update (replace-cross).
  • The problem is largely unexplored in tabular data.

Entities

Institutions

  • arXiv

Sources