ARTFEED — Contemporary Art Intelligence

Provable Training Data Identification for Large Language Models

ai-technology · 2026-08-10

A novel technique known as Provable Training Data Identification (PTDI) has been introduced to reliably pinpoint the training data of extensive models with statistical assurances. This method redefines training data identification as a set-level inference challenge, allowing for rigorous control over false identification rates. PTDI calculates conformal p-values for each data instance by utilizing known unseen data and introduces a Jackknife-corrected Beta boundary (JKBB) estimator to gauge the proportion of training data within the test set. By employing the Benjamini-Hochberg (BH) procedure on the scaled p-values, it identifies a group of data points with regulated error rates. This research tackles the shortcomings of current instance-wise identification techniques that lack error rate oversight, essential for copyright disputes, privacy assessments, and equitable evaluations. The paper can be found on arXiv with the identifier 2510.09717.

Key facts

  • PTDI is a distribution-free approach for training data identification.
  • It provides provable and strict false identification rate control.
  • The method uses conformal p-values and a Jackknife-corrected Beta boundary estimator.
  • It applies the Benjamini-Hochberg procedure to scaled p-values.
  • The approach is relevant for copyright litigation, privacy auditing, and fair evaluation.
  • The paper is available on arXiv (2510.09717).
  • The method addresses limitations of instance-wise identification methods.
  • It formalizes training data identification as a set-level inference problem.

Entities

Institutions

  • arXiv

Sources