ARTFEED — Contemporary Art Intelligence

New Framework Audits Data Provenance in LLM Fine-tuning

ai-technology · 2026-08-04

A new framework called Distribution Provenance Audit (DPA) has been proposed to audit data intellectual property (IP) infringement in fine-tuning large language models (LLMs). The proliferation of customized LLMs poses risks of unauthorized fine-tuning on proprietary data. Existing audit techniques require intervention during data preparation or training and are fragile under malicious obfuscations like data paraphrasing and knowledge distillation. DPA is a post-hoc framework that works under black-box and malicious settings. It is based on the insight that regardless of fine-tuning tactics, maintaining utility forces the model to preserve the intersection of semantic substance and lexical form. DPA captures this as intrinsic distributional fingerprints. The framework formulates the audit as a statistical testing problem, leveraging the intrinsic distributional fingerprints to detect unauthorized use. The paper is available on arXiv with identifier 2608.02154.

Key facts

  • The framework is called Distribution Provenance Audit (DPA).
  • It addresses data IP infringement in LLM fine-tuning.
  • Existing audit techniques require intervention during data preparation or training.
  • DPA works post-hoc and under black-box settings.
  • It is robust against malicious obfuscations like paraphrasing and knowledge distillation.
  • The framework uses intrinsic distributional fingerprints.
  • The paper is on arXiv with ID 2608.02154.
  • The announcement type is 'new'.

Entities

Institutions

  • arXiv

Sources