Stemma: A Black-Box Method for LLM Provenance Testing
A recent publication on arXiv presents Stemma, a novel black-box technique for assessing the provenance of large language models (LLMs). This method overcomes the shortcomings of current techniques that depend on response-level traits, which may change during adaptation or deployment. By mapping open-ended outputs into a finite decision space, Stemma creates induced decision regions, minimizing surface-form variations. This approach redefines provenance testing to focus on the inheritance of decision regions. Empirical findings indicate that related models maintain source-induced regions more effectively than unrelated ones. Stemma also establishes stability, robustness, and specificity as complementary principles for selecting probes, ensuring reliable fingerprinting.
Key facts
- Paper published on arXiv with ID 2607.25880.
- LLM provenance testing determines if a suspect LLM belongs to the same lineage as a source.
- Existing black-box methods infer provenance from response-level characteristics.
- Response-level characteristics may shift under adaptation or deployment.
- Stemma introduces induced decision regions by mapping open-ended outputs into a finite decision space.
- Provenance testing is reframed as measuring inheritance of decision regions.
- Empirical analysis shows source-induced regions are preserved more strongly in related models.
- Stemma uses stability, robustness, and specificity as probe-selection principles.
Entities
Institutions
- arXiv