Self-Poisoning in Adaptive OOD Detection: A Sharp-Threshold Theory
A theoretical investigation into adaptive out-of-distribution (OOD) detectors has uncovered a verifiable dynamic principle that governs self-poisoning. The study conceptualizes memory bank impurity as a generalized Pólya urn, demonstrating almost-sure convergence towards a mean-field equilibrium. Stability is dictated by the slope of a reproduction number: when below one, impurity is harmless; when above one, the bank becomes entirely poisoned, leading to detector failure. The observed admission kernel is affine (R² ≥ 0.996), with a slope slightly under one across all encoder families, suggesting a near-critical design. In 96 scenarios, the anticipated threshold aligns with actual collapse, with ungated dictionaries experiencing a loss of up to 0.163 AUROC. A certified admission gate, which accesses only a frozen reserve, disrupts the feedback loop and mitigates transitions at any contamination level, including adversarial cases, while managing false positives.
Key facts
- Adaptive OOD detectors update a memory bank from unlabelled stream
- Adaptation obeys a provable dynamical law
- Bank impurity modeled as generalized Pólya urn
- Almost-sure convergence to mean-field equilibrium
- Slope acts as reproduction number
- Below one: impurity benign; above one: full poisoning
- Measured admission kernel affine with R² ≥ 0.996
- Slope just below one in every encoder family
- Predicted threshold matches empirical collapse across 96 settings
- Ungated dictionaries lose up to 0.163 AUROC
- Certified admission gate severs feedback loop
- Gate removes transition at any contamination rate, even adversarial
- Gate controls false positives
Entities
—