New Signal Detects Spurious Correlations in Trained Models
A new arXiv paper (2608.05419) introduces a method to identify samples affected by spurious correlations in machine learning models without requiring group annotations. The authors demonstrate that after model convergence, when loss no longer distinguishes between populations, a perturbation-based signal can effectively separate samples consistent with the spurious correlation from those that are not. The method involves applying a fixed perturbation to the inputs of a converged model and observing that predictions for samples fitting the spurious correlation are more stable, while predictions for other samples flip more frequently. This approach requires only two forward passes per training sample, making it computationally efficient. The paper addresses a critical limitation in existing methods that rely on early training signals and require hyperparameter tuning with group-labeled validation data. The findings have implications for improving model robustness and fairness in real-world applications where spurious correlations are prevalent.
Key facts
- Paper arXiv:2608.05419
- Method identifies spuriously correlated samples without group annotations
- Signal available after convergence
- Requires two forward passes per training sample
- Perturbation flips predictions of non-spurious samples more often
- Addresses limitations of early training signals
- No hyperparameter tuning needed
- Improves model robustness and fairness
Entities
Institutions
- arXiv