Knowledge Distillation Has Asymmetric Effects on Bias in Small Language Models
A recent study published on arXiv (2607.28639) indicates that knowledge distillation in compact instruction-tuned language models has uneven impacts on bias. In unambiguous tasks (BBQ-disambig), response-based distillation from the Gemma-2-9B model enhances context-following, reducing the context-overriding error rate from 44% to 24% for the most biased baseline (SmolLM2-1.7B-Instruct). Conversely, in ambiguous tasks (BBQ-ambig), the same process undermines per-item refusal calibration, with 15% of items receiving stereotype responses instead of the correct abstentions. This trend is mirrored in another student model (OLMo-2-1B-Instruct), showing an 8% silence-loss and 89% of new bias attributed to filled-silence. Across 28 configurations, silence-loss and filled-silence show no correlation (Spearman ρ=0.19, n.s.), suggesting distinct mechanisms at play. These results question the notion that distillation merely transfers knowledge and underscore the importance of evaluating bias in distilled models.
Key facts
- Paper arXiv:2607.28639
- Knowledge distillation has asymmetric effects on bias
- On unambiguous tasks (BBQ-disambig), distillation improves context-following
- For SmolLM2-1.7B-Instruct, context-overriding error rate drops from 44% to 24%
- On ambiguous tasks (BBQ-ambig), distillation destroys refusal calibration
- 15% of items where baseline abstained now receive stereotype answers
- Pattern reproduces on OLMo-2-1B-Instruct with silence-loss of 8%
- Filled-silence accounts for 89% of new bias
- Silence-loss and filled-silence magnitudes are uncorrelated (Spearman ρ=0.19)
- Effects arise from distinct mechanisms
Entities
Institutions
- arXiv