ARTFEED — Contemporary Art Intelligence

Knowledge Distillation Has Asymmetric Effects on Bias in Small Language Models

ai-technology · 2026-08-03

A recent study published on arXiv (2607.28639) indicates that knowledge distillation in compact instruction-tuned language models has uneven impacts on bias. In unambiguous tasks (BBQ-disambig), response-based distillation from the Gemma-2-9B model enhances context-following, reducing the context-overriding error rate from 44% to 24% for the most biased baseline (SmolLM2-1.7B-Instruct). Conversely, in ambiguous tasks (BBQ-ambig), the same process undermines per-item refusal calibration, with 15% of items receiving stereotype responses instead of the correct abstentions. This trend is mirrored in another student model (OLMo-2-1B-Instruct), showing an 8% silence-loss and 89% of new bias attributed to filled-silence. Across 28 configurations, silence-loss and filled-silence show no correlation (Spearman ρ=0.19, n.s.), suggesting distinct mechanisms at play. These results question the notion that distillation merely transfers knowledge and underscore the importance of evaluating bias in distilled models.

Key facts

  • Paper arXiv:2607.28639
  • Knowledge distillation has asymmetric effects on bias
  • On unambiguous tasks (BBQ-disambig), distillation improves context-following
  • For SmolLM2-1.7B-Instruct, context-overriding error rate drops from 44% to 24%
  • On ambiguous tasks (BBQ-ambig), distillation destroys refusal calibration
  • 15% of items where baseline abstained now receive stereotype answers
  • Pattern reproduces on OLMo-2-1B-Instruct with silence-loss of 8%
  • Filled-silence accounts for 89% of new bias
  • Silence-loss and filled-silence magnitudes are uncorrelated (Spearman ρ=0.19)
  • Effects arise from distinct mechanisms

Entities

Institutions

  • arXiv

Sources