Activation Source Selection Significantly Affects Steering in Language Models
A new study on arXiv (2607.25270) investigates activation source selection in steering language models. Activation steering modifies model behavior by adding vectors to hidden states at inference, but the choice of source activations is often overlooked. The research shows across three instruction-tuned models and four task families that changing the source context and readout policy substantially alters steering success. Effective steering signals come from execution-boundary states where the model is about to produce target behavior, not merely from desired behavior appearing in source text. This pre-/post-realization distinction explains why answer-based sources sometimes work.
Key facts
- arXiv paper 2607.25270 studies activation source selection in steering language models.
- Activation steering adds vectors or features to hidden states at inference time.
- Source choice involves source context and activation readout policy.
- Three instruction-tuned models and four steering task families were tested.
- Changing source activations substantially changes steering success.
- Effective steering signals come from execution-boundary states.
- Execution-boundary states are where the model is about to produce or continue target behavior.
- Pre-/post-realization distinction explains why answer-based sources sometimes work.
Entities
Institutions
- arXiv