Qwen3-32B Temporal Preference Steering via Contrastive Activation Addition
A novel technique has been created by researchers to influence the temporal preferences of large language models, particularly Qwen3-32B, through contrastive activation addition. The findings, published as arXiv preprint 2608.03892, reveal linear representations of temporal horizons within the model's residual stream. By training contrastive linear probes on teacher-forced temporal-choice responses, a distinction between short-term and long-term preferences was discovered. When applying contrastive activation-addition steering to binary temporal-choice tasks, out-of-distribution monetary intertemporal choices, and the TravelPlanner benchmark, notable bidirectional preference shifts were achieved. The key takeaway is that simple linear probes can identify temporal-horizon directions, enabling adjustments in the model's indifference threshold regarding smaller-sooner versus larger-later rewards, which has significant implications for AI alignment and time-related decision-making.
Key facts
- The study focuses on Qwen3-32B, a large language model.
- Contrastive linear probes are trained on teacher-forced temporal-choice answers.
- A short-term versus long-term direction is identified in the residual stream.
- Contrastive activation-addition steering is evaluated on multiple tasks.
- The method induces large, bidirectional preference changes.
- Out-of-distribution monetary choice tasks show shifts in indifference thresholds.
- The paper is available on arXiv with ID 2608.03892.
- The research is relevant to AI alignment and temporal decision-making.
Entities
Institutions
- arXiv