Activation Steering Enhances Role-Conditioned Agents in Social Simulations
A new arXiv preprint (2608.00023) introduces an activation-steering screening workflow for role-conditioned language model agents in social simulations. The method involves defining a role profile, extracting a role-specific direction, sweeping four steering coefficients, evaluating role-profile alignment, and passing or flagging candidate configurations. Applied to OLMo-3-7B-Instruct with a mixed 275-role inventory and 228 role-agnostic questions, the approach uses GPT-4.1-mini for prompted role references and judging. Role-specific directions achieved higher judged role-profile alignment (mean overall score 63.2) compared to an assistant-axis directional control (41.1) across the tested grid, while preserving high lexical diversity. The role-level screen is the main practical output, enabling pre-screening of agents before population deployment.
Key facts
- The preprint is arXiv:2608.00023, announced as a cross-type.
- The workflow includes defining a role profile, extracting a role-specific direction, sweeping four steering coefficients, evaluating role-profile alignment, and passing or flagging configurations.
- The model used is OLMo-3-7B-Instruct.
- The inventory includes 275 roles and 228 role-agnostic questions.
- GPT-4.1-mini is used for prompted role references and as judges.
- Role-specific directions scored 63.2 vs 41.1 for the control in mean overall alignment.
- Role-specific directions preserve high lexical diversity, while the control drops at larger coefficients.
- The role-level screen is the main practical output.
Entities
Institutions
- arXiv