Causal Steering of Language Identity in LLMs: A New Study
There's a new preprint on arXiv, numbered 2608.12334, that dives into how language identity works in Large Language Models (LLMs). It examines if we can decode this identity from hidden states or if it can be influenced through a specific activation direction. The researchers analyzed various models, like Qwen 3.5-2B and Llama-3.2-1B-Instruct, using causal interventions. They ran experiments on over 1.26 million generations with the FLORES-200 dataset, finding that guiding these models in certain directions effectively switches languages, whether it's from English to Chinese or English to Spanish. Interestingly, random changes didn’t really affect the outcome. Their analysis revealed that language commitment is quite specific, suggesting we can control language identity in LLMs, which could improve multilingual capabilities.
Key facts
- Preprint arXiv:2608.12334 investigates causal control of language identity in LLMs.
- Models studied: Qwen 3.5-2B and Llama-3.2-1B-Instruct.
- Experiments conducted on FLORES-200 dataset with 1.26 million generations.
- PCA-derived 'language axes' used for steering and ablation.
- Steering reliably forces language switching in cross-script (English to Chinese) and same-script (English to Spanish) settings.
- Random perturbations of equal magnitude have virtually no effect.
- Layerwise analysis shows language commitment is localized and language-pair-dependent.
- Study asks whether language identity is linearly decodable or causally controllable.
Entities
Institutions
- arXiv