Layer-wise Positional Bias in Short-Context Language Modeling
A recent study published on arXiv (2601.04098) presents a framework for layer conductance aimed at assessing positional bias in transformer language models. The findings indicate that recency bias consistently rises across the layers, whereas primacy bias is less pronounced and decreases. Utilizing a sliding-window approach for short-context next-word prediction allowed the researchers to separate model-internal behavior from task and context-window influences. The layer-wise positional importance profiles obtained remain consistent across various texts and lexical arrangements, validating their reflection of the model's internal structure. This research elucidates the evolution of these profiles with depth, contributing to a better understanding of how input positions influence predictions at each layer, which may enhance language model performance and interpretability.
Key facts
- Paper: arXiv:2601.04098
- Introduces a layer conductance framework
- Applied to short-context next-word prediction
- Uses a sliding-window design
- Finds recency bias increases monotonically across layers
- Primacy bias is subtle and diminishes
- Profiles are stable across diverse texts and lexical scrambling
- Confirms profiles reflect model-internal structure
Entities
Institutions
- arXiv