LLM Conflict Task Reveals Distinct Processing Pathways in Gemma-2 and Pythia Models
A new study introduces a verbal-only conflict task for large language models (LLMs), revealing distinct processing pathways that handle congruent and incongruent conditions. The research, available on arXiv (2608.11510), tests Gemma-2-2B and six Pythia models (410M to 12B parameters) using a prompt stem that elicits a default same-color completion, with an explicit rule either agreeing (congruent) or conflicting (incongruent) with that completion. The results show strong default same-color tendencies across all models, and six of seven models exhibit robust congruency effects. Through causal attribution analysis, attention analysis, and attention ablations, the authors identified two distinct pathways: one involving short-range attention to a superficial color cue, preferentially activated in the congruent condition, and another involving long-range attention to the rule prefix, active in the incongruent condition. This work extends the study of congruency effects, traditionally explored in psychology and neuroscience via Stroop and flanker tasks, to the domain of LLMs, offering mechanistic insights into how these models process conflicting information. The findings contribute to a deeper understanding of the internal mechanisms of LLMs, potentially informing future model design and evaluation.
Key facts
- The study introduces a verbal-only LLM conflict task.
- The task uses a prompt stem that elicits a default same-color completion.
- An explicit rule either agrees (congruent) or conflicts (incongruent) with the completion.
- Gemma-2-2B and six Pythia models (410M to 12B parameters) were tested.
- All models showed strong default same-color tendencies.
- Six of seven models showed strong congruency effects.
- Causal attribution analysis, attention analysis, and attention ablations were used.
- Two distinct processing pathways were identified: short-range attention to color cue (congruent) and long-range attention to rule prefix (incongruent).
Entities
—