ARTFEED — Contemporary Art Intelligence

Gradient-Based Explanations of LLMs Align with Brain Activity

ai-technology · 2026-08-07

A recent study published on arXiv (2502.14671) explores the connection between Large Language Model (LLM) representations and brain activity during language processing. Researchers aimed to determine if explainable AI (XAI) could shed light on this connection by employing attribution methods to assess the influence of each input word on the LLM's next-word predictions. They subsequently utilized these insights to forecast fMRI data from participants engaged in narrative listening. The results indicated that gradient-based attribution methods align closely with brain activity, providing unique variance beyond acoustic and word-rate factors, and surpassing internal representations in early auditory areas. By applying conductance, the study revealed that early layers exhibit heightened sensitivity to word types and preferential alignment with auditory regions, while later layers correspond with higher-order language areas. These findings imply that gradient-based explanations illuminate aspects of neural language processing overlooked by internal representations, presenting a novel approach to understanding the brain-language connection.

Key facts

  • Study from arXiv:2502.14671
  • Uses explainable AI (XAI) to test LLM-brain alignment
  • Attribution methods quantify word contributions to next-word predictions
  • Predicts fMRI data from participants listening to narratives
  • Gradient-based attribution methods align robustly with brain activity
  • Contribute unique variance beyond acoustic and word-rate confounds
  • Outperform internal representations in early auditory regions
  • Conductance extends attribution from words to individual layers
  • Early layers show greater word-type sensitivity and align with auditory regions
  • Later layers align with higher-order language regions

Entities

Institutions

  • arXiv

Sources