ARTFEED — Contemporary Art Intelligence

Training Data Granularity Determines Parametric Modularity in LLMs

ai-technology · 2026-08-13

A recent preprint on arXiv (2608.10214) explores the presence of domain-specific parametric shells within large language models. These shells consist of neuron groups that, when removed, impair performance in specific domains without affecting others. The research employs a consistent causal approach across three model families (ranging from 1.5B to 7B parameters) and eight domains, analyzing two levels of domain granularity. At the academic subject level, no neurons surpass 60% domain selectivity among 939,008 FFN neurons, with flat causal damage matrices, despite achieving over 85% accuracy in domain identity. In contrast, at the language and modality level, 0.65–1.14% of neurons exceed 60% selectivity, with damage matrices showing near-perfect diagonal patterns. Masking code-selective neurons results in a 16–24 percentage point drop in mathematical reasoning accuracy across all models, while masking neurons related to Spanish or Chinese maintains performance at random levels. The results indicate that parametric modularity varies by granularity, revealing modular structures in coarse-grained domains but not in fine-grained academic subjects. This study offers insights into the existence of domain-specific neural circuits in LLMs, which may influence interpretability and targeted interventions. The research is conducted by a team of scholars and is available on the arXiv preprint server.

Key facts

  • Study examines parametric modularity in LLMs across two domain granularities.
  • Three model families (1.5B to 7B parameters) and eight domains are tested.
  • At academic subject level, no neurons exceed 60% selectivity across 939,008 FFN neurons.
  • Domain identity is linearly decodable above 85% accuracy.
  • At language/modality level, 0.65–1.14% of neurons exceed 60% selectivity.
  • Damage matrices are near-perfectly diagonal with ratios up to 595:1.
  • Shell neuron sets are essentially disjoint (IoU < 0.003).
  • Masking code-selective neurons reduces math reasoning by 16–24 percentage points.
  • Masking Spanish or Chinese neurons leaves math reasoning at or below random.
  • Paper is available on arXiv with ID 2608.10214.

Entities

Institutions

  • arXiv

Sources