Closed-Form Bound on Cross-Layer Interaction in Weight-Space Ablation
A recent publication on arXiv (2608.03629) expands upon a theoretical finding regarding the agreement of activation patching and weight-space ablation in transformer architectures. The initial companion study examined a simplified model where conditional computation is added through a residual stream, deriving a precise first-order interaction formula for a specific case involving two architecturally dependent carriers: an attention head and its corresponding layer's normalization-MLP composition. This finding was limited to a single residual block and was only evaluated on small transformers using a synthetic task. The new paper broadens these findings by demonstrating that the interaction from ablating carriers across multiple layers can be decomposed into same-block terms for each affected layer, along with a cross-layer remainder. Additionally, it precisely isolates this remainder for two layers as a double integral of a specific expression. Authored by researchers, the paper is theoretical and does not include empirical tests on actual pretrained models, despite its title suggesting such a test; the abstract does not mention any empirical evaluation.
Key facts
- Paper extends theoretical result on activation patching vs. weight-space ablation.
- Original result confined to single residual block and small transformers on synthetic task.
- New result decomposes interaction from ablating carriers spanning several layers into same-block terms plus cross-layer remainder.
- Cross-layer remainder isolated exactly for two layers as a double integral.
- Paper posted on arXiv with identifier 2608.03629.
- No empirical test on real pretrained model described in abstract.
- Theoretical work, no specific institutions or locations mentioned.
- Authors not named in the provided content.
Entities
Institutions
- arXiv