Interior Interpretability: Contraction and Propagation in Transformers
A recent study presents the concept of interior interpretability, offering a propagation-based view on how internal models are structured, specifically applied to tabular Transformers through attention rollout. The researchers define rollout as a row-stochastic operator that captures attention-driven propagation among feature tokens. Utilizing classical Doeblin–Dobrushin contraction theory, they demonstrate that a rollout operator with a minimal Dobrushin coefficient closely resembles a rank-one stochastic matrix, where the shared row is defined by its normalized column sums. This finding provides a structural understanding of the associated rollout propagation profile. In Transformers designed for predicting metabolomic age, the observed rollout contraction increases with depth, suggesting enhanced internal organization.
Key facts
- Interior interpretability is a propagation-based perspective on internal model organization.
- It is instantiated for tabular Transformers using attention rollout.
- Rollout is interpreted as a row-stochastic operator encoding attention-mediated propagation.
- Doeblin–Dobrushin contraction theory is applied to rollout operators.
- A small Dobrushin coefficient implies closeness to a rank-one stochastic matrix.
- The common row of the rank-one matrix is determined by normalized column sums.
- In Transformers for metabolomic age prediction, rollout contraction strengthens with depth.
Entities
Institutions
- arXiv