Effective Alignment Dimension Predicts Width Scaling in Residual Networks
A new theoretical study has introduced the effective alignment dimension, which quantifies how activation gradients behave in normalized residual networks. The researchers determined the mean and variance of the inner product between training and test gradients, establishing an upper limit on the chance of misalignment that depends on the effective alignment dimension and sample size. This limit requires only finite second moments and a nonzero population gradient, without needing assumptions about covariance or growth rates. Their findings fit into a framework for expanding residuals in train-test scenarios, presenting a high-probability condition for beneficial directions in unseen data. This research addresses gaps in width-scaling theories that typically ignore insights from finite training data.
Key facts
- Introduces effective alignment dimension for residual networks
- Derives exact mean and variance of gradient inner products
- Provides finite-sample upper bound on misalignment probability
- Bound depends only on effective alignment dimension and effective sample size
- Requires finite second moments and nonzero population gradient
- No covariance spectral assumptions or prescribed width-growth rates
- Integrates certificate into train-test residual-expansion framework
- Addresses generalization of expansion directions from finite data
Entities
—