ARTFEED — Contemporary Art Intelligence

Hyperball Optimizers' Advantage Remains Unclear

other · 2026-07-27

A recent study published on arXiv questions the presumed benefits of Hyperball-style optimizers in scale-invariant deep networks. The researchers introduce an angular effective learning rate that incorporates the angle of parameter updates, as well as the norms of parameters and updates. They demonstrate that the traditional norm-based approach is merely a specific instance in cases of orthogonality. By breaking down updates into radial and tangential parts, they discover that radial updates have minimal direct impact on the angular effective learning rate, which fails to clarify why MuonH experiences slower convergence compared to MuonWD during the initial training phases.

Key facts

  • Paper title: Hyperball May Not Be a Free Lunch
  • arXiv ID: 2607.22444
  • Announce type: cross
  • Focuses on scale-invariant deep networks
  • Derives angular effective learning rate
  • Shows norm-based measure is a special case under orthogonality
  • Decomposes optimizer updates into radial and tangential components
  • Radial component has limited direct effect on angular effective learning rate

Entities

Institutions

  • arXiv

Sources