Graded Large Language Models: An Algebraic Framework
A new paper on arXiv introduces Graded Large Language Models (GLLMs), an algebraic framework that adds a grading structure to transformer representation spaces. This grading propagates through embeddings, self-attention, and training objectives, extending graded neural networks and graded transformers to autoregressive language models. The framework preserves expressive power, asymptotic computational complexity, and inference cost. The geometric picture is based on geometric invariant theory, with benefits expressed via a Kempf–Ness functional on the grading torus. Grades that improve upon uniform architecture form an open convex cone, decided by a Hilbert–Mumford-type criterion. Optimal grades are given by a closed-form coincidence point of two moment maps.
Key facts
- Paper arXiv:2607.22757 introduces Graded Large Language Models (GLLMs).
- GLLMs equip transformer representation space with a grading.
- Grading propagates through embeddings, self-attention, and training objective.
- Extends graded neural networks and graded transformers to autoregressive models.
- Preserves expressive power, asymptotic complexity, and inference cost.
- Geometric picture based on geometric invariant theory.
- Benefit expressed by Kempf–Ness functional on grading torus.
- Optimal grades given by closed-form coincidence point of two moment maps.
Entities
Institutions
- arXiv