Raven: A New Linear-Time Sequence Model with Sparse Memory Routing
Researchers introduce Raven, a linear-time sequence model that improves long-context recall by maintaining a fixed set of memory slots and updating only a selected subset via learned, input-dependent routing. This approach mitigates interference from dense state updates in state-space models (SSMs) and linear Transformers, while avoiding the hard eviction of sliding-window attention (SWA). Raven achieves high recall by preserving long-range content without the computational cost of full attention. The model is detailed in a preprint on arXiv (2607.25357).
Key facts
- Raven is a linear-time sequence model.
- It uses a fixed set of memory slots.
- Updates only a selected subset via learned, input-dependent routing.
- Mitigates interference from dense state updates in SSMs.
- Avoids hard eviction of sliding-window attention.
- Preserves long-range content.
- Described in arXiv preprint 2607.25357.
- Announce type: cross.
Entities
Institutions
- arXiv