Single-Block Spatio-Temporal Transformer for Multi-Entity Reasoning
A new arXiv paper (2607.23077) proposes a structured spatio-temporal transformer block that explicitly models spatial, temporal, and cross interactions among entities in a single stage, reducing the need for deep stacking. The approach uses parallel spatial and temporal self-attention, bidirectional cross-attention, and learnable gated fusion. Evaluated on video-based group activity recognition, it achieves competitive performance with lower computational cost. The paper revisits multi-entity temporal modeling from a structural perspective, decomposing dynamics into three interaction types.
Key facts
- arXiv paper ID: 2607.23077
- Proposes a single-block spatio-temporal transformer
- Models spatial, temporal, and cross interactions explicitly
- Uses parallel spatial and temporal self-attention
- Includes bidirectional cross-attention and learnable gated fusion
- Reduces need for deep stacking of layers
- Evaluated on video-based group activity recognition
- Aims to lower computational cost while maintaining performance
Entities
Institutions
- arXiv