ARTFEED — Contemporary Art Intelligence

Single-Block Spatio-Temporal Transformer for Multi-Entity Reasoning

other · 2026-07-29

A new arXiv paper (2607.23077) proposes a structured spatio-temporal transformer block that explicitly models spatial, temporal, and cross interactions among entities in a single stage, reducing the need for deep stacking. The approach uses parallel spatial and temporal self-attention, bidirectional cross-attention, and learnable gated fusion. Evaluated on video-based group activity recognition, it achieves competitive performance with lower computational cost. The paper revisits multi-entity temporal modeling from a structural perspective, decomposing dynamics into three interaction types.

Key facts

  • arXiv paper ID: 2607.23077
  • Proposes a single-block spatio-temporal transformer
  • Models spatial, temporal, and cross interactions explicitly
  • Uses parallel spatial and temporal self-attention
  • Includes bidirectional cross-attention and learnable gated fusion
  • Reduces need for deep stacking of layers
  • Evaluated on video-based group activity recognition
  • Aims to lower computational cost while maintaining performance

Entities

Institutions

  • arXiv

Sources