ARTFEED — Contemporary Art Intelligence

HSTGFormer: A Graph-Enhanced Transformer for 3D Human Pose Estimation

other · 2026-08-13

A novel approach for estimating 3D human poses, called HSTGFormer, has been revealed in a paper on arXiv (ID: 2608.12187). This method tackles the shortcomings of current transformer-based models that separate spatial and temporal reasoning, which can diminish the interconnectedness of human motion and reduce frame-level structural data before temporal analysis. HSTGFormer redefines spatial-temporal reasoning through localized coupled graph aggregation at joint-time nodes. It features a Hyper Spatial-Temporal Graph (HSTG) that breaks down global reasoning into localized receptive fields around joint-time nodes by expanding per-frame skeleton graphs into temporal neighborhoods. This allows for structure-aware coupled reasoning while maintaining local motion details. An adaptive dual-branch mechanism is also included. The full paper can be accessed at https://arxiv.org/abs/2608.12187.

Key facts

  • HSTGFormer is a graph-enhanced Transformer framework for 3D human pose estimation.
  • It addresses limitations of existing transformer methods that separate spatial and temporal reasoning.
  • The method introduces a Hyper Spatial-Temporal Graph (HSTG) for localized coupled graph aggregation.
  • HSTG decomposes global spatial-temporal reasoning into local receptive fields around joint-time nodes.
  • It extends per-frame skeleton graphs into temporal neighborhoods to preserve structural motion information.
  • The paper is available on arXiv with ID 2608.12187.
  • The approach aims to improve unified spatial-temporal interdependencies in human motion.
  • The paper is a cross-announcement, indicating it was previously announced in another venue.

Entities

Institutions

  • arXiv

Sources