ARTFEED — Contemporary Art Intelligence

Depth-Guided Video Object Counting in Crowded Scenes

ai-technology · 2026-08-07

A novel technique for counting items in densely populated video scenes utilizing depth data has been unveiled. The Depth-Guided Detector (DG-Det) combines depth signals with multi-scale RGB-D cross-attention and clear occlusion forecasting to enhance spatial comprehension and reliable detection. This method tackles the shortcomings of current RGB-only approaches in congested and obstructed environments. Additionally, a comprehensive de-duplication system removes redundant counts across frames. The team has also made available a new RGB-D Video Object Counting dataset featuring depth data and various object categories for each sequence. Tests indicate a 62.01% decrease in counting inaccuracies.

Key facts

  • Method: Depth-Guided Detector (DG-Det)
  • Integrates depth cues with multi-scale RGB-D cross-attention
  • Includes explicit occlusion prediction
  • Unified de-duplication framework for cross-frame redundancy
  • New RGB-D Video Object Counting dataset released
  • Dataset features depth information and multiple object categories per sequence
  • Achieves 62.01% reduction in counting errors
  • Paper available on arXiv:2608.06236

Entities

Institutions

  • arXiv

Sources