Depth-Guided Video Object Counting in Crowded Scenes
A novel technique for counting items in densely populated video scenes utilizing depth data has been unveiled. The Depth-Guided Detector (DG-Det) combines depth signals with multi-scale RGB-D cross-attention and clear occlusion forecasting to enhance spatial comprehension and reliable detection. This method tackles the shortcomings of current RGB-only approaches in congested and obstructed environments. Additionally, a comprehensive de-duplication system removes redundant counts across frames. The team has also made available a new RGB-D Video Object Counting dataset featuring depth data and various object categories for each sequence. Tests indicate a 62.01% decrease in counting inaccuracies.
Key facts
- Method: Depth-Guided Detector (DG-Det)
- Integrates depth cues with multi-scale RGB-D cross-attention
- Includes explicit occlusion prediction
- Unified de-duplication framework for cross-frame redundancy
- New RGB-D Video Object Counting dataset released
- Dataset features depth information and multiple object categories per sequence
- Achieves 62.01% reduction in counting errors
- Paper available on arXiv:2608.06236
Entities
Institutions
- arXiv