DSMM-TBIR: New Benchmark for Multi-Match Text-Based Image Retrieval in Surveillance
A recent study published on arXiv (2608.10524) presents a specialized data engine for multi-match text-based image retrieval (DSMM-TBIR) along with a benchmark named SecMM-TBIR. This initiative aims to tackle the shortcomings of current TBIR benchmarks, particularly in fields like surveillance. The authors, in their paper titled 'Rethinking Text-Based Image Retrieval in Specific Domain,' critique existing benchmarks for assuming a one-to-one match between queries and images, which does not align with practical situations where a single query may relate to multiple images. To address this, they developed SecMM-TBIR, featuring 50,000 surveillance images and 200 detailed queries. The study also notes that traditional contrastive learning in these contexts leads to significant false negatives, adversely affecting model performance. The DSMM-TBIR engine is proposed to create such benchmarks, emphasizing the importance of multi-match assessments in domain-specific TBIR. This research is pertinent to computer vision, information retrieval, and AI, especially in security and surveillance applications.
Key facts
- arXiv paper 2608.10524 introduces DSMM-TBIR data engine and SecMM-TBIR benchmark.
- SecMM-TBIR contains 50,000 surveillance images and 200 comprehensive queries.
- Existing TBIR benchmarks assume single-match between query and images.
- The paper addresses multi-match scenarios in specific domains like surveillance.
- Vanilla contrastive learning suffers from false negatives in specific domains.
- The benchmark is designed to evaluate TBIR systems in practical surveillance settings.
- The paper is titled 'Rethinking Text-Based Image Retrieval in Specific Domain'.
- The research is relevant to computer vision and AI applications in security.
Entities
Institutions
- arXiv