ARTFEED — Contemporary Art Intelligence

DSMM-TBIR: New Benchmark for Multi-Match Text-Based Image Retrieval in Surveillance

ai-technology · 2026-08-13

A recent study published on arXiv (2608.10524) presents a specialized data engine for multi-match text-based image retrieval (DSMM-TBIR) along with a benchmark named SecMM-TBIR. This initiative aims to tackle the shortcomings of current TBIR benchmarks, particularly in fields like surveillance. The authors, in their paper titled 'Rethinking Text-Based Image Retrieval in Specific Domain,' critique existing benchmarks for assuming a one-to-one match between queries and images, which does not align with practical situations where a single query may relate to multiple images. To address this, they developed SecMM-TBIR, featuring 50,000 surveillance images and 200 detailed queries. The study also notes that traditional contrastive learning in these contexts leads to significant false negatives, adversely affecting model performance. The DSMM-TBIR engine is proposed to create such benchmarks, emphasizing the importance of multi-match assessments in domain-specific TBIR. This research is pertinent to computer vision, information retrieval, and AI, especially in security and surveillance applications.

Key facts

  • arXiv paper 2608.10524 introduces DSMM-TBIR data engine and SecMM-TBIR benchmark.
  • SecMM-TBIR contains 50,000 surveillance images and 200 comprehensive queries.
  • Existing TBIR benchmarks assume single-match between query and images.
  • The paper addresses multi-match scenarios in specific domains like surveillance.
  • Vanilla contrastive learning suffers from false negatives in specific domains.
  • The benchmark is designed to evaluate TBIR systems in practical surveillance settings.
  • The paper is titled 'Rethinking Text-Based Image Retrieval in Specific Domain'.
  • The research is relevant to computer vision and AI applications in security.

Entities

Institutions

  • arXiv

Sources