Vision-Language Model Achieves 98.54% Accuracy in Dengue Mosquito Detection
A recent study published on arXiv (2608.12677) introduces a vision-language framework that integrates YOLO and Contrastive Language-Image Pre-training (CLIP) to classify mosquito flight frames, differentiating between uninfected mosquitoes and those infected with Dengue virus serotype 2 (DENV2). Initially, YOLO is utilized to extract mosquito regions from their surroundings. Subsequently, visual features from video frames are matched with relevant textual prompts in a unified embedding space. The multimodal model underwent fine-tuning via supervised bidirectional contrastive learning and was assessed through frame-level image-text similarity classification, achieving an impressive accuracy of 98.54%. This research tackles the difficulties of identifying infection-related behavioral changes in mosquitoes, which are often hindered by their small size and erratic movements, as well as environmental factors. Its implications for public health could enhance automated monitoring of disease vectors, facilitating early dengue outbreak detection and management.
Key facts
- The paper is published on arXiv with ID 2608.12677.
- The framework uses YOLO and CLIP.
- It classifies mosquito flight frames as uninfected or DENV2-infected.
- YOLO isolates mosquito regions from the background.
- Visual features are aligned with textual prompts in a shared embedding space.
- The model was fine-tuned using supervised bidirectional contrastive learning.
- Evaluation was done via frame-level image-text similarity-based classification.
- The method achieved 98.54% accuracy.
Entities
Institutions
- arXiv