EgoCross Challenge at CVPR 2026: Cross-Domain Egocentric Video QA
The inaugural EgoCross Challenge took place during the Third EgoVis Workshop at CVPR 2026, unveiling a benchmark for egocentric video question answering across different domains. This challenge assesses the ability of multimodal large language models to extend their generalization beyond typical everyday situations by utilizing first-person videos from four specific areas: surgery, industrial assembly, extreme sports, and animal viewpoints. Each evaluation consists of an egocentric video, a corresponding question, and four possible answers for the model to choose from. The technical report outlines the challenge's tasks, benchmark resources, and two official Codabench tracks: the Source-Limited Track, which confines participants to the official baseline model and a limited support set, and the Open-Source Track, allowing more flexibility in model and training data selection while banning manual target-domain training data construction. The challenge seeks to enhance AI comprehension in specialized visual contexts, impacting sectors such as medical training, industrial automation, and wildlife research.
Key facts
- EgoCross is a cross-domain egocentric video question answering benchmark.
- The first EgoCross Challenge was hosted at the Third EgoVis Workshop at CVPR 2026.
- The benchmark evaluates multimodal large language models on first-person videos from four domains: surgery, industrial assembly, extreme sports, and animal perspectives.
- Each test example includes a video clip, a question, and four candidate answers.
- The challenge has two official Codabench tracks: Source-Limited and Open-Source.
- The Source-Limited Track restricts participants to the official baseline model and a small support set.
- The Open-Source Track allows broader model and data choices but prohibits manual construction of target-domain training data.
- The technical report introduces the challenge task and benchmark resources.
Entities
Institutions
- EgoVis Workshop
- CVPR
- Codabench