Search papers, labs, and topics across Lattice.
The EgoCross Challenge at EgoVis 2026 established a novel benchmark for evaluating the generalization capabilities of multimodal large language models in cross-domain egocentric video question answering. Participants tackled first-person video clips across diverse domains such as surgery and extreme sports, with over 1,500 submissions highlighting the competitive landscape. The challenge's structured tracks鈥擲ource-Limited and Open-Source鈥攑rovided a framework for assessing model performance while ensuring a fair evaluation environment, culminating in a comprehensive leaderboard and insights into winning strategies.
Over 1,500 submissions revealed stark differences in model performance across diverse domains, highlighting the challenges of generalizing egocentric video understanding.
EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scenarios. The first EgoCross Challenge was hosted at the Third EgoVis Workshop at CVPR 2026 and evaluated models on first-person videos from four target domains: surgery, industrial assembly, extreme sports, and animal perspectives. Each test example consists of an egocentric video clip, a question, and four candidate answers, from which the model must select the correct option. This technical report introduces the challenge task, benchmark resources, and two official Codabench tracks. The Source-Limited Track restricts participants to the official baseline model and a small support set, whereas the Open-Source Track permits broader choices of models and training data under rules that prohibit the manual construction of target-domain training data. In total, the challenge received more than 1,500 submissions from over 130 participants, with 19 teams participating in the Open-Source Track and 38 teams in the Source-Limited Track. We further present the official leaderboard results and summarize the winning solutions from both tracks. We hope that this report will serve as a useful technical reference for advancing cross-domain egocentric video understanding. All resources, including the challenge data, baseline implementation, and code released by the winning teams, are made publicly available.