collaborators

32 papers

cs.CV2026

The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering

Yuqian Fu, Tianwen Qian, Yanjun Li +30

EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scena…

cs.CV2026

SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video

Weili Guan, Haoyu Zhang, Meng Liu +3

Visual-spatial understanding, defined as the ability to infer object relationships and scene layouts from visual inputs, is fundamental to downstream tasks such as robotic navigati…

cs.MM2026

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning

Yibo Lyu, Rui Shao, Gongwei Chen +3

As multimedia content expands, the demand for unified multimodal retrieval (UMR) in real-world applications increases. Recent work leverages multimodal large language models (MLLMs…

cs.CV2026

R^3: Composed Video Retrieval via Reasoning-Guided Recalling and Re-ranking

Zixu Li, Yupeng Hu, Zhiheng Fu +3

The CoVR-R challenge evaluates composed video retrieval, where a system must retrieve a target video from a large gallery given a reference video and a textual edit instruction. Th…

cs.CV2026

EgoAdapt: A Multi-Scene Egocentric Adaptation Method for CVPR 2026 HD-EPIC VQA Challenge

Zhiwei Chen, Yupeng Hu, Zixu Li +4

This technical report presents our solution, EgoAdapt (Egocentric Adaptation via Category, Calibration, and Consistency), to the CVPR 2026 HD-EPIC VQA challenge. HD-EPIC evaluates…

cs.CV2026

EgoAction: Egocentric Action Composition with Reliability-Aware Temporal Fusion for the EPIC-KITCHENS Action Detection Challenge at CVPR 2026

Zhiheng Fu, Zixu Li, Zhiwei Chen +4

The EPIC-KITCHENS-100 Action Detection challenge evaluates whether a model can localize the start and end of each action in long untrimmed egocentric videos and assign the correspo…