33 papers
The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering
Yuqian Fu, Tianwen Qian, Yanjun Li +30
EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scena…
ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding
Mingkang Dong, Muxin Pu, Jie Li +8
ObjectStream introduces a training‑free method that extracts latent objects from frozen Video‑LLM representations and uses them as persistent memory anchors to improve streaming vi…
DVPSFormer: Efficient Online Depth-aware Video Panoptic Segmentation for Autonomous Driving
Yung-Hsu Yang, Luigi Piccinelli, Siyuan Li +8
DVPSFormer is an online architecture that jointly estimates metric depth, semantic segmentation, and instance trajectories for autonomous driving by using explicit scene discretiza…
InterEdit: Navigating Text-Guided 3D Dyadic Human Motion Editing
Yebin Yang, Di Wen, Lei Qi +10
Text-guided 3D motion editing has seen success in single-person scenarios, but its extension to multi-person settings is less explored due to limited paired data and the complexity…
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning
Kailing Li, Qi'ao Xu, Tianwen Qian +3
Embodied Visual Reasoning (EVR) seeks to follow complex, free-form instructions based on egocentric video, enabling semantic understanding and spatiotemporal reasoning in dynamic e…
AVA: Attentive VLM Agent for Mastering StarCraft II
Weiyu Ma, Yuqian Fu, Zecheng Zhang +2
We introduce AVACraft, a multimodal StarCraft II benchmark supporting both Multi-Agent Reinforcement Learning (MARL) and Vision-Language Model (VLM) paradigms. Unlike SMAC-family e…