6 papers · 1 filter
The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering
Yuqian Fu, Tianwen Qian, Yanjun Li +30
EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scena…
The Right Inference Strategy Is All You Need: Nearly Training-Free Domain-Wise Inference for EgoCross Challenge
Leyi Wu, Yifan Zhao, Jinjie Zhang +2
EgoCross evaluates multimodal large language models on egocentric video question answering under substantial domain shift, where test videos come from surgery, industrial assembly,…
RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes
Leyi Wu, Yifan Zhao, Jinjie Zhang +11
Vision-Language Models (VLMs) have shown strong visual understanding and are increasingly deployed in embodied AI systems, where reliable perception under real conditions is essent…
Find, Fix, Reason: Context Repair for Video Reasoning
Haojian Huang, Chuanyu Qin, Yinchuan Li +1
Reinforcement learning has advanced video reasoning in large multi-modal models, yet dominant pipelines either rely on on-policy self-exploration, which plateaus at the model's kno…
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators
Jiazhou Zhou, Yucheng Chen, Hongyang Li +4
Multimodal Large Language Models (MLLMs) have achieved remarkable success, yet they remain prone to perception-related hallucinations in fine-grained tasks. This vulnerability aris…
LongLive: Real-time Interactive Long Video Generation
Shuai Yang, Wei Huang, Ruihang Chu +9
We present LongLive, a frame-level autoregressive (AR) framework for real-time and interactive long video generation. Long video generation presents challenges in both efficiency a…