9 papers
The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering
Yuqian Fu, Tianwen Qian, Yanjun Li +30
EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scena…
AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding
Qize Yu, Jiadi You, Yuran Wang +10
Vision-Language-Action (VLA) models leverage the rich world knowledge of pretrained vision-language models (VLMs) to enable instruction-following robotic manipulation. However, the…
Policy and World Modeling Co-Training for Language Agents
Ning Lu, Baijiong Lin, Shengcai Liu +9
Reinforcement learning (RL) improves large language model (LLM) agents by teaching them which actions lead to high rewards, but provides little supervision on what those actions do…
The Right Inference Strategy Is All You Need: Nearly Training-Free Domain-Wise Inference for EgoCross Challenge
Leyi Wu, Yifan Zhao, Jinjie Zhang +2
EgoCross evaluates multimodal large language models on egocentric video question answering under substantial domain shift, where test videos come from surgery, industrial assembly,…
RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes
Leyi Wu, Yifan Zhao, Jinjie Zhang +11
Vision-Language Models (VLMs) have shown strong visual understanding and are increasingly deployed in embodied AI systems, where reliable perception under real conditions is essent…
Affordance Agent Harness: Verification-Gated Skill Orchestration
Haojian Huang, Jiahao Shi, Yinchuan Li +1
Affordance grounding requires identifying where and how an agent should interact in open-world scenes, where actionable regions are often small, occluded, reflective, and visually…