11 papers
RAE-NWM: Navigation World Model in Dense Visual Representation Space
Mingkun Zhang, Wangtian Shen, Fan Zhang +3
Visual navigation requires agents to reach goals in complex environments through perception and planning. World models address this task by simulating action-conditioned state tran…
OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics
Mingxian Lin, Shengju Qian, Yuqi Liu +9
Vision-language model (VLM) agents are increasingly deployed in interactive game environments. Yet game benchmarks for VLM agents typically report a single first-attempt score per…
Your Model Already Knows: Attention-Guided Safety Filter for Vision-Language-Action Models
Seongbin Park, Fan Zhang, Baharan Mirzasoleiman +2
Vision-Language-Action (VLA) models have demonstrated impressive end-to-end performance across a variety of robotic manipulation tasks. However, these policies offer no guarantees…
ProbeAct: Probe-Guided Training-Free Failure Recovery in Vision-Language-Action Models
Fan Zhang, Seongbin Park, Baharan Mirzasoleiman +2
Vision-Language-Action (VLA) models demonstrate strong perfor-1 mance on language-conditioned robotic manipulation within their training dis-2 tribution, yet their generalization c…
MERGE: Guided Vision-Language Models for Multi-Actor Event Reasoning and Grounding in Human-Robot Interaction
Joerg Deigmoeller, Nakul Agarwal, Stephan Hasler +8
We introduce MERGE, a system for situational grounding of actors, objects, and events in dynamic human-robot group interactions. Effective collaboration in such settings requires c…
Generation of Real-time Robotic Emotional Expressions Learning from Human Demonstration in Mixed Reality
Chao Wang, Michael Gienger, Fan Zhang
Expressive behaviors in robots are critical for effectively conveying their emotional states during interactions with humans. In this work, we present a framework that autonomously…