11 papers · 1 filter
DreamWAM: Beyond RGB Future Prediction for World Action Models
Shanglin Yuan, Weiheng Zhao, Xin Shi +6
World Action Models (WAMs) learn action-relevant representations by predicting how the observed world will evolve. Most existing WAMs define this future in RGB space, where task-re…
EA-Nav: Learning Safe Visual Navigation Policies with Embodiment Awareness
Jialu Zhang, Yong Du, Xianda Guo +6
Cross-embodiment navigation is a key challenge in embodied intelligence. Due to differences in embodiment, the same visual observation may imply different actions for different age…
SkillNav: Score-Level Skill Intervention for Zero-Shot Object Goal Navigation
Ruijie Sang, Yiqun Duan, Pinhan Fu +3
Vision-Language Model (VLM) agents have advanced zero-shot object-goal navigation, yet single-frame reasoning leaves them without the cross-step behavioral awareness an embodied na…
TrustVLA: Mechanism-Guided Inference-Time Defense Against Vision-Language-Action Backdoors
Pinhan Fu, Xianda Guo, Xuetao Li +5
Vision-Language-Action (VLA) models are deployed through pipelines that end users cannot audit, and a poisoned VLA can behave normally on clean observations while a small visual tr…
Humanoid-OmniOcc: Stereo-Based Full-View Occupancy Dataset for Embodied AI
Xianda Guo, Bohao Zhang, Chenwei Huang +6
Occupancy prediction at voxel-level granularity is essential for safe robotic navigation and interaction in complex environments. Existing occupancy datasets, however, are predomin…
ReactSim-Bench: Benchmarking Reactive Behavior World Model Simulation in Autonomous Driving
Zhiyuan Zhang, Yanlun Peng, Jianing Zhang +7
Reactive capability is a key property of data-driven behavior world model simulators for autonomous driving simulation systems. With this capability, simulated world agents can res…