From the 1 of 13 linked papers with an AI index.
1 citations · 1 across the 11 of their papers we have counts for
6 papers · 1 filter
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting
Juyi Lin, Amir Taherin, Arash Akbari +11
Recent large-scale Vision Language Action (VLA) models have shown superior performance in robotic manipulation tasks guided by natural language. However, current VLA models suffer…
GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning
Fengyi Wu, Yifei Dong, Yilong Dai +7
We introduce Goal-Conditioned Visual Navigation Instruction Generation (GoViG), a new task that aims to generate contextually coherent navigation instructions solely from egocentri…
RoboStereo: Dual-Tower 4D Embodied World Models for Unified Policy Optimization
Ruicheng Zhang, Guangyu Chen, Zunnan Xu +5
Scalable Embodied AI faces fundamental constraints due to prohibitive costs and safety risks of real-world interaction. While Embodied World Models (EWMs) offer promise through ima…
Securing the Skies: A Comprehensive Survey on Anti-UAV Methods, Benchmarking, and Future Directions
Yifei Dong, Fengyi Wu, Sanjian Zhang +8
Unmanned Aerial Vehicles (UAVs) are indispensable for infrastructure inspection, surveillance, and related tasks, yet they also introduce critical security challenges. This survey…
Language-Conditioned World Modeling for Visual Navigation
Yifei Dong, Fengyi Wu, Yilong Dai +10
We study language-conditioned visual navigation (LCVN), in which an embodied agent is asked to follow a natural language instruction based only on an initial egocentric observation…
Universal Visuo-Tactile Video Understanding for Embodied Interaction
Yifan Xie, Mingyang Li, Shoujie Li +5
Tactile perception is essential for embodied agents to understand physical attributes of objects that cannot be determined through visual inspection alone. While existing approache…