17 papers
What Matters for Latent Actions in Robot Learning
Xizhou Bu, Qingda Hu, Lei Zhou +13
Latent Action Models (LAMs) have emerged as a promising paradigm for enabling robot learning to leverage large-scale unlabeled videos through latent actions that serve as compact s…
Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments
Xianhui Meng, Zirui Song, Yuchen Zhang +10
Large Language Models (LLMs) have demonstrated remarkable capabilities in 3D indoor synthesis for Manhattan environments. However, existing methods often fail to capture plausible…
FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation
Lingfeng Zhang, Zeying Gong, Xiaoshuai Hao +7
Vision-and-language navigation (VLN) in continuous environments requires an agent to ground instructions in egocentric observations while maintaining spatial understanding across l…
OneVLA: A Unified Framework for Embodied Tasks
Lingfeng Zhang, Xiaoshuai Hao, Yingbo Tang +10
Navigation and manipulation are fundamental capabilities of embodied intelligence, enabling robots to interpret natural language commands and interact physically with their surroun…
Beyond Binary Success: A Diagnostic Meta-Evaluation Framework for Fine-Grained Manipulation
He-Yang Xu, Pengyuan Zhang, Zongyuan Ge +5
Fine-grained manipulation marks a regime where global scene context no longer suffices, and success hinges on the tight coupling of local attribute grounding, high-fidelity spatial…
Data-Asymmetric Latent Imagination and Reranking for 3D Robotic Imitation Learning
Lianghao Luo, Xizhou Bu, Ruyan Liu +5
Robotic imitation learning typically assumes access to optimal demonstrations, yet real-world data collection often yields suboptimal, exploratory, or even failed trajectories. Dis…