7 papers
TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models
Ziheng Liu, Quantao Yang
Vision-language-action (VLA) models are commonly adapted to downstream manipulation tasks via supervised fine-tuning (SFT) or online reinforcement learning (RL) post-training. SFT…
Learning to Localize Reference Trajectories in Image-Space for Visual Navigation
Finn Lukas Busch, Matti Vahs, Quantao Yang +4
We present LoTIS, a model for visual navigation that provides robot-agnostic image-space guidance by localizing a reference RGB trajectory in the robot's current view, without requ…
FUS3DMaps: Scalable and Accurate Open-Vocabulary Semantic Mapping by 3D Fusion of Voxel- and Instance-Level Layers
Timon Homberger, Finn Lukas Busch, Jesús Gerardo Ortega Peimbert +2
Open-vocabulary semantic mapping enables robots to spatially ground previously unseen concepts without requiring predefined class sets. Current training-free methods commonly rely…
DIV-Nav: Open-Vocabulary Spatial Relationships for Multi-Object Navigation
Jesús Ortega-Peimbert, Finn Lukas Busch, Timon Homberger +2
Advances in open-vocabulary semantic mapping and object navigation have enabled robots to perform an informed search of their environment for an arbitrary object. However, such zer…
ViSA-Flow: Accelerating Robot Skill Learning via Large-Scale Video Semantic Action Flow
Changhe Chen, Quantao Yang, Xiaohao Xu +2
One of the central challenges preventing robots from acquiring complex manipulation skills is the prohibitive cost of collecting large-scale robot demonstrations. In contrast, huma…
Diffusion Trajectory-guided Policy for Long-horizon Robot Manipulation
Shichao Fan, Quantao Yang, Yajie Liu +4
Recently, Vision-Language-Action models (VLA) have advanced robot imitation learning, but high data collection costs and limited demonstrations hinder generalization and current im…