6 papers
Towards Safe Mobility: A Unified Transportation Foundation Model enabled by Open-Ended Vision-Language Dataset
Wenhui Huang, Songyan Zhang, Collister Chua +4
Urban transportation systems face growing safety challenges that require scalable intelligence for emerging smart mobility infrastructures. While recent advances in foundation mode…
Inference-Time Enhancement of Generative Robot Policies via Predictive World Modeling
Han Qi, Haocheng Yin, Aris Zhu +2
We present Generative Predictive Control (GPC), an inference-time method for improving pretrained behavior-cloning policies without retraining. GPC augments a frozen diffusion poli…
Compose by Focus: Scene Graph-based Atomic Skills
Han Qi, Changhe Chen, Heng Yang
A key requirement for generalist robots is compositional generalization - the ability to combine atomic skills to solve complex, long-horizon tasks. While prior work has primarily…
Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models
Rishabh Tiwari, Aditya Tomar, Udbhav Bamba +5
Process Reward Models (PRMs) are rapidly becoming the backbone of LLM reasoning pipelines, yet we demonstrate that state-of-the-art PRMs are systematically exploitable under advers…
MoTVLA: A Vision-Language-Action Model with Unified Fast-Slow Reasoning
Wenhui Huang, Changhe Chen, Han Qi +3
Integrating visual-language instructions into visuomotor policies is gaining momentum in robot learning for enhancing open-world generalization. Despite promising advances, existin…
Control-oriented Clustering of Visual Latent Representation
Han Qi, Haocheng Yin, Heng Yang
We initiate a study of the geometry of the visual representation space -- the information channel from the vision encoder to the action decoder -- in an image-based control pipelin…