collaborators

6 papers

cs.CV2026

Towards Safe Mobility: A Unified Transportation Foundation Model enabled by Open-Ended Vision-Language Dataset

Wenhui Huang, Songyan Zhang, Collister Chua +4

Urban transportation systems face growing safety challenges that require scalable intelligence for emerging smart mobility infrastructures. While recent advances in foundation mode…

cs.RO2026

Inference-Time Enhancement of Generative Robot Policies via Predictive World Modeling

Han Qi, Haocheng Yin, Aris Zhu +2

We present Generative Predictive Control (GPC), an inference-time method for improving pretrained behavior-cloning policies without retraining. GPC augments a frozen diffusion poli…

cs.RO2026

Compose by Focus: Scene Graph-based Atomic Skills

Han Qi, Changhe Chen, Heng Yang

A key requirement for generalist robots is compositional generalization - the ability to combine atomic skills to solve complex, long-horizon tasks. While prior work has primarily…

cs.LG2026

Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models

Rishabh Tiwari, Aditya Tomar, Udbhav Bamba +5

Process Reward Models (PRMs) are rapidly becoming the backbone of LLM reasoning pipelines, yet we demonstrate that state-of-the-art PRMs are systematically exploitable under advers…

cs.RO2025

MoTVLA: A Vision-Language-Action Model with Unified Fast-Slow Reasoning

Wenhui Huang, Changhe Chen, Han Qi +3

Integrating visual-language instructions into visuomotor policies is gaining momentum in robot learning for enhancing open-world generalization. Despite promising advances, existin…

cs.LG2025

Control-oriented Clustering of Visual Latent Representation

Han Qi, Haocheng Yin, Heng Yang

We initiate a study of the geometry of the visual representation space -- the information channel from the vision encoder to the action decoder -- in an image-based control pipelin…