collaborators

16 papers

cs.RO2026

EA-Nav: Learning Safe Visual Navigation Policies with Embodiment Awareness

Jialu Zhang, Yong Du, Xianda Guo +6

Cross-embodiment navigation is a key challenge in embodied intelligence. Due to differences in embodiment, the same visual observation may imply different actions for different age…

cs.RO2026

SkillNav: Score-Level Skill Intervention for Zero-Shot Object Goal Navigation

Ruijie Sang, Yiqun Duan, Pinhan Fu +3

Vision-Language Model (VLM) agents have advanced zero-shot object-goal navigation, yet single-frame reasoning leaves them without the cross-step behavioral awareness an embodied na…

cs.CV2026

X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras

Heng Zhou, Shuhong Liu, Yonghao He +6

X-Lens is a compact feed‑forward model that estimates metric depth in real time from a mix of calibrated fisheye and pinhole camera views using geometry‑aware calibration tokens an…

cs.RO2026

TrustVLA: Mechanism-Guided Inference-Time Defense Against Vision-Language-Action Backdoors

Pinhan Fu, Xianda Guo, Xuetao Li +5

The paper introduces TrustVLA, an inference-time defense that detects and mitigates visual backdoor triggers in vision‑language‑action models by monitoring epistemic uncertainty an…

cs.CV2026

TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation

Ziqian Wang, Yonghao He, Licheng Yang +6

Simulation provides a low-cost, scalable pathway to large-scale robotic manipulation data collection. However, existing 3D scene generation methods can rarely be applied directly t…

cs.RO2026

dVLA-RL: Reinforcement Learning over Denoising Trajectories for Discrete Diffusion Vision-Language-Action Models

Yuhao Wu, Yitian Liu, Weijie Shen +13

Vision-Language-Action (VLA) models have established a powerful paradigm for generalist robotic manipulation by grounding control into the semantic reasoning of VLMs. Prevailing ar…