collaborators

5 papers

cs.AI2026

DriveCache: Action-Aware Caching for Driving World Model Inference

Jianchun Yang, Jian Liang, Xianda Guo +5

Driving video generation models support autonomous-driving development by predicting controllable future scenes for simulation, planning evaluation, and offline data generation. Di…

cs.RO2026

SkillNav: Score-Level Skill Intervention for Zero-Shot Object Goal Navigation

Ruijie Sang, Yiqun Duan, Pinhan Fu +3

Vision-Language Model (VLM) agents have advanced zero-shot object-goal navigation, yet single-frame reasoning leaves them without the cross-step behavioral awareness an embodied na…

cs.RO2026

TrustVLA: Mechanism-Guided Inference-Time Defense Against Vision-Language-Action Backdoors

Pinhan Fu, Xianda Guo, Xuetao Li +5

Vision-Language-Action (VLA) models are deployed through pipelines that end users cannot audit, and a poisoned VLA can behave normally on clean observations while a small visual tr…

cs.CV2026

StereoFactory: A Unified Merging Framework for Robust Stereo Matching

Xianda Guo, Pinhan Fu, Ruilin Wang +3

Stereo matching has advanced through foundation models trained on large-scale datasets, yet this paradigm suffers from a scalability bottleneck: incorporating new data requires cos…

cs.RO2026

When Attention Betrays: Erasing Backdoor Attacks in Robotic Policies by Reconstructing Visual Tokens

Xuetao Li, Pinhan Fu, Wenke Huang +7

Downstream fine-tuning of vision-language-action (VLA) models enhances robotics, yet exposes the pipeline to backdoor risks. Attackers can pretrain VLAs on poisoned data to implant…