5 papers
DriveCache: Action-Aware Caching for Driving World Model Inference
Jianchun Yang, Jian Liang, Xianda Guo +5
Driving video generation models support autonomous-driving development by predicting controllable future scenes for simulation, planning evaluation, and offline data generation. Di…
SkillNav: Score-Level Skill Intervention for Zero-Shot Object Goal Navigation
Ruijie Sang, Yiqun Duan, Pinhan Fu +3
Vision-Language Model (VLM) agents have advanced zero-shot object-goal navigation, yet single-frame reasoning leaves them without the cross-step behavioral awareness an embodied na…
TrustVLA: Mechanism-Guided Inference-Time Defense Against Vision-Language-Action Backdoors
Pinhan Fu, Xianda Guo, Xuetao Li +5
Vision-Language-Action (VLA) models are deployed through pipelines that end users cannot audit, and a poisoned VLA can behave normally on clean observations while a small visual tr…
StereoFactory: A Unified Merging Framework for Robust Stereo Matching
Xianda Guo, Pinhan Fu, Ruilin Wang +3
Stereo matching has advanced through foundation models trained on large-scale datasets, yet this paradigm suffers from a scalability bottleneck: incorporating new data requires cos…
When Attention Betrays: Erasing Backdoor Attacks in Robotic Policies by Reconstructing Visual Tokens
Xuetao Li, Pinhan Fu, Wenke Huang +7
Downstream fine-tuning of vision-language-action (VLA) models enhances robotics, yet exposes the pipeline to backdoor risks. Attackers can pretrain VLAs on poisoned data to implant…