works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.RO2026

SkillNav: Score-Level Skill Intervention for Zero-Shot Object Goal Navigation

Ruijie Sang, Yiqun Duan, Pinhan Fu +3

Vision-Language Model (VLM) agents have advanced zero-shot object-goal navigation, yet single-frame reasoning leaves them without the cross-step behavioral awareness an embodied na…

cs.RO2026

TrustVLA: Mechanism-Guided Inference-Time Defense Against Vision-Language-Action Backdoors

Pinhan Fu, Xianda Guo, Xuetao Li +5

The paper introduces TrustVLA, an inference-time defense that detects and mitigates visual backdoor triggers in vision‑language‑action models by monitoring epistemic uncertainty an…

cs.RO2026

Humanoid-OmniOcc: Stereo-Based Full-View Occupancy Dataset for Embodied AI

Xianda Guo, Bohao Zhang, Chenwei Huang +6

Occupancy prediction at voxel-level granularity is essential for safe robotic navigation and interaction in complex environments. Existing occupancy datasets, however, are predomin…

cs.CV2026

StereoFactory: A Unified Merging Framework for Robust Stereo Matching

Xianda Guo, Pinhan Fu, Ruilin Wang +3

Stereo matching has advanced through foundation models trained on large-scale datasets, yet this paradigm suffers from a scalability bottleneck: incorporating new data requires cos…

cs.CV2026

ROVR-Open-Dataset: A Large-Scale Depth Dataset for Autonomous Driving

Xianda Guo, Ruijun Zhang, Yiqun Duan +9

Depth estimation is a fundamental component of spatial perception for autonomous driving and other unmanned systems operating in open urban environments. Existing depth datasets su…

cs.CV2025

Stereo Anything: Unifying Zero-shot Stereo Matching with Large-Scale Mixed Data

Xianda Guo, Chenming Zhang, Youmin Zhang +8

Stereo matching serves as a cornerstone in 3D vision, aiming to establish pixel-wise correspondences between stereo image pairs for depth recovery. Despite remarkable progress driv…