collaborators

7 papers

cs.AI2026

DriveCache: Action-Aware Caching for Driving World Model Inference

Jianchun Yang, Jian Liang, Xianda Guo +5

Driving video generation models support autonomous-driving development by predicting controllable future scenes for simulation, planning evaluation, and offline data generation. Di…

cs.CV2026

InstructVVT: Instruction-Driven Video Virtual Try-On without Auxiliary Spatial Priors

Dingbao Shao, Song Wu, Xinyu Chen +18

Video virtual try-on is a highly constrained editing task requiring the precise replacement of a target person's clothing while strictly preserving the original video's spatial str…

cs.CV2026

Visual Geometry Foundation-Aware Gaussians for Single-Frame Surround-View Driving Reconstruction

Junhong Lin, Jinlong Wang, Xianda Guo +6

Single-frame surround-view reconstruction faces severe geometric instability and rendering artifacts due to minimal inter-camera overlap. While existing methods rely on complex dec…

cs.CV2026

Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning

Qianlong Yang, Bowen Ye, Xianda Guo +4

Despite the progress of multimodal large language models (MLLMs), they continue to exhibit deficiencies in visual perception. Following visual instruction tuning, internal MLLM rep…

cs.CV2026

VGOcc: Learning Visual-Geometric Gaussians for Vision-Centric 3D Driving Occupancy Prediction

Junhong Lin, Xianda Guo, Kangli Wang +4

Vision-only occupancy prediction requires recovering a semantic 3D occupancy field from calibrated surround-view images, where each view provides observations with ambiguous depth…

cs.CV2026

Bridging 3D Gaussians and Semantic Occupancy for Comprehensive Open-Vocabulary Scene Understanding from Unposed Images

Hu Zhu, Bohan Li, Xianda Guo +5

Comprehensive 3D scene understanding from sparse, unposed images requires a model to recover renderable geometry, open-vocabulary semantics, and free/occupied 3D space without rely…