collaborators

12 papers

cs.CV2026

UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving

Zhexiao Xiong, Xin Ye, Burhan Yaman +5

World models have become central to autonomous driving, where accurate scene understanding and future prediction are crucial for safe control. Recent work has explored using vision…

cs.CV2026

ActWorld: From Explorable to Interactive World Model via Action-Aware Memory

Zhexiao Xiong, Yizhi Song, Hao Kang +11

Interactive world models aim to simulate environment dynamics under real-time user actions. However, their action vocabulary is largely confined to navigation: most actions corresp…

cs.CV2026

Track2View: 4D-Consistent Camera-Controlled Video Generation via Paired 3D Point Tracks

Feng Qiao, Zhaochong An, Zhexiao Xiong +2

Re-rendering an existing video from a novel camera viewpoint requires the output to follow the prescribed camera trajectory while preserving the appearance and dynamics of the orig…

cs.CV2026

MCPDepth: Omnidirectional Depth Estimation via Stereo Matching from Multi-Cylindrical Panoramas

Feng Qiao, Zhexiao Xiong, Xinge Zhu +3

Omnidirectional depth estimation presents a significant challenge due to the inherent distortions in panoramic images. Despite notable advancements, the impact of projection method…

cs.CV2026

GenOpticalFlow: A Generative Approach to Unsupervised Optical Flow Learning

Yixuan Luo, Feng Qiao, Zhexiao Xiong +2

Optical flow estimation is a fundamental problem in computer vision, yet the reliance on expensive ground-truth annotations limits the scalability of supervised approaches. Althoug…

cs.CV2026

Reconstruction Matters: Learning Geometry-Aligned BEV Representation through 3D Gaussian Splatting

Yiren Lu, Xin Ye, Burhaneddin Yaman +4

Bird's-Eye-View (BEV) perception serves as a cornerstone for autonomous driving, offering a unified spatial representation that fuses surrounding-view images to enable reasoning fo…