works on

From the 1 of 12 linked papers with an AI index.

collaborators

12 papers

cs.CV2026

FoundationGeo: Learning Spatial Pixel-Wise Fields for Monocular Metric Geometry

Muxin Liu, Xiaoyang Lyu, Tianhe Ren +7

FoundationGeo is a two‑stage framework that first learns an affine‑invariant geometry model from a large multi‑domain dataset, then refines metric depth using lightweight pixel‑wis…

cs.CV2026

DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving

Chen Shi, Jinrui Xu, Shaoshuai Shi +3

Pretrained foundation models have become an important basis for end-to-end autonomous driving. In contrast to vision-language models pretrained primarily on static image-text pairs…

cs.CV2026

Stabilizing Streaming Video Geometry via Dynamic Feature Normalization

Xiaoyang Lyu, Muxin Liu, Xiaoshan Wu +5

Consistent 3D geometry estimation from streaming RGB input is crucial for real-world applications such as autonomous driving, embodied AI, and large-scale reconstruction. While mod…

cs.CV2026

GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation

Jingjing Qian, Boyao Han, Chen Shi +4

Vision-Language-Action (VLA) models achieve strong generalization in robotic manipulation but remain largely reactive and 2D-centric, making them unreliable in tasks that require p…

cs.CV2026

JiSAM: Alleviate Labeling Burden and Corner Case Problems in Autonomous Driving via Minimal Real-World Data

Runjian Chen, Wenqi Shao, Bo Zhang +3

Deep-learning-based autonomous driving (AD) perception introduces a promising picture for safe and environment-friendly transportation. However, the over-reliance on real labeled d…

cs.CV2026

LIVE: Long-horizon Interactive Video World Modeling

Junchao Huang, Ziyang Ye, Xinting Hu +5

Autoregressive video world models predict future visual observations conditioned on actions. While effective over short horizons, these models often struggle with long-horizon gene…