2 papers
cs.CV2026
SV-WAM: An Efficient Surround-View World-Action Model for End-to-End Autonomous Driving
Jinyang Wang, Shiwei Li, Junjian Wang +12
World models (WMs) have demonstrated strong potential for end-to-end autonomous driving by learning predictive representations of future scene dynamics. However, generating future…
cs.CV2026
SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts
Haida Feng, Hao Wei, Haolin Wang +3
Recent Multimodal Large Language Models (MLLMs) struggle to bridge the representational gap between 2D semantic understanding and 3D spatial geometry. Existing 3D-aware models eith…