4 papers · 1 filter
DLWM: Diverse Latent World Models for Efficient Multimodal Reasoning
David Huang, Lianlei Shan
Reasoning capabilities of multimodal large language models (MLLMs) have improved considerably in recent years. Existing approaches typically rely on explicit chain-of-thought or co…
TurboVGGT: Fast Visual Geometry Reconstruction with Adaptive Alternating Attention
David Huang, Guile Wu, Chengjie Huang +2
Recent feed-forward 3D reconstruction methods, such as visual geometry transformers, have substantially advanced the traditional per-scene optimization paradigm by enabling effecti…
MoVieDrive: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer
Guile Wu, David Huang, Dongfeng Bai +1
Urban scene synthesis with video generation models has recently shown great potential for autonomous driving. Existing video generation approaches to autonomous driving primarily f…
Language and Geometry Grounded Sparse Voxel Representations for Holistic Scene Understanding
Guile Wu, David Huang, Bingbing Liu +1
Existing 3D open-vocabulary scene understanding methods mostly emphasize distilling language features from 2D foundation models into 3D feature fields, but largely overlook the syn…