5 papers
DLWM: Diverse Latent World Models for Efficient Multimodal Reasoning
David Huang, Lianlei Shan
Reasoning capabilities of multimodal large language models (MLLMs) have improved considerably in recent years. Existing approaches typically rely on explicit chain-of-thought or co…
TurboVGGT: Fast Visual Geometry Reconstruction with Adaptive Alternating Attention
David Huang, Guile Wu, Chengjie Huang +2
Recent feed-forward 3D reconstruction methods, such as visual geometry transformers, have substantially advanced the traditional per-scene optimization paradigm by enabling effecti…
MoVieDrive: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer
Guile Wu, David Huang, Dongfeng Bai +1
Urban scene synthesis with video generation models has recently shown great potential for autonomous driving. Existing video generation approaches to autonomous driving primarily f…
Language and Geometry Grounded Sparse Voxel Representations for Holistic Scene Understanding
Guile Wu, David Huang, Bingbing Liu +1
Existing 3D open-vocabulary scene understanding methods mostly emphasize distilling language features from 2D foundation models into 3D feature fields, but largely overlook the syn…
Robo-DM: Data Management For Large Robot Datasets
Kaiyuan Chen, Letian Fu, David Huang +9
Recent results suggest that very large datasets of teleoperated robot demonstrations can be used to train transformer-based models that have the potential to generalize to new scen…