8 papers
AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects
Danrui Li, Jiahao Zhang, Bernhard Egger +4
Assembling objects from parts requires understanding multimodal instructions, linking them to 3D components, and predicting physically plausible 6-DoF motions for each assembly ste…
LASER: Layer-wise Scale Alignment for Training-Free Streaming 4D Reconstruction
Tianye Ding, Yiming Xie, Yiqing Liang +3
Recent feed-forward reconstruction models like VGGT and achieve impressive reconstruction quality but cannot process streaming videos due to quadratic memory complexity, lim…
Understanding Dynamic Compute Allocation in Recurrent Transformers
Ibraheem Muhammad Moosa, Suhas Lohit, Ye Wang +2
Token-level adaptive computation seeks to reduce inference cost by allocating more computation to harder tokens and less to easier ones. However, prior work is primarily evaluated…
A Probability-guided Sampler for Neural Implicit Surface Rendering
Gonçalo Dias Pais, Valter Piedade, Moitreya Chatterjee +2
Several variants of Neural Radiance Fields (NeRFs) have significantly improved the accuracy of synthesized images and surface reconstruction of 3D scenes/objects. In all of these m…
Programmatic Video Prediction Using Large Language Models
Hao Tang, Kevin Ellis, Suhas Lohit +2
The task of estimating the world model describing the dynamics of a real world process assumes immense importance for anticipating and preparing for future outcomes. For applicatio…
UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing
Yung-Hsuan Lai, Janek Ebbers, Yu-Chiang Frank Wang +3
Audio-Visual Video Parsing (AVVP) entails the challenging task of localizing both uni-modal events (i.e., those occurring exclusively in either the visual or acoustic modality of a…