8 papers
MilliVid: Hierarchical Latents for Long-Range Consistency in Video Generation
Ishaan Preetam Chandratreya, David Charatan, Basile Van Hoorick +4
Video generative models have become increasingly powerful, but long-range consistency remains challenging to achieve because even a few dozen frames require impractically long tran…
Reconstruction by Generation: 3D Multi-Object Scene Reconstruction from Sparse Observations
Andrii Zadaianchuk, Leonardo Barcellona, Lennard Schuenemann +7
Accurately reconstructing complex full multi-object scenes from sparse observations remains a core challenge in computer vision and a key step toward scalable and reliable simulati…
AnyView: Synthesizing Any Novel View in Dynamic Scenes
Basile Van Hoorick, Dian Chen, Shun Iwase +7
Modern generative video models excel at producing convincing, high-quality outputs, but struggle to maintain multi-view and spatiotemporal consistency in highly dynamic real-world…
Steerable Scene Generation with Post Training and Inference-Time Search
Nicholas Pfaff, Hongkai Dai, Sergey Zakharov +2
Training robots in simulation requires diverse 3D scenes that reflect the specific challenges of downstream tasks. However, scenes that satisfy strict task requirements, such as hi…
OmniShape: Zero-Shot Multi-Hypothesis Shape and Pose Estimation in the Real World
Katherine Liu, Sergey Zakharov, Dian Chen +4
We would like to estimate the pose and full shape of an object from a single observation, without assuming known 3D model or category. In this work, we propose OmniShape, the first…
A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation
TRI LBM Team, Jose Barreiros, Andrew Beaulieu +79
Robot manipulation has seen tremendous progress in recent years, with imitation learning policies enabling successful performance of dexterous and hard-to-model tasks. Concurrently…