4 papers · 1 filter
MilliVid: Hierarchical Latents for Long-Range Consistency in Video Generation
Ishaan Preetam Chandratreya, David Charatan, Basile Van Hoorick +4
Video generative models have become increasingly powerful, but long-range consistency remains challenging to achieve because even a few dozen frames require impractically long tran…
Understanding Multi-View Transformers
Michal Stary, Julien Gaubil, Ayush Tewari +1
Multi-view transformers such as DUSt3R are revolutionizing 3D vision by solving 3D tasks in a feed-forward manner. However, contrary to previous optimization-based pipelines, the i…
FlowMap: High-Quality Camera Poses, Intrinsics, and Depth via Gradient Descent
Cameron Smith, David Charatan, Ayush Tewari +1
This paper introduces FlowMap, an end-to-end differentiable method that solves for precise camera poses, camera intrinsics, and per-frame dense depth of a video sequence. Our metho…
Approaching human 3D shape perception with neurally mappable models
Thomas P. O'Connell, Tyler Bonnen, Yoni Friedman +4
Humans effortlessly infer the 3D shape of objects. What computations underlie this ability? Although various computational models have been proposed, none of them capture the human…