2.1k citations · 2.5k across the 13 of their papers we have counts for
8 papers · 1 filter
Imagen Video: High Definition Video Generation with Diffusion Models
Jonathan Ho, William Chan, Chitwan Saharia +8
We present Imagen Video, a text-conditional video generation system based on a cascade of video diffusion models. Given a text prompt, Imagen Video generates high definition videos…
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
Chitwan Saharia, William Chan, Saurabh Saxena +11
We present Imagen, a text-to-image diffusion model with an unprecedented degree of photorealism and a deep level of language understanding. Imagen builds on the power of large tran…
Kubric: A scalable dataset generator
Klaus Greff, Francois Belletti, Lucas Beyer +32
Data is the driving force of machine learning, with the amount and quality of training data often being more important for the performance of a system than architecture and trainin…
Unsupervised part representation by Flow Capsules
Sara Sabour, Andrea Tagliasacchi, Soroosh Yazdani +2
Capsule networks aim to parse images into a hierarchy of objects, parts and relations. While promising, they remain limited by an inability to learn effective low level part descri…
Walking on Thin Air: Environment-Free Physics-based Markerless Motion Capture
Micha Livne, Leonid Sigal, Marcus A. Brubaker +1
We propose a generative approach to physics-based motion capture. Unlike prior attempts to incorporate physics into tracking that assume the subject and scene geometry are calibrat…
Hierarchical Video Understanding
Farzaneh Mahdisoltani, Roland Memisevic, David Fleet
We introduce a hierarchical architecture for video understanding that exploits the structure of real world actions by capturing targets at different levels of granularity. We desig…