activity
20122023
most citedDeep Convolutional Inverse Graphics Network

747 citations · 2.2k across the 86 of their papers we have counts for

collaborators
Showing cs.CVShow all

55 papers · 1 filter

cs.CV2023

Physion++: Evaluating Physical Scene Understanding that Requires Online Inference of Different Physical Properties

Hsiao-Yu Tung, Mingyu Ding, Zhenfang Chen +6

General physical scene understanding requires more than simply localizing and recognizing objects -- it requires knowledge that objects can have different latent properties (e.g.,…

cs.CV2023

Diffusion with Forward Models: Solving Stochastic Inverse Problems Without Direct Supervision

Ayush Tewari, Tianwei Yin, George Cazenavette +5

Denoising diffusion models are a powerful type of generative models used to capture complex distributions of real-world signals. However, their applicability is limited to scenario…

cs.CV2023

Unsupervised Compositional Concepts Discovery with Text-to-Image Generative Models

Nan Liu, Yilun Du, Shuang Li +2

Text-to-image generative models have enabled high-resolution image synthesis across different domains, but require users to specify the content they wish to generate. In this paper…

cs.CV2023

3D-IntPhys: Towards More Generalized 3D-grounded Visual Intuitive Physics under Challenging Scenes

Haotian Xue, Antonio Torralba, Joshua B. Tenenbaum +3

Given a visual scene, humans have strong intuitions about how a scene can evolve over time under given actions. The intuition, often termed visual intuitive physics, is a critical…

cs.CV20236 cited

Embodied Concept Learner: Self-supervised Learning of Concepts and Mapping through Instruction Following

Mingyu Ding, Yan Xu, Zhenfang Chen +4

Humans, even at a very early age, can learn visual concepts and understand geometry and layout through active interaction with the environment, and generalize their compositions to…

cs.CV2023

Visual Dependency Transformers: Dependency Tree Emerges from Reversed Attention

Mingyu Ding, Yikang Shen, Lijie Fan +5

Humans possess a versatile mechanism for extracting structured representations of our visual world. When looking at an image, we can decompose the scene into entities and their par…