activity
20162019
most citedMarrNet: 3D Shape Reconstruction via 2.5D Sketches

237 citations · 559 across the 5 of their papers we have counts for

collaborators

6 papers

cs.RO201917 cited

DensePhysNet: Learning Dense Physical Object Representations via Multi-step Dynamic Interactions

Zhenjia Xu, Jiajun Wu, Andy Zeng +2

We study the problem of learning physical object representations for robot manipulation. Understanding object physics is critical for successful object manipulation, but also chall…

cs.CV2019143 cited

The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision

Jiayuan Mao, Chuang Gan, Pushmeet Kohli +2

We propose the Neuro-Symbolic Concept Learner (NS-CL), a model that learns visual concepts, words, and semantic parsing of sentences without explicit supervision on any of them; in…

cs.CV2017

Learning Sight from Sound: Ambient Sound Provides Supervision for Visual Learning

Andrew Owens, Jiajun Wu, Josh H. McDermott +2

The sound of crashing waves, the roar of fast-moving cars -- sound conveys important information about the objects in our surroundings. In this work, we show that ambient sounds ca…

cs.CV201719 cited

Attention Clusters: Purely Attention Based Local Feature Integration for Video Classification

Xiang Long, Chuang Gan, Gerard de Melo +3

Recently, substantial research effort has focused on how to apply CNNs or RNNs to better extract temporal patterns from videos, so as to improve the accuracy of video classificatio…

cs.CV2017237 cited

MarrNet: 3D Shape Reconstruction via 2.5D Sketches

Jiajun Wu, Yifan Wang, Tianfan Xue +3

3D object reconstruction from a single image is a highly under-determined problem, requiring strong prior knowledge of plausible 3D shapes. This introduces challenges for learning-…

cs.CV2016143 cited

Visual Dynamics: Probabilistic Future Frame Synthesis via Cross Convolutional Networks

Tianfan Xue, Jiajun Wu, Katherine L. Bouman +1

We study the problem of synthesizing a number of likely future frames from a single input image. In contrast to traditional methods, which have tackled this problem in a determinis…