230 citations · 1.1k across the 45 of their papers we have counts for
38 papers · 1 filter
Greedy Hierarchical Variational Autoencoders for Large-Scale Video Prediction
Bohan Wu, Suraj Nair, Roberto Martin-Martin +2
A video prediction model that generalizes to diverse scenes would enable intelligent agents such as robots to perform a variety of tasks via planning with the model. However, while…
Learning Physical Graph Representations from Visual Scenes
Daniel M. Bear, Chaofei Fan, Damian Mrowca +8
Convolutional Neural Networks (CNNs) have proved exceptional at learning representations for visual object categorization. However, CNNs do not explicitly encode objects, parts, an…
Towards Fairer Datasets: Filtering and Balancing the Distribution of the People Subtree in the ImageNet Hierarchy
Kaiyu Yang, Klint Qinami, Li Fei-Fei +2
Computer vision technology is being used by many but remains representative of only a few. People have reported misbehavior of computer vision models, including offensive predictio…
Action Genome: Actions as Composition of Spatio-temporal Scene Graphs
Jingwei Ji, Ranjay Krishna, Li Fei-Fei +1
Action recognition has typically treated actions and activities as monolithic events that occur in videos. However, there is evidence from Cognitive Science and Neuroscience that p…
Deep Bayesian Active Learning for Multiple Correct Outputs
Khaled Jedoui, Ranjay Krishna, Michael Bernstein +1
Typical active learning strategies are designed for tasks, such as classification, with the assumption that the output space is mutually exclusive. The assumption that these tasks…
6-PACK: Category-level 6D Pose Tracker with Anchor-Based Keypoints
Chen Wang, Roberto Martín-Martín, Danfei Xu +5
We present 6-PACK, a deep learning approach to category-level 6D object pose tracking on RGB-D data. Our method tracks in real-time novel object instances of known object categorie…