630 citations · 2.3k across the 46 of their papers we have counts for
21 papers · 1 filter
Towards Practical Multi-Object Manipulation using Relational Reinforcement Learning
Richard Li, Allan Jabri, Trevor Darrell +1
Learning robotic manipulation tasks using reinforcement learning with sparse rewards is currently impractical due to the outrageous data requirements. Many practical tasks require…
Something-Else: Compositional Action Recognition with Spatial-Temporal Interaction Networks
Joanna Materzynska, Tete Xiao, Roei Herzig +3
Human action is naturally compositional: humans can easily recognize and perform actions with objects that are different from those used in training demonstrations. In this paper,…
Learning Canonical Representations for Scene Graph to Image Generation
Roei Herzig, Amir Bar, Huijuan Xu +3
Generating realistic images of complex visual scenes becomes challenging when one wishes to control the structure of the generated images. Previous approaches showed that scenes wi…
Semantic Bottleneck Scene Generation
Samaneh Azadi, Michael Tschannen, Eric Tzeng +3
Coupling the high-fidelity generation capabilities of label-conditional image synthesis methods with the flexibility of unconditional generative models, we propose a semantic bottl…
Iterative Answer Prediction with Pointer-Augmented Multimodal Transformers for TextVQA
Ronghang Hu, Amanpreet Singh, Trevor Darrell +1
Many visual scenes contain text that carries crucial information, and it is thus essential to understand text in images for downstream reasoning tasks. For example, a deep water la…
Plan Arithmetic: Compositional Plan Vectors for Multi-Task Control
Coline Devin, Daniel Geng, Pieter Abbeel +2
Autonomous agents situated in real-world environments must be able to master large repertoires of skills. While a single short skill can be learned quickly, it would be impractical…