237 citations · 1.2k across the 69 of their papers we have counts for
15 papers · 1 filter
Look, Listen, and Act: Towards Audio-Visual Embodied Navigation
Chuang Gan, Yiwei Zhang, Jiajun Wu +2
A crucial ability of mobile intelligent agents is to integrate the evidence from multiple sensory inputs in an environment and to make a sequence of actions to reach their goals. I…
Accurate Vision-based Manipulation through Contact Reasoning
Alina Kloss, Maria Bauza, Jiajun Wu +3
Planning contact interactions is one of the core challenges of many robotic tasks. Optimizing contact locations while taking dynamics into account is computationally costly and, in…
Entity Abstraction in Visual Model-Based Reinforcement Learning
Rishi Veerapaneni, John D. Co-Reyes, Michael Chang +5
This paper tests the hypothesis that modeling a scene in terms of entities and their local interactions, as opposed to modeling the scene globally, provides a significant benefit i…
Learning Compositional Koopman Operators for Model-Based Control
Yunzhu Li, Hao He, Jiajun Wu +2
Finding an embedding space for a linear approximation of a nonlinear dynamical system enables efficient system identification and control synthesis. The Koopman operator theory lay…
CLEVRER: CoLlision Events for Video REpresentation and Reasoning
Kexin Yi, Chuang Gan, Yunzhu Li +4
The ability to reason about temporal and causal events from videos lies at the core of human intelligence. Most video reasoning benchmarks, however, focus on pattern recognition fr…
Program-Guided Image Manipulators
Jiayuan Mao, Xiuming Zhang, Yikai Li +3
Humans are capable of building holistic representations for images at various levels, from local objects, to pairwise relations, to global structures. The interpretation of structu…