10 citations · 18 across the 2 of their papers we have counts for
5 papers
The ThreeDWorld Transport Challenge: A Visually Guided Task-and-Motion Planning Benchmark for Physically Realistic Embodied AI
Chuang Gan, Siyuan Zhou, Jeremy Schwartz +8
We introduce a visually-guided and physics-driven task-and-motion planning benchmark, which we call the ThreeDWorld Transport Challenge. In this challenge, an embodied agent equipp…
Untangling in Invariant Speech Recognition
Cory Stephenson, Jenelle Feather, Suchismita Padhy +4
Encouraged by the success of deep neural networks on a variety of visual tasks, much theoretical and experimental work has been aimed at understanding and interpreting how vision n…
Self-Supervised Audio-Visual Co-Segmentation
Andrew Rouditchenko, Hang Zhao, Chuang Gan +2
Segmenting objects in images and separating sound sources in audio are challenging tasks, in part because traditional approaches require large amounts of labeled data. In this pape…
The Sound of Pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko +3
We introduce PixelPlayer, a system that, by leveraging large amounts of unlabeled videos, learns to locate image regions which produce sounds and separate the input sounds into a s…
Learning Sight from Sound: Ambient Sound Provides Supervision for Visual Learning
Andrew Owens, Jiajun Wu, Josh H. McDermott +2
The sound of crashing waves, the roar of fast-moving cars -- sound conveys important information about the objects in our surroundings. In this work, we show that ambient sounds ca…