10 citations · 18 across the 2 of their papers we have counts for
4 papers · 1 filter
The ThreeDWorld Transport Challenge: A Visually Guided Task-and-Motion Planning Benchmark for Physically Realistic Embodied AI
Chuang Gan, Siyuan Zhou, Jeremy Schwartz +8
We introduce a visually-guided and physics-driven task-and-motion planning benchmark, which we call the ThreeDWorld Transport Challenge. In this challenge, an embodied agent equipp…
Self-Supervised Audio-Visual Co-Segmentation
Andrew Rouditchenko, Hang Zhao, Chuang Gan +2
Segmenting objects in images and separating sound sources in audio are challenging tasks, in part because traditional approaches require large amounts of labeled data. In this pape…
The Sound of Pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko +3
We introduce PixelPlayer, a system that, by leveraging large amounts of unlabeled videos, learns to locate image regions which produce sounds and separate the input sounds into a s…
Learning Sight from Sound: Ambient Sound Provides Supervision for Visual Learning
Andrew Owens, Jiajun Wu, Josh H. McDermott +2
The sound of crashing waves, the roar of fast-moving cars -- sound conveys important information about the objects in our surroundings. In this work, we show that ambient sounds ca…