activity
20182022
most citedVisualVoice: Audio-Visual Speech Separation with Cross-Modal Consistency

11 citations · 43 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV20223 cited

ObjectFolder 2.0: A Multisensory Object Dataset for Sim2Real Transfer

Ruohan Gao, Zilin Si, Yen-Yu Chang +5

Objects play a crucial role in our everyday activities. Though multisensory object-centric learning has shown great potential lately, the modeling of objects in prior work is rathe…

cs.CV202111 cited

VisualVoice: Audio-Visual Speech Separation with Cross-Modal Consistency

Ruohan Gao, Kristen Grauman

We introduce a new approach for audio-visual speech separation. Given a video, the goal is to extract the speech associated with a face in spite of simultaneous background sounds a…

cs.CV2020

Learning to Set Waypoints for Audio-Visual Navigation

Changan Chen, Sagnik Majumder, Ziad Al-Halah +3

In audio-visual navigation, an agent intelligently travels through a complex, unmapped 3D environment using both sights and sounds to find a sound source (e.g., a phone ringing in…

cs.CV202010 cited

VisualEchoes: Spatial Image Representation Learning through Echolocation

Ruohan Gao, Changan Chen, Ziad Al-Halah +2

Several animal species (e.g., bats, dolphins, and whales) and even visually impaired humans have the remarkable ability to perform echolocation: a biological sonar used to perceive…

cs.CV2019

Listen to Look: Action Recognition by Previewing Audio

Ruohan Gao, Tae-Hyun Oh, Kristen Grauman +1

In the face of the video data deluge, today's expensive clip-level classifiers are increasingly impractical. We propose a framework for efficient action recognition in untrimmed vi…

cs.CV2019

Co-Separating Sounds of Visual Objects

Ruohan Gao, Kristen Grauman

Learning how objects sound from video is challenging, since they often heavily overlap in a single audio channel. Current methods for visually-guided audio source separation sidest…