12 citations · 19 across the 8 of their papers we have counts for
10 papers
Audio-Visual Scene-Aware Dialog and Reasoning using Audio-Visual Transformers with Joint Student-Teacher Learning
Ankit P. Shah, Shijie Geng, Peng Gao +5
In previous work, we have proposed the Audio-Visual Scene-Aware Dialog (AVSD) task, collected an AVSD dataset, developed AVSD technologies, and hosted an AVSD challenge track at bo…
Advancing Momentum Pseudo-Labeling with Conformer and Initialization Strategy
Yosuke Higuchi, Niko Moritz, Jonathan Le Roux +1
Pseudo-labeling (PL), a semi-supervised learning (SSL) method where a seed model performs self-training using pseudo-labels generated from untranscribed speech, has been shown to e…
Leveraging Low-Distortion Target Estimates for Improved Speech Enhancement
Zhong-Qiu Wang, Gordon Wichern, Jonathan Le Roux
A promising approach for multi-microphone speech separation involves two deep neural networks (DNN), where the predicted target speech from the first DNN is used to compute signal…
Visual Scene Graphs for Audio Source Separation
Moitreya Chatterjee, Jonathan Le Roux, Narendra Ahuja +1
State-of-the-art approaches for visually-guided audio source separation typically assume sources that have characteristic sounds, such as musical instruments. These approaches ofte…
Convolutive Prediction for Reverberant Speech Separation
Zhong-Qiu Wang, Gordon Wichern, Jonathan Le Roux
We investigate the effectiveness of convolutive prediction, a novel formulation of linear prediction for speech dereverberation, for speaker separation in reverberant conditions. T…
Optimizing Latency for Online Video CaptioningUsing Audio-Visual Transformers
Chiori Hori, Takaaki Hori, Jonathan Le Roux
Video captioning is an essential technology to understand scenes and describe events in natural language. To apply it to real-time monitoring, a system needs not only to describe e…