23 citations · 24 across the 3 of their papers we have counts for
3 papers
Scenario-Aware Audio-Visual TF-GridNet for Target Speech Extraction
Zexu Pan, Gordon Wichern, Yoshiki Masuyama +4
Target speech extraction aims to extract, based on a given conditioning cue, a target speech signal that is corrupted by interfering sources, such as noise or competing speakers. B…
Generation or Replication: Auscultating Audio Latent Diffusion Models
Dimitrios Bralios, Gordon Wichern, François G. Germain +4
The introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how…
Attention-Based Multimodal Fusion for Video Description
Chiori Hori, Takaaki Hori, Teng-Yok Lee +3
Currently successful methods for video description are based on encoder-decoder sentence generation using recur-rent neural networks (RNNs). Recent work has shown the advantage of…