1 citations · 1 across the 3 of their papers we have counts for
Showing eess.ASShow all
3 papers · 1 filter
eess.AS2023
SLMGAN: Exploiting Speech Language Model Representations for Unsupervised Zero-Shot Voice Conversion in GANs
Yinghao Aaron Li, Cong Han, Nima Mesgarani
In recent years, large-scale pre-trained speech language models (SLMs) have demonstrated remarkable advancements in various generative speech modeling applications, such as text-to…
eess.AS2023
Online Binaural Speech Separation of Moving Speakers With a Wavesplit Network
Cong Han, Nima Mesgarani
Binaural speech separation in real-world scenarios often involves moving speakers. Most current speech separation methods use utterance-level permutation invariant training (u-PIT)…
eess.AS2023★ 1 cited
Improved Decoding of Attentional Selection in Multi-Talker Environments with Self-Supervised Learned Speech Representation
Cong Han, Vishal Choudhari, Yinghao Aaron Li +1
Auditory attention decoding (AAD) is a technique used to identify and amplify the talker that a listener is focused on in a noisy environment. This is done by comparing the listene…