30 citations · 47 across the 8 of their papers we have counts for
11 papers · 1 filter
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
Yinghao Aaron Li, Xilin Jiang, Cong Han +1
The rapid development of large-scale text-to-speech (TTS) models has led to significant advancements in modeling diverse speaker prosody and voices. However, these models often fac…
Dual-path Mamba: Short and Long-term Bidirectional Selective Structured State Space Models for Speech Separation
Xilin Jiang, Cong Han, Nima Mesgarani
Transformers have been the most successful architecture for various speech modeling tasks, including speech separation. However, the self-attention mechanism in transformers with q…
Multi-Channel Speech Denoising for Machine Ears
Cong Han, E. Merve Kaya, Kyle Hoefer +2
This work describes a speech denoising system for machine ears that aims to improve speech intelligibility and the overall listening experience in noisy environments. We recorded a…
Dual-Path Modeling for Long Recording Speech Separation in Meetings
Chenda Li, Zhuo Chen, Yi Luo +6
The continuous speech separation (CSS) is a task to separate the speech sources from a long, partially overlapped recording, which involves a varying number of speakers. A straight…
Continuous Speech Separation Using Speaker Inventory for Long Multi-talker Recording
Cong Han, Yi Luo, Chenda Li +8
Leveraging additional speaker information to facilitate speech separation has received increasing attention in recent years. Recent research includes extracting target speech by us…
Group Communication with Context Codec for Lightweight Source Separation
Yi Luo, Cong Han, Nima Mesgarani
Despite the recent progress on neural network architectures for speech separation, the balance between the model size, model complexity and model performance is still an important…