activity
20172026
most citedFaSNet: Low-latency Adaptive Beamforming for Multi-microphone Audio Processing

30 citations · 61 across the 20 of their papers we have counts for

collaborators
Showing eess.ASShow all

15 papers · 1 filter

eess.AS2025

MeanFlow-TSE: One-Step Generative Target Speaker Extraction with Mean Flow

Riki Shimizu, Xilin Jiang, Nima Mesgarani

Target speaker extraction (TSE) aims to isolate a desired speaker's voice from a multi-speaker mixture using auxiliary information such as a reference utterance. Although recent ad…

eess.AS2025

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis

Yinghao Aaron Li, Xilin Jiang, Fei Tao +4

Diffusion-based text-to-speech (TTS) systems have made remarkable progress in zero-shot speech synthesis, yet optimizing all components for perceptual metrics remains challenging.…

eess.AS2025

Exploring Finetuned Audio-LLM on Heart Murmur Features

Adrian Florea, Xilin Jiang, Nima Mesgarani +1

Large language models (LLMs) for audio have excelled in recognizing and analyzing human speech, music, and environmental sounds. However, their potential for understanding other ty…

eess.AS2024

StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion

Yinghao Aaron Li, Xilin Jiang, Cong Han +1

The rapid development of large-scale text-to-speech (TTS) models has led to significant advancements in modeling diverse speaker prosody and voices. However, these models often fac…

eess.AS202411 cited

Dual-path Mamba: Short and Long-term Bidirectional Selective Structured State Space Models for Speech Separation

Xilin Jiang, Cong Han, Nima Mesgarani

Transformers have been the most successful architecture for various speech modeling tasks, including speech separation. However, the self-attention mechanism in transformers with q…

eess.AS20206 cited

Continuous Speech Separation Using Speaker Inventory for Long Multi-talker Recording

Cong Han, Yi Luo, Chenda Li +8

Leveraging additional speaker information to facilitate speech separation has received increasing attention in recent years. Recent research includes extracting target speech by us…