activity
20242026
collaborators
Showing eess.ASShow all

12 papers · 1 filter

eess.AS2026

DiaScriber: A Speech LLM for Joint Diarization and Transcription in Multi-Speaker Scenarios

Bingshen Mu, Xian Shi, Xiong Wang +6

Multi-speaker automatic speech recognition (MSASR) aims to jointly predict content transcriptions, speaker identities, and timestamps, thereby addressing the key question of "who s…

eess.AS2026

Audio-Mind: An Auditable Agentic Framework for Audio Understanding

Yucheng Wang, Jing Peng, Hanqi Li +6

Audio agents extend large audio-language models (LALMs) by decomposing audio questions into tool calls, intermediate evidence, and iterative reasoning steps. However, as LALMs beco…

eess.AS2026

Detect, Attend and Extract: Keyword Guided Target Speaker Extraction

Haoyu Li, Yu Xi, Yidi Jiang +5

Target speaker extraction (TSE) aims to extract the speech of a target speaker from mixtures containing multiple competing speakers. Conventional TSE systems predominantly rely on…

eess.AS2025

Time-Layer Adaptive Alignment for Speaker Similarity in Flow-Matching Based Zero-Shot TTS

Haoyu Li, Mingyang Han, Yu Xi +8

Flow-Matching (FM)-based zero-shot text-to-speech (TTS) systems exhibit high-quality speech synthesis and robust generalization capabilities. However, the speaker representation ab…

eess.AS2025

MFA-KWS: Effective Keyword Spotting with Multi-head Frame-asynchronous Decoding

Yu Xi, Haoyu Li, Xiaoyu Gu +2

Keyword spotting (KWS) is essential for voice-driven applications, demanding both accuracy and efficiency. Traditional ASR-based KWS methods, such as greedy and beam search, explor…

eess.AS2024

Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario

Wen Wen, Qiang Zhou, Yu Xi +3

In multi-speaker scenarios, leveraging spatial features is essential for enhancing target speech. While with limited microphone arrays, developing a compact multi-channel speech en…