activity
20242026
collaborators

12 papers

eess.AS2026

DiaScriber: A Speech LLM for Joint Diarization and Transcription in Multi-Speaker Scenarios

Bingshen Mu, Xian Shi, Xiong Wang +6

Multi-speaker automatic speech recognition (MSASR) aims to jointly predict content transcriptions, speaker identities, and timestamps, thereby addressing the key question of "who s…

eess.AS2026

Detect, Attend and Extract: Keyword Guided Target Speaker Extraction

Haoyu Li, Yu Xi, Yidi Jiang +5

Target speaker extraction (TSE) aims to extract the speech of a target speaker from mixtures containing multiple competing speakers. Conventional TSE systems predominantly rely on…

cs.SD2026

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

Bohan Li, Shi Lian, Hankun Wang +6

Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-quality waveforms. Existing speech tokenize…

eess.AS2026

Audio-Mind: An Auditable Agentic Framework for Audio Understanding

Yucheng Wang, Jing Peng, Hanqi Li +6

Audio agents extend large audio-language models (LALMs) by decomposing audio questions into tool calls, intermediate evidence, and iterative reasoning steps. However, as LALMs beco…

eess.AS2026

Time-Layer Adaptive Alignment for Speaker Similarity in Flow-Matching Based Zero-Shot TTS

Haoyu Li, Mingyang Han, Yu Xi +8

Flow-Matching (FM)-based zero-shot text-to-speech (TTS) systems exhibit high-quality speech synthesis and robust generalization capabilities. However, the speaker representation ab…

eess.AS2025

A Survey on Speech Large Language Models for Understanding

Jing Peng, Yucheng Wang, Bohan Li +9

Speech understanding is essential for interpreting the diverse forms of information embedded in spoken language, including linguistic, paralinguistic, and non-linguistic cues that…