17 papers
dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model
Hankun Wang, Bohan Li, Shi Lian +6
Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural language provides a flexible interface fo…
SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings
Shuai Wang, Zihan Qian, Ke Zhang +9
The paper presents the REAL‑TSE Challenge, a benchmark for extracting a target speaker’s voice from real conversational recordings in Mandarin and English, with both online low‑lat…
Detect, Attend and Extract: Keyword Guided Target Speaker Extraction
Haoyu Li, Yu Xi, Yidi Jiang +5
Target speaker extraction (TSE) aims to extract the speech of a target speaker from mixtures containing multiple competing speakers. Conventional TSE systems predominantly rely on…
HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding
Bohan Li, Shi Lian, Hankun Wang +6
Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-quality waveforms. Existing speech tokenize…
G-STAR: End-to-End Global Speaker-Tracking Attributed Recognition
Jing Peng, Ziyi Chen, Haoyu Li +7
We study timestamped speaker-attributed automatic speech recognition (SA-ASR) for long-form, multi-party speech with overlap. In this setting, chunk-wise inference must preserve me…
Audio-Mind: An Auditable Agentic Framework for Audio Understanding
Yucheng Wang, Jing Peng, Hanqi Li +6
Audio agents extend large audio-language models (LALMs) by decomposing audio questions into tool calls, intermediate evidence, and iterative reasoning steps. However, as LALMs beco…