activity
20242026
collaborators

19 papers

cs.SD2026

dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model

Hankun Wang, Bohan Li, Shi Lian +6

Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural language provides a flexible interface fo…

cs.SD2026

RAS: a Reliability Oriented Metric for Automatic Speech Recognition

Wenbin Huang, Yuhang Qiu, Bohan Li +5

Automatic speech recognition systems often produce confident yet incorrect transcriptions under noisy or ambiguous conditions, which can be misleading for both users and downstream…

cs.SD2026

dots.tts Technical Report

Shi Lian, Changtao Li, Bohan Li +6

We present dotstts, a 2B-parameter continuous autoregressive text-to-speech (TTS) foundation model that models speech in a continuous latent space. Compared with existing contin…

eess.AS2026

Read What You Hear: Reference-Free Hypotheses Evaluation with Acoustic Discrepancy

Zhihan Li, Hankun Wang, Yiwei Guo +3

Automatic speech recognition systems commonly rely on reference transcriptions for evaluation, while reference-free approaches often depend on internal confidence estimation or aux…

cs.SD2026

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

Bohan Li, Shi Lian, Hankun Wang +6

Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-quality waveforms. Existing speech tokenize…

eess.AS2026

Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios

Changhao Pan, Rui Yang, Han Wang +12

Recent advances in speech generation have enabled high-fidelity synthesis, yet systematic evaluation of models under long-context conditions remains largely underexplored. A compre…