Showing 2025Show all
2 papers · 1 filter
cs.SD2025
Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations
Yichen Han, Xiaoyang Hao, Keming Chen +25
Text-to-speech (TTS) synthesis has seen renewed progress under the discrete modeling paradigm. Existing autoregressive approaches often rely on single-codebook representations, whi…
cs.SD2025
FLAM: Frame-Wise Language-Audio Modeling
Yusong Wu, Christos Tsirigotis, Ke Chen +5
Recent multi-modal audio-language models (ALMs) excel at text-audio retrieval but struggle with frame-wise audio understanding. Prior works use temporal-aware labels or unsupervise…