3 citations · 4 across the 10 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching
Hui Wang, Shujie Liu, Lingwei Meng +9
To advance continuous-valued token modeling and temporal-coherence enforcement, we propose FELLE, an autoregressive model that integrates language modeling with token-wise flow mat…
cs.CL2025
Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning
Yexing Du, Youcheng Pan, Ziyang Ma +7
Multimodal Large Language Models (MLLMs) have achieved significant success in Speech-to-Text Translation (S2TT) tasks. While most existing research has focused on English-centric t…
cs.CL2025
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
Mingyu Cui, Yifan Yang, Jiajun Deng +7
Self-supervised learning (SSL) based discrete speech representations are highly compact and domain adaptable. In this paper, SSL discrete speech features extracted from WavLM model…