4 papers
MDM-ASR: Bridging Accuracy and Efficiency in ASR with Diffusion-Based Non-Autoregressive Decoding
Hao Yen, Pin-Jui Ku, Ante JukiÄ +1
In sequence-to-sequence Transformer ASR, autoregressive (AR) models achieve strong accuracy but suffer from slow decoding, while non-autoregressive (NAR) models enable parallel dec…
Discrete Diffusion for Generative Modeling of Text-Aligned Speech Tokens
Pin-Jui Ku, He Huang, Jean-Marie Lemercier +3
This paper introduces a discrete diffusion model (DDM) framework for text-aligned speech tokenization and reconstruction. By replacing the auto-regressive speech decoder with a dis…
A Bottom-up Framework with Language-universal Speech Attribute Modeling for Syllable-based ASR
Hao Yen, Pin-Jui Ku, Sabato Marco Siniscalchi +1
We propose a bottom-up framework for automatic speech recognition (ASR) in syllable-based languages by unifying language-universal articulatory attribute modeling with syllable-lev…
An Investigation on Combining Geometry and Consistency Constraints into Phase Estimation for Speech Enhancement
Chun-Wei Ho, Pin-Jui Ku, Hao Yen +3
We propose a novel iterative phase estimation framework, termed multi-source Griffin-Lim algorithm (MSGLA), for speech enhancement (SE) under additive noise conditions. The core id…