4 papers
MDM-ASR: Bridging Accuracy and Efficiency in ASR with Diffusion-Based Non-Autoregressive Decoding
Hao Yen, Pin-Jui Ku, Ante JukiÄ +1
In sequence-to-sequence Transformer ASR, autoregressive (AR) models achieve strong accuracy but suffer from slow decoding, while non-autoregressive (NAR) models enable parallel dec…
A Bottom-up Framework with Language-universal Speech Attribute Modeling for Syllable-based ASR
Hao Yen, Pin-Jui Ku, Sabato Marco Siniscalchi +1
We propose a bottom-up framework for automatic speech recognition (ASR) in syllable-based languages by unifying language-universal articulatory attribute modeling with syllable-lev…
An Investigation on Combining Geometry and Consistency Constraints into Phase Estimation for Speech Enhancement
Chun-Wei Ho, Pin-Jui Ku, Hao Yen +3
We propose a novel iterative phase estimation framework, termed multi-source Griffin-Lim algorithm (MSGLA), for speech enhancement (SE) under additive noise conditions. The core id…
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
Hao Yen, Shaoshi Ling, Guoli Ye
We propose a novel approach to end-to-end automatic speech recognition (ASR) to achieve efficient speech in-context learning (SICL) for (i) long-form speech decoding, (ii) test-tim…