3 papers
cs.SD2025
A Non-autoregressive Model for Joint STT and TTS
Vishal Sunder, Brian Kingsbury, George Saon +5
In this paper, we take a step towards jointly modeling automatic speech recognition (STT) and speech synthesis (TTS) in a fully non-autoregressive way. We develop a novel multimoda…
cs.LG2025
Improving Transducer-Based Spoken Language Understanding with Self-Conditioned CTC and Knowledge Transfer
Vishal Sunder, Eric Fosler-Lussier
In this paper, we propose to improve end-to-end (E2E) spoken language understand (SLU) in an RNN transducer model (RNN-T) by incorporating a joint self-conditioned CTC automatic sp…
eess.AS2023
End-to-End real time tracking of children's reading with pointer network
Vishal Sunder, Beulah Karrolla, Eric Fosler-Lussier
In this work, we explore how a real time reading tracker can be built efficiently for children's voices. While previously proposed reading trackers focused on ASR-based cascaded ap…