5 citations · 5 across the 3 of their papers we have counts for
3 papers
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
Sathvik Udupa, Shinji Watanabe, Petr Schwarz +1
Accurate, low-latency endpointing is crucial for effective spoken dialogue systems. While traditional endpointers often rely on spectrum-based audio features, this work proposes re…
Speaking rate attention-based duration prediction for speed control TTS
Jesuraj Bandekar, Sathvik Udupa, Abhayjeet Singh +6
With the advent of high-quality speech synthesis, there is a lot of interest in controlling various prosodic attributes of speech. Speaking rate is an essential attribute towards m…
Model Adaptation for ASR in low-resource Indian Languages
Abhayjeet Singh, Arjun Singh Mehta, Ashish Khuraishi K S +17
Automatic speech recognition (ASR) performance has improved drastically in recent years, mainly enabled by self-supervised learning (SSL) based acoustic models such as wav2vec2 and…