1 citations · 1 across the 3 of their papers we have counts for
7 papers
Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
Xinlu He, Swayambhu Nath Ray, Harish Mallidi +5
Unified architectures in multimodal large language models (MLLM) have shown promise in handling diverse tasks within a single framework. In the text-to-speech (TTS) task, current M…
Unified Modeling of Multi-Domain Multi-Device ASR Systems
Soumyajit Mitra, Swayambhu Nath Ray, Bharat Padi +6
Modern Automatic Speech Recognition (ASR) systems often use a portfolio of domain-specific models in order to get high accuracy for distinct user utterance types across different d…
Improving RNN-T ASR Performance with Date-Time and Location Awareness
Swayambhu Nath Ray, Soumyajit Mitra, Raghavendra Bilgi +1
In this paper, we explore the benefits of incorporating context into a Recurrent Neural Network (RNN-T) based Automatic Speech Recognition (ASR) model to improve the speech recogni…
Listen with Intent: Improving Speech Recognition with Audio-to-Intent Front-End
Swayambhu Nath Ray, Minhua Wu, Anirudh Raju +7
Comprehending the overall intent of an utterance helps a listener recognize the individual words spoken. Inspired by this fact, we perform a novel study of the impact of explicitly…
Timestamping Documents and Beliefs
Swayambhu Nath Ray
Most of the textual information available to us are temporally variable. In a world where information is dynamic, time-stamping them is a very important task. Documents are a good…
Dating Documents using Graph Convolution Networks
Shikhar Vashishth, Shib Sankar Dasgupta, Swayambhu Nath Ray +1
Document date is essential for many important tasks, such as document retrieval, summarization, event detection, etc. While existing approaches for these tasks assume accurate know…