activity
20202026
most citedAudio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

4 citations · 5 across the 9 of their papers we have counts for

collaborators
Showing cs.SDShow all

6 papers · 1 filter

cs.SD2024

Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment

Paarth Neekhara, Shehzeen Hussain, Subhankar Ghosh +4

Large Language Model (LLM) based text-to-speech (TTS) systems have demonstrated remarkable capabilities in handling large speech datasets and generating natural speech for new spea…

cs.SD2024★ 4 cited

Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

Zhifeng Kong, Arushi Goel, Rohan Badlani +3

Augmenting large language models (LLMs) to understand audio -- including non-speech sounds and non-verbal speech -- is critically important for diverse real-world applications of L…

cs.SD2024

Scaling NVIDIA's Multi-speaker Multi-lingual TTS Systems with Zero-Shot TTS to Indic Languages

Akshit Arora, Rohan Badlani, Sungwon Kim +2

In this paper, we describe the TTS models developed by NVIDIA for the MMITS-VC (Multi-speaker, Multi-lingual Indic TTS with Voice Cloning) 2024 Challenge. In Tracks 1 and 2, we uti…

cs.SD2023

VANI: Very-lightweight Accent-controllable TTS for Native and Non-native speakers with Identity Preservation

Rohan Badlani, Akshit Arora, Subhankar Ghosh +5

We introduce VANI, a very lightweight multi-lingual accent controllable speech synthesis system. Our model builds upon disentanglement strategies proposed in RADMMM and supports ex…

cs.SD2023★ 1 cited

Multilingual Multiaccented Multispeaker TTS with RADTTS

Rohan Badlani, Rafael Valle, Kevin J. Shih +3

We work to create a multilingual speech synthesis system which can generate speech with the proper accent while retaining the characteristics of an individual voice. This is challe…

cs.SD2021

One TTS Alignment To Rule Them All

Rohan Badlani, Adrian Łancucki, Kevin J. Shih +3

Speech-to-text alignment is a critical component of neural textto-speech (TTS) models. Autoregressive TTS models typically use an attention mechanism to learn these alignments on-l…