activity
20172026
most citedUnsupervised Image-to-Image Translation with Generative Adversarial Networks

70 citations · 149 across the 19 of their papers we have counts for

collaborators
Showing 2025Show all

5 papers · 1 filter

cs.AI2025

Align2Speak: Improving TTS for Low Resource Languages via ASR-Guided Online Preference Optimization

Shehzeen Hussain, Paarth Neekhara, Xuesong Yang +7

Developing high-quality text-to-speech (TTS) systems for low-resource languages is challenging due to the scarcity of paired text and speech data. In contrast, automatic speech rec…

eess.AS2025

Frame-Stacked Local Transformers For Efficient Multi-Codebook Speech Generation

Roy Fejgin, Paarth Neekhara, Xuesong Yang +6

Speech generation models based on large language models (LLMs) typically operate on discrete acoustic codes, which differ fundamentally from text tokens due to their multicodebook…

eess.AS2025

NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference

Edresson Casanova, Paarth Neekhara, Ryan Langman +6

Large Language Models (LLMs) have significantly advanced audio processing by leveraging audio codecs to discretize audio into tokens, enabling the application of language modeling…

eess.AS2025

HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset

Ryan Langman, Xuesong Yang, Paarth Neekhara +4

This paper introduces HiFiTTS-2, a large-scale speech dataset designed for high-bandwidth speech synthesis. The dataset is derived from LibriVox audiobooks, and contains approximat…

cs.SD2025

Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance

Shehzeen Hussain, Paarth Neekhara, Xuesong Yang +6

While autoregressive speech token generation models produce speech with remarkable variety and naturalness, their inherent lack of controllability often results in issues such as h…