23 citations · 65 across the 11 of their papers we have counts for
20 papers
VoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents
Edresson Casanova, Jaehyeon Kim, Mariana Graterol Fuenmayor +17
Spoken dialogue is a natural form of human--computer interaction, yet most speech language models remain limited to turn-based operation and lack real-time adaptability, such as us…
Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space
Thanh V. T. Tran, Ngoc-Son Nguyen, Luong Tran +4
Video-to-audio (V2A) generation aims to synthesize realistic audio that is both semantically consistent with and temporally synchronized to a silent video. Despite recent progress,…
MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form data
Subhankar Ghosh, Jason Li, Paarth Neekhara +4
Neural Text-to-Speech (TTS) systems achieve remarkable quality on short utterances but long-form speech generation shows prosodic drift, speaker inconsistencies and sentence bounda…
Align2Speak: Improving TTS for Low Resource Languages via ASR-Guided Online Preference Optimization
Shehzeen Hussain, Paarth Neekhara, Xuesong Yang +7
Developing high-quality text-to-speech (TTS) systems for low-resource languages is challenging due to the scarcity of paired text and speech data. In contrast, automatic speech rec…
Frame-Stacked Local Transformers For Efficient Multi-Codebook Speech Generation
Roy Fejgin, Paarth Neekhara, Xuesong Yang +6
Speech generation models based on large language models (LLMs) typically operate on discrete acoustic codes, which differ fundamentally from text tokens due to their multicodebook…
NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
Edresson Casanova, Paarth Neekhara, Ryan Langman +6
Large Language Models (LLMs) have significantly advanced audio processing by leveraging audio codecs to discretize audio into tokens, enabling the application of language modeling…