activity
20232026
most citedFreeSVC: Towards Zero-shot Multilingual Singing Voice Conversion

3 citations · 5 across the 10 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2026

VoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents

Edresson Casanova, Jaehyeon Kim, Mariana Graterol Fuenmayor +17

Spoken dialogue is a natural form of human--computer interaction, yet most speech language models remain limited to turn-based operation and lack real-time adaptability, such as us…

eess.AS2025

Frame-Stacked Local Transformers For Efficient Multi-Codebook Speech Generation

Roy Fejgin, Paarth Neekhara, Xuesong Yang +6

Speech generation models based on large language models (LLMs) typically operate on discrete acoustic codes, which differ fundamentally from text tokens due to their multicodebook…

eess.AS2025

NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference

Edresson Casanova, Paarth Neekhara, Ryan Langman +6

Large Language Models (LLMs) have significantly advanced audio processing by leveraging audio codecs to discretize audio into tokens, enabling the application of language modeling…

eess.AS2025

HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset

Ryan Langman, Xuesong Yang, Paarth Neekhara +4

This paper introduces HiFiTTS-2, a large-scale speech dataset designed for high-bandwidth speech synthesis. The dataset is derived from LibriVox audiobooks, and contains approximat…

eess.AS2024

Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference

Edresson Casanova, Ryan Langman, Paarth Neekhara +5

Large language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modelin…

eess.AS2023

CML-TTS A Multilingual Dataset for Speech Synthesis in Low-Resource Languages

Frederico S. Oliveira, Edresson Casanova, Arnaldo Cândido Júnior +2

In this paper, we present CML-TTS, a recursive acronym for CML-Multi-Lingual-TTS, a new Text-to-Speech (TTS) dataset developed at the Center of Excellence in Artificial Intelligenc…