10 citations · 10 across the 2 of their papers we have counts for
3 papers
Factorized RVQ-GAN For Disentangled Speech Tokenization
Sameer Khurana, Dominik Klement, Antoine Laurent +13
We propose Hierarchical Audio Codec (HAC), a unified neural speech codec that factorizes its bottleneck into three linguistic levels-acoustic, phonetic, and lexical-within a single…
LFAR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models
Santiago Cuervo, Adel Moumen, Yanis Labrak +5
Text-pretrained language models (LMs) encode rich world knowledge, but adapting them to process and generate perceptual modalities such as audio and images while effectively levera…
Direct Text to Speech Translation System using Acoustic Units
Victoria Mingote, Pablo Gimeno, Luis Vicente +3
This paper proposes a direct text to speech translation system using discrete acoustic units. This framework employs text in different source languages as input to generate speech…