4 papers
Continuous Audio Language Models
Simon Rouard, Manu Orsini, Axel Roebel +2
Audio Language Models (ALM) have emerged as the dominant paradigm for speech and music generation by representing audio as sequences of discrete tokens. Yet, unlike text tokens, wh…
Lina-Speech: Gated Linear Attention and Initial-State Tuning for Multi-Sample Prompting Text-To-Speech Synthesis
Théodor Lemerle, Téo Guichoux, Axel Roebel +1
Neural codec language models, built on transformer architecture, have revolutionized text-to-speech (TTS) synthesis, excelling in voice cloning by treating it as a prefix continuat…
PitchFlower: A flow-based neural audio codec with pitch controllability
Diego Torres, Axel Roebel, Nicolas Obin
We present PitchFlower, a flow-based neural audio codec with explicit pitch controllability. Our approach enforces disentanglement through a simple perturbation: during training, F…
Fast-VGAN: Lightweight Voice Conversion with Explicit Control of F0 and Duration Parameters
Mathilde Abrassart, Nicolas Obin, Axel Roebel
Precise control over speech characteristics, such as pitch, duration, and speech rate, remains a significant challenge in the field of voice conversion. The ability to manipulate p…