4 papers
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
Téo Guichoux, Théodor Lemerle, Shivam Mehta +5
Human communication is multimodal, with speech and gestures tightly coupled, yet most computational methods for generating speech and gestures synthesize them sequentially, weakeni…
Lina-Speech: Gated Linear Attention and Initial-State Tuning for Multi-Sample Prompting Text-To-Speech Synthesis
Théodor Lemerle, Téo Guichoux, Axel Roebel +1
Neural codec language models, built on transformer architecture, have revolutionized text-to-speech (TTS) synthesis, excelling in voice cloning by treating it as a prefix continuat…
PitchFlower: A flow-based neural audio codec with pitch controllability
Diego Torres, Axel Roebel, Nicolas Obin
We present PitchFlower, a flow-based neural audio codec with explicit pitch controllability. Our approach enforces disentanglement through a simple perturbation: during training, F…
Learning Relationships Between Separate Audio Tracks for Creative Applications
Balthazar Bujard, Jérôme Nika, Fédéric Bevilacqua +1
This paper presents the first step in a research project situated within the field of musical agents. The objective is to achieve, through training, the tuning of the desired music…