3 papers
cs.SD2025
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
Téo Guichoux, Théodor Lemerle, Shivam Mehta +5
Human communication is multimodal, with speech and gestures tightly coupled, yet most computational methods for generating speech and gestures synthesize them sequentially, weakeni…
eess.AS2024
Lina-Speech: Gated Linear Attention and Initial-State Tuning for Multi-Sample Prompting Text-To-Speech Synthesis
Théodor Lemerle, Téo Guichoux, Axel Roebel +1
Neural codec language models, built on transformer architecture, have revolutionized text-to-speech (TTS) synthesis, excelling in voice cloning by treating it as a prefix continuat…
eess.AS2024
Small-E: Small Language Model with Linear Attention for Efficient Speech Synthesis
Théodor Lemerle, Nicolas Obin, Axel Roebel
Recent advancements in text-to-speech (TTS) powered by language models have showcased remarkable capabilities in achieving naturalness and zero-shot voice cloning. Notably, the dec…