Showing cs.SDShow all
2 papers · 1 filter
cs.SD2026
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
Téo Guichoux, Théodor Lemerle, Shivam Mehta +5
Human communication is multimodal, with speech and gestures tightly coupled, yet most computational methods for generating speech and gestures synthesize them sequentially, weakeni…
cs.SD2025
Learning Relationships Between Separate Audio Tracks for Creative Applications
Balthazar Bujard, Jérôme Nika, Fédéric Bevilacqua +1
This paper presents the first step in a research project situated within the field of musical agents. The objective is to achieve, through training, the tuning of the desired music…