12 papers
The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction Dataset
Rajmund Nagy, Silvia Arellano García, Hendric Voss +4
This preprint presents the results of the fourth GENEA Challenge, a large-scale human evaluation of five speech-driven gesture-generation systems trained by participating teams on…
Causality for Tabular Data Synthesis: A High-Order Structure Causal Benchmark Framework
Zineb Senane, Axel Karlsson, Lele Cao +6
Existing evaluations of tabular synthesis models rely primarily on low-order statistics and downstream task performance, leaving multivariate causal relationships that go beyond pa…
Speak Your Mind: The Speech Continuation Task as a Probe of Voice-Based Model Bias
Shree Harsha Bokkahalli Satish, Harm Lameris, Olivier Perrotin +2
Speech Continuation (SC) is the task of generating a coherent extension of a spoken prompt while preserving both semantic context and speaker identity. Because SC is constrained to…
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
Téo Guichoux, Théodor Lemerle, Shivam Mehta +5
Human communication is multimodal, with speech and gestures tightly coupled, yet most computational methods for generating speech and gestures synthesize them sequentially, weakeni…
VoXtream2: Full-stream TTS with dynamic speaking rate control
Nikita Torgashov, Gustav Eje Henter, Gabriel Skantze
Full-stream text-to-speech (TTS) for interactive systems must start speaking with minimal delay while remaining controllable as text arrives incrementally. We present VoXtream2, a…
HiFi-Glot: High-Fidelity Neural Formant Synthesis with Differentiable Resonant Filters
Yicheng Gu, Pablo Pérez Zarazaga, Chaoren Wang +4
Formant synthesis aims to generate speech with controllable formant structures, enabling precise control of vocal resonance and phonetic features. However, while existing formant s…