4 papers
Speech-DRAME: A Framework for Human-Aligned Benchmarks in Speech Role-Play
Jiatong Shi, Jionghao Han, Yichen Lu +6
Role-play has become a key testbed for generative models, expanding from text-only dialogue to multimodal interaction. Extending role-play to speech captures prosody, emotion, and…
Sequential Contrastive Audio-Visual Learning
Ioannis Tsiamas, Santiago Pascual, Chunghsin Yeh +1
Contrastive learning has emerged as a powerful technique in audio-visual representation learning, leveraging the natural co-occurrence of audio and visual modalities in webscale vi…
Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
Xiaoyu Liu, Xu Li, Joan Serrà +1
Speech restoration aims at restoring full-band speech with high quality and intelligibility, considering a diverse set of distortions. MaskSR is a recently proposed generative mode…
Masked Generative Video-to-Audio Transformers with Enhanced Synchronicity
Santiago Pascual, Chunghsin Yeh, Ioannis Tsiamas +1
Video-to-audio (V2A) generation leverages visual-only video features to render plausible sounds that match the scene. Importantly, the generated sound onsets should match the visua…