1 citations · 1 across the 3 of their papers we have counts for
4 papers · 1 filter
Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
Xiaoyu Liu, Xu Li, Joan Serrà +1
Speech restoration aims at restoring full-band speech with high quality and intelligibility, considering a diverse set of distortions. MaskSR is a recently proposed generative mode…
Masked Generative Video-to-Audio Transformers with Enhanced Synchronicity
Santiago Pascual, Chunghsin Yeh, Ioannis Tsiamas +1
Video-to-audio (V2A) generation leverages visual-only video features to render plausible sounds that match the scene. Importantly, the generated sound onsets should match the visua…
Sequential Contrastive Audio-Visual Learning
Ioannis Tsiamas, Santiago Pascual, Chunghsin Yeh +1
Contrastive learning has emerged as a powerful technique in audio-visual representation learning, leveraging the natural co-occurrence of audio and visual modalities in webscale vi…
GASS: Generalizing Audio Source Separation with Large-scale Data
Jordi Pons, Xiaoyu Liu, Santiago Pascual +1
Universal source separation targets at separating the audio sources of an arbitrary mix, removing the constraint to operate on a specific domain like speech or music. Yet, the pote…