4 citations · 7 across the 7 of their papers we have counts for
Showing cs.SDShow all
2 papers · 1 filter
cs.SD2024
Self-supervised Pretraining for Robust Personalized Voice Activity Detection in Adverse Conditions
Holger Severin Bovbjerg, Jesper Jensen, Jan Østergaard +1
In this paper, we propose the use of self-supervised pretraining on a large unlabelled data set to improve the performance of a personalized voice activity detection (VAD) model in…
cs.SD2023★ 1 cited
Speech inpainting: Context-based speech synthesis guided by video
Juan F. Montesinos, Daniel Michelsanti, Gloria Haro +2
Audio and visual modalities are inherently connected in speech signals: lip movements and facial expressions are correlated with speech sounds. This motivates studies that incorpor…