54 citations · 166 across the 33 of their papers we have counts for
65 papers
Joint Minimum Processing Beamforming and Near-end Listening Enhancement
Andreas J. Fuglsig, Jesper Jensen, Zheng-Hua Tan +3
We consider speech enhancement for signals picked up in one noisy environment that must be rendered to a listener in another noisy environment. For both far-end noise reduction and…
Masked Autoencoders with Multi-Window Local-Global Attention Are Better Audio Learners
Sarthak Yadav, Sergios Theodoridis, Lars Kai Hansen +1
In this work, we propose a Multi-Window Masked Autoencoder (MW-MAE) fitted with a novel Multi-Window Multi-Head Attention (MW-MHA) module that facilitates the modelling of local-gl…
Speech inpainting: Context-based speech synthesis guided by video
Juan F. Montesinos, Daniel Michelsanti, Gloria Haro +2
Audio and visual modalities are inherently connected in speech signals: lip movements and facial expressions are correlated with speech sounds. This motivates studies that incorpor…
PAC-Bayesian bounds for learning LTI-ss systems with input from empirical loss
Deividas Eringis, John Leth, Zheng-Hua Tan +2
In this paper we derive a Probably Approxilmately Correct(PAC)-Bayesian error bound for linear time-invariant (LTI) stochastic dynamical systems with inputs. Such bounds are widesp…
PAC-Bayesian-Like Error Bound for a Class of Linear Time-Invariant Stochastic State-Space Models
Deividas Eringis, John Leth, Zheng-Hua Tan +2
In this paper we derive a PAC-Bayesian-Like error bound for a class of stochastic dynamical systems with inputs, namely, for linear time-invariant stochastic state-space models (st…
Leveraging Domain Features for Detecting Adversarial Attacks Against Deep Speech Recognition in Noise
Christian Heider Nielsen, Zheng-Hua Tan
In recent years, significant progress has been made in deep model-based automatic speech recognition (ASR), leading to its widespread deployment in the real world. At the same time…