activity
20162023
most citedA Hybrid RNN-HMM Approach for Weakly Supervised Temporal Action Segmentation

99 citations · 236 across the 12 of their papers we have counts for

collaborators
Showing 2022Show all

7 papers · 1 filter

cs.SD2022

End-to-End Binaural Speech Synthesis

Wen Chin Huang, Dejan Markovic, Alexander Richard +2

In this work, we present an end-to-end binaural speech synthesis system that combines a low-bitrate audio codec with a powerful binaural decoder that is capable of accurate speech…

cs.CV2022★ 50 cited

Multiface: A Dataset for Neural Face Rendering

Cheng-hsin Wuu, Ningyuan Zheng, Scott Ardisson +28

Photorealistic avatars of human faces have come a long way in recent years, yet research along this area is limited by a lack of publicly available, high-quality datasets covering…

cs.SD2022

Implicit Neural Spatial Filtering for Multichannel Source Separation in the Waveform Domain

Dejan Markovic, Alexandre Defossez, Alexander Richard

We present a single-stage casual waveform-to-waveform multichannel model that can separate moving sound sources based on their broad spatial locations in a dynamic acoustic scene.…

cs.CV2022

Audio-Visual Speech Codecs: Rethinking Audio-Visual Speech Enhancement by Re-Synthesis

Karren Yang, Dejan Markovic, Steven Krenn +2

Since facial actions such as lip movements contain significant information about speech content, it is not surprising that audio-visual speech enhancement methods are more accurate…

cs.CV2022

LiP-Flow: Learning Inference-time Priors for Codec Avatars via Normalizing Flows in Latent Space

Emre Aksan, Shugao Ma, Akin Caliskan +5

Neural face avatars that are trained from multi-view data captured in camera domes can produce photo-realistic 3D reconstructions. However, at inference time, they must be driven b…

eess.AS2022

Conditional Diffusion Probabilistic Model for Speech Enhancement

Yen-Ju Lu, Zhong-Qiu Wang, Shinji Watanabe +3

Speech enhancement is a critical component of many user-oriented audio applications, yet current systems still suffer from distorted and unnatural outputs. While generative models…