261 citations · 280 across the 2 of their papers we have counts for
5 papers
CDPAM: Contrastive learning for perceptual audio similarity
Pranay Manocha, Zeyu Jin, Richard Zhang +1
Many speech processing methods based on deep learning require an automatic and differentiable audio metric for the loss function. The DPAM approach of Manocha et al. learns a full-…
HiFi-GAN: High-Fidelity Denoising and Dereverberation Based on Speech Deep Features in Adversarial Networks
Jiaqi Su, Zeyu Jin, Adam Finkelstein
Real-world audio recordings are often degraded by factors such as noise, reverberation, and equalization distortion. This paper introduces HiFi-GAN, a deep learning method to trans…
A Differentiable Perceptual Audio Metric Learned from Just Noticeable Differences
Pranay Manocha, Adam Finkelstein, Richard Zhang +3
Many audio processing tasks require perceptual assessment. The ``gold standard`` of obtaining human judgments is time-consuming, expensive, and cannot be used as an optimization cr…
Text-based Editing of Talking-head Video
Ohad Fried, Ayush Tewari, Michael Zollhöfer +7
Editing talking-head video to change the speech content or to remove filler words is challenging. We propose a novel method to edit talking-head video based on its transcript to pr…
TurkerGaze: Crowdsourcing Saliency with Webcam based Eye Tracking
Pingmei Xu, Krista A Ehinger, Yinda Zhang +3
Traditional eye tracking requires specialized hardware, which means collecting gaze data from many observers is expensive, tedious and slow. Therefore, existing saliency prediction…