6 papers
SSNAPS: Audio-Visual Separation of Speech and Background Noise with Diffusion Inverse Sampling
Yochai Yemini, Yoav Ellinson, Rami Ben-Ari +2
This paper addresses the challenge of audio-visual single-microphone speech separation and enhancement in the presence of real-world environmental noise. Our approach is based on g…
On the Usefulness of Diffusion-Based Room Impulse Response Interpolation to Microphone Array Processing
Sagi Della Torre, Mirco Pezzoli, Fabio Antonacci +1
Room Impulse Responses estimation is a fundamental problem in spatial audio processing and speech enhancement. In this paper, we build upon our previously introduced diffusion-base…
Spectral or spatial? Leveraging both for speaker extraction in challenging data conditions
Aviad Eisenberg, Sharon Gannot, Shlomo E. Chazan
This paper presents a robust multi-channel speaker extraction algorithm designed to handle inaccuracies in reference information. While existing approaches often rely solely on eit…
Diffusion-Based Unsupervised Audio-Visual Speech Separation in Noisy Environments with Noise Prior
Yochai Yemini, Rami Ben-Ari, Sharon Gannot +1
In this paper, we address the problem of single-microphone speech separation in the presence of ambient noise. We propose a generative unsupervised technique that directly models b…
Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes
Neta Glazer, David Chernin, Idan Achituve +2
Recent advancements in Text-to-Speech (TTS) models, particularly in voice cloning, have intensified the demand for adaptable and efficient deepfake detection methods. As TTS system…
End-to-End Multi-Microphone Speaker Extraction Using Relative Transfer Functions
Aviad Eisenberg, Sharon Gannot, Shlomo E. Chazan
This paper introduces a multi-microphone method for extracting a desired speaker from a mixture involving multiple speakers and directional noise in a reverberant environment. In t…