5 papers
AMDM-SE: Attention-based Multichannel Diffusion Model for Speech Enhancement
Renana Opochinsky, Sharon Gannot
Diffusion models have recently achieved impressive results in reconstructing images from noisy inputs, and similar ideas have been applied to speech enhancement by treating time-fr…
Socially Pertinent Robots in Gerontological Healthcare
Xavier Alameda-Pineda, Angus Addlesee, Daniel Hernández GarcÃa +41
Despite the many recent achievements in developing and deploying social robotics, there are still many underexplored environments and applications for which systematic evaluation o…
Transient Noise Removal via Diffusion-based Speech Inpainting
Mordehay Moradi, Sharon Gannot
In this paper, we present PGDI, a diffusion-based speech inpainting framework for restoring missing or severely corrupted speech segments. Unlike previous methods that struggle wit…
Multi-Microphone and Multi-Modal Emotion Recognition in Reverberant Environment
Ohad Cohen, Gershon Hazan, Sharon Gannot
This paper presents a Multi-modal Emotion Recognition (MER) system designed to enhance emotion recognition accuracy in challenging acoustic conditions. Our approach combines a modi…
Audio-Visual Approach For Multimodal Concurrent Speaker Detection
Amit Eliav, Sharon Gannot
Concurrent Speaker Detection (CSD), the task of identifying active speakers and their overlaps in an audio signal, is essential for various audio applications, including meeting tr…