17 citations · 20 across the 4 of their papers we have counts for
5 papers
One model to enhance them all: array geometry agnostic multi-channel personalized speech enhancement
Hassan Taherian, Sefik Emre Eskimez, Takuya Yoshioka +3
With the recent surge of video conferencing tools usage, providing high-quality speech signals and accurate captions have become essential to conduct day-to-day business or connect…
Personalized Speech Enhancement: New Models and Comprehensive Evaluation
Sefik Emre Eskimez, Takuya Yoshioka, Huaming Wang +3
Personalized speech enhancement (PSE) models utilize additional cues, such as speaker embeddings like d-vectors, to remove background noise and interfering speech in real-time and…
Human Listening and Live Captioning: Multi-Task Training for Speech Enhancement
Sefik Emre Eskimez, Xiaofei Wang, Min Tang +5
With the surge of online meetings, it has become more critical than ever to provide high-quality speech audio and live captioning under various noise conditions. However, most mona…
Advances in Online Audio-Visual Meeting Transcription
Takuya Yoshioka, Igor Abramovski, Cem Aksoylar +23
This paper describes a system that generates speaker-annotated transcripts of meetings by using a microphone array and a 360-degree camera. The hallmark of the system is its abilit…
Cracking the cocktail party problem by multi-beam deep attractor network
Zhuo Chen, Jinyu Li, Xiong Xiao +4
While recent progresses in neural network approaches to single-channel speech separation, or more generally the cocktail party problem, achieved significant improvement, their perf…