activity
20182022
most citedImproving Readability for Automatic Speech Recognition Transcription

19 citations · 30 across the 12 of their papers we have counts for

collaborators

13 papers

eess.AS2022

Speech separation with large-scale self-supervised learning

Zhuo Chen, Naoyuki Kanda, Jian Wu +6

Self-supervised learning (SSL) methods such as WavLM have shown promising speech separation (SS) results in small-scale simulation-based experiments. In this work, we extend the ex…

eess.AS2022

Breaking the trade-off in personalized speech enhancement with cross-task knowledge distillation

Hassan Taherian, Sefik Emre Eskimez, Takuya Yoshioka

Personalized speech enhancement (PSE) models achieve promising results compared with unconditional speech enhancement models due to their ability to remove interfering speech in ad…

eess.AS2022

Leveraging Real Conversational Data for Multi-Channel Continuous Speech Separation

Xiaofei Wang, Dongmei Wang, Naoyuki Kanda +2

Existing multi-channel continuous speech separation (CSS) models are heavily dependent on supervised data - either simulated data which causes data mismatch between the training an…

eess.AS2022

Fast Real-time Personalized Speech Enhancement: End-to-End Enhancement Network (E3Net) and Knowledge Distillation

Manthan Thakker, Sefik Emre Eskimez, Takuya Yoshioka +1

This paper investigates how to improve the runtime speed of personalized speech enhancement (PSE) networks while maintaining the model quality. Our approach includes two aspects: a…

eess.AS2021

One model to enhance them all: array geometry agnostic multi-channel personalized speech enhancement

Hassan Taherian, Sefik Emre Eskimez, Takuya Yoshioka +3

With the recent surge of video conferencing tools usage, providing high-quality speech signals and accurate captions have become essential to conduct day-to-day business or connect…

eess.AS20211 cited

Personalized Speech Enhancement: New Models and Comprehensive Evaluation

Sefik Emre Eskimez, Takuya Yoshioka, Huaming Wang +3

Personalized speech enhancement (PSE) models utilize additional cues, such as speaker embeddings like d-vectors, to remove background noise and interfering speech in real-time and…