activity
20192022
most citedAsteroid: the PyTorch-based audio source separation toolkit for researchers

51 citations · 52 across the 5 of their papers we have counts for

collaborators

6 papers

eess.AS2022

Speech separation with large-scale self-supervised learning

Zhuo Chen, Naoyuki Kanda, Jian Wu +6

Self-supervised learning (SSL) methods such as WavLM have shown promising speech separation (SS) results in small-scale simulation-based experiments. In this work, we extend the ex…

eess.AS20221 cited

Simulating realistic speech overlaps improves multi-talker ASR

Muqiao Yang, Naoyuki Kanda, Xiaofei Wang +5

Multi-talker automatic speech recognition (ASR) has been studied to generate transcriptions of natural conversation including overlapping speech of multiple speakers. Due to the di…

eess.AS202051 cited

Asteroid: the PyTorch-based audio source separation toolkit for researchers

Manuel Pariente, Samuele Cornell, Joris Cosentino +11

This paper describes Asteroid, the PyTorch-based audio source separation toolkit for researchers. Inspired by the most successful neural source separation systems, it provides all…

eess.AS2019

The Speed Submission to DIHARD II: Contributions & Lessons Learned

Md Sahidullah, Jose Patino, Samuele Cornell +11

This paper describes the speaker diarization systems developed for the Second DIHARD Speech Diarization Challenge (DIHARD II) by the Speed team. Besides describing the system, whic…

eess.AS2019

SLOGD: Speaker LOcation Guided Deflation approach to speech separation

Sunit Sivasankaran, Emmanuel Vincent, Dominique Fohr

Speech separation is the process of separating multiple speakers from an audio recording. In this work we propose to separate the sources using a Speaker LOcalization Guided Deflat…

eess.AS2019

Analyzing the impact of speaker localization errors on speech separation for automatic speech recognition

Sunit Sivasankaran, Emmaneul Vincent, Dominique Fohr

We investigate the effect of speaker localization on the performance of speech recognition systems in a multispeaker, multichannel environment. Given the speaker location informati…