activity
20192022
most citedGraph-PIT: Generalized permutation invariant training for continuous separation of arbitrary numbers of speakers

20 citations · 32 across the 6 of their papers we have counts for

collaborators

8 papers

eess.AS2022

MMS-MSG: A Multi-purpose Multi-Speaker Mixture Signal Generator

Tobias Cord-Landwehr, Thilo von Neumann, Christoph Boeddeker +1

The scope of speech enhancement has changed from a monolithic view of single, independent tasks, to a joint processing of complex conversational speech recordings. Training and eva…

eess.AS20222 cited

A Meeting Transcription System for an Ad-Hoc Acoustic Sensor Network

Tobias Gburrek, Christoph Boeddeker, Thilo von Neumann +3

We propose a system that transcribes the conversation of a typical meeting scenario that is captured by a set of initially unsynchronized microphone arrays at unknown positions. It…

cs.SD20221 cited

An Initialization Scheme for Meeting Separation with Spatial Mixture Models

Christoph Boeddeker, Tobias Cord-Landwehr, Thilo von Neumann +1

Spatial mixture model (SMM) supported acoustic beamforming has been extensively used for the separation of simultaneously active speakers. However, it has hardly been considered fo…

eess.AS20216 cited

Speeding Up Permutation Invariant Training for Source Separation

Thilo von Neumann, Christoph Boeddeker, Keisuke Kinoshita +2

Permutation invariant training (PIT) is a widely used training criterion for neural network-based source separation, used for both utterance-level separation with utterance-level P…

eess.AS202120 cited

Graph-PIT: Generalized permutation invariant training for continuous separation of arbitrary numbers of speakers

Thilo von Neumann, Keisuke Kinoshita, Christoph Boeddeker +2

Automatic transcription of meetings requires handling of overlapped speech, which calls for continuous speech separation (CSS) systems. The uPIT criterion was proposed for utteranc…

eess.AS20203 cited

Multi-path RNN for hierarchical modeling of long sequential data and its application to speaker stream separation

Keisuke Kinoshita, Thilo von Neumann, Marc Delcroix +2

Recently, the source separation performance was greatly improved by time-domain audio source separation based on dual-path recurrent neural network (DPRNN). DPRNN is a simple but e…