activity
20192024
most citedReal-time Speaker counting in a cocktail party scenario using Attention-guided Convolutional Neural Network

1 citations · 1 across the 5 of their papers we have counts for

collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS2024

CoVoMix: Advancing Zero-Shot Speech Generation for Human-like Multi-talker Conversations

Leying Zhang, Yao Qian, Long Zhou +9

Recent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech. However, dialogue generation, along w…

eess.AS2021

Speaker conditioning of acoustic models using affine transformation for multi-speaker speech recognition

Midia Yousefi, John H. L. Hanse

This study addresses the problem of single-channel Automatic Speech Recognition of a target speaker within an overlap speech scenario. In the proposed method, the hidden representa…

eess.AS20211 cited

Real-time Speaker counting in a cocktail party scenario using Attention-guided Convolutional Neural Network

Midia Yousefi, John H. L. Hansen

Most current speech technology systems are designed to operate well even in the presence of multiple active speakers. However, most solutions assume that the number of co-current s…

eess.AS2020

Frame-based overlapping speech detection using Convolutional Neural Networks

Midia Yousefi, John H. L. Hansen

Naturalistic speech recordings usually contain speech signals from multiple speakers. This phenomenon can degrade the performance of speech technologies due to the complexity of tr…

eess.AS2019

Probabilistic Permutation Invariant Training for Speech Separation

Midia Yousefi, Soheil Khorram, John H. L. Hansen

Single-microphone, speaker-independent speech separation is normally performed through two steps: (i) separating the specific speech sources, and (ii) determining the best output-l…