activity
20192026
most citedOn End-to-end Multi-channel Time Domain Speech Separation in Reverberant Environments

49 citations · 72 across the 24 of their papers we have counts for

collaborators
Showing eess.ASShow all

15 papers · 1 filter

eess.AS2026

Text-Prompted CLAP: Learning Text-Conditioned Audio Representations via Contrastive Learning

Mohan Li, Rama Doddipatla, Philip C. Woodland

Contrastive Language-Audio Pretraining (CLAP) aligns text and audio in a shared embedding space, but encoding each modality independently limits its ability to model cross-modal se…

eess.AS2024

WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding

Mohan Li, Cong-Thanh Do, Simon Keizer +3

Speech large language models (speech-LLMs) integrate speech and text-based foundation models to provide a unified framework for handling a wide range of downstream tasks. In this p…

eess.AS2024

Prompting Whisper for QA-driven Zero-shot End-to-end Spoken Language Understanding

Mohan Li, Simon Keizer, Rama Doddipatla

Zero-shot spoken language understanding (SLU) enables systems to comprehend user utterances in new domains without prior exposure to training data. Recent studies often rely on lar…

eess.AS2024

Geodesic interpolation of frame-wise speaker embeddings for the diarization of meeting scenarios

Tobias Cord-Landwehr, Christoph Boeddeker, Cătălin Zorilă +2

We propose a modified teacher-student training for the extraction of frame-wise speaker embeddings that allows for an effective diarization of meeting scenarios containing partiall…

eess.AS2023

A Teacher-Student approach for extracting informative speaker embeddings from speech mixtures

Tobias Cord-Landwehr, Christoph Boeddeker, Cătălin Zorilă +2

We introduce a monaural neural speaker embeddings extractor that computes an embedding for each speaker present in a speech mixture. To allow for supervised training, a teacher-stu…

eess.AS2023

Frame-wise and overlap-robust speaker embeddings for meeting diarization

Tobias Cord-Landwehr, Christoph Boeddeker, Cătălin Zorilă +2

Using a Teacher-Student training approach we developed a speaker embedding extraction system that outputs embeddings at frame rate. Given this high temporal resolution and the fact…