49 citations · 72 across the 24 of their papers we have counts for
15 papers · 1 filter
Text-Prompted CLAP: Learning Text-Conditioned Audio Representations via Contrastive Learning
Mohan Li, Rama Doddipatla, Philip C. Woodland
Contrastive Language-Audio Pretraining (CLAP) aligns text and audio in a shared embedding space, but encoding each modality independently limits its ability to model cross-modal se…
WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding
Mohan Li, Cong-Thanh Do, Simon Keizer +3
Speech large language models (speech-LLMs) integrate speech and text-based foundation models to provide a unified framework for handling a wide range of downstream tasks. In this p…
Prompting Whisper for QA-driven Zero-shot End-to-end Spoken Language Understanding
Mohan Li, Simon Keizer, Rama Doddipatla
Zero-shot spoken language understanding (SLU) enables systems to comprehend user utterances in new domains without prior exposure to training data. Recent studies often rely on lar…
Geodesic interpolation of frame-wise speaker embeddings for the diarization of meeting scenarios
Tobias Cord-Landwehr, Christoph Boeddeker, Cătălin Zorilă +2
We propose a modified teacher-student training for the extraction of frame-wise speaker embeddings that allows for an effective diarization of meeting scenarios containing partiall…
A Teacher-Student approach for extracting informative speaker embeddings from speech mixtures
Tobias Cord-Landwehr, Christoph Boeddeker, Cătălin Zorilă +2
We introduce a monaural neural speaker embeddings extractor that computes an embedding for each speaker present in a speech mixture. To allow for supervised training, a teacher-stu…
Frame-wise and overlap-robust speaker embeddings for meeting diarization
Tobias Cord-Landwehr, Christoph Boeddeker, Cătălin Zorilă +2
Using a Teacher-Student training approach we developed a speaker embedding extraction system that outputs embeddings at frame rate. Given this high temporal resolution and the fact…