5 papers · 1 filter
Balancing ASR and diarization in end-to-end LLMs for multi-talker speech recognition
Naijun Zheng, Yuke Lin, Sanli Tian +4
Multi-talker speech recognition is often addressed by combining automatic speech recognition (ASR) and speaker diarization in a pipeline system. Recently, LLM-based approaches have…
The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge
Yuke Lin, Ming Cheng, Ze Li +1
We present the DKU system for Task 2 of the MLC-SLM Challenge, which aims to perform multi-speaker automatic speech recognition directly from raw audio without Oracle speaker label…
Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection and Representation
Ming Cheng, Yuke Lin, Ming Li
This paper proposes a novel Sequence-to-Sequence Neural Diarization (S2SND) framework to perform online and offline speaker diarization. It is developed from the sequence-to-sequen…
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
Yuke Lin, Ming Cheng, Ze Li +2
Multi-speaker automatic speech recognition (MS-ASR) faces significant challenges in transcribing overlapped speech, a task critical for applications like meeting transcription and…
The Database and Benchmark for the Source Speaker Tracing Challenge 2024
Ze Li, Yuke Lin, Tian Yao +6
Voice conversion (VC) systems can transform audio to mimic another speaker's voice, thereby attacking speaker verification (SV) systems. However, ongoing studies on source speaker…