collaborators

5 papers

eess.AS2026

M2S-AVSR: Modality-aware Multi-view Self-supervised Representation for Robust Audio-Visual Speech Recognition

Fei Su, Cancan Li, Ming Li +1

Audio-Visual Speech Recognition (AVSR) enhances speech recognition robustness by leveraging visual cues, while real-world scenarios remain challenging due to viewpoint variation, a…

eess.AS2026

Enhancing Speaker Verification with w2v-BERT 2.0 and Knowledge Distillation guided Structured Pruning

Ze Li, Ming Cheng, Ming Li

Large-scale self-supervised Pre-Trained Models (PTMs) have shown significant improvements in the speaker verification (SV) task by providing rich feature representations. In this p…

eess.AS2025

The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge

Yuke Lin, Ming Cheng, Ze Li +1

We present the DKU system for Task 2 of the MLC-SLM Challenge, which aims to perform multi-speaker automatic speech recognition directly from raw audio without Oracle speaker label…

eess.AS2025

Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection and Representation

Ming Cheng, Yuke Lin, Ming Li

This paper proposes a novel Sequence-to-Sequence Neural Diarization (S2SND) framework to perform online and offline speaker diarization. It is developed from the sequence-to-sequen…

eess.AS2025

Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models

Yuke Lin, Ming Cheng, Ze Li +2

Multi-speaker automatic speech recognition (MS-ASR) faces significant challenges in transcribing overlapped speech, a task critical for applications like meeting transcription and…