most citedWhisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models

1 citations · 1 across the 3 of their papers we have counts for

collaborators

5 papers

eess.AS2025

Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing

Mengqi Wang, Zhan Liu, Zengrui Jin +3

Diffusion-based large language models (DLLMs) have recently attracted growing interest as an alternative to autoregressive decoders. In this work, we present an empirical study on…

cs.SD2025

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification

Yiyang Zhao, Shuai Wang, Guangzhi Sun +4

Short-utterance speaker verification presents significant challenges due to the limited information in brief speech segments, which can undermine accuracy and reliability. Recently…

cs.SD2025

Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition and Translation

Qiuming Zhao, Guangzhi Sun, Chao Zhang

Language diversity presents a significant challenge in speech-to-text (S2T) tasks, such as automatic speech recognition and translation. Traditional multi-lingual multi-task traini…

eess.AS2024

SOT Triggered Neural Clustering for Speaker Attributed ASR

Xianrui Zheng, Guangzhi Sun, Chao Zhang +1

This paper introduces a novel approach to speaker-attributed ASR transcription using a neural clustering method. With a parallel processing mechanism, diarisation and ASR can be ap…

cs.SD20241 cited

Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models

Yiyang Zhao, Shuai Wang, Guangzhi Sun +4

In this paper, Whisper, a large-scale pre-trained model for automatic speech recognition, is proposed to apply to speaker verification. A partial multi-scale feature aggregation (P…