activity
20242026
most citedWhisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models

1 citations · 1 across the 5 of their papers we have counts for

collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD2026

HoliAntiSpoof: Audio LLM for Holistic Speech Anti-Spoofing

Xuenan Xu, Yiming Ren, Liwei Liu +5

Recent advances in speech synthesis and editing have made speech spoofing increasingly challenging. However, most existing methods treat spoofing as binary classification, overlook…

cs.SD2026

MMEDIT: A Unified Framework for Multi-Type Audio Editing via Audio Language Model

Ye Tao, Wen Wu, Chao Zhang +3

Text-guided audio editing aims to modify specific acoustic events while strictly preserving non-target content. Despite recent progress, existing approaches remain fundamentally li…

cs.SD2025

Can Audio Large Language Models Verify Speaker Identity?

Yiming Ren, Xuenan Xu, Baoxiang Li +2

This paper investigates adapting Audio Large Language Models (ALLMs) for speaker verification (SV). We reformulate SV as an audio question-answering task and conduct comprehensive…

cs.SD2025

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification

Yiyang Zhao, Shuai Wang, Guangzhi Sun +4

Short-utterance speaker verification presents significant challenges due to the limited information in brief speech segments, which can undermine accuracy and reliability. Recently…

cs.SD2024★ 1 cited

Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models

Yiyang Zhao, Shuai Wang, Guangzhi Sun +4

In this paper, Whisper, a large-scale pre-trained model for automatic speech recognition, is proposed to apply to speaker verification. A partial multi-scale feature aggregation (P…