activity
20242026
most citedESPnet-SpeechLM: An Open Speech Language Model Toolkit

1 citations · 1 across the 7 of their papers we have counts for

collaborators
Showing 2025Show all

7 papers · 1 filter

eess.AS2025

BUT Systems for WildSpoof Challenge: SASV in the Wild

Junyi Peng, Jin Li, Johan Rohdin +3

This paper presents the BUT submission to the WildSpoof Challenge, focusing on the Spoofing-robust Automatic Speaker Verification (SASV) track. We propose a SASV framework designed…

eess.AS2025

BUT Systems for Environmental Sound Deepfake Detection in the ESDD 2026 Challenge

Junyi Peng, Lin Zhang, Jin Li +2

This paper describes the BUT submission to the ESDD 2026 Challenge, specifically focusing on Track 1: Environmental Sound Deepfake Detection with Unseen Generators. To address the…

eess.AS2025

Hybrid Pruning: In-Situ Compression of Self-Supervised Speech Models for Speaker Verification and Anti-Spoofing

Junyi Peng, Lin Zhang, Jiangyu Han +5

Although large-scale self-supervised learning (SSL) models like WavLM have achieved state-of-the-art performance in speech processing, their significant size impedes deployment on…

eess.AS2025

Analysis of ABC Frontend Audio Systems for the NIST-SRE24

Sara Barahona, Anna Silnova, Ladislav Mošner +14

We present a comprehensive analysis of the embedding extractors (frontends) developed by the ABC team for the audio track of NIST SRE 2024. We follow the two scenarios imposed by N…

cs.CL2025

TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models

Junyi Peng, Takanori Ashihara, Marc Delcroix +4

Self-supervised learning (SSL) models have significantly advanced speech processing tasks, and several benchmarks have been proposed to validate their effectiveness. However, previ…

eess.AS2025

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

Helin Wang, Jiarui Hai, Dongchao Yang +7

Target Speech Extraction (TSE) aims to isolate a target speaker's voice from a mixture of multiple speakers by leveraging speaker-specific cues, typically provided as auxiliary aud…