NewEvery arXiv paper, its researchers & institutions — mapped.
papers

Publications (47)

eess.AS2025

ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech

Xin Wang, Héctor Delgado, Hemlata Tak +26

eess.AS2021

Closing the Gap Between Time-Domain Multi-Channel Speech Enhancement on Real and Simulation Conditions

Wangyou Zhang, Jing Shi, Chenda Li +2

eess.AS2025

Text-To-Speech Synthesis In The Wild

Jee-weon Jung, Wangyou Zhang, Soumi Maiti +11

cs.CL2024

Towards Robust Speech Representation Learning for Thousands of Languages

William Chen, Wangyou Zhang, Yifan Peng +7

eess.AS2020

ESPnet-se: end-to-end speech enhancement and separation toolkit designed for asr integration

Chenda Li, Jing Shi, Wangyou Zhang +8

eess.AS2020

Recent Developments on ESPnet Toolkit Boosted by Conformer

Pengcheng Guo, Florian Boyer, Xuankai Chang +12

eess.AS2025

P.808 Multilingual Speech Enhancement Testing: Approach and Results of URGENT 2025 Challenge

Marvin Sach, Yihui Fu, Kohei Saijo +9

eess.AS2020

The 2020 ESPnet update: new features, broadened applications, performance improvements, and future plans

Shinji Watanabe, Florian Boyer, Xuankai Chang +12

eess.AS2025

Less is More: Data Curation Matters in Scaling Speech Enhancement

Chenda Li, Wangyou Zhang, Wei Wang +10

eess.AS2019

MIMO-SPEECH: End-to-End Multi-Channel Multi-Speaker Speech Recognition

Xuankai Chang, Wangyou Zhang, Yanmin Qian +2

eess.AS2024

Scale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling

Leying Zhang, Wangyou Zhang, Chenda Li +1

eess.AS2021

Separating Long-Form Speech with Group-Wise Permutation Invariant Training

Wangyou Zhang, Zhuo Chen, Naoyuki Kanda +8

eess.AS2020

End-to-End Multi-speaker Speech Recognition with Transformer

Xuankai Chang, Wangyou Zhang, Yanmin Qian +2

eess.AS2025

URGENT-PK: Perceptually-Aligned Ranking Model Designed for Speech Enhancement Competition

Jiahe Wang, Chenda Li, Wei Wang +11

eess.AS2022

End-to-End Multi-speaker ASR with Independent Vector Analysis

Robin Scheibler, Wangyou Zhang, Xuankai Chang +2

cs.SD2025

VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music

Jiatong Shi, Hye-jin Shim, Jinchuan Tian +14

eess.AS2024

Beyond Performance Plateaus: A Comprehensive Study on Scalability in Speech Enhancement

Wangyou Zhang, Kohei Saijo, Jee-weon Jung +3

eess.AS2020

End-to-End Far-Field Speech Recognition with Unified Dereverberation and Beamforming

Wangyou Zhang, Aswin Shanmugam Subramanian, Xuankai Chang +2

cs.SD2023

Exploring the Integration of Speech Separation and Recognition with Self-Supervised Learning Representation

Yoshiki Masuyama, Xuankai Chang, Wangyou Zhang +5

eess.AS2025

MeanSE: Efficient Generative Speech Enhancement with Mean Flows

Jiahe Wang, Hongyu Wang, Wei Wang +5

eess.AS2025

Interspeech 2025 URGENT Speech Enhancement Challenge

Kohei Saijo, Wangyou Zhang, Samuele Cornell +9

cs.SD2026

TF-MoE: Time-Frequency Mixture-of-Experts for Efficient Speech Separation

Qinzhe Hu, Chenda Li, Wangyou Zhang +3

eess.AS2023

A Single Speech Enhancement Model Unifying Dereverberation, Denoising, Speaker Counting, Separation, and Extraction

Kohei Saijo, Wangyou Zhang, Zhong-Qiu Wang +3

eess.AS2024

URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement

Wangyou Zhang, Robin Scheibler, Kohei Saijo +9

eess.AS2021

End-to-End Dereverberation, Beamforming, and Speech Recognition with Improved Numerical Stability and Advanced Frontend

Wangyou Zhang, Christoph Boeddeker, Shinji Watanabe +7

eess.AS2024

Toward Universal Speech Enhancement for Diverse Input Conditions

Wangyou Zhang, Kohei Saijo, Zhong-Qiu Wang +2

cs.SD2025

Improving Speech Enhancement with Multi-Metric Supervision from Learned Quality Assessment

Wei Wang, Wangyou Zhang, Chenda Li +3

cs.CL2024

SpeechComposer: Unifying Multiple Speech Tasks with Prompt Composition

Yihan Wu, Soumi Maiti, Yifan Peng +6

eess.AS2026

Towards Array-Invariant Speech Enhancement via Geometry-Aware Dynamic Convolution

Zhenglong Liu, Wangyou Zhang, Chenda Li +1

eess.AS2024

Improving Design of Input Condition Invariant Speech Enhancement

Wangyou Zhang, Jee-weon Jung, Shinji Watanabe +1

eess.AS2026

ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era

Masao Someki, Alexander Polok, Carlos Carvalho +14

cs.SD2026

UrgentMOS: Unified Multi-Metric and Preference Learning for Robust Speech Quality Assessment

Wei Wang, Wangyou Zhang, Chenda Li +12

eess.AS2026

Representation-Regularized Convolutional Audio Transformer for Audio Understanding

Bing Han, Chushu Zhou, Yifan Yang +4

eess.AS2022

ESPnet-SE++: Speech Enhancement for Robust Speech Recognition, Translation, and Understanding

Yen-Ju Lu, Xuankai Chang, Chenda Li +10

cs.CL2023

Reproducing Whisper-Style Training Using an Open-Source Toolkit and Publicly Available Data

Yifan Peng, Jinchuan Tian, Brian Yan +13

eess.AS2025

Lessons Learned from the URGENT 2024 Speech Enhancement Challenge

Wangyou Zhang, Kohei Saijo, Samuele Cornell +10

cs.SD2025

SpoofCeleb: Speech Deepfake Detection and SASV In The Wild

Jee-weon Jung, Yihan Wu, Xin Wang +11

cs.CL2023

Joint Prediction and Denoising for Large-scale Multilingual Self-supervised Learning

William Chen, Jiatong Shi, Brian Yan +6

cs.SD2026

On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation

Changhao Cheng, Wei Wang, Wangyou Zhang +4

eess.AS2023

Weakly-Supervised Speech Pre-training: A Case Study on Target Speech Recognition

Wangyou Zhang, Yanmin Qian

cs.SD2025

Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction

Leying Zhang, Wangyou Zhang, Zhengyang Chen +1

eess.AS2022

Towards Low-distortion Multi-channel Speech Enhancement: The ESPNet-SE Submission to The L3DAS22 Challenge

Yen-Ju Lu, Samuele Cornell, Xuankai Chang +5

cs.SD2021

Convolutive Transfer Function Invariant SDR training criteria for Multi-Channel Reverberant Speech Separation

Christoph Boeddeker, Wangyou Zhang, Tomohiro Nakatani +6

cs.SD2025

PURE Codec: Progressive Unfolding of Residual Entropy for Speech Codec Learning

Jiatong Shi, Haoran Wang, William Chen +4

cs.SD2024

ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models

Jee-weon Jung, Wangyou Zhang, Jiatong Shi +5

cs.CL2019

A Comparative Study on Transformer vs RNN in Speech Applications

Shigeki Karita, Nanxin Chen, Tomoki Hayashi +10

eess.AS2026

ICASSP 2026 URGENT Speech Enhancement Challenge

Chenda Li, Wei Wang, Marvin Sach +8