activity
20182025
most citedRecent Developments on ESPnet Toolkit Boosted by Conformer

40 citations · 126 across the 32 of their papers we have counts for

collaborators
Showing eess.ASShow all

15 papers · 1 filter

eess.AS2025

Summary on The Multilingual Conversational Speech Language Model Challenge: Datasets, Tasks, Baselines, and Methods

Bingshen Mu, Pengcheng Guo, Zhaokai Sun +8

This paper summarizes the Interspeech2025 Multilingual Conversational Speech Language Model (MLC-SLM) challenge, which aims to advance the exploration of building effective multili…

eess.AS2024

NPU-NTU System for Voice Privacy 2024 Challenge

Jixun Yao, Nikita Kuzmin, Qing Wang +6

Speaker anonymization is an effective privacy protection solution that conceals the speaker's identity while preserving the linguistic content and paralinguistic information of the…

eess.AS2024

MUSA: Multi-lingual Speaker Anonymization via Serial Disentanglement

Jixun Yao, Qing Wang, Pengcheng Guo +4

Speaker anonymization is an effective privacy protection solution designed to conceal the speaker's identity while preserving the linguistic content and para-linguistic information…

eess.AS2024

Distinctive and Natural Speaker Anonymization via Singular Value Transformation-assisted Matrix

Jixun Yao, Qing Wang, Pengcheng Guo +2

Speaker anonymization is an effective privacy protection solution that aims to conceal the speaker's identity while preserving the naturalness and distinctiveness of the original s…

eess.AS20241 cited

The NPU-ASLP-LiAuto System Description for Visual Speech Recognition in CNVSRC 2023

He Wang, Pengcheng Guo, Wei Chen +2

This paper delineates the visual speech recognition (VSR) system introduced by the NPU-ASLP-LiAuto (Team 237) in the first Chinese Continuous Visual Speech Recognition Challenge (C…

eess.AS2023

TranUSR: Phoneme-to-word Transcoder Based Unified Speech Representation Learning for Cross-lingual Speech Recognition

Hongfei Xue, Qijie Shao, Peikun Chen +3

UniSpeech has achieved superior performance in cross-lingual automatic speech recognition (ASR) by explicitly aligning latent representations to phoneme units using multi-task self…