activity
20212026
most citedCross-lingual Text Classification with Heterogeneous Graph Neural Network

2 citations · 3 across the 5 of their papers we have counts for

collaborators

7 papers

cs.SD2026

FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech

Shuoyi Zhou, Yixuan Zhou, Peiji Yang +4

Controllable text-to-speech (TTS) has become a key research focus. However, methods based on either reference speech or text descriptions lack flexibility and precise control, and…

eess.AS2026

TellWhisper: Tell Whisper Who Speaks When

Yifan Hu, Peiji Yang, Zhisheng Wang +2

Multi-speaker automatic speech recognition (MASR) aims to predict ''who spoke when and what'' from multi-speaker speech, a key technology for multi-party dialogue understanding. Ho…

cs.SD2025

HPSU: A Benchmark for Human-Level Perception in Real-World Spoken Speech Understanding

Chen Li, Peiji Yang, Yicheng Zhong +5

Recent advances in Speech Large Language Models (Speech LLMs) have led to great progress in speech understanding tasks such as Automatic Speech Recognition (ASR) and Speech Emotion…

cs.SD2025

Multi-Reward GRPO for Stable and Prosodic Single-Codebook TTS LLMs at Scale

Yicheng Zhong, Peiji Yang, Zhisheng Wang

Recent advances in Large Language Models (LLMs) have transformed text-to-speech (TTS) synthesis, inspiring autoregressive frameworks that represent speech as sequences of discrete…

cs.SD2024

Optimizing Neural Speech Codec for Low-Bitrate Compression via Multi-Scale Encoding

Peiji Yang, Fengping Wang, Yicheng Zhong +2

Neural speech codecs have demonstrated their ability to compress high-quality speech and audio by converting them into discrete token representations. Most existing methods utilize…

cs.SD20221 cited

AccentSpeech: Learning Accent from Crowd-sourced Data for Target Speaker TTS with Accents

Yongmao Zhang, Zhichao Wang, Peiji Yang +3

Learning accent from crowd-sourced data is a feasible way to achieve a target speaker TTS system that can synthesize accent speech. To this end, there are two challenging problems…