activity
20222025
most citedMSAC: Multiple Speech Attribute Control Method for Reliable Speech Emotion Recognition

3 citations · 4 across the 9 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2025

Audiobook-CC: Controllable Long-context Speech Generation for Multicast Audiobook

Min Liu, JingJing Yin, Xiang Zhang +6

Existing text-to-speech systems predominantly focus on single-sentence synthesis and lack adequate contextual modeling as well as fine-grained performance control capabilities for…

eess.AS2023

PP-MeT: a Real-world Personalized Prompt based Meeting Transcription System

Xiang Lyu, Yuhang Cao, Qing Wang +5

Speaker-attributed automatic speech recognition (SA-ASR) improves the accuracy and applicability of multi-speaker ASR systems in real-world scenarios by assigning speaker labels to…

eess.AS2023★ 1 cited

PromptVC: Flexible Stylistic Voice Conversion in Latent Space Driven by Natural Language Prompts

Jixun Yao, Yuguang Yang, Yi Lei +7

Style voice conversion aims to transform the style of source speech to a desired style according to real-world application demands. However, the current style voice conversion appr…

eess.AS2023

HYBRIDFORMER: improving SqueezeFormer with hybrid attention and NSR mechanism

Yuguang Yang, Yu Pan, Jingjing Yin +3

SqueezeFormer has recently shown impressive performance in automatic speech recognition (ASR). However, its inference speed suffers from the quadratic complexity of softmax-attenti…

eess.AS2022

LMEC: Learnable Multiplicative Absolute Position Embedding Based Conformer for Speech Recognition

Yuguang Yang, Yu Pan, Jingjing Yin +1

This paper proposes a Learnable Multiplicative absolute position Embedding based Conformer (LMEC). It contains a kernelized linear attention (LA) module called LMLA to solve the ti…

eess.AS2022

The USTC-Ximalaya system for the ICASSP 2022 multi-channel multi-party meeting transcription (M2MeT) challenge

Maokui He, Xiang Lv, Weilin Zhou +8

We propose two improvements to target-speaker voice activity detection (TS-VAD), the core component in our proposed speaker diarization system that was submitted to the 2022 Multi-…