268 citations · 349 across the 37 of their papers we have counts for
4 papers · 1 filter
AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling
Jiacheng Shi, Hongfei Du, Xinyuan Song +3
Neural speech codecs provide discrete representations for speech language models, but emotional cues are often degraded during quantization. Existing codecs mainly optimize acousti…
Advanced Long-Content Speech Recognition With Factorized Neural Transducer
Xun Gong, Yu Wu, Jinyu Li +4
In this paper, we propose two novel approaches, which integrate long-content information into the factorized neural transducer (FNT) based architecture in both non-streaming (refer…
The Microsoft System for VoxCeleb Speaker Recognition Challenge 2022
Gang Liu, Tianyan Zhou, Yong Zhao +4
In this report, we describe our submitted system for track 2 of the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC-22). We fuse a variety of good-performing models ranging fro…
Speech Pre-training with Acoustic Piece
Shuo Ren, Shujie Liu, Yu Wu +2
Previous speech pre-training methods, such as wav2vec2.0 and HuBERT, pre-train a Transformer encoder to learn deep representations from audio data, with objectives predicting eithe…