18 citations · 101 across the 75 of their papers we have counts for
12 papers · 1 filter
emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Ziyang Ma, Zhisheng Zheng, Jiaxin Ye +4
We propose emotion2vec, a universal speech emotion representation model. emotion2vec is pre-trained on open-source unlabeled emotion data through self-supervised online distillatio…
Hourglass-AVSR: Down-Up Sampling-based Computational Efficiency Model for Audio-Visual Speech Recognition
Fan Yu, Haoxu Wang, Ziyang Ma +1
Recently audio-visual speech recognition (AVSR), which better leverages video modality as additional information to extend automatic speech recognition (ASR), has shown promising r…
LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
Zhihao Du, Jiaming Wang, Qian Chen +12
Generative Pre-trained Transformer (GPT) models have achieved remarkable performance on various natural language processing tasks, and have shown great potential as backbones for a…
Fast-HuBERT: An Efficient Training Framework for Self-Supervised Speech Representation Learning
Guanrou Yang, Ziyang Ma, Zhisheng Zheng +3
Recent years have witnessed significant advancements in self-supervised learning (SSL) methods for speech-processing tasks. Various speech-based SSL models have been developed and…
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
Ziyang Ma, Wen Wu, Zhisheng Zheng +4
In this paper, we explored how to boost speech emotion recognition (SER) with the state-of-the-art speech pre-trained model (PTM), data2vec, text generation technique, GPT-4, and s…
Towards Universal Speech Discrete Tokens: A Case Study for ASR and TTS
Yifan Yang, Feiyu Shen, Chenpeng Du +4
Self-supervised learning (SSL) proficiency in speech-related tasks has driven research into utilizing discrete tokens for speech tasks like recognition and translation, which offer…