activity
20212026
most citedLauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

18 citations · 101 across the 75 of their papers we have counts for

collaborators
Showing 2023Show all

12 papers · 1 filter

cs.CL2023★ 6 cited

emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Ziyang Ma, Zhisheng Zheng, Jiaxin Ye +4

We propose emotion2vec, a universal speech emotion representation model. emotion2vec is pre-trained on open-source unlabeled emotion data through self-supervised online distillatio…

cs.SD2023

Hourglass-AVSR: Down-Up Sampling-based Computational Efficiency Model for Audio-Visual Speech Recognition

Fan Yu, Haoxu Wang, Ziyang Ma +1

Recently audio-visual speech recognition (AVSR), which better leverages video modality as additional information to extend automatic speech recognition (ASR), has shown promising r…

cs.SD2023★ 18 cited

LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Zhihao Du, Jiaming Wang, Qian Chen +12

Generative Pre-trained Transformer (GPT) models have achieved remarkable performance on various natural language processing tasks, and have shown great potential as backbones for a…

cs.CL2023

Fast-HuBERT: An Efficient Training Framework for Self-Supervised Speech Representation Learning

Guanrou Yang, Ziyang Ma, Zhisheng Zheng +3

Recent years have witnessed significant advancements in self-supervised learning (SSL) methods for speech-processing tasks. Various speech-based SSL models have been developed and…

cs.CL2023★ 1 cited

Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition

Ziyang Ma, Wen Wu, Zhisheng Zheng +4

In this paper, we explored how to boost speech emotion recognition (SER) with the state-of-the-art speech pre-trained model (PTM), data2vec, text generation technique, GPT-4, and s…

eess.AS2023

Towards Universal Speech Discrete Tokens: A Case Study for ASR and TTS

Yifan Yang, Feiyu Shen, Chenpeng Du +4

Self-supervised learning (SSL) proficiency in speech-related tasks has driven research into utilizing discrete tokens for speech tasks like recognition and translation, which offer…