Showing eess.ASShow all
2 papers · 1 filter
eess.AS2024
NEST-RQ: Next Token Prediction for Speech Self-Supervised Pre-Training
Minglun Han, Ye Bai, Chen Shen +6
Speech self-supervised pre-training can effectively improve the performance of downstream tasks. However, previous self-supervised learning (SSL) methods for speech, such as HuBERT…
eess.AS2023
VILAS: Exploring the Effects of Vision and Language Context in Automatic Speech Recognition
Ziyi Ni, Minglun Han, Feilong Chen +4
Enhancing automatic speech recognition (ASR) performance by leveraging additional multimodal information has shown promising results in previous studies. However, most of these wor…