4 papers
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
Mingyu Cui, Yifan Yang, Jiajun Deng +7
Self-supervised learning (SSL) based discrete speech representations are highly compact and domain adaptable. In this paper, SSL discrete speech features extracted from WavLM model…
VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
Jianheng Zhuo, Yifan Yang, Yiwen Shao +4
Automatic speech recognition (ASR) has made remarkable progress but heavily relies on large-scale labeled data, which is scarce for low-resource languages like Vietnamese. While ex…
k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
Yifan Yang, Jianheng Zhuo, Zengrui Jin +9
Self-supervised learning (SSL) has achieved great success in speech-related tasks. While Transformer and Conformer architectures have dominated SSL backbones, encoders like Zipform…
A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario
Zheshu Song, Ziyang Ma, Yifan Yang +2
Large Language Models (LLMs) have showcased exceptional performance across diverse NLP tasks, and their integration with speech encoder is rapidly emerging as a dominant trend in t…