22 citations · 23 across the 5 of their papers we have counts for
6 papers
NEST-RQ: Next Token Prediction for Speech Self-Supervised Pre-Training
Minglun Han, Ye Bai, Chen Shen +6
Speech self-supervised pre-training can effectively improve the performance of downstream tasks. However, previous self-supervised learning (SSL) methods for speech, such as HuBERT…
X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages
Feilong Chen, Minglun Han, Haozhi Zhao +4
Large language models (LLMs) have demonstrated remarkable language abilities. GPT-4, based on advanced LLMs, exhibits extraordinary multimodal capabilities beyond previous visual l…
Matching-based Term Semantics Pre-training for Spoken Patient Query Understanding
Zefa Hu, Xiuyi Chen, Haoran Wu +5
Medical Slot Filling (MSF) task aims to convert medical queries into structured information, playing an essential role in diagnosis dialogue systems. However, the lack of sufficien…
Complex Dynamic Neurons Improved Spiking Transformer Network for Efficient Automatic Speech Recognition
Minglun Han, Qingyu Wang, Tielin Zhang +3
The spiking neural network (SNN) using leaky-integrated-and-fire (LIF) neurons has been commonly used in automatic speech recognition (ASR) tasks. However, the LIF neuron is still…
Improving End-to-End Contextual Speech Recognition with Fine-Grained Contextual Knowledge Selection
Minglun Han, Linhao Dong, Zhenlin Liang +4
Nowadays, most methods in end-to-end contextual speech recognition bias the recognition process towards contextual knowledge. Since all-neural contextual biasing methods rely on ph…
CIF-based Collaborative Decoding for End-to-end Contextual Speech Recognition
Minglun Han, Linhao Dong, Shiyu Zhou +1
End-to-end (E2E) models have achieved promising results on multiple speech recognition benchmarks, and shown the potential to become the mainstream. However, the unified structure…