activity
20202023
most citedX-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

22 citations · 23 across the 5 of their papers we have counts for

collaborators

6 papers

eess.AS2024

NEST-RQ: Next Token Prediction for Speech Self-Supervised Pre-Training

Minglun Han, Ye Bai, Chen Shen +6

Speech self-supervised pre-training can effectively improve the performance of downstream tasks. However, previous self-supervised learning (SSL) methods for speech, such as HuBERT…

cs.CL202322 cited

X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Feilong Chen, Minglun Han, Haozhi Zhao +4

Large language models (LLMs) have demonstrated remarkable language abilities. GPT-4, based on advanced LLMs, exhibits extraordinary multimodal capabilities beyond previous visual l…

cs.CL2023

Matching-based Term Semantics Pre-training for Spoken Patient Query Understanding

Zefa Hu, Xiuyi Chen, Haoran Wu +5

Medical Slot Filling (MSF) task aims to convert medical queries into structured information, playing an essential role in diagnosis dialogue systems. However, the lack of sufficien…

cs.NE20231 cited

Complex Dynamic Neurons Improved Spiking Transformer Network for Efficient Automatic Speech Recognition

Minglun Han, Qingyu Wang, Tielin Zhang +3

The spiking neural network (SNN) using leaky-integrated-and-fire (LIF) neurons has been commonly used in automatic speech recognition (ASR) tasks. However, the LIF neuron is still…

cs.CL2022

Improving End-to-End Contextual Speech Recognition with Fine-Grained Contextual Knowledge Selection

Minglun Han, Linhao Dong, Zhenlin Liang +4

Nowadays, most methods in end-to-end contextual speech recognition bias the recognition process towards contextual knowledge. Since all-neural contextual biasing methods rely on ph…

cs.CL2020

CIF-based Collaborative Decoding for End-to-end Contextual Speech Recognition

Minglun Han, Linhao Dong, Shiyu Zhou +1

End-to-end (E2E) models have achieved promising results on multiple speech recognition benchmarks, and shown the potential to become the mainstream. However, the unified structure…