22 citations · 29 across the 6 of their papers we have counts for
4 papers · 1 filter
X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages
Feilong Chen, Minglun Han, Haozhi Zhao +4
Large language models (LLMs) have demonstrated remarkable language abilities. GPT-4, based on advanced LLMs, exhibits extraordinary multimodal capabilities beyond previous visual l…
Matching-based Term Semantics Pre-training for Spoken Patient Query Understanding
Zefa Hu, Xiuyi Chen, Haoran Wu +5
Medical Slot Filling (MSF) task aims to convert medical queries into structured information, playing an essential role in diagnosis dialogue systems. However, the lack of sufficien…
Improving End-to-End Contextual Speech Recognition with Fine-Grained Contextual Knowledge Selection
Minglun Han, Linhao Dong, Zhenlin Liang +4
Nowadays, most methods in end-to-end contextual speech recognition bias the recognition process towards contextual knowledge. Since all-neural contextual biasing methods rely on ph…
CIF-based Collaborative Decoding for End-to-end Contextual Speech Recognition
Minglun Han, Linhao Dong, Shiyu Zhou +1
End-to-end (E2E) models have achieved promising results on multiple speech recognition benchmarks, and shown the potential to become the mainstream. However, the unified structure…