most citedIdeal-LLM: Integrating Dual Encoders and Language-Adapted LLM for Multilingual Speech-to-Text

1 citations · 2 across the 5 of their papers we have counts for

collaborators

5 papers

cs.SD20251 cited

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Xuelong Geng, Kun Wei, Qijie Shao +18

Large Language Models (LLMs) have made significant progress in various downstream tasks, inspiring the development of Speech Understanding Language Models (SULMs) to enable compreh…

cs.SD2025

DQ-Data2vec: Decoupling Quantization for Multilingual Speech Recognition

Qijie Shao, Linhao Dong, Kun Wei +2

Data2vec is a self-supervised learning (SSL) approach that employs a teacher-student architecture for contextual representation learning via masked prediction, demonstrating remark…

cs.SD2025

CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech Recognition

He Wang, Xucheng Wan, Naijun Zheng +4

Code-switching automatic speech recognition (ASR) aims to transcribe speech that contains two or more languages accurately. To better capture language-specific speech representatio…

eess.AS20241 cited

Ideal-LLM: Integrating Dual Encoders and Language-Adapted LLM for Multilingual Speech-to-Text

Hongfei Xue, Wei Ren, Xuelong Geng +6

Integrating audio encoders with LLMs through connectors has enabled these models to process and comprehend audio modalities, significantly enhancing speech-to-text tasks, including…

eess.AS2024

MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition

Bingshen Mu, Yangze Li, Qijie Shao +5

Despite notable advancements in automatic speech recognition (ASR), performance tends to degrade when faced with adverse conditions. Generative error correction (GER) leverages the…