papers

Publications (6)

cs.SD2024

Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection

Haoyu Wang, Guoqiang Hu, Guodong Lin +2

As a robust and large-scale multilingual speech recognition model, Whisper has demonstrated impressive results in many low-resource and out-of-distribution scenarios. However, its…

cs.CL2025

Dolphin: A Large-Scale Automatic Speech Recognition Model for Eastern Languages

Yangyang Meng, Jinpeng Li, Guodong Lin +7

This report introduces Dolphin, a large-scale multilingual automatic speech recognition (ASR) model that extends the Whisper architecture to support a wider range of languages. Our…

cs.CL2026

Unified Gradient Projection: Language-Balanced Continual Learning for Multilingual Low-Resource ASR

Ziang Ren, Guodong Lin, Yuchen Ai +2

The paper introduces Unified Gradient Projection (UGP), a method that uses language‑balanced replay gradients to constrain parameter updates, reducing dominant‑language bias and ca…

#continual learning#multilingual asr#low-resource languages#gradient projection
cs.SD2026

Enhancing Multilingual LLM-based ASR with Mixture of Experts and Dynamic Downsampling

Guodong Lin, Ziqi Chen, Yuxiang Fu +2

The rapid progress of large language models (LLMs) has opened up a new frontier for automatic speech recognition (ASR), making their effective integration a critical and challengin…

eess.AS2026

GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

Yujie Tu, Yifan Yang, Tianrui Wang +36

While modern ASR systems achieve low error rates on high-resource benchmarks, such performance often overestimates real-world robustness. Existing evaluations address challenges in…

cs.CL2026

Dolphin-CN-Dialect: Where Chinese Dialects Matter

Yangyang Meng, Huihang Zhong, Guodong Lin +6

We present Dolphin-CN-Dialect, a streaming-capable ASR model with a focus on Chinese and dialect-rich scenarios. Compared to the previous version, Dolphin-CN-Dialect introduces sub…