Publications (6)
Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection
Haoyu Wang, Guoqiang Hu, Guodong Lin +2
As a robust and large-scale multilingual speech recognition model, Whisper has demonstrated impressive results in many low-resource and out-of-distribution scenarios. However, its…
Dolphin: A Large-Scale Automatic Speech Recognition Model for Eastern Languages
Yangyang Meng, Jinpeng Li, Guodong Lin +7
This report introduces Dolphin, a large-scale multilingual automatic speech recognition (ASR) model that extends the Whisper architecture to support a wider range of languages. Our…
Unified Gradient Projection: Language-Balanced Continual Learning for Multilingual Low-Resource ASR
Ziang Ren, Guodong Lin, Yuchen Ai +2
The paper introduces Unified Gradient Projection (UGP), a method that uses language‑balanced replay gradients to constrain parameter updates, reducing dominant‑language bias and ca…
Enhancing Multilingual LLM-based ASR with Mixture of Experts and Dynamic Downsampling
Guodong Lin, Ziqi Chen, Yuxiang Fu +2
The rapid progress of large language models (LLMs) has opened up a new frontier for automatic speech recognition (ASR), making their effective integration a critical and challengin…
GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark
Yujie Tu, Yifan Yang, Tianrui Wang +36
While modern ASR systems achieve low error rates on high-resource benchmarks, such performance often overestimates real-world robustness. Existing evaluations address challenges in…
Dolphin-CN-Dialect: Where Chinese Dialects Matter
Yangyang Meng, Huihang Zhong, Guodong Lin +6
We present Dolphin-CN-Dialect, a streaming-capable ASR model with a focus on Chinese and dialect-rich scenarios. Compared to the previous version, Dolphin-CN-Dialect introduces sub…