20 citations · 39 across the 8 of their papers we have counts for
8 papers
Qwen-Audio-3.0-ASR Technical Report
Chuanmeng Bian, Daren Chen, Peixin Chen +42
In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model scaling, and deep integra…
SA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASR
Zhiyun Fan, Linhao Dong, Jun Zhang +2
Multi-talker automatic speech recognition plays a crucial role in scenarios involving multi-party interactions, such as meetings and conversations. Due to its inherent complexity,…
Language-specific Acoustic Boundary Learning for Mandarin-English Code-switching Speech Recognition
Zhiyun Fan, Linhao Dong, Chen Shen +4
Code-switching speech recognition (CSSR) transcribes speech that switches between multiple languages or dialects within a single sentence. The main challenge in this task is that d…
Token-level Speaker Change Detection Using Speaker Difference and Speech Content via Continuous Integrate-and-fire
Zhiyun Fan, Zhenlin Liang, Linhao Dong +6
In multi-talker scenarios such as meetings and conversations, speech processing systems are usually required to segment the audio and then transcribe each segmentation. These two s…
Sequence-level Speaker Change Detection with Difference-based Continuous Integrate-and-fire
Zhiyun Fan, Linhao Dong, Meng Cai +2
Speaker change detection is an important task in multi-party interactions such as meetings and conversations. In this paper, we address the speaker change detection task from the p…
Exploring wav2vec 2.0 on speaker verification and language identification
Zhiyun Fan, Meng Li, Shiyu Zhou +1
Wav2vec 2.0 is a recently proposed self-supervised framework for speech representation learning. It follows a two-stage training process of pre-training and fine-tuning, and perfor…