activity
20242026
most citedCode-switching Speech Recognition Under the Lens: Model- and Data-Centric Perspectives

1 citations · 1 across the 4 of their papers we have counts for

collaborators
Showing eess.ASShow all

11 papers · 1 filter

eess.AS2026

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action

Haoyang Zhang, Jun Chen, Donghang Wu +13

Recent advances in spoken dialogue language models have shifted from turn-based to full-duplex designs, where the model continuously listens to the user while generating responses.…

eess.AS2026

Step-Audio-R1.5 Technical Report

Yuxin Zhang, Xiangyu Tony Zhang, Daijiao Liu +16

Recent advancements in large audio language models have extended Chain-of-Thought (CoT) reasoning into the auditory domain, enabling models to tackle increasingly complex acoustic…

eess.AS2026

The WER Trap: Shattering the Illusion of Unified Tokens in Speech Language Models

Xiangyu Zhang, Yuxin Li, Haoyang Zhang +5

The pursuit of a "unified" discrete token for both speech understanding and generation has led the Speech Language Model (SLM) community to heavily rely on Word Error Rate (WER) --…

eess.AS20261 cited

Code-switching Speech Recognition Under the Lens: Model- and Data-Centric Perspectives

Hexin Liu, Haoyang Zhang, Qiquan Zhang +4

Code-switching automatic speech recognition (CS-ASR) presents unique challenges due to language confusion introduced by spontaneous intra-sentence switching and accent bias that bl…

eess.AS2025

Aligning Speech to Languages to Enhance Code-switching Speech Recognition

Hexin Liu, Xiangyu Zhang, Haoyang Zhang +4

Code-switching (CS) refers to the switching of languages within a speech signal and results in language confusion for automatic speech recognition (ASR). To address language confus…

eess.AS2025

Long-Context Modeling Networks for Monaural Speech Enhancement: A Comparative Study

Qiquan Zhang, Moran Chen, Zeyang Song +3

Advanced long-context modeling backbone networks, such as Transformer, Conformer, and Mamba, have demonstrated state-of-the-art performance in speech enhancement. However, a system…