Showing eess.ASShow all
3 papers · 1 filter
eess.AS2026
Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models
Longhao Li, Hongjie Chen, Zehan Li +5
Recent advances in reasoning models have driven significant progress in text and multimodal domains, yet audio reasoning remains relatively limited. Only a few Large Audio Language…
eess.AS2025
Bridging the Gap between Continuous and Informative Discrete Representations by Random Product Quantization
Xueqing Li, Hao Ma, Zehan Li +8
Self-supervised learning (SSL) has become a core technique in speech processing, but the high dimensionality of its representations makes discretization essential for improving eff…
eess.AS2025
Multilingual Speech Recognition Using Discrete Tokens with a Two-step Training Strategy
Zehan Li, Yan Yang, Xueqing Li +3
Pre-trained models, especially self-supervised learning (SSL) models, have demonstrated impressive results in automatic speech recognition (ASR) task. While most applications of SS…