works on

From the 1 of 41 linked papers with an AI index.

activity
20242026
most citedSLAM-LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing

3 citations · 5 across the 21 of their papers we have counts for

collaborators
Showing 2025Show all

13 papers · 1 filter

cs.AI2025

Step-Audio-R1 Technical Report

Fei Tian, Xiangyu Tony Zhang, Yuxin Zhang +14

Recent advances in reasoning models have demonstrated remarkable success in text and vision domains through extended chain-of-thought deliberation. However, a perplexing phenomenon…

eess.AS2025

Aligning Speech to Languages to Enhance Code-switching Speech Recognition

Hexin Liu, Xiangyu Zhang, Haoyang Zhang +4

Code-switching (CS) refers to the switching of languages within a speech signal and results in language confusion for automatic speech recognition (ASR). To address language confus…

eess.AS2025

MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models

Yayue Deng, Guoqiang Hu, Haiyang Sun +6

Spoken Dialogue Models (SDMs) have advanced rapidly, yet their ability to sustain genuinely interactive multi-turn conversations remains underexplored, as most benchmarks focus on…

cs.SD2025

A correlation-permutation approach for speech-music encoders model merging

Fabian Ritter-Gutierrez, Yi-Cheng Lin, Jeremy H. M Wong +3

Creating a unified speech and music model requires expensive pre-training. Model merging can instead create an unified audio model with minimal computational expense. However, dire…

eess.AS2025

Speechless: Speech Instruction Training Without Speech for Low Resource Languages

Alan Dao, Dinh Bach Vu, Huy Hoang Ha +6

The rapid growth of voice assistants powered by large language models (LLM) has highlighted a need for speech instruction data to train these systems. Despite the abundance of spee…

eess.AS2025

From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based Methodology

Haoyang Li, Yuchen Hu, Chen Chen +3

Deep neural network (DNN)-based speech enhancement (SE) usually uses conventional activation functions, which lack the expressiveness to capture complex multiscale structures neede…