works on

From the 1 of 6 linked papers with an AI index.

activity
20242026
collaborators

6 papers

cs.SD2026

Qwen-Music Technical Report

Jin Xu, Kangdi Wang, Ruibin Yuan +24

Qwen-Music is a large language model‑based system that generates high‑fidelity songs with vocals from text prompts or re‑imagines existing tracks, using a semantic token representa…

cs.CV2025

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Chaoyou Fu, Haojia Lin, Xiong Wang +13

Recent Multimodal Large Language Models (MLLMs) have typically focused on integrating visual and textual modalities, with less emphasis placed on the role of speech in enhancing in…

cs.SD2025

Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR

Longhao Li, Yangze Li, Hongfei Xue +4

CTC-based streaming ASR has gained significant attention in real-world applications but faces two main challenges: accuracy degradation in small chunks and token emission latency.…

cs.SD2025

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Xuelong Geng, Kun Wei, Qijie Shao +18

Large Language Models (LLMs) have made significant progress in various downstream tasks, inspiring the development of Speech Understanding Language Models (SULMs) to enable compreh…

cs.SD2024

Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Xiong Wang, Yangze Li, Chaoyou Fu +5

Rapidly developing large language models (LLMs) have brought tremendous intelligent applications. Especially, the GPT-4o's excellent duplex speech interaction ability has brought i…

cs.SD2024

Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets

Xuelong Geng, Tianyi Xu, Kun Wei +9

Large Language Models (LLMs) have demonstrated unparalleled effectiveness in various NLP tasks, and integrating LLMs with automatic speech recognition (ASR) is becoming a mainstrea…