From the 1 of 7 linked papers with an AI index.
7 papers
Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model
Zhiwei Lin, Jun Chen, Boshi Tang +5
Text-controlled symbolic music generation has recently gained research attention due to its versatile, flexible and straightforward approach to music composition. However, previous…
Video = World + Event Stream
Lianghua Huang, Zhi-Fan Wu, Yupeng Shi +24
The paper introduces Wan-Streamer v0.3, a model that treats video as a combination of a persistent world and a dynamic event stream, enabling real-time multimodal audio‑visual inte…
Balancing ASR and diarization in end-to-end LLMs for multi-talker speech recognition
Naijun Zheng, Yuke Lin, Sanli Tian +4
Multi-talker speech recognition is often addressed by combining automatic speech recognition (ASR) and speaker diarization in a pipeline system. Recently, LLM-based approaches have…
LeVo: High-Quality Song Generation with Multi-Preference Alignment
Shun Lei, Yaoxun Xu, Zhiwei Lin +10
Recent advances in large language models (LLMs) and audio language models have significantly improved music generation, particularly in lyrics-to-song generation. However, existing…
VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents
Weihao Wu, Liang Cao, Xinyu Wu +4
Recent significant advancements in Large Language Models (LLMs) have greatly propelled the development of Role-Playing Conversational Agents (RPCAs). These systems aim to create im…
MuCodec: Ultra Low-Bitrate Music Codec
Yaoxun Xu, Hangting Chen, Jianwei Yu +5
Music codecs are a vital aspect of audio codec research, and ultra low-bitrate compression holds significant importance for music transmission and generation. Due to the complexity…