3 papers
cs.MM2025
Multi Agents Semantic Emotion Aligned Music to Image Generation with Music Derived Captions
Junchang Shi, Gang Li
When people listen to music, they often experience rich visual imagery. We aim to externalize this inner imagery by generating images conditioned on music. We propose MESA MIG, a m…
cs.SD2025
GLAP: General contrastive audio-text pretraining across domains and languages
Heinrich Dinkel, Zhiyong Yan, Tianzi Wang +7
Contrastive Language Audio Pretraining (CLAP) is a widely-used method to bridge the gap between audio and text domains. Current CLAP methods enable sound and music retrieval in Eng…
cs.SD2025
Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering
Gang Li, Jizhong Liu, Heinrich Dinkel +3
Recently, reinforcement learning (RL) has been shown to greatly enhance the reasoning capabilities of large language models (LLMs), and RL-based approaches have been progressively…