9 papers
Break-the-Beat! Controllable MIDI-to-Drum Audio Synthesis
Shuyang Cui, Zhi Zhong, Qiyu Wu +9
Current methods for creating drum loop audio in digital music production, such as using one-shot samples or resampling, often demand non-trivial efforts of creators. While recent g…
VIRTUE: Visual-Interactive Text-Image Universal Embedder
Wei-Yao Wang, Kazuya Tateishi, Qiyu Wu +2
Multimodal representation learning models have demonstrated successful operation across complex tasks, and the integration of vision-language models (VLMs) has further enabled embe…
LLM2Fx-Tools: Tool Calling For Music Post-Production
Seungheon Doh, Junghyun Koo, Marco A. MartÃnez-RamÃrez +5
This paper introduces LLM2Fx-Tools, a multimodal tool-calling framework that generates executable sequences of audio effects (Fx-chain) for music post-production. LLM2Fx-Tools uses…
Mining the Gold: Student-AI Chat Logs as Rich Sources for Automated Knowledge Gap Detection
Quanzhi Fu, Qiyu Wu, Dan Williams
With the significant increase in enrollment in computing-related programs over the past 20 years, lecture sizes have grown correspondingly. In large lectures, instructors face chal…
Remedy-R: Generative Reasoning for Machine Translation Evaluation without Error Annotations
Shaomu Tan, Ryosuke Mitani, Ritvik Choudhary +3
Over the years, automatic MT metrics have hillclimbed benchmarks and presented strong and sometimes human-level agreement with human ratings. Yet they remain black-box, offering li…
MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval
Qiyu Wu, Shuyang Cui, Satoshi Hayakawa +3
Multimodal retrieval, which seeks to retrieve relevant content across modalities such as text or image, supports applications from AI search to contents production. Despite the suc…