5 papers
Omni-Interactive Universal Embedder
Wei-Yao Wang, Kazuya Tateishi, Shuyang Cui +4
Multimodal representation learning has been shifting from traditional two-tower architectures to large language model (LLM)-based embedders due to their strong instruction-followin…
MMAudio-LABEL: Audio Event Labeling via Audio Generation for Silent Video
Kazuya Tateishi, Akira Takahashi, Atsuo Hiroe +3
Recent advances in multimodal generation have enabled high-quality audio generation from silent videos. Practical applications, such as sound production, demand not only the genera…
VIRTUE: Visual-Interactive Text-Image Universal Embedder
Wei-Yao Wang, Kazuya Tateishi, Qiyu Wu +2
Multimodal representation learning models have demonstrated successful operation across complex tasks, and the integration of vision-language models (VLMs) has further enabled embe…
Extending Audio Masked Autoencoders Toward Audio Restoration
Zhi Zhong, Hao Shi, Masato Hirano +5
Audio classification and restoration are among major downstream tasks in audio signal processing. However, restoration derives less of a benefit from pretrained models compared to…
An Attention-based Approach to Hierarchical Multi-label Music Instrument Classification
Zhi Zhong, Masato Hirano, Kazuki Shimada +3
Although music is typically multi-label, many works have studied hierarchical music tagging with simplified settings such as single-label data. Moreover, there lacks a framework to…