1 paper
Suhwan Choi, Kyu Won Kim, Myungjoo Kang
We introduce Multimodal Matching based on Valence and Arousal (MMVA), a tri-modal encoder framework designed to capture emotional content across images, music, and musical captions…