1 paper
Ayuto Tsutsumi, Kohei Tanaka, Sayaka Shiota
In this paper, we propose a submission to the x-to-audio alignment (XACLE) challenge. The goal is to predict semantic alignment of a given general audio and text pair. The proposed…