1 citations · 1 across the 1 of their papers we have counts for
Showing eess.ASShow all
2 papers · 1 filter
eess.AS2025
Listening without Looking: Modality Bias in Audio-Visual Captioning
Yuchi Ishikawa, Toranosuke Manabe, Tatsuya Komatsu +1
Audio-visual captioning aims to generate holistic scene descriptions by jointly modeling sound and vision. While recent methods have improved performance through sophisticated moda…
eess.AS2024
Pre-training with Synthetic Patterns for Audio
Yuchi Ishikawa, Tatsuya Komatsu, Yoshimitsu Aoki
In this paper, we propose to pre-train audio encoders using synthetic patterns instead of real audio data. Our proposed framework consists of two key elements. The first one is Mas…