2 citations · 3 across the 5 of their papers we have counts for
5 papers
Can CLIP Help Sound Source Localization?
Sooyoung Park, Arda Senocak, Joon Son Chung
Large-scale pre-trained image-text models demonstrate remarkable versatility across diverse tasks, benefiting from their robust representational capabilities and effective multimod…
Sound Source Localization is All about Cross-Modal Alignment
Arda Senocak, Hyeonggon Ryu, Junsik Kim +3
Humans can easily perceive the direction of sound sources in a visual scene, termed sound source localization. Recent studies on learning-based sound source localization have mainl…
FlexiAST: Flexibility is What AST Needs
Jiu Feng, Mehmet Hamza Erol, Joon Son Chung +1
The objective of this work is to give patch-size flexibility to Audio Spectrogram Transformers (AST). Recent advancements in ASTs have shown superior performance in various audio-b…
Hindi as a Second Language: Improving Visually Grounded Speech with Semantically Similar Samples
Hyeonggon Ryu, Arda Senocak, In So Kweon +1
The objective of this work is to explore the learning of visually grounded speech models (VGS) from multilingual perspective. Bilingual VGS models are generally trained with an equ…
Sound to Visual Scene Generation by Audio-to-Visual Latent Alignment
Kim Sung-Bin, Arda Senocak, Hyunwoo Ha +2
How does audio describe the world around us? In this paper, we propose a method for generating an image of a scene from sound. Our method addresses the challenges of dealing with t…