2 citations · 7 across the 9 of their papers we have counts for
4 papers · 1 filter
Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment
Kim Sung-Bin, Arda Senocak, Hyunwoo Ha +1
How does audio describe the world around us? In this work, we propose a method for generating images of visual scenes from diverse in-the-wild sounds. This cross-modal generation t…
Can CLIP Help Sound Source Localization?
Sooyoung Park, Arda Senocak, Joon Son Chung
Large-scale pre-trained image-text models demonstrate remarkable versatility across diverse tasks, benefiting from their robust representational capabilities and effective multimod…
Sound Source Localization is All about Cross-Modal Alignment
Arda Senocak, Hyeonggon Ryu, Junsik Kim +3
Humans can easily perceive the direction of sound sources in a visual scene, termed sound source localization. Recent studies on learning-based sound source localization have mainl…
Sound to Visual Scene Generation by Audio-to-Visual Latent Alignment
Kim Sung-Bin, Arda Senocak, Hyunwoo Ha +2
How does audio describe the world around us? In this paper, we propose a method for generating an image of a scene from sound. Our method addresses the challenges of dealing with t…