most citedSound to Visual Scene Generation by Audio-to-Visual Latent Alignment

2 citations · 3 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV2023

Can CLIP Help Sound Source Localization?

Sooyoung Park, Arda Senocak, Joon Son Chung

Large-scale pre-trained image-text models demonstrate remarkable versatility across diverse tasks, benefiting from their robust representational capabilities and effective multimod…

cs.CV2023

Sound Source Localization is All about Cross-Modal Alignment

Arda Senocak, Hyeonggon Ryu, Junsik Kim +3

Humans can easily perceive the direction of sound sources in a visual scene, termed sound source localization. Recent studies on learning-based sound source localization have mainl…

cs.SD2023

FlexiAST: Flexibility is What AST Needs

Jiu Feng, Mehmet Hamza Erol, Joon Son Chung +1

The objective of this work is to give patch-size flexibility to Audio Spectrogram Transformers (AST). Recent advancements in ASTs have shown superior performance in various audio-b…

cs.CL20231 cited

Hindi as a Second Language: Improving Visually Grounded Speech with Semantically Similar Samples

Hyeonggon Ryu, Arda Senocak, In So Kweon +1

The objective of this work is to explore the learning of visually grounded speech models (VGS) from multilingual perspective. Bilingual VGS models are generally trained with an equ…

cs.CV20232 cited

Sound to Visual Scene Generation by Audio-to-Visual Latent Alignment

Kim Sung-Bin, Arda Senocak, Hyunwoo Ha +2

How does audio describe the world around us? In this paper, we propose a method for generating an image of a scene from sound. Our method addresses the challenges of dealing with t…