1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.MM2024
Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment
Arda Senocak, Hyeonggon Ryu, Junsik Kim +3
Recent studies on learning-based sound source localization have mainly focused on the localization performance perspective. However, prior work and existing benchmarks overlook a c…
cs.CV2023
Sound Source Localization is All about Cross-Modal Alignment
Arda Senocak, Hyeonggon Ryu, Junsik Kim +3
Humans can easily perceive the direction of sound sources in a visual scene, termed sound source localization. Recent studies on learning-based sound source localization have mainl…
cs.CL2023★ 1 cited
Hindi as a Second Language: Improving Visually Grounded Speech with Semantically Similar Samples
Hyeonggon Ryu, Arda Senocak, In So Kweon +1
The objective of this work is to explore the learning of visually grounded speech models (VGS) from multilingual perspective. Bilingual VGS models are generally trained with an equ…