2 papers
cs.CV2022
Audio-Visual Fusion Layers for Event Type Aware Video Recognition
Arda Senocak, Junsik Kim, Tae-Hyun Oh +3
Human brain is continuously inundated with the multisensory information and their complex interactions coming from the outside world at any given moment. Such information is automa…
cs.CV2022
Learning Sound Localization Better From Semantically Similar Samples
Arda Senocak, Hyeonggon Ryu, Junsik Kim +1
The objective of this work is to localize the sound sources in visual scenes. Existing audio-visual works employ contrastive learning by assigning corresponding audio-visual pairs…