4 papers
Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence
Han Hu, Dongheng Lin, Yuqi Hou +3
Localising multiple sound sources in visual scenes remains a fundamental challenge in multimodal perception due to an inherent circular dependency: separating mixed audio requires…
Gaze-Anchored Social Net: Decoding Implicit Relations via Joint Modeling
Yuqi Hou, Zhuo Chen, Han Hu +3
Human gaze does more than point to visual targets; it serves as a subtle indicator of social intent within static images, whereas standard models typically process individuals inde…
What Time Tells Us? An Explorative Study of Time Awareness Learned from Static Images
Dongheng Lin, Han Hu, Jianbo Jiao
Time becomes visible through illumination changes in what we see. Inspired by this, in this paper we explore the potential to learn time awareness from static images, trying to ans…
Audio-Visual Separation with Hierarchical Fusion and Representation Alignment
Han Hu, Dongheng Lin, Qiming Huang +3
Self-supervised audio-visual source separation leverages natural correlations between audio and vision modalities to separate mixed audio signals. In this work, we first systematic…