activity
20232026
collaborators

5 papers

cs.MM2026

Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence

Han Hu, Dongheng Lin, Yuqi Hou +3

Localising multiple sound sources in visual scenes remains a fundamental challenge in multimodal perception due to an inherent circular dependency: separating mixed audio requires…

cs.CV2026

Gaze-Anchored Social Net: Decoding Implicit Relations via Joint Modeling

Yuqi Hou, Zhuo Chen, Han Hu +3

Human gaze does more than point to visual targets; it serves as a subtle indicator of social intent within static images, whereas standard models typically process individuals inde…

cs.MM2025

Audio-Visual Separation with Hierarchical Fusion and Representation Alignment

Han Hu, Dongheng Lin, Qiming Huang +3

Self-supervised audio-visual source separation leverages natural correlations between audio and vision modalities to separate mixed audio signals. In this work, we first systematic…

cs.CV2024

Few Exemplar-Based General Medical Image Segmentation via Domain-Aware Selective Adaptation

Chen Xu, Qiming Huang, Yuqi Hou +4

Medical image segmentation poses challenges due to domain gaps, data modality variations, and dependency on domain knowledge or experts, especially for low- and middle-income count…

cs.CV2023

Multi-Modal Gaze Following in Conversational Scenarios

Yuqi Hou, Zhongqun Zhang, Nora Horanyi +3

Gaze following estimates gaze targets of in-scene person by understanding human behavior and scene information. Existing methods usually analyze scene images for gaze following. Ho…