9 papers
Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence
Han Hu, Dongheng Lin, Yuqi Hou +3
Localising multiple sound sources in visual scenes remains a fundamental challenge in multimodal perception due to an inherent circular dependency: separating mixed audio requires…
PeCA: Palette Context Assisted Inference for Test-Time Paint-Bucket Colourisation on Animation Videos
Dongheng Lin, Jianbo Jiao
In animation production, paint-bucket colourisation for hand-drawn animation is a labour-intensive procedure that assigns each enclosed region in line sketches a colour from refere…
Gaze-Anchored Social Net: Decoding Implicit Relations via Joint Modeling
Yuqi Hou, Zhuo Chen, Han Hu +3
Human gaze does more than point to visual targets; it serves as a subtle indicator of social intent within static images, whereas standard models typically process individuals inde…
FlowSeg: Dynamic Semantic Guidance for LLM-Conditioned Segmentation
Zekang Zhang, Guangyu Gao, Youyun Tang +7
LLM-conditioned segmentation has recently advanced rapidly by coupling large language models with iterative mask generation frameworks. However, we identify a persistent failure mo…
Structure-Aware Feature Rectification with Region Adjacency Graphs for Training-Free Open-Vocabulary Semantic Segmentation
Qiming Huang, Hao Ai, Jianbo Jiao
Benefiting from the inductive biases learned from large-scale datasets, open-vocabulary semantic segmentation (OVSS) leverages the power of vision-language models, such as CLIP, to…
What Time Tells Us? An Explorative Study of Time Awareness Learned from Static Images
Dongheng Lin, Han Hu, Jianbo Jiao
Time becomes visible through illumination changes in what we see. Inspired by this, in this paper we explore the potential to learn time awareness from static images, trying to ans…