collaborators

5 papers

cs.MM2026

Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence

Han Hu, Dongheng Lin, Yuqi Hou +3

Localising multiple sound sources in visual scenes remains a fundamental challenge in multimodal perception due to an inherent circular dependency: separating mixed audio requires…

cs.CV2026

PeCA: Palette Context Assisted Inference for Test-Time Paint-Bucket Colourisation on Animation Videos

Dongheng Lin, Jianbo Jiao

In animation production, paint-bucket colourisation for hand-drawn animation is a labour-intensive procedure that assigns each enclosed region in line sketches a colour from refere…

cs.CV2025

What Time Tells Us? An Explorative Study of Time Awareness Learned from Static Images

Dongheng Lin, Han Hu, Jianbo Jiao

Time becomes visible through illumination changes in what we see. Inspired by this, in this paper we explore the potential to learn time awareness from static images, trying to ans…

cs.CV2025

A Unified Reasoning Framework for Holistic Zero-Shot Video Anomaly Analysis

Dongheng Lin, Mengxue Qu, Kunyang Han +3

Most video-anomaly research stops at frame-wise detection, offering little insight into why an event is abnormal, typically outputting only frame-wise anomaly scores without spatia…

cs.MM2025

Audio-Visual Separation with Hierarchical Fusion and Representation Alignment

Han Hu, Dongheng Lin, Qiming Huang +3

Self-supervised audio-visual source separation leverages natural correlations between audio and vision modalities to separate mixed audio signals. In this work, we first systematic…