6 papers
Distilling Image Prototypes for Guided Test-Time Adaptation
Liwen Wang, Xingbo Dong, Iman Yi Liao +4
Test-Time Adaptation (TTA) enhances the robustness of models against distribution shifts but faces two critical challenges: error accumulation from noisy pseudo-labels and catastro…
LightAVSeg: Lightweight Audio-Visual Segmentation
Qing Zhong, Guodong Ding, Lingqiao Liu +3
Audio-Visual Segmentation (AVS) targets pixel level localization of sounding emitting objects in videos. However, existing models rely on dense cross-modal attention with quadratic…
Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation
Bingrui Zhao, Lin Yuanbo Wu, Xiangtian Fan +5
Referring Video Object Segmentation (RVOS) aims to segment an object of interest throughout a video based on a language description. The prominent challenge lies in aligning static…
A Novel Local Focusing Mechanism for Deepfake Detection Generalization
Mingliang Li, Lin Yuanbo Wu, Changhong Liu +1
The rapid advancement of deepfake generation techniques has intensified the need for robust and generalizable detection methods. Existing approaches based on reconstruction learnin…
Blended Latent Diffusion under Attention Control for Real-World Video Editing
Deyin Liu, Lin Yuanbo Wu, Xianghua Xie
Due to lack of fully publicly available text-to-video models, current video editing methods tend to build on pre-trained text-to-image generation models, however, they still face g…
Towards Efficient Pixel Labeling for Industrial Anomaly Detection and Localization
Hanxi Li, Jingqi Wu, Lin Yuanbo Wu +3
In the realm of practical Anomaly Detection (AD) tasks, manual labeling of anomalous pixels proves to be a costly endeavor. Consequently, many AD methods are crafted as one-class c…