32 citations · 122 across the 20 of their papers we have counts for
24 papers · 1 filter
Scalable Video Object Segmentation with Simplified Framework
Qiangqiang Wu, Tianyu Yang, Wei WU +1
The current popular methods for video object segmentation (VOS) implement feature matching through several hand-crafted modules that separately perform feature extraction and match…
Edit Temporal-Consistent Videos with Image Diffusion Model
Yuanzhi Wang, Yong Li, Xiaoya Zhang +4
Large-scale text-to-image (T2I) diffusion models have been extended for text-guided video editing, yielding impressive zero-shot video editing performance. Nonetheless, the generat…
Human Attention-Guided Explainable Artificial Intelligence for Computer Vision Models
Guoyang Liu, Jindi Zhang, Antoni B. Chan +1
We examined whether embedding human attention knowledge into saliency-based explainable AI (XAI) methods for computer vision models could enhance their plausibility and faithfulnes…
ODAM: Gradient-based instance-specific visual explanations for object detection
Chenyang Zhao, Antoni B. Chan
We propose the gradient-weighted Object Detector Activation Maps (ODAM), a visualized explanation technique for interpreting the predictions of object detectors. Utilizing the grad…
DropMAE: Learning Representations via Masked Autoencoders with Spatial-Attention Dropout for Temporal Matching Tasks
Qiangqiang Wu, Tianyu Yang, Ziquan Liu +3
This paper studies masked autoencoder (MAE) video pre-training for various temporal matching-based downstream tasks, i.e., object-level tracking tasks including video object tracki…
An Empirical Study on Distribution Shift Robustness From the Perspective of Pre-Training and Data Augmentation
Ziquan Liu, Yi Xu, Yuanhong Xu +5
The performance of machine learning models under distribution shift has been the focus of the community in recent years. Most of current methods have been proposed to improve the r…