16 papers · 1 filter
Did Models Learn Sufficiently? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation
Yannan Chen, Ruoyu Chen, Wei Wang +6
Current visual models often make predictions based on a limited set of discriminative visual cues. As a result, they may become unreliable when the distribution shifts or when thes…
Domain Adaptive Object Detection via Dual-Stream Bilevel-Cycle Optimization
Yannan Chen, Wei Wang, Wenqiang Wang +5
Cycle self-training (CST) breaks the shared classifier assumption of the standard self-training framework, which is effective for unsupervised domain adaptation and exploits unlabe…
CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
Qiming Li, Zekai Ye, Xiaocheng Feng +9
Although Large Vision-Language Models (LVLMs) have demonstrated remarkable performance on downstream tasks, they frequently produce contents that deviate from visual information, l…
Interactive Tracking: A Human-in-the-Loop Paradigm with Memory-Augmented Adaptation
Yuqing Huang, Guotian Zeng, Zhenqiao Yuan +4
Existing visual trackers mainly operate in a non-interactive, fire-and-forget manner, making them impractical for real-world scenarios that require human-in-the-loop adaptation. To…
Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining
Weijun Zhuang, Yuqing Huang, Weikang Meng +5
Large-scale video-language pretraining enables strong generalization across multimodal tasks but often incurs prohibitive computational costs. Although recent advances in masked vi…
Prototype Perturbation for Relaxing Alignment Constraints in Backward-Compatible Learning
Zikun Zhou, Yushuai Sun, Wenjie Pei +2
The traditional paradigm to update retrieval models requires re-computing the embeddings of the gallery data, a time-consuming and computationally intensive process known as backfi…