13 papers
Did Models Learn Sufficiently? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation
Yannan Chen, Ruoyu Chen, Wei Wang +6
Current visual models often make predictions based on a limited set of discriminative visual cues. As a result, they may become unreliable when the distribution shifts or when thes…
Domain Adaptive Object Detection via Dual-Stream Bilevel-Cycle Optimization
Yannan Chen, Wei Wang, Wenqiang Wang +5
Cycle self-training (CST) breaks the shared classifier assumption of the standard self-training framework, which is effective for unsupervised domain adaptation and exploits unlabe…
NormDirection: Restoring the Missing Query Norm in Vision Linear Attention
Weikang Meng, Yadan Luo, Liangyu Huo +4
Linear attention mitigates the quadratic complexity of softmax attention but suffers from a critical loss of expressiveness. We identify two primary causes: (1) The normalization o…
CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
Qiming Li, Zekai Ye, Xiaocheng Feng +9
Although Large Vision-Language Models (LVLMs) have demonstrated remarkable performance on downstream tasks, they frequently produce contents that deviate from visual information, l…
MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages
Yexing Du, Kaiyuan Liu, Youcheng Pan +7
Multimodal Large Language Models (MLLMs) have achieved great success in Speech-to-Text Translation (S2TT) tasks. However, current research is constrained by two key challenges: lan…
Interactive Tracking: A Human-in-the-Loop Paradigm with Memory-Augmented Adaptation
Yuqing Huang, Guotian Zeng, Zhenqiao Yuan +4
Existing visual trackers mainly operate in a non-interactive, fire-and-forget manner, making them impractical for real-world scenarios that require human-in-the-loop adaptation. To…