4 papers
CIVA: Critic-Induced Value-Subspace Attacks on Visual World-Model Agents
Jiancheng Wang, Mingli Zhu, Tong Zhang +4
Visual world-model agents such as DreamerV3 act through a recurrent latent state rather than a single observation, which weakens frame-wise observation attacks and makes their pert…
Consistent Evidence, Robust Recognition: Faithful Attribution Regularization under Geometric Transformations
Xianghao Jiao, Ruoyu Chen, Wei Wang +6
Attribution methods are widely used to characterize the evidence underlying model predictions, yet their potential to improve model behavior remains underexplored. Attribution inco…
Domain Adaptive Object Detection via Dual-Stream Bilevel-Cycle Optimization
Yannan Chen, Wei Wang, Wenqiang Wang +5
Cycle self-training (CST) breaks the shared classifier assumption of the standard self-training framework, which is effective for unsupervised domain adaptation and exploits unlabe…
Vision Enhancing LLMs: Empowering Multimodal Knowledge Storage and Sharing in LLMs
Yunxin Li, Zhenyu Liu, Baotian Hu +4
Recent advancements in multimodal large language models (MLLMs) have achieved significant multimodal generation capabilities, akin to GPT-4. These models predominantly map visual i…