2 papers
cs.CV2026
Learning What Not to Learn: Adversarial Disentangled Prompt Tuning for Robust Vision-Language Models
Yang Chen, Zhan Zhuang, Yanbin Wei +3
While adversarial prompt tuning can enhance robustness of vision-language models efficiently, we find that existing methods aggravate robust generalization overfitting on seen clas…
cs.LG2026
Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR
Chaoli Mou, Zhan Zhuang, Xinning Chen +1
Reinforcement Learning with Verifiable Rewards (RLVR) has become a key approach for improving the reasoning abilities of large language models. However, widely used critic-free alg…