2 papers
cs.CV2026
Diffusion-based Cumulative Adversarial Purification for Vision Language Models
Jia Fu, Yongtao Wu, Yihang Chen +5
Vision Language Models (VLMs) have shown remarkable capabilities in multimodal understanding, yet their susceptibility to adversarial perturbations poses a significant threat to th…
cs.LG2025
Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees
Yongtao Wu, Luca Viano, Yihang Chen +4
Reinforcement Learning from Human Feedback (RLHF) has been highly successful in aligning large language models with human preferences. While prevalent methods like DPO have demonst…