2 papers
cs.LG2025
Accelerating RLHF Training with Reward Variance Increase
Zonglin Yang, Zhexuan Gu, Houduo Qi +1
Reinforcement learning from human feedback (RLHF) is an essential technique for ensuring that large language models (LLMs) are aligned with human values and preferences during the…
cs.CV2025
ZeroPur: Succinct Training-Free Adversarial Purification
Erhu Liu, Zonglin Yang, Bo Liu +4
Adversarial purification is a kind of defense technique that can defend against various unseen adversarial attacks without modifying the victim classifier. Existing methods often d…