3 papers
cs.LG2026
BandPO: Bridging Trust Regions and Ratio Clipping via Probability-Aware Bounds for LLM Reinforcement Learning
Yuan Li, Bo Wang, Yufei Gao +4
Proximal constraints are fundamental to the stability of the Large Language Model reinforcement learning. While the canonical clipping mechanism in PPO serves as an efficient surro…
cs.LG2026
When Priors Backfire: On the Vulnerability of Unlearnable Examples to Pretraining
Zhihao Li, Gezheng Xu, Jiale Cai +5
Unlearnable Examples (UEs) serve as a data protection strategy that generates imperceptible perturbations to mislead models into learning spurious correlations instead of underlyin…
cs.LG2025
FedOne: Query-Efficient Federated Learning for Black-box Discrete Prompt Learning
Ganyu Wang, Jinjie Fang, Maxwell J. Yin +5
Black-Box Discrete Prompt Learning is a prompt-tuning method that optimizes discrete prompts without accessing model parameters or gradients, making the prompt tuning on a cloud-ba…