2 papers
cs.LG2026
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment
Zhanyu Liu, Qingguo Hu, Ante Wang +5
Reinforcement Learning with Verifiable Reward (RLVR) has proven effective for training reasoning-oriented large language models, but existing methods largely assume high-resource s…
cs.LG2026
FedCARE: Federated Unlearning with Conflict-Aware Projection and Relearning-Resistant Recovery
Yue Li, Mingmin Chu, Xilei Yang +5
Federated learning (FL) enables collaborative model training without centralizing raw data, but privacy regulations such as the right to be forgotten require FL systems to remove t…