2 papers
cs.CL2026
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients
Mingwei Xu, Hao Fang
Reinforcement learning with verifiable rewards (RLVR), due to the deterministic verification, becomes a dominant paradigm for enhancing the reasoning ability of large language mode…
cs.LG2025
Analytical Survey of Learning with Low-Resource Data: From Analysis to Investigation
Xiaofeng Cao, Mingwei Xu, Xin Yu +8
Learning with high-resource data has demonstrated substantial success in artificial intelligence (AI); however, the costs associated with data annotation and model training remain…