2 papers
cs.AI2026
Mitigating Distribution Sharpening in Math RLVR via Distribution-Aligned Hint Synthesis and Backward Hint Annealing
Pei-Xi Xie, Che-Yu Lin, Cheng-Lin Yang
Reinforcement learning with verifiable rewards (RLVR) can improve low- reasoning accuracy while narrowing solution coverage on challenging math questions, and pass@1 gains do no…
cs.CL2024
Lightweight Contenders: Navigating Semi-Supervised Text Mining through Peer Collaboration and Self Transcendence
Qianren Mao, Weifeng Jiang, Junnan Liu +5
The semi-supervised learning (SSL) strategy in lightweight models requires reducing annotated samples and facilitating cost-effective inference. However, the constraint on model pa…