Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
TTPO: Test-Time Policy Optimization
Aozhe Wang, Zhengxi Lu, Jianze Wang +8
Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large l…
cs.CL2026
SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning
Jianze Wang, Kunwang Zheng, Ying Liu +5
Test-time reinforcement learning (TTRL) enables language models to self-evolve at inference time without labeled feedback. Existing methods rely on answer voting and therefore do n…
cs.CL2024
MCSD: An Efficient Language Model with Diverse Fusion
Hua Yang, Duohai Li, Shiman Li
Transformers excel in Natural Language Processing (NLP) due to their prowess in capturing long-term dependencies but suffer from exponential resource consumption with increasing se…