From the 1 of 16 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning
Ziyan Liu, Xueda Shen, Yuzhe Gu +7
Large Reasoning Models (LRMs) have achieved remarkable progress thanks to Reinforcement Learning with Verifiable Rewards (RLVR) on Chain-of-Thoughts (CoTs). However, since long CoT…
cs.AI2025
RL in the Wild: Characterizing RLVR Training in LLM Deployment
Jiecheng Zhou, Qinghao Hu, Yuyang Jin +7
Large Language Models (LLMs) are now widely used across many domains. With their rapid development, Reinforcement Learning with Verifiable Rewards (RLVR) has surged in recent month…