3 citations · 3 across the 8 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Step-wise Rubric Rewards for LLM Reasoning
Weichu Xie, Haozhe Zhao, Wenpu Liu +15
Reinforcement Learning with Verifiable Rewards (RLVR) is widely used to improve reasoning in large language models, but rewards only final-answer correctness with no supervision ov…
cs.LG2026
LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning
Bowen Ping, Zijun Chen, Tingfeng Hui +4
Reinforcement Learning (RL) has emerged as a critical driver for enhancing the reasoning capabilities of Large Language Models (LLMs). While recent advancements have focused on rew…