Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
PRIME: A Process-Outcome Alignment Benchmark for Verifiable Reasoning in Mathematics and Engineering
Xiangfeng Wang, Hangyu Guo, Yanlin Lai +11
While model-based verifiers are essential for scaling Reinforcement Learning with Verifiable Rewards (RLVR), current outcome-centric verification paradigms primarily focus on the c…
cs.CL2026
R-Align: Enhancing Generative Reward Models through Rationale-Centric Meta-Judging
Yanlin Lai, Mitt Huang, Hangyu Guo +11
Reinforcement Learning from Human Feedback (RLHF) remains indispensable for aligning large language models (LLMs) in subjective domains. To enhance robustness, recent work shifts t…
cs.CL2026
ECR: Manifold-Guided Semantic Cues for Compact Language Models
Chung-Wei Victor Yuan
Compact models often lose the structure of their embedding space. The issue shows up when the capacity is tight or the data spans several languages. Such collapse makes it difficul…