3 papers
cs.CL2026
When a Data Artifact Isn't a Shortcut: Causal Auditing of Synthetic RLVR Corpora
Esther Xin
Several recent pipelines build RLVR training data by masking a span of real corpus text and asking a language model to invent plausible wrong answers around it. The correct option…
cs.LG2026
Are Verifier Errors Independent Within a GRPO Group? Evidence from Qwen2.5 Rollouts
Esther Xin
Group-based reinforcement learning with verifiable rewards (RLVR) scoresmultiple completions per prompt using automatic verifiers. Analysesbased on independent verifier errors may…
cs.CL2026
Where the Verifier Fails: A Category-Level Audit of Reward Signals in RLVR
Esther Xin
Reinforcement learning with verifiable rewards (RLVR) and standard benchmark evaluation both rely on an automatic verifier that turns a free text answer into a binary reward. Prior…