2 papers
cs.LG2026
When the Reward Suite Is Leaky: A Preregistered Causal Contrast of Natural Verifier False Positives in RLVR
Chuyifei Zhang
The test suites used as RLVR rewards for code have natural false positives: per-task, persistent, asymmetric errors that accept the same wrong programs every time they appear, unli…
cs.CL2026
Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks
Chuyifei Zhang, Hongyu Cui, Xiaowen Huang +1
Position-controlled evaluation is standard for retrieval tasks such as Needle-in-a-Haystack and RULER, but mainstream reasoning benchmarks do not control positional placement of ta…