audit 1causal evaluation 1code generation 1reinforcement learning 1reward design 1test suite leakage 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.LG2026
When the Reward Suite Is Leaky: A Preregistered Causal Contrast of Natural Verifier False Positives in RLVR
Chuyifei Zhang
The paper investigates how test suites used as rewards for reinforcement‑learning‑based code generation can contain systematic false positives, and shows that hardening the suite r…
cs.CL2026
Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks
Chuyifei Zhang, Hongyu Cui, Xiaowen Huang +1
Position-controlled evaluation is standard for retrieval tasks such as Needle-in-a-Haystack and RULER, but mainstream reasoning benchmarks do not control positional placement of ta…