3 papers
cs.LG2026
Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Tests
Thanawat Lodkaew, Johannes Ackermann, Soichiro Nishimori +3
A growing failure mode in agent evaluation and training is that models can achieve high evaluation scores by exploiting shortcuts instead of solving the intended task, producing de…
cs.LG2026
CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting
Takashi Ishida, Thanawat Lodkaew, Ikko Yamane
Publishing a large language model (LLM) benchmark (especially its ground-truth answers) on the Internet risks contaminating future LLMs and enabling evaluation gaming: it may be un…
cs.LG2025
On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
Soichiro Nishimori, Yu-Jie Zhang, Thanawat Lodkaew +1
Optimizing policies based on human preferences is key to aligning language models with human intent. This work focuses on reward modeling, a core component in reinforcement learnin…