2 papers
cs.LG2026
Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Tests
Thanawat Lodkaew, Johannes Ackermann, Soichiro Nishimori +3
A growing failure mode in agent evaluation and training is that models can achieve high evaluation scores by exploiting shortcuts instead of solving the intended task, producing de…
cs.LG2025
On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
Soichiro Nishimori, Yu-Jie Zhang, Thanawat Lodkaew +1
Optimizing policies based on human preferences is key to aligning language models with human intent. This work focuses on reward modeling, a core component in reinforcement learnin…