Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
Amit Roth, Ankur Samanta, Matan Halevy +2
Aligning autonomous agents with human intent remains a central challenge in modern AI. A key manifestation of this challenge is reward hacking, whereby agents appear successful und…
cs.LG2025
Language model developers should report train-test overlap
Andy K Zhang, Kevin Klyman, Yifan Mai +4
Language models are extensively evaluated, but correctly interpreting evaluation results requires knowledge of train-test overlap which refers to the extent to which the language m…