4 papers
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
Amit Roth, Ankur Samanta, Matan Halevy +2
Aligning autonomous agents with human intent remains a central challenge in modern AI. A key manifestation of this challenge is reward hacking, whereby agents appear successful und…
Language model developers should report train-test overlap
Andy K Zhang, Kevin Klyman, Yifan Mai +4
Language models are extensively evaluated, but correctly interpreting evaluation results requires knowledge of train-test overlap which refers to the extent to which the language m…
FormulaOne: Measuring the Depth of Algorithmic Reasoning Beyond Competitive Programming
Gal Beniamini, Yuval Dor, Alon Vinnikov +10
Frontier AI models demonstrate formidable breadth of knowledge. But how close are they to true human -- or superhuman -- expertise? Genuine experts can tackle the hardest problems…
Tradeoffs Between Alignment and Helpfulness in Language Models with Steering Methods
Yotam Wolf, Noam Wies, Dorin Shteyman +3
Language model alignment has become an important component of AI safety, allowing safe interactions between humans and language models, by enhancing desired behaviors and inhibitin…