1 paper · 1 filter
Pradyumna Shyama Prasad, Meiri Anto, Leon Eshuijs +3
LLM agents are increasingly used to run autonomous ML experiments, iterating on target metrics with little human oversight. Prior work has documented reward hacking in these enviro…