3 papers
cs.AI2026
Reward Hacking in Language Model Agents: Revisiting AI Safety Gridworlds
Ãmer Veysel ÃaÄatan, Xuandong Zhao
Reward hacking, where AI systems exploit misspecified objectives to achieve high reward without satisfying intended goals, remains a central challenge in AI safety. Yet most known…
cs.LG2026
Clipping-Free Policy Optimization for Large Language Models
Ãmer Veysel ÃaÄatan, BarıŠAkgün, Gözde Gül Åahin +1
Reinforcement learning has become central to post-training large language models, yet dominant algorithms rely on clipping mechanisms that introduce optimization issues at scale, i…
cs.CL2025
MMTEB: Massive Multilingual Text Embedding Benchmark
Kenneth Enevoldsen, Isaac Chung, Imene Kerboua +83
Text embeddings are typically evaluated on a limited set of tasks, which are constrained by language, domain, and task diversity. To address these limitations and provide a more co…