3 papers
cs.AI2026
Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks
Amit Roth, Ivan Bercovich, Yonathan Efroni
As agents grow more capable and autonomous, their tendency to reward hack, satisfying a task's checks while violating its intent, becomes an increasingly important failure mode. Me…
cs.LG2026
Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias
Ofek I. Cohen, Lior Shani, Aviv Rosenberg +3
Many organizations aim to adapt language models for internal use, both to improve performance on domain-specific tasks and to address privacy concerns around sensitive data. Howeve…
cs.AI2026
BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation
Ankur Samanta, Akshayaa Magesh, Tal Lancewicki +7
Large language models (LLMs) are typically deployed in multi-turn conversations, where each turn provides new evidence that should reduce epistemic uncertainty about their environm…