Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks
Amit Roth, Ivan Bercovich, Yonathan Efroni
As agents grow more capable and autonomous, their tendency to reward hack, satisfying a task's checks while violating its intent, becomes an increasingly important failure mode. Me…
cs.AI2026
BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation
Ankur Samanta, Akshayaa Magesh, Tal Lancewicki +7
Large language models (LLMs) are typically deployed in multi-turn conversations, where each turn provides new evidence that should reduce epistemic uncertainty about their environm…