3 papers
cs.AI2026
Agentic Uncertainty Reveals Agentic Overconfidence
Jean Kaddour, Srijan Patel, Gbètondji Dovonon +3
Can AI agents predict whether they will succeed at a task? We study agentic uncertainty by eliciting success probability estimates before, during, and after task execution. All res…
cs.LG2024
When Can Proxies Improve the Sample Complexity of Preference Learning?
Yuchen Zhu, Daniel Augusto de Souza, Zhengyan Shi +4
We address the problem of reward hacking, where maximising a proxy reward does not necessarily increase the true reward. This is a key concern for Large Language Models (LLMs), as…
cs.LG2024
An Auditing Test To Detect Behavioral Shift in Language Models
Leo Richter, Xuanli He, Pasquale Minervini +1
As language models (LMs) approach human-level performance, a comprehensive understanding of their behavior becomes crucial. This includes evaluating capabilities, biases, task perf…