From the 1 of 4 linked papers with an AI index.
4 papers
Plausible Deniability Guarantees for Whistleblowers
Leo Richter, Matt J. Kusner
The paper proposes formal privacy guarantees for whistleblowers by applying per-report (0,δ)-differential privacy to audit selection transcripts, using a reduction to private conti…
Agentic Uncertainty Reveals Agentic Overconfidence
Jean Kaddour, Srijan Patel, Gbètondji Dovonon +3
Can AI agents predict whether they will succeed at a task? We study agentic uncertainty by eliciting success probability estimates before, during, and after task execution. All res…
An Auditing Test To Detect Behavioral Shift in Language Models
Leo Richter, Xuanli He, Pasquale Minervini +1
As language models (LMs) approach human-level performance, a comprehensive understanding of their behavior becomes crucial. This includes evaluating capabilities, biases, task perf…
When Can Proxies Improve the Sample Complexity of Preference Learning?
Yuchen Zhu, Daniel Augusto de Souza, Zhengyan Shi +4
We address the problem of reward hacking, where maximising a proxy reward does not necessarily increase the true reward. This is a key concern for Large Language Models (LLMs), as…