7 citations · 10 across the 3 of their papers we have counts for
3 papers
Dense Reward for Free in Reinforcement Learning from Human Feedback
Alex J. Chan, Hao Sun, Samuel Holt +1
Reinforcement Learning from Human Feedback (RLHF) has been credited as the key advance that has allowed Large Language Models (LLMs) to effectively follow instructions and produce…
Optimising Human-AI Collaboration by Learning Convincing Explanations
Alex J. Chan, Alihan Huyuk, Mihaela van der Schaar
Machine learning models are being increasingly deployed to take, or assist in taking, complicated and high-impact decisions, from quasi-autonomous vehicles to clinical decision sup…
How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions
Lorenzo Pacchiardi, Alex J. Chan, Sören Mindermann +5
Large language models (LLMs) can "lie", which we define as outputting false statements despite "knowing" the truth in a demonstrable sense. LLMs might "lie", for example, when inst…