581 citations · 1.6k across the 57 of their papers we have counts for
4 papers · 1 filter
How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions
Lorenzo Pacchiardi, Alex J. Chan, Sören Mindermann +5
Large language models (LLMs) can "lie", which we define as outputting false statements despite "knowing" the truth in a demonstrable sense. LLMs might "lie", for example, when inst…
In-Context Learning Learns Label Relationships but Is Not Conventional Learning
Jannik Kossen, Yarin Gal, Tom Rainforth
The predictions of Large Language Models (LLMs) on downstream tasks often improve significantly when including examples of the input--label relationship in the context. However, th…
Revisiting Automated Prompting: Are We Actually Doing Better?
Yulin Zhou, Yiren Zhao, Ilia Shumailov +2
Current literature demonstrates that Large Language Models (LLMs) are great few-shot learners, and prompting significantly increases their performance on a range of downstream task…
Wat zei je? Detecting Out-of-Distribution Translations with Variational Transformers
Tim Z. Xiao, Aidan N. Gomez, Yarin Gal
We detect out-of-training-distribution sentences in Neural Machine Translation using the Bayesian Deep Learning equivalent of Transformer models. For this we develop a new measure…