3 citations · 7 across the 8 of their papers we have counts for
4 papers · 1 filter
Evaluating the Goal-Directedness of Large Language Models
Tom Everitt, Cristina Garbacea, Alexis Bellot +4
To what extent do LLMs use their capabilities towards their given goal? We take this as a measure of their goal-directedness. We evaluate goal-directedness on tasks that require in…
Mapping AI Benchmark Data to Quantitative Risk Estimates Through Expert Elicitation
Malcolm Murray, Henry Papadatos, Otter Quarks +2
The literature and multiple experts point to many potential risks from large language models (LLMs), but there are still very few direct measurements of the actual harms posed. AI…
A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management
Simeon Campos, Henry Papadatos, Fabien Roger +3
The recent development of powerful AI systems has highlighted the need for robust risk management frameworks in the AI industry. Although companies have begun to implement safety f…
Linear Probe Penalties Reduce LLM Sycophancy
Henry Papadatos, Rachel Freedman
Large language models (LLMs) are often sycophantic, prioritizing agreement with their users over accurate or objective statements. This problematic behavior becomes more pronounced…