14 citations · 14 across the 2 of their papers we have counts for
3 papers
Probing Persona-Dependent Preferences in Language Models
Oscar Gilg, Pierre Beckmann, Daniel Paleka +1
Large language models (LLMs) can be said to have preferences: they reliably pick certain tasks and outputs over others, and preferences shaped by post-training and system prompts a…
Refusal in Language Models Is Mediated by a Single Direction
Andy Arditi, Oscar Obeso, Aaquib Syed +4
Conversational large language models are fine-tuned for both instruction-following and safety, resulting in models that obey benign requests but refuse harmful ones. While this ref…
ARB: Advanced Reasoning Benchmark for Large Language Models
Tomohiro Sawada, Daniel Paleka, Alexander Havrilla +6
Large Language Models (LLMs) have demonstrated remarkable performance on various quantitative reasoning and knowledge benchmarks. However, many of these benchmarks are losing utili…