5 citations · 6 across the 5 of their papers we have counts for
4 papers · 1 filter
Applying Refusal-Vector Ablation to Llama 3.1 70B Agents
Simon Lermen, Mateusz Dziemian, Govind Pimpale
Recently, language models like Llama 3.1 Instruct have become increasingly capable of agentic behavior, enabling them to perform tasks requiring short-term planning and tool use. I…
Exploring the Robustness of Model-Graded Evaluations and Automated Interpretability
Simon Lermen, Ondřej Kvapil
There has been increasing interest in evaluations of language models for a variety of risks and characteristics. Evaluations relying on natural language understanding for grading c…
BadLlama: cheaply removing safety fine-tuning from Llama 2-Chat 13B
Pranav Gade, Simon Lermen, Charlie Rogers-Smith +1
Llama 2-Chat is a collection of large language models that Meta developed and released to the public. While Meta fine-tuned Llama 2-Chat to refuse to output harmful content, we hyp…
Evaluating Shutdown Avoidance of Language Models in Textual Scenarios
Teun van der Weij, Simon Lermen, Leon lang
Recently, there has been an increase in interest in evaluating large language models for emergent and dangerous capabilities. Importantly, agents could reason that in some scenario…