Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
Applying Refusal-Vector Ablation to Llama 3.1 70B Agents
Simon Lermen, Mateusz Dziemian, Govind Pimpale
Recently, language models like Llama 3.1 Instruct have become increasingly capable of agentic behavior, enabling them to perform tasks requiring short-term planning and tool use. I…
cs.CL2024
BadLlama: cheaply removing safety fine-tuning from Llama 2-Chat 13B
Pranav Gade, Simon Lermen, Charlie Rogers-Smith +1
Llama 2-Chat is a collection of large language models that Meta developed and released to the public. While Meta fine-tuned Llama 2-Chat to refuse to output harmful content, we hyp…