3 papers
cs.AI2025
Representation Engineering for Large-Language Models: Survey and Research Challenges
Lukasz Bartoszcze, Sarthak Munshi, Bryan Sukidi +6
Large-language models are capable of completing a variety of tasks, but remain unpredictable and intractable. Representation engineering seeks to resolve this problem through a new…
cs.CL2024
Representation Noising: A Defence Mechanism Against Harmful Finetuning
Domenic Rosati, Jan Wehner, Kai Williams +7
Releasing open-source large language models (LLMs) presents a dual-use risk since bad actors can easily fine-tune these models for harmful purposes. Even without the open release o…
cs.CL2024
Immunization against harmful fine-tuning attacks
Domenic Rosati, Jan Wehner, Kai Williams +4
Large Language Models (LLMs) are often trained with safety guards intended to prevent harmful text generation. However, such safety training can be removed by fine-tuning the LLM o…