5 papers · 1 filter
Interpreting Learned Feedback Patterns in Large Language Models
Luke Marks, Amir Abdullah, Clement Neo +4
Reinforcement learning from human feedback (RLHF) is widely used to train large language models (LLMs). However, it is unclear whether LLMs accurately learn the underlying preferen…
Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
Minseon Kim, Jin Myung Kwak, Lama Alssum +5
Fine-tuning language models is commonly believed to inevitably harm their safety, i.e., refusing to respond to harmful user requests, even when using harmless datasets, thus requir…
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
Michael Lan, Philip Torr, Austin Meek +3
The Universality Hypothesis in large language models (LLMs) claims that different models converge towards similar concept representations in their latent spaces. Providing evidence…
Open Problems in Machine Unlearning for AI Safety
Fazl Barez, Tingchen Fu, Ameya Prabhu +16
As AI systems become more capable, widely deployed, and increasingly autonomous in critical areas such as cybersecurity, biological research, and healthcare, ensuring their safety…
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
Luke Marks, Alasdair Paren, David Krueger +1
Sparse Autoencoders (SAEs) have shown promise in improving the interpretability of neural network activations, but can learn features that are not features of the input, limiting t…