collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2025

Interpreting Learned Feedback Patterns in Large Language Models

Luke Marks, Amir Abdullah, Clement Neo +4

Reinforcement learning from human feedback (RLHF) is widely used to train large language models (LLMs). However, it is unclear whether LLMs accurately learn the underlying preferen…

cs.LG2025

Rethinking Safety in LLM Fine-tuning: An Optimization Perspective

Minseon Kim, Jin Myung Kwak, Lama Alssum +5

Fine-tuning language models is commonly believed to inevitably harm their safety, i.e., refusing to respond to harmful user requests, even when using harmless datasets, thus requir…

cs.LG2025

Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders

Michael Lan, Philip Torr, Austin Meek +3

The Universality Hypothesis in large language models (LLMs) claims that different models converge towards similar concept representations in their latent spaces. Providing evidence…

cs.LG2025

Open Problems in Machine Unlearning for AI Safety

Fazl Barez, Tingchen Fu, Ameya Prabhu +16

As AI systems become more capable, widely deployed, and increasingly autonomous in critical areas such as cybersecurity, biological research, and healthcare, ensuring their safety…

cs.LG2024

Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders

Luke Marks, Alasdair Paren, David Krueger +1

Sparse Autoencoders (SAEs) have shown promise in improving the interpretability of neural network activations, but can learn features that are not features of the input, limiting t…