collaborators

6 papers

cs.LG2025

Interpreting Learned Feedback Patterns in Large Language Models

Luke Marks, Amir Abdullah, Clement Neo +4

Reinforcement learning from human feedback (RLHF) is widely used to train large language models (LLMs). However, it is unclear whether LLMs accurately learn the underlying preferen…

cs.LG2025

Rethinking Safety in LLM Fine-tuning: An Optimization Perspective

Minseon Kim, Jin Myung Kwak, Lama Alssum +5

Fine-tuning language models is commonly believed to inevitably harm their safety, i.e., refusing to respond to harmful user requests, even when using harmless datasets, thus requir…

cs.CR2025

PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning

Tingchen Fu, Mrinank Sharma, Philip Torr +3

Preference learning is a central component for aligning current LLMs, but this process can be vulnerable to data poisoning attacks. To address this concern, we introduce PoisonBenc…

cs.LG2025

Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders

Michael Lan, Philip Torr, Austin Meek +3

The Universality Hypothesis in large language models (LLMs) claims that different models converge towards similar concept representations in their latent spaces. Providing evidence…

cs.CV2025

Towards Interpreting Visual Information Processing in Vision-Language Models

Clement Neo, Luke Ong, Philip Torr +3

Vision-Language Models (VLMs) are powerful tools for processing and understanding text and images. We study the processing of visual tokens in the language model component of LLaVA…

cs.LG2025

Open Problems in Machine Unlearning for AI Safety

Fazl Barez, Tingchen Fu, Ameya Prabhu +16

As AI systems become more capable, widely deployed, and increasingly autonomous in critical areas such as cybersecurity, biological research, and healthcare, ensuring their safety…