2 papers
cs.CL2025
Flash Interpretability: Decoding Specialised Feature Neurons in Large Language Models with the LM-Head
Harry J Davies
Large Language Models (LLMs) typically have billions of parameters and are thus often difficult to interpret in their operation. In this work, we demonstrate that it is possible to…
cs.CL2024
Targeted Angular Reversal of Weights (TARS) for Knowledge Removal in Large Language Models
Harry J. Davies, Giorgos Iacovides, Danilo P. Mandic
The sheer scale of data required to train modern large language models (LLMs) poses significant risks, as models are likely to gain knowledge of sensitive topics such as bio-securi…