4 papers
Extracting Paragraphs from LLM Token Activations
Nicholas Pochinkov, Angelo Benoit, Lovkush Agarwal +2
Generative large language models (LLMs) excel in natural language processing tasks, yet their inner workings remain underexplored beyond token-level predictions. This study investi…
Modularity in Transformers: Investigating Neuron Separability & Specialization
Nicholas Pochinkov, Thomas Jones, Mohammed Rashidur Rahman
Transformer models are increasingly prevalent in various applications, yet our understanding of their internal workings remains limited. This paper investigates the modularity and…
Investigating Neuron Ablation in Attention Heads: The Case for Peak Activation Centering
Nicholas Pochinkov, Ben Pasero, Skylar Shibayama
The use of transformer-based models is growing rapidly throughout society. With this growth, it is important to understand how they work, and in particular, how the attention mecha…
Dissecting Language Models: Machine Unlearning via Selective Pruning
Nicholas Pochinkov, Nandi Schoots
Understanding and shaping the behaviour of Large Language Models (LLMs) is increasingly important as applications become more powerful and more frequently adopted. This paper intro…