16 citations · 31 across the 14 of their papers we have counts for
Showing 2025 · cs.CLShow all
2 papers · 2 filters
cs.CL2025
Steering Language Models with Weight Arithmetic
Constanza Fierro, Fabien Roger
Providing high-quality feedback to Large Language Models (LLMs) on a diverse training distribution can be difficult and expensive, and providing feedback only on a narrow distribut…
cs.CL2025
Mechanistic Interpretability Needs Philosophy
Iwan Williams, Ninell Oldenburg, Ruchira Dhar +6
Mechanistic interpretability (MI) aims to explain how neural networks work by uncovering their underlying mechanisms. As the field grows in influence, it is increasingly important…