2 papers
cs.LG2025
Priors in Time: Missing Inductive Biases for Language Model Interpretability
Ekdeep Singh Lubana, Can Rager, Sai Sumedh R. Hindupur +13
Recovering meaningful concepts from language model activations is a central aim of interpretability. While existing feature extraction methods aim to identify concepts that are ind…
cs.LG2025
Towards Interpretable Soft Prompts
Oam Patel, Jason Wang, Nikhil Shivakumar Nayak +2
Soft prompts have been popularized as a cheap and easy way to improve task-specific LLM performance beyond few-shot prompts. Despite their origin as an automated prompting method,…