2 papers
cs.LG2026
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
Narmeen Oozeer, Luke Marks, Shreyans Jain +2
Controlling multiple behavioral attributes in large language models (LLMs) at inference time is a challenging problem due to interference between attributes and the limitations of…
cs.CL2024
Scalable Influence and Fact Tracing for Large Language Model Pretraining
Tyler A. Chang, Dheeraj Rajagopal, Tolga Bolukbasi +2
Training data attribution (TDA) methods aim to attribute model outputs back to specific training examples, and the application of these methods to large language model (LLM) output…