2 papers
cs.LG2026
Prototype Language Models
Dan Ley, Giang Nguyen, Himabindu Lakkaraju +1
Knowing which training examples drive outputs is fundamental to auditing, correcting, and understanding language models, yet for modern LLMs this remains expensive, approximate, an…
cs.LG2024
Generalized Group Data Attribution
Dan Ley, Suraj Srinivas, Shichang Zhang +2
Data Attribution (DA) methods quantify the influence of individual training data points on model outputs and have broad applications such as explainability, data selection, and noi…