3 papers
cs.CL2024
Removing Spurious Correlation from Neural Network Interpretations
Milad Fotouhi, Mohammad Taha Bahadori, Oluwaseyi Feyisetan +2
The existing algorithms for identification of neurons responsible for undesired and harmful behaviors do not consider the effects of confounders such as topic of the conversation.…
cs.CL2024
Fast Training Dataset Attribution via In-Context Learning
Milad Fotouhi, Mohammad Taha Bahadori, Oluwaseyi Feyisetan +2
We investigate the use of in-context learning and prompt engineering to estimate the contributions of training data in the outputs of instruction-tuned large language models (LLMs)…
stat.ME2024
Multiply-Robust Causal Change Attribution
Victor Quintas-Martinez, Mohammad Taha Bahadori, Eduardo Santiago +3
Comparing two samples of data, we observe a change in the distribution of an outcome variable. In the presence of multiple explanatory variables, how much of the change can be expl…