4 papers
Natively Unlearnable Large Language Models
Gaurav R. Ghosal, Pratyush Maini, Aditi Raghunathan
Unlearning aims to remove the influence of specific training data sources, but this has proved challenging because the contributions of different sources are entangled within the m…
Early Data Exposure Improves Robustness to Subsequent Fine-Tuning
Lawrence Feng, Gaurav R. Ghosal, Jacob Mitchell Springer +2
How can we train models whose post-trained capabilities survive subsequent fine-tuning? Rather than focusing on downstream interventions to mitigate forgetting of upstream capabili…
Can LLMs Reconcile Knowledge Conflicts in Counterfactual Reasoning
Khurram Yamin, Gaurav Ghosal, Bryan Wilder
Large Language Models have been shown to contain extensive world knowledge in their parameters, enabling impressive performance on many knowledge intensive tasks. However, when dep…
Failure Modes of LLMs for Causal Reasoning on Narratives
Khurram Yamin, Shantanu Gupta, Gaurav R. Ghosal +2
The ability to robustly identify causal relationships is essential for autonomous decision-making and adaptation to novel scenarios. However, accurately inferring causal structure…