3 papers
cs.LG2026
Weight Decay Improves Language Model Plasticity
Tessa Han, Sebastian Bordt, Hanlin Zhang +1
Large language models are typically trained in two broad phases: pretraining to produce a base model, followed by further training to improve downstream performance. However, hyper…
cs.LG2025
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
Shichang Zhang, Tessa Han, Usha Bhalla +1
The increasing complexity of AI systems has made understanding their behavior critical. Numerous interpretability methods have been developed to attribute model behavior to three k…
cs.LG2025
The Disagreement Problem in Explainable Machine Learning: A Practitioner's Perspective
Satyapriya Krishna, Tessa Han, Alex Gu +3
As various post hoc explanation methods are increasingly being leveraged to explain complex models in high-stakes settings, it becomes critical to develop a deeper understanding of…