6 papers
Weight Decay Improves Language Model Plasticity
Tessa Han, Sebastian Bordt, Hanlin Zhang +1
Large language models are typically trained in two broad phases: pretraining to produce a base model, followed by further training to improve downstream performance. However, hyper…
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
Shichang Zhang, Tessa Han, Usha Bhalla +1
The increasing complexity of AI systems has made understanding their behavior critical. Numerous interpretability methods have been developed to attribute model behavior to three k…
The Disagreement Problem in Explainable Machine Learning: A Practitioner's Perspective
Satyapriya Krishna, Tessa Han, Alex Gu +3
As various post hoc explanation methods are increasingly being leveraged to explain complex models in high-stakes settings, it becomes critical to develop a deeper understanding of…
MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models
Tessa Han, Aounon Kumar, Chirag Agarwal +1
As large language models (LLMs) develop increasingly sophisticated capabilities and find applications in medical settings, it becomes important to assess their medical safety due t…
Hevelius Report: Visualizing Web-Based Mobility Test Data For Clinical Decision and Learning Support
Hongjin Lin, Tessa Han, Krzysztof Z. Gajos +1
Hevelius, a web-based computer mouse test, measures arm movement and has been shown to accurately evaluate severity for patients with Parkinson's disease and ataxias. A Hevelius se…
Characterizing Data Point Vulnerability via Average-Case Robustness
Tessa Han, Suraj Srinivas, Himabindu Lakkaraju
Studying the robustness of machine learning models is important to ensure consistent model behaviour across real-world settings. To this end, adversarial robustness is a standard f…