activity
20242026
collaborators

5 papers

cs.LG2026

Position: Explainability Research Must Prioritize Foundations over Ad-hoc Methods

Michal Moshkovitz, Suraj Srinivas, Lesia Semenova +7

Despite the proliferation of Explainable AI (XAI) techniques -- from feature attributions to sparse autoencoders -- explanations rarely influence real-world workflows. In practice,…

cs.LG2025

The Disagreement Problem in Explainable Machine Learning: A Practitioner's Perspective

Satyapriya Krishna, Tessa Han, Alex Gu +3

As various post hoc explanation methods are increasingly being leveraged to explain complex models in high-stakes settings, it becomes critical to develop a deeper understanding of…

cs.CL2025

On the Impact of Fine-Tuning on Chain-of-Thought Reasoning

Elita Lobo, Chirag Agarwal, Himabindu Lakkaraju

Large language models have emerged as powerful tools for general intelligence, showcasing advanced natural language processing capabilities that find applications across diverse do…

cs.CL2025

Certifying LLM Safety against Adversarial Prompting

Aounon Kumar, Chirag Agarwal, Suraj Srinivas +3

Large language models (LLMs) are vulnerable to adversarial attacks that add malicious tokens to an input prompt to bypass the safety guardrails of an LLM and cause it to produce ha…

cs.CL2024

More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness

Aaron J. Li, Satyapriya Krishna, Himabindu Lakkaraju

The trustworthiness of Large Language Models (LLMs) refers to the extent to which their outputs are reliable, safe, and ethically aligned, and it has become a crucial consideration…