activity
20132026
most citedBridging Mode Connectivity in Loss Landscapes and Adversarial Robustness

33 citations · 96 across the 25 of their papers we have counts for

collaborators
Showing 2024Show all

5 papers · 1 filter

cs.LG2024

Final-Model-Only Data Attribution with a Unifying View of Gradient-Based Methods

Dennis Wei, Inkit Padhi, Soumya Ghosh +3

Training data attribution (TDA) is concerned with understanding model behavior in terms of the training data. This paper draws attention to the common setting where one has access…

cs.CL20243 cited

Evaluating the Prompt Steerability of Large Language Models

Erik Miehling, Michael Desmond, Karthikeyan Natesan Ramamurthy +5

Building pluralistic AI requires designing models that are able to be shaped to represent a wide range of value systems and cultures. Achieving this requires first being able to ev…

cs.LG2024

Identifying Sub-networks in Neural Networks via Functionally Similar Representations

Tian Gao, Amit Dhurandhar, Karthikeyan Natesan Ramamurthy +1

Providing human-understandable insights into the inner workings of neural networks is an important step toward achieving more explainable and trustworthy AI. Existing approaches to…

cs.LG2024

Programming Refusal with Conditional Activation Steering

Bruce W. Lee, Inkit Padhi, Karthikeyan Natesan Ramamurthy +4

LLMs have shown remarkable capabilities, but precisely controlling their response behavior remains challenging. Existing activation steering methods alter LLM behavior indiscrimina…

cs.CL2024

Value Alignment from Unstructured Text

Inkit Padhi, Karthikeyan Natesan Ramamurthy, Prasanna Sattigeri +3

Aligning large language models (LLMs) to value systems has emerged as a significant area of research within the fields of AI and NLP. Currently, this alignment process relies on th…