activity
20242026
most citedAuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors

2 citations · 2 across the 4 of their papers we have counts for

collaborators

8 papers

cs.CL2026

How Frontier LLMs Adapt to Neurodivergence Context: A Measurement Framework for Surface vs. Structural Change in System-Prompted Responses

Ishan Gupta, Pavlo Buryi

We examine if frontier chat-based large language models (LLMs) adjust their outputs based on neurodivergence (ND) context in system prompts and describe the nature of these adjustm…

cs.CL20262 cited

AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors

Abhay Sheshadri, Aidan Ewart, Kai Fronsdal +5

We introduce AuditBench, an alignment auditing benchmark. AuditBench consists of 56 language models with implanted hidden behaviors. Each model has one of 14 concerning behaviors--…

cs.LG2025

Scaling Laws and In-Context Learning: A Unified Theoretical Framework

Sushant Mehta, Ishan Gupta

In-context learning (ICL) enables large language models to adapt to new tasks from demonstrations without parameter updates. Despite extensive empirical studies, a principled under…

cs.LG2025

Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed

Isha Gupta, Rylan Schaeffer, Joshua Kazdan +2

The field of adversarial robustness has long established that adversarial examples can successfully transfer between image classifiers and that text jailbreaks can successfully tra…

cs.LG2025

Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track

Rylan Schaeffer, Joshua Kazdan, Yegor Denisov-Blanch +11

Science progresses by iteratively advancing and correcting humanity's understanding of the world. In machine learning (ML) research, rapid advancements have led to an explosion of…

cs.LG2025

"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models

Isha Gupta, David Khachaturov, Robert Mullins

The rise of multimodal large language models has introduced innovative human-machine interaction paradigms but also significant challenges in machine learning safety. Audio-Languag…