activity
20242026
most citedAuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors

2 citations · 2 across the 5 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2025

Scaling Laws and In-Context Learning: A Unified Theoretical Framework

Sushant Mehta, Ishan Gupta

In-context learning (ICL) enables large language models to adapt to new tasks from demonstrations without parameter updates. Despite extensive empirical studies, a principled under…

cs.LG2025

Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed

Isha Gupta, Rylan Schaeffer, Joshua Kazdan +2

The field of adversarial robustness has long established that adversarial examples can successfully transfer between image classifiers and that text jailbreaks can successfully tra…

cs.LG2025

"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models

Isha Gupta, David Khachaturov, Robert Mullins

The rise of multimodal large language models has introduced innovative human-machine interaction paradigms but also significant challenges in machine learning safety. Audio-Languag…

cs.LG2025

Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track

Rylan Schaeffer, Joshua Kazdan, Yegor Denisov-Blanch +11

Science progresses by iteratively advancing and correcting humanity's understanding of the world. In machine learning (ML) research, rapid advancements have led to an explosion of…

cs.LG2024

CIMRL: Combining IMitation and Reinforcement Learning for Safe Autonomous Driving

Jonathan Booher, Khashayar Rohanimanesh, Junhong Xu +5

Modern approaches to autonomous driving rely heavily on learned components trained with large amounts of human driving data via imitation learning. However, these methods require l…

cs.LG2024

Fragile Giants: Understanding the Susceptibility of Models to Subpopulation Attacks

Isha Gupta, Hidde Lycklama, Emanuel Opel +2

As machine learning models become increasingly complex, concerns about their robustness and trustworthiness have become more pressing. A critical vulnerability of these models is d…