From the 1 of 5 linked papers with an AI index.
5 papers
Bad Memory: Evaluating Prompt Injection Risks from Memory in Agentic Systems
Soham Gadgil, David Alexander, Sai Sunku +1
The paper investigates how malicious instructions embedded in persistent memory files can be used to launch prompt injection attacks on agentic AI systems, evaluating several large…
Ensembling Sparse Autoencoders
Soham Gadgil, Chris Lin, Su-In Lee
Sparse autoencoders (SAEs) are used to decompose neural network activations into human-interpretable features. Typically, features learned by a single SAE are used for downstream a…
SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) Models
Mingyu Lu, Soham Gadgil, Chris Lin +2
As Text-to-Image (T2I) diffusion models are increasingly used in real-world creative workflows, a principled framework for valuing contributors who provide a collection of data is…
Where to Steer: Input-Dependent Layer Selection for Steering Improves LLM Alignment
Soham Gadgil, Chris Lin, Su-In Lee
Steering vectors have emerged as a lightweight and effective approach for aligning large language models (LLMs) at inference time, enabling modulation over model behaviors by shift…
Reflections from Research Roundtables at the Conference on Health, Inference, and Learning (CHIL) 2025
Emily Alsentzer, Marie-Laure Charpignon, Bill Chen +90
The 6th Annual Conference on Health, Inference, and Learning (CHIL 2025), hosted by the Association for Health Learning and Inference (AHLI), was held in person on June 25-27, 2025…