activity
20242026
collaborators

7 papers

cs.AI2026

Understanding Annotator Safety Policy with Interpretability

Alex Oesterling, Donghao Ren, Yannick Assogba +4

Safety policies define what constitutes safe and unsafe AI outputs, guiding data annotation and model development. However, annotation disagreement is pervasive and can stem from m…

cs.CL2026

Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability

Usha Bhalla, Alex Oesterling, Claudio Mayrink Verdun +2

Translating the internal representations and computations of models into concepts that humans can understand is a key goal of interpretability. While recent dictionary learning met…

cs.LG2025

Inference-Time Reward Hacking in Large Language Models

Hadi Khalaf, Claudio Mayrink Verdun, Alex Oesterling +2

A common paradigm to improve the performance of large language models is optimizing for a reward model. Reward models assign a numerical score to an LLM's output that indicates, fo…

cs.CV2025

Multi-Group Proportional Representation for Text-to-Image Models

Sangwon Jung, Alex Oesterling, Claudio Mayrink Verdun +3

Text-to-image (T2I) generative models can create vivid, realistic images from textual descriptions. As these models proliferate, they expose new concerns about their ability to rep…

cs.IT2025

Soft Best-of-n Sampling for Model Alignment

Claudio Mayrink Verdun, Alex Oesterling, Himabindu Lakkaraju +1

Best-of- (BoN) sampling is a practical approach for aligning language model outputs with human preferences without expensive fine-tuning. BoN sampling is performed by generating…

cs.LG2024

Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)

Usha Bhalla, Alex Oesterling, Suraj Srinivas +2

CLIP embeddings have demonstrated remarkable performance across a wide range of multimodal applications. However, these high-dimensional, dense vector representations are not easil…