papers
Publications (3)
cs.CV2021
CheXbreak: Misclassification Identification for Deep Learning Models Interpreting Chest X-rays
Emma Chen, Andy Kim, Rayan Krishnan +3
A major obstacle to the integration of deep learning models for chest x-ray interpretation into clinical settings is the lack of understanding of their failure modes. In this work,…
cs.CR2026
ToxScreen: Detecting Whether an LLM Has Been Poisoned
Anthony Hughes, Nicole Xing, Collin Francel +2
The paper introduces ToxScreen, a benchmark of backdoored large language models, and evaluates methods for recovering hidden triggers under realistic defender constraints, finding…
#large language models#backdoor detection#adversarial attacks#model security
cs.CL2026
Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
Cameron Tice, Puria Radmard, Samuel Ratnam +3
Pretraining corpora contain extensive discourse about AI systems, yet the causal influence of this discourse on downstream alignment remains poorly understood. If prevailing descri…