From the 2 of 59 linked papers with an AI index.
59 papers
CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment
Chengxiao Wang, Enyi Jiang, Xiaojing Liao +1
Improving the safety of large language models (LLMs) often comes at the expense of utility, as globally applied safety tuning may affect model responses to both harmful and benign…
What to Forget in Unlearning? Forget Set Curation for Language Models
Animesh Jha, Arpandeep Khatua, Youssef Allouah +1
Machine unlearning aims to remove targeted data or behaviors from a trained model without retraining from scratch. Yet most evaluations assume that the examples to forget are alrea…
Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation
Yuta Kobayashi, Pradyun Ramesh, Muhammad Ahmed Chaudhry +5
Vision-Language Models (VLMs) for radiology report generation are typically trained on retrospective clinical reports, which suffer from omission noise: clinically present findings…
The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context
Yanzhe Zhang, Sanmi Koyejo, Diyi Yang
The paper shows that while large language models seem robust to irrelevant context when measured by overall accuracy, adding even meaningless pseudo‑words can cause prediction flip…
Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack
Cristian Trout, Sanmi Koyejo, Sasha Romanosky +34
The paper proposes a comprehensive AI insurance framework to price and manage risks from the emerging AI agent economy, outlining an eight‑component stack for data collection, mode…
Auditing of Unlearning Algorithms
Sahasrajit Sarmasarkar, Anastasia Koloskova, Sanmi Koyejo
Evaluating whether unlearning algorithms truly remove training data influence remains an open challenge. We propose a practical auditor that computes data-dependent lower bounds on…