collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

When Should We Introduce Safety Interventions During Pretraining?

Dylan Sam, Sachin Goyal, Pratyush Maini +2

Prior work has shown that safety interventions applied during pretraining, such as removing and rephrasing harmful content, can substantially improve the robustness of the resultin…

cs.LG2025

Evaluating Language Model Reasoning about Confidential Information

Dylan Sam, Alexander Robey, Andy Zou +2

As language models are increasingly deployed as autonomous agents in high-stakes settings, ensuring that they reliably follow user-defined rules has become a critical safety concer…

cs.LG2025

Command-V: Pasting LLM Behaviors via Activation Profiles

Barry Wang, Avi Schwarzschild, Alexander Robey +4

Retrofitting large language models (LLMs) with new behaviors typically requires full finetuning or distillation-costly steps that must be repeated for every architecture. In this w…

cs.LG2025

Existing Large Language Model Unlearning Evaluations Are Inconclusive

Zhili Feng, Yixuan Even Xu, Alexander Robey +5

Machine unlearning aims to remove sensitive or undesired data from large language models. However, recent studies suggest that unlearning is often shallow, claiming that removed kn…

cs.LG2025

Safety Pretraining: Toward the Next Generation of Safe AI

Pratyush Maini, Sachin Goyal, Dylan Sam +7

As large language models (LLMs) are increasingly deployed in high-stakes settings, the risk of generating harmful or toxic content remains a central challenge. Post-hoc alignment m…