19 papers
Asymmetric Collapse in Model Merging: When Refusal Over- writes Recognition
Aarnav Choudhary, Matheus Fonseca Rocha, Jiwon Seo +2
Model merging is often used to combine capabilities from separately fine-tuned models without additional training, but it is unclear whether standard merging methods preserve multi…
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
Jasmine Brazilek, Maheep Chaudhary, Zoe Lu +1
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failu…
Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillation
Uwe König, Hamza Kazmi, Ruizhe Li +1
Distillation of a language model intended to transfer benign behavior to a student model may also transfer undesirable characteristics, if they are present in the teacher model, a…
In-Context Environments Induce Evaluation-Awareness in Language Models
Maheep Chaudhary
Humans often become more self-aware under threat, yet can lose self-awareness when absorbed in a task; we hypothesize that language models exhibit environment-dependent \textit{eva…
LLM Scheming Inversely Scales with Pretraining Language Coverage
Nathan Truong, Aryan Panda, Rayming Ye +2
With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent work has empirically demonstrated in-con…
Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?
Jeanmely Rojas Nunez, Viraj Sawant, Nathan Allen +4
Fine-tuning large language models (LLMs) frequently induces catastrophic forgetting of prior capabilities. Recent work has shown that reinforcement learning (RL) retains prior capa…