collaborators

19 papers

cs.AI2026

Asymmetric Collapse in Model Merging: When Refusal Over- writes Recognition

Aarnav Choudhary, Matheus Fonseca Rocha, Jiwon Seo +2

Model merging is often used to combine capabilities from separately fine-tuned models without additional training, but it is unclear whether standard merging methods preserve multi…

cs.MA2026

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation

Jasmine Brazilek, Maheep Chaudhary, Zoe Lu +1

Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failu…

cs.LG2026

Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillation

Uwe König, Hamza Kazmi, Ruizhe Li +1

Distillation of a language model intended to transfer benign behavior to a student model may also transfer undesirable characteristics, if they are present in the teacher model, a…

cs.AI2026

In-Context Environments Induce Evaluation-Awareness in Language Models

Maheep Chaudhary

Humans often become more self-aware under threat, yet can lose self-awareness when absorbed in a task; we hypothesize that language models exhibit environment-dependent \textit{eva…

cs.AI2026

LLM Scheming Inversely Scales with Pretraining Language Coverage

Nathan Truong, Aryan Panda, Rayming Ye +2

With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent work has empirically demonstrated in-con…

cs.LG2026

Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?

Jeanmely Rojas Nunez, Viraj Sawant, Nathan Allen +4

Fine-tuning large language models (LLMs) frequently induces catastrophic forgetting of prior capabilities. Recent work has shown that reinforcement learning (RL) retains prior capa…