collaborators

11 papers

cs.CL2026

Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs

Nikita Afonin, Nikita Andriianov, Vahagn Hovhannisyan +9

Recent work has shown that narrow finetuning can produce broadly misaligned LLMs, a phenomenon termed emergent misalignment (EM). While concerning, these findings were limited to f…

cs.SI2026

@GrokSet: multi-party Human-LLM Interactions in Social Media

Matteo Migliarini, Berat Ercevik, Oluwagbemike Olowe +5

Large Language Models (LLMs) are increasingly deployed as active participants on public social media platforms, yet their behavior in these unconstrained social environments remain…

cs.LG2026

A Few Bad Neurons: Isolating and Surgically Correcting Sycophancy

Claire O'Brien, Jessica Seto, Dristi Roy +6

Behavioral alignment in large language models (LLMs) is often achieved through broad fine-tuning, which can result in undesired side effects like distributional shift and low inter…

cs.LG2025

Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs

Isha Chaturvedi, Anjana Nair, Yushen Li +5

We introduce Contrastive Region Masking (CRM), a training free diagnostic that reveals how multimodal large language models (MLLMs) depend on specific visual regions at each step o…

cs.CR2025

SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought

Shourya Batra, Pierce Tillman, Samarth Gaggar +6

As Large Language Models (LLMs) evolve into personal assistants with access to sensitive user data, they face a critical privacy challenge: while prior work has addressed output-le…

cs.CL2025

Modeling and Predicting Multi-Turn Answer Instability in Large Language Models

Jiahang He, Rishi Ramachandran, Neel Ramachandran +5

As large language models (LLMs) are adopted in an increasingly wide range of applications, user-model interactions have grown in both frequency and scale. Consequently, research ha…