8 papers · 1 filter
Asymmetric Collapse in Model Merging: When Refusal Over- writes Recognition
Aarnav Choudhary, Matheus Fonseca Rocha, Jiwon Seo +2
Model merging is often used to combine capabilities from separately fine-tuned models without additional training, but it is unclear whether standard merging methods preserve multi…
In-Context Environments Induce Evaluation-Awareness in Language Models
Maheep Chaudhary
Humans often become more self-aware under threat, yet can lose self-awareness when absorbed in a task; we hypothesize that language models exhibit environment-dependent \textit{eva…
LLM Scheming Inversely Scales with Pretraining Language Coverage
Nathan Truong, Aryan Panda, Rayming Ye +2
With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent work has empirically demonstrated in-con…
Playing Devil's Advocate: Off-the-Shelf Persona Vectors Rival Targeted Steering for Sycophancy
Ishaan Kelkar, Nebras Alam, Vikram Kakaria +3
We study the effect of different persona on \textbf{sycophancy}: model's agreement with users even when the user is incorrect. The standard mitigation, Contrastive Activation Addit…
Why Limit the Residual Stream to Layers and Not Tokens? Persistent Memory for Continuous Latent Reasoning
Mujtaba Farhan, Maheep Chaudhary
Large language models (LLMs) have demonstrated remarkable reasoning abilities on mathematical and multi-hop planning tasks. The CoCoNuT (Chain of Continuous Thought) paradigm~\cite…
Optimizing Chain-of-Thought Confidence via Topological and Dirichlet Risk Analysis
Abhishek More, Anthony Zhang, Nicole Bonilla +4
Chain-of-thought (CoT) prompting enables Large Language Models to solve complex problems, but deploying these models safely requires reliable confidence estimates, a capability whe…