collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

A Few Bad Neurons: Isolating and Surgically Correcting Sycophancy

Claire O'Brien, Jessica Seto, Dristi Roy +6

Behavioral alignment in large language models (LLMs) is often achieved through broad fine-tuning, which can result in undesired side effects like distributional shift and low inter…

cs.LG2025

Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs

Isha Chaturvedi, Anjana Nair, Yushen Li +5

We introduce Contrastive Region Masking (CRM), a training free diagnostic that reveals how multimodal large language models (MLLMs) depend on specific visual regions at each step o…

cs.LG2025

Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits

Dev Patel, Gabrielle Gervacio, Diekola Raimi +5

Large Language Models require substantial computational resources for inference, posing deployment challenges. While dynamic pruning offers superior efficiency over static methods…

cs.LG2025

Amortized Latent Steering: Low-Cost Alternative to Test-Time Optimization

Nathan Egbuna, Saatvik Gaur, Sunishchal Dev +2

Test-time optimization remains impractical at scale due to prohibitive inference costs--techniques like iterative refinement and multi-step verification can require

cs.LG2025

Shared Parameter Subspaces and Cross-Task Linearity in Emergently Misaligned Behavior

Daniel Aarao Reis Arturi, Eric Zhang, Andrew Ansah +3

Recent work has discovered that large language models can develop broadly misaligned behaviors after being fine-tuned on narrowly harmful datasets, a phenomenon known as emergent m…