10 papers
Conditional Optimal Bridge for Riemannian Activation Steering
Seyed Arshan Dalili, Ajay Narayanan Sridhar, Vijaykrishnan Narayanan +1
Activation steering offers a lightweight alternative to fine-tuning for controlling large language models at inference time. While many existing methods implicitly optimize a log-d…
Correlation-Aware Contextual Bandits with Surrogate Rewards for LLM Routing
Ajay Narayanan Sridhar, Ronak Singh, Mehrdad Mahdavi +1
We study contextual bandit problems with correlated arms and access to surrogate reward signals produced by a machine learning model, motivated by applications such as large langua…
Curvature-Guided Module Localization for Low-Rank Detoxification of Backdoored Large Language Models
Arash Raftari, Mehrdad Mahdavi, Nathan Blackthorn +1
Backdoor attacks pose a serious threat to large language models (LLMs) by causing otherwise benign systems to produce attacker-specified malicious behavior when a hidden trigger is…
Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability
Seyed Arshan Dalili, Mehrdad Mahdavi
Sparse Autoencoders (SAEs) are widely used for mechanistic interpretability in large language models, yet their formulation assigns each latent feature a single decoder direction,…
Merge before Forget: A Single LoRA Continual Learning via Continual Merging
Fuli Qiao, Mehrdad Mahdavi
Parameter-efficient continual learning has emerged as a promising approach for large language models (LLMs) to mitigate catastrophic forgetting while enabling adaptation to new tas…
Model Merging via Multi-Teacher Knowledge Distillation
Seyed Arshan Dalili, Mehrdad Mahdavi
Model merging has emerged as a lightweight alternative to joint multi-task learning (MTL), yet the generalization properties of merged models remain largely unexplored. Establishin…