collaborators

10 papers

cs.LG2026

Conditional Optimal Bridge for Riemannian Activation Steering

Seyed Arshan Dalili, Ajay Narayanan Sridhar, Vijaykrishnan Narayanan +1

Activation steering offers a lightweight alternative to fine-tuning for controlling large language models at inference time. While many existing methods implicitly optimize a log-d…

cs.LG2026

Correlation-Aware Contextual Bandits with Surrogate Rewards for LLM Routing

Ajay Narayanan Sridhar, Ronak Singh, Mehrdad Mahdavi +1

We study contextual bandit problems with correlated arms and access to surrogate reward signals produced by a machine learning model, motivated by applications such as large langua…

cs.CR2026

Curvature-Guided Module Localization for Low-Rank Detoxification of Backdoored Large Language Models

Arash Raftari, Mehrdad Mahdavi, Nathan Blackthorn +1

Backdoor attacks pose a serious threat to large language models (LLMs) by causing otherwise benign systems to produce attacker-specified malicious behavior when a hidden trigger is…

cs.LG2026

Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability

Seyed Arshan Dalili, Mehrdad Mahdavi

Sparse Autoencoders (SAEs) are widely used for mechanistic interpretability in large language models, yet their formulation assigns each latent feature a single decoder direction,…

cs.LG2025

Merge before Forget: A Single LoRA Continual Learning via Continual Merging

Fuli Qiao, Mehrdad Mahdavi

Parameter-efficient continual learning has emerged as a promising approach for large language models (LLMs) to mitigate catastrophic forgetting while enabling adaptation to new tas…

cs.LG2025

Model Merging via Multi-Teacher Knowledge Distillation

Seyed Arshan Dalili, Mehrdad Mahdavi

Model merging has emerged as a lightweight alternative to joint multi-task learning (MTL), yet the generalization properties of merged models remain largely unexplored. Establishin…