3 papers
cs.CL2026
Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Models
Sanskar Pandey, Ruhaan Chopra, Angkul Puniya +1
Large language models internalize a structural trade-off between truthfulness and obsequious flattery, emerging from reward optimization that conflates helpfulness with polite subm…
cs.AI2026
A Mechanistic Investigation of Supervised Fine Tuning
Ruhaan Chopra
The cosine similarity between a large language model's hidden activations before and after Supervised Fine-Tuning (SFT) remains very high. This, at first glance, suggests that SFT…
cs.AI2025
Hecto: Modular Sparse Experts for Adaptive and Interpretable Reasoning
Sanskar Pandey, Ruhaan Chopra, Saad Murtaza Bhat +1
Mixture-of-Experts (MoE) models enable conditional computation by routing inputs to specialized experts, but these experts rely on identical inductive biases, thus limiting represe…