4 papers
One Lens, Many Worlds : A Capability-Typed Interface for World-Model Interpretability
Bhavith Chandra Challagundla, Sanskar Pandey, Param Thakkar +7
World models are now built on substantially different computational substrates. Latent recurrent state-space models such as PlaNet and the Dreamer family compress observations into…
Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Models
Sanskar Pandey, Ruhaan Chopra, Angkul Puniya +1
Large language models internalize a structural trade-off between truthfulness and obsequious flattery, emerging from reward optimization that conflates helpfulness with polite subm…
Hecto: Modular Sparse Experts for Adaptive and Interpretable Reasoning
Sanskar Pandey, Ruhaan Chopra, Saad Murtaza Bhat +1
Mixture-of-Experts (MoE) models enable conditional computation by routing inputs to specialized experts, but these experts rely on identical inductive biases, thus limiting represe…
Chronocept: Instilling a Sense of Time in Machines
Krish Goel, Sanskar Pandey, KS Mahadevan +2
Human cognition is deeply intertwined with a sense of time, known as Chronoception. This sense allows us to judge how long facts remain valid and when knowledge becomes outdated. D…