collaborators

8 papers

cs.CL2026

Reasoning Fine-Tuning Induces Persistent Latent Policy States

Abir Harrasse, Michael Lan, Hunar Batra +2

Reasoning-specialized language models show large performance gains over base models, yet the internal changes responsible for improved multi-step reasoning remain poorly understood…

cs.SE2026

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

Chaithanya Bandi, Razvan-Gabriel Dumitru, Ben Hertzberg +20

The Model Context Protocol (MCP) is emerging as a standard interface through which large language model (LLM) agents discover and invoke external tools. However, existing MCP evalu…

cs.CY2026

Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing

Michael Lan, Narmeen Fatimah Oozeer, Chaithanya Bandi +4

While mechanistic interpretability (MI) has produced important insights into neural network internals, the field has yet to establish a standardized system to audit experiments. As…

cs.CL2026

Debate, Deliberate, Decide (D3): A Cost-Aware Adversarial Framework for Reliable and Interpretable LLM Evaluation

Abir Harrasse, Chaithanya Bandi, Hari Bandi

The evaluation of Large Language Models (LLMs) remains challenging due to inconsistency, bias, and the absence of transparent decision criteria in automated judging. We present Deb…

cs.AI2025

ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents

Manasi Sharma, Chen Bo Calvin Zhang, Chaithanya Bandi +13

Deep Research (DR) is an emerging agent application that leverages large language models (LLMs) to address open-ended queries. It requires the integration of several capabilities,…

cs.AI2025

Position: Require Frontier AI Labs To Release Small "Analog" Models

Shriyash Upadhyay, Chaithanya Bandi, Narmeen Oozeer +1

Recent proposals for regulating frontier AI models have sparked concerns about the cost of safety regulation, and most such regulations have been shelved due to the safety-innovati…