1 citations · 1 across the 4 of their papers we have counts for
9 papers
Reasoning Fine-Tuning Induces Persistent Latent Policy States
Abir Harrasse, Michael Lan, Hunar Batra +2
Reasoning-specialized language models show large performance gains over base models, yet the internal changes responsible for improved multi-step reasoning remain poorly understood…
ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
Manasi Sharma, Chen Bo Calvin Zhang, Chaithanya Bandi +13
Deep Research (DR) is an emerging agent application that leverages large language models (LLMs) to address open-ended queries. It requires the integration of several capabilities,…
Position: Require Frontier AI Labs To Release Small "Analog" Models
Shriyash Upadhyay, Chaithanya Bandi, Narmeen Oozeer +1
Recent proposals for regulating frontier AI models have sparked concerns about the cost of safety regulation, and most such regulations have been shelved due to the safety-innovati…
Beyond Monoliths: Expert Orchestration for More Capable, Democratic, and Safe Language Models
Philip Quirke, Narmeen Oozeer, Chaithanya Bandi +8
This position paper argues that the prevailing trajectory toward ever larger, more expensive generalist foundation models controlled by a handful of companies limits innovation and…
REL: Working out is all you need
Toby Simonds, Jey Han Lau, Chaithanya Bandi
Recent developments, particularly OpenAI's O1 model, have demonstrated the remarkable potential of Large Language Models (LLMs) for complex reasoning tasks. Through analysis of O1'…
Debate, Deliberate, Decide (D3): A Cost-Aware Adversarial Framework for Reliable and Interpretable LLM Evaluation
Abir Harrasse, Chaithanya Bandi, Hari Bandi
The evaluation of Large Language Models (LLMs) remains challenging due to inconsistency, bias, and the absence of transparent decision criteria in automated judging. We present Deb…