From the 1 of 14 linked papers with an AI index.
14 papers
(Im)Paired Programming: Coding Agents Improve Productivity but Harm Understanding
Nishant Balepur, Connor Baumler, Valerie Chen +3
The study finds that AI coding assistants help students finish a programming task faster but reduce their understanding of the code, making it harder for them to extend it later.
DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute
Nishant Balepur, Malachi Hamada, Varsha Kishore +9
Scientific Deep Research (DR) agents answer user queries by synthesizing research papers into multi-section reports. User feedback can improve their utility, but existing protocols…
Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
Nishant Balepur, Malachi Hamada, Varsha Kishore +7
Deep Research (DR) systems help researchers cope with ballooning publishing counts. Such tools synthesize scientific papers to answer research queries, but lack understanding of th…
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
Nishant Balepur, Bhavya Rajasekaran, Jane Oh +7
Multiple-choice question answering (MCQA) is standard in NLP, but benchmarks lack rigorous quality control. We present BenchMarker, an education-inspired toolkit using LLM judges t…
No Single Best Model for Diversity: Learning a Router for Sample Diversity
Yuhan Liu, Fangyuan Xu, Vishakh Padmakumar +2
When posed with prompts that permit a large number of valid answers, comprehensively generating them is the first step towards satisfying a wide range of users. In this paper, we s…
Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models
Ayush Rajesh Jhaveri, Anthony GX-Chen, Ilia Sucholutsky +1
Confirmation bias, the tendency to seek evidence that supports rather than challenges one's belief, hinders one's reasoning ability. We examine whether large language models (LLMs)…