activity
20182026
most citedMisleading Failures of Partial-input Baselines

6 citations · 6 across the 9 of their papers we have counts for

collaborators
Showing cs.CLShow all

19 papers · 1 filter

cs.CL2026

VibeJam: A User Study Platform for Web Development with Agents

Nishant Balepur, Connor Baumler, Valerie Chen +3

Programming with AI is increasingly agentic, users prompt LLMs to directly edit their code and review the changes, with adoption growing especially for web development tasks. Despi…

cs.CL2026

Check The Scoreboard: An Analysis of Scoring Schemes on Multiple-Choice Evaluation

Nishant Balepur, Paiheng Xu, Wei Ai +3

Multiple-choice question answering (MCQA) benchmarks in NLP use number-right scoring (accuracy), but in educational testing, the scoring scheme, the combination of the response mod…

cs.CL2026

(Im)Paired Programming: Coding Agents Improve Productivity but Harm Understanding

Nishant Balepur, Connor Baumler, Valerie Chen +3

Coding agents (e.g., Cursor) improve developer productivity by optimizing task completion, but shifting users from writing code to prompting and reviewing may harm their understand…

cs.CL2026

Measuring User's Mental Models of Speech Translation in Human-AI Collaboration

HyoJung Han, Nishant Balepur, Jordan Boyd-Graber +1

Millions of people use machine translation (MT) tools daily, yet little is known about their perception of what systems can and cannot do. This paper studies users' mental models o…

cs.CL2026

DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute

Nishant Balepur, Malachi Hamada, Varsha Kishore +9

Scientific Deep Research (DR) agents answer user queries by synthesizing research papers into multi-section reports. User feedback can improve their utility, but existing protocols…

cs.CL2026

Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users

Nishant Balepur, Malachi Hamada, Varsha Kishore +7

Deep Research (DR) systems help researchers cope with ballooning publishing counts. Such tools synthesize scientific papers to answer research queries, but lack understanding of th…