9 papers
VibeJam: A User Study Platform for Web Development with Agents
Nishant Balepur, Connor Baumler, Valerie Chen +3
Programming with AI is increasingly agentic, users prompt LLMs to directly edit their code and review the changes, with adoption growing especially for web development tasks. Despi…
Check The Scoreboard: An Analysis of Scoring Schemes on Multiple-Choice Evaluation
Nishant Balepur, Paiheng Xu, Wei Ai +3
Multiple-choice question answering (MCQA) benchmarks in NLP use number-right scoring (accuracy), but in educational testing, the scoring scheme, the combination of the response mod…
Language Models Encode the Contextual Truth of Propositions
Rupak Sarkar, Pritika Ramu, Rachel Rudinger
Prior work has shown that LLMs encode the truth of factual propositions along linear directions in activation space. It's unclear how these representations extend to contextual tru…
Sycophancy Undermines Epistemic Vigilance in Cooperative Vision-Language Tasks
Rupak Sarkar, Neha Srikanth, Saloni Gupta +3
To maintain common ground in cooperative conversation, humans iteratively update their beliefs as conversation participants share new information; participants who are epistemicall…
(Im)Paired Programming: Coding Agents Improve Productivity but Harm Understanding
Nishant Balepur, Connor Baumler, Valerie Chen +3
Coding agents (e.g., Cursor) improve developer productivity by optimizing task completion, but shifting users from writing code to prompting and reviewing may harm their understand…
On the Role of Citations in Preference Data
Yu Hou, Hal Daumé, Rachel Rudinger +1
Many NLP tasks require systems to provide attribution in their outputs--i.e. citations to grounding sources. Attribution serves as a bulwark against model hallucination and as a me…