961 citations · 1.8k across the 37 of their papers we have counts for
4 papers · 1 filter
ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues
Shanda Li, Qiuhong Anna Wei, Jingwu Tang +5
Reproducing research results from papers and released code is central to scientific progress. Existing works have introduced benchmarks to evaluate whether LLM agents can assist wi…
Completion Collaboration: Scaling Collaborative Effort with Agents
Shannon Zejiang Shen, Valerie Chen, Ken Gu +11
Current evaluations of agents remain centered around one-shot task completion, failing to account for the inherently iterative and collaborative nature of many real-world problems,…
ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data
Junhong Shen, Atishay Jain, Zedian Xiao +4
Large Language Model (LLM) agents are rapidly improving to handle increasingly complex web-based tasks. Most of these agents rely on general-purpose, proprietary models like GPT-4…
Do LLMs exhibit human-like response biases? A case study in survey design
Lindia Tjuatja, Valerie Chen, Sherry Tongshuang Wu +2
As large language models (LLMs) become more capable, there is growing excitement about the possibility of using LLMs as proxies for humans in real-world tasks where subjective labe…