3 citations · 3 across the 4 of their papers we have counts for
4 papers
Copilot Arena: A Platform for Code LLM Evaluation in the Wild
Wayne Chi, Valerie Chen, Anastasios Nikolas Angelopoulos +7
Evaluating in-the-wild coding capabilities of large language models (LLMs) is a challenging endeavor with no clear solution. We introduce Copilot Arena, a platform to collect user…
FeedbackLogs: Recording and Incorporating Stakeholder Feedback into Machine Learning Pipelines
Matthew Barker, Emma Kallina, Dhananjay Ashok +6
Even though machine learning (ML) pipelines affect an increasing array of stakeholders, there is little work on how input from stakeholders is recorded and incorporated. We propose…
A Case Study on Designing Evaluations of ML Explanations with Simulated User Studies
Ada Martin, Valerie Chen, Sérgio Jesus +1
When conducting user studies to ascertain the usefulness of model explanations in aiding human decision-making, it is important to use real-world use cases, data, and users. Howeve…
Assisting Human Decisions in Document Matching
Joon Sik Kim, Valerie Chen, Danish Pruthi +2
Many practical applications, ranging from paper-reviewer assignment in peer review to job-applicant matching for hiring, require human decision makers to identify relevant matches…