5 papers
SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations
Shuaiqi Wang, Aadyaa Maddi, Zinan Lin +1
Today, tool-calling agents are commonly evaluated or tested on static datasets of execution traces, including input commands, agent responses, and associated tool calls. However, i…
Smooth Partial Lotteries for Stable Randomized Selection
Alexander Goldberg, Giulia Fanti, Nihar B. Shah
Competitive selection processes, from scientific funding to admissions and hiring, use evaluations to score candidates, and eventually choose a subset of them based on those scores…
A Principled Approach to Randomized Selection under Uncertainty: Applications to Peer Review and Grant Funding
Alexander Goldberg, Giulia Fanti, Nihar B. Shah
Many decision-making processes involve evaluating and then selecting items; examples include scientific peer review, job hiring, school admissions, and investment decisions. The ev…
Benchmarking Fraud Detectors on Private Graph Data
Alexander Goldberg, Giulia Fanti, Nihar Shah +1
We introduce the novel problem of benchmarking fraud detectors on private graph-structured data. Currently, many types of fraud are managed in part by automated detection algorithm…
Privacy Requirements and Realities of Digital Public Goods
Geetika Gopi, Aadyaa Maddi, Omkhar Arasaratnam +1
In the international development community, the term "digital public goods" is used to describe open-source digital products (e.g., software, datasets) that aim to address the Unit…