4 papers
E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing
Shuvom Sadhuka, Drew Prinster, Clara Fannjiang +4
Agentic AI systems execute a sequence of actions, such as reasoning steps or tool calls, in response to a user prompt. To evaluate the success of their trajectories, researchers ha…
A Bayesian Model for Multi-stage Censoring
Shuvom Sadhuka, Sophia Lin, Bonnie Berger +1
Many sequential decision settings in healthcare feature funnel structures characterized by a series of stages, such as screenings or evaluations, where the number of patients who a…
Urban Incident Prediction with Graph Neural Networks: Integrating Government Ratings and Crowdsourced Reports
Sidhika Balachandar, Shuvom Sadhuka, Bonnie Berger +2
Graph neural networks (GNNs) are widely used in urban spatiotemporal forecasting, such as predicting infrastructure problems. In this setting, government officials wish to know in…
Evaluating multiple models using labeled and unlabeled data
Divya Shanmugam, Shuvom Sadhuka, Manish Raghavan +3
It remains difficult to evaluate machine learning classifiers in the absence of a large, labeled dataset. While labeled data can be prohibitively expensive or impossible to obtain,…