13 papers
Truthful Calibration Measures for Sequential Prediction
Anagha Gokul, Jason Hartline, Lunjia Hu +2
Calibration requires probabilistic reports to be conditionally unbiased and reliably interpretable as probabilities. A calibration measure assigns numerical error to miscalibrated…
Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics
Shengwei Xu, Yuxuan Lu, Yifan Wu +2
Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing it with a reference response.…
A Perfectly Truthful Calibration Measure
Jason Hartline, Lunjia Hu, Yifan Wu
Calibration requires that predictions are conditionally unbiased and, therefore, reliably interpretable as probabilities. A calibration measure quantifies how far a predictor is fr…
ComplLLM: Fine-tuning LLMs to Discover Complementary Signals for Decision-making
Ziyang Guo, Yifan Wu, Jason Hartline +2
Multi-agent decision pipelines can outperform single agent workflows when complementarity holds, i.e., different agents bring unique information to the table to inform a final deci…
Clarification of `Algorithmic Collusion without Threats'
Jason Hartline
This brief note clarifies that the scenario described in Arunachaleswaran et al. (2025) -- titled `Algorithmic Collusion without Threats' -- is not one of collusion, but one where…
Explaining and Improving Information Complementarities in Multi-Agent Decision-making
Ziyang Guo, Yifan Wu, Jason Hartline +1
Multiple agents are increasingly combined to make decisions with the expectation of achieving complementary performance, where the decisions they make together outperform those mad…