10 papers
Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics
Shengwei Xu, Yuxuan Lu, Yifan Wu +2
Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing it with a reference response.…
To Reason or Not to: Selective Chain-of-Thought in Medical Question Answering
Zaifu Zhan, Min Zeng, Shuang Zhou +6
Objective: To improve the efficiency of medical question answering (MedQA) with large language models (LLMs) by avoiding unnecessary reasoning while maintaining accuracy. Methods:…
ComplLLM: Fine-tuning LLMs to Discover Complementary Signals for Decision-making
Ziyang Guo, Yifan Wu, Jason Hartline +2
Multi-agent decision pipelines can outperform single agent workflows when complementarity holds, i.e., different agents bring unique information to the table to inform a final deci…
Explaining and Improving Information Complementarities in Multi-Agent Decision-making
Ziyang Guo, Yifan Wu, Jason Hartline +1
Multiple agents are increasingly combined to make decisions with the expectation of achieving complementary performance, where the decisions they make together outperform those mad…
ElicitationGPT: Text Elicitation Mechanisms via Language Models
Yifan Wu, Jason Hartline
Scoring rules evaluate probabilistic forecasts of an unknown state against the realized state and are a fundamental building block in the incentivized elicitation of information. T…
A Decision Theoretic Framework for Measuring AI Reliance
Ziyang Guo, Yifan Wu, Jason Hartline +1
Humans frequently make decisions with the aid of artificially intelligent (AI) systems. A common pattern is for the AI to recommend an action to the human who retains control over…