collaborators

10 papers

cs.AI2026

Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics

Shengwei Xu, Yuxuan Lu, Yifan Wu +2

Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing it with a reference response.…

cs.CL2026

To Reason or Not to: Selective Chain-of-Thought in Medical Question Answering

Zaifu Zhan, Min Zeng, Shuang Zhou +6

Objective: To improve the efficiency of medical question answering (MedQA) with large language models (LLMs) by avoiding unnecessary reasoning while maintaining accuracy. Methods:…

cs.AI2026

ComplLLM: Fine-tuning LLMs to Discover Complementary Signals for Decision-making

Ziyang Guo, Yifan Wu, Jason Hartline +2

Multi-agent decision pipelines can outperform single agent workflows when complementarity holds, i.e., different agents bring unique information to the table to inform a final deci…

cs.AI2026

Explaining and Improving Information Complementarities in Multi-Agent Decision-making

Ziyang Guo, Yifan Wu, Jason Hartline +1

Multiple agents are increasingly combined to make decisions with the expectation of achieving complementary performance, where the decisions they make together outperform those mad…

cs.AI2025

ElicitationGPT: Text Elicitation Mechanisms via Language Models

Yifan Wu, Jason Hartline

Scoring rules evaluate probabilistic forecasts of an unknown state against the realized state and are a fundamental building block in the incentivized elicitation of information. T…

cs.AI2025

A Decision Theoretic Framework for Measuring AI Reliance

Ziyang Guo, Yifan Wu, Jason Hartline +1

Humans frequently make decisions with the aid of artificially intelligent (AI) systems. A common pattern is for the AI to recommend an action to the human who retains control over…