4 papers
PieArena: Ranking and Profiling Language Agents in Realistic Negotiation Scenarios
Chris Zhu, Sasha Cui, Will Sanok Dufallo +4
We present an in-depth evaluation of LLMs' ability to negotiate, a central business task requiring strategic reasoning, theory of mind, and economic value creation. To do so, we in…
Bias-Corrected Data Synthesis for Imbalanced Learning
Pengfei Lyu, Zhengchi Ma, Linjun Zhang +1
Imbalanced data, where the positive samples represent only a small proportion compared to the negative samples, makes it challenging for classification problems to balance the fals…
Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals
Zihan Dong, Zhixian Zhang, Yang Zhou +3
Evaluating mathematical reasoning in LLMs is constrained by limited benchmark sizes and inherent model stochasticity, yielding high-variance accuracy estimates and unstable ranking…
Unified Inference Framework for Single and Multi-Player Performative Prediction: Method and Asymptotic Optimality
Zhixian Zhang, Xiaotian Hou, Linjun Zhang
Performative prediction characterizes environments where predictive models alter the very data distributions they aim to forecast, triggering complex feedback loops. While prior re…