5 papers · 1 filter
WorldReasoner: Evaluating Whether Language Model Agents Forecast Events with Valid Reasoning
Yizhou Chi, Eric Chamoun, Zifeng Ding +1
Forecasting real-world events requires language-model agents to reason under uncertainty from incomplete, time-bounded information. Yet evaluating whether agents genuinely forecast…
Demystifying Multi-Agent Debate: The Role of Confidence and Diversity
Xiaochen Zhu, Caiqi Zhang, Yizhou Chi +3
Multi-agent debate (MAD) is widely used to improve large language model (LLM) performance through test-time scaling, yet recent work shows that vanilla MAD often underperforms simp…
SciPaths: Forecasting Pathways to Scientific Discovery
Eric Chamoun, Yizhou Chi, Yulong Chen +4
Scientific progress depends on sequences of enabling contributions, yet existing AI4Science benchmarks largely focus on citation prediction, literature retrieval, or idea generatio…
THOUGHTSCULPT: Reasoning with Intermediate Revision and Search
Yizhou Chi, Kevin Yang, Dan Klein
We present THOUGHTSCULPT, a general reasoning and search method for tasks with outputs that can be decomposed into components. THOUGHTSCULPT explores a search tree of potential sol…
AMONGAGENTS: Evaluating Large Language Models in the Interactive Text-Based Social Deduction Game
Yizhou Chi, Lingjun Mao, Zineng Tang
Strategic social deduction games serve as valuable testbeds for evaluating the understanding and inference skills of language models, offering crucial insights into social science,…