4 papers
Agentic Time Machine as an Infrastructure for Future-Event Forecasting
Jingyi Chai, Bingyang Zheng, Xiangrui Liu +5
Forecasting future events is a critical challenge for large language model (LLM) agents, spanning domains from elections and monetary policy to financial markets. However, evaluati…
CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive Engagement
Hong Qian, Yuanhao Liu, Zihan Zhou +7
While LLM-based agents excel at individual tasks, effective collaboration with realistic human partners remains challenging. Most of the existing conversation-level collaborative s…
SetupX: Can LLM Agents Learn from Past Failures in Functionality-Correct Code Repository Setup?
Zihang Zhou, Ziqian Ren, Yukai Wu +7
Functionality-correct repository setup aims to configure execution environments (e.g., dependencies, build scripts) to successfully execute a repository's documented features. It p…
Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies
Zirui Tang, Xuanhe Zhou, Yumou Liu +19
Workspace learning requires AI agents to identify, reason over, exploit, and update explicit and implicit dependencies among heterogeneous files in a worker's workspace, enabling t…