most citedFutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction

1 citations · 1 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CL2026

NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents

Jingzhe Ding, Shengda Long, Changxin Pu +46

Recent advances in coding agents suggest rapid progress toward autonomous software development, yet existing benchmarks fail to rigorously evaluate the long-horizon capabilities re…

cs.CL2025

DiscoX: Benchmarking Discourse-Level Translation task in Expert Domains

Xiying Zhao, Zhoufutu Wen, Zhixuan Chen +20

The evaluation of discourse-level translation in expert domains remains inadequate, despite its centrality to knowledge dissemination and cross-lingual scholarly communication. Whi…

cs.MA2025

Multi-Agent Medical Decision Consensus Matrix System: An Intelligent Collaborative Framework for Oncology MDT Consultations

Xudong Han, Xianglun Gao, Xiaoyi Qu +1

Multidisciplinary team (MDT) consultations are the gold standard for cancer care decision-making, yet current practice lacks structured mechanisms for quantifying consensus and ens…

cs.AI2025

LPFQA: A Long-Tail Professional Forum-based Benchmark for LLM Evaluation

Liya Zhu, Peizhuang Cong, Jingzhe Ding +17

Large Language Models (LLMs) perform well on standard reasoning and question-answering benchmarks, yet such evaluations often fail to capture their ability to handle long-tail, exp…

cs.LG2025

FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and Reasoning

Liang Hu, Jianpeng Jiao, Jiashuo Liu +20

Search has emerged as core infrastructure for LLM-based agents and is widely viewed as critical on the path toward more general intelligence. Finance is a particularly demanding pr…

cs.AI20251 cited

FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction

Zhiyuan Zeng, Jiashuo Liu, Siyuan Chen +28

Future prediction is a complex task for LLM agents, requiring a high level of analytical thinking, information gathering, contextual understanding, and decision-making under uncert…