most citedFIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights

1 citations · 1 across the 1 of their papers we have counts for

collaborators

13 papers

cs.AI20261 cited

FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights

Zhen Wang, Fan Bai, Zhongyan Luo +9

Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery r…

cs.AI2026

General Agentic Planning Through Simulative Reasoning with World Models

Mingkai Deng, Jinyu Hou, Zhiting Hu +1

What does it mean to plan? Current agentic systems, whether scaffolded workflows or end-to-end policies, rely on reactive decision-making: selecting the next action via a fixed pro…

cs.CL2026

CocoaBench: Evaluating Unified Digital Agents in the Wild

CocoaBench Team, Shibo Hao, Zhining Zhang +29

LLM agents now perform strongly in software engineering, deep research, GUI automation, and various other applications, while recent agent scaffolds and models are increasingly int…

cs.CV2026

World Reasoning Arena

PAN Team, Qiyue Gao, Kun Zhou +15

World models (WMs) are intended to serve as internal simulators of the real world that enable agents to understand, anticipate, and act upon complex environments. Existing WM bench…

cs.LG2026

IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RL

Zhoujun Cheng, Yutao Xie, Yuxiao Qu +12

While scaling laws guide compute allocation for LLM pre-training, analogous prescriptions for reinforcement learning (RL) post-training of large language models (LLMs) remain poorl…

cs.AI2026

scPilot: Large Language Model Reasoning Toward Automated Single-Cell Analysis and Discovery

Yiming Gao, Zhen Wang, Jefferson Chen +8

We present scPilot, the first systematic framework to practice omics-native reasoning: a large language model (LLM) converses in natural language while directly inspecting single-c…