most citedFIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights

1 citations · 1 across the 1 of their papers we have counts for

collaborators

9 papers

cs.AI20261 cited

FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights

Zhen Wang, Fan Bai, Zhongyan Luo +9

Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery r…

cs.CL2026

CocoaBench: Evaluating Unified Digital Agents in the Wild

CocoaBench Team, Shibo Hao, Zhining Zhang +29

LLM agents now perform strongly in software engineering, deep research, GUI automation, and various other applications, while recent agent scaffolds and models are increasingly int…

cs.CV2026

World Reasoning Arena

PAN Team, Qiyue Gao, Kun Zhou +15

World models (WMs) are intended to serve as internal simulators of the real world that enable agents to understand, anticipate, and act upon complex environments. Existing WM bench…

cond-mat.mtrl-sci2026

TritonDFT: Automating DFT with a Multi-Agent Framework

Zhengding Hu, Kuntal Talit, Zhen Wang +8

Density Functional Theory (DFT) is a cornerstone of materials science, yet executing DFT in practice requires coordinating a complex, multi-step workflow. Existing tools and LLM-ba…

cs.AI2026

Learning Modal-Mixed Chain-of-Thought Reasoning with Latent Embeddings

Yifei Shao, Kun Zhou, Ziming Xu +5

We study how to extend chain-of-thought (CoT) beyond language to better handle multimodal reasoning. While CoT helps LLMs and VLMs articulate intermediate steps, its text-only form…

cs.AI2025

Auto-scaling Continuous Memory for GUI Agent

Wenyi Wu, Kun Zhou, Ruoxin Yuan +4

We study how to endow GUI agents with scalable memory that help generalize across unfamiliar interfaces and long-horizon tasks. Prior GUI agents compress past trajectories into tex…