1 citations · 1 across the 1 of their papers we have counts for
9 papers
FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights
Zhen Wang, Fan Bai, Zhongyan Luo +9
Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery r…
CocoaBench: Evaluating Unified Digital Agents in the Wild
CocoaBench Team, Shibo Hao, Zhining Zhang +29
LLM agents now perform strongly in software engineering, deep research, GUI automation, and various other applications, while recent agent scaffolds and models are increasingly int…
World Reasoning Arena
PAN Team, Qiyue Gao, Kun Zhou +15
World models (WMs) are intended to serve as internal simulators of the real world that enable agents to understand, anticipate, and act upon complex environments. Existing WM bench…
TritonDFT: Automating DFT with a Multi-Agent Framework
Zhengding Hu, Kuntal Talit, Zhen Wang +8
Density Functional Theory (DFT) is a cornerstone of materials science, yet executing DFT in practice requires coordinating a complex, multi-step workflow. Existing tools and LLM-ba…
Learning Modal-Mixed Chain-of-Thought Reasoning with Latent Embeddings
Yifei Shao, Kun Zhou, Ziming Xu +5
We study how to extend chain-of-thought (CoT) beyond language to better handle multimodal reasoning. While CoT helps LLMs and VLMs articulate intermediate steps, its text-only form…
Auto-scaling Continuous Memory for GUI Agent
Wenyi Wu, Kun Zhou, Ruoxin Yuan +4
We study how to endow GUI agents with scalable memory that help generalize across unfamiliar interfaces and long-horizon tasks. Prior GUI agents compress past trajectories into tex…