most citedMulti-Agent Teams Hold Experts Back

2 citations · 2 across the 4 of their papers we have counts for

collaborators
Showing cs.AIShow all

9 papers · 1 filter

cs.AI2026

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing

Jiacheng Miao, Jin Mu, Guanhua Chen +1

Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate this process, as they can ins…

cs.AI2026

The Agentic Garden of Forking Paths

Jiacheng Miao, Jonathan K Pritchard, James Zou

Empirical research rarely admits a unique analysis. Different analytical choices can lead to different conclusions from the same data, yet these hidden forking paths are difficult…

cs.AI2026

Cost-of-Pass: An Economic Framework for Evaluating Language Models

Mehmet Hamza Erol, Batu El, Mirac Suzgun +2

Widespread adoption of AI systems hinges on their ability to generate economic value that outweighs their inference costs. Evaluating this tradeoff requires metrics accounting for…

cs.AI2026

What LLMs Think When You Don't Tell Them What to Think About?

Yongchan Kwon, James Zou

Characterizing the behavior of large language models (LLMs) across diverse settings is critical for reliable monitoring and AI safety. However, most existing analyses rely on topic…

cs.AI2025

Paper2Agent: Reimagining Research Papers As Interactive and Reliable AI Agents

Jiacheng Miao, Joe R. Davis, Yaohui Zhang +2

We introduce Paper2Agent, an automated framework that converts research papers into AI agents. Paper2Agent transforms research output from passive artifacts into active systems tha…

cs.AI2025

Inefficiencies of Meta Agents for Agent Design

Batu El, Mert Yuksekgonul, James Zou

Recent works began to automate the design of agentic systems using meta-agents that propose and iteratively refine new agent architectures. In this paper, we examine three key chal…