2 citations · 2 across the 4 of their papers we have counts for
9 papers · 1 filter
Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing
Jiacheng Miao, Jin Mu, Guanhua Chen +1
Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate this process, as they can ins…
The Agentic Garden of Forking Paths
Jiacheng Miao, Jonathan K Pritchard, James Zou
Empirical research rarely admits a unique analysis. Different analytical choices can lead to different conclusions from the same data, yet these hidden forking paths are difficult…
Cost-of-Pass: An Economic Framework for Evaluating Language Models
Mehmet Hamza Erol, Batu El, Mirac Suzgun +2
Widespread adoption of AI systems hinges on their ability to generate economic value that outweighs their inference costs. Evaluating this tradeoff requires metrics accounting for…
What LLMs Think When You Don't Tell Them What to Think About?
Yongchan Kwon, James Zou
Characterizing the behavior of large language models (LLMs) across diverse settings is critical for reliable monitoring and AI safety. However, most existing analyses rely on topic…
Paper2Agent: Reimagining Research Papers As Interactive and Reliable AI Agents
Jiacheng Miao, Joe R. Davis, Yaohui Zhang +2
We introduce Paper2Agent, an automated framework that converts research papers into AI agents. Paper2Agent transforms research output from passive artifacts into active systems tha…
Inefficiencies of Meta Agents for Agent Design
Batu El, Mert Yuksekgonul, James Zou
Recent works began to automate the design of agentic systems using meta-agents that propose and iteratively refine new agent architectures. In this paper, we examine three key chal…