17 papers
Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing
Jiacheng Miao, Jin Mu, Guanhua Chen +1
Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate this process, as they can ins…
The Agentic Garden of Forking Paths
Jiacheng Miao, Jonathan K Pritchard, James Zou
Empirical research rarely admits a unique analysis. Different analytical choices can lead to different conclusions from the same data, yet these hidden forking paths are difficult…
Multi-Agent Teams Hold Experts Back
Aneesh Pappu, Batu El, Hancheng Cao +4
Multi-agent LLM systems are increasingly deployed as autonomous collaborators, where agents interact freely rather than execute fixed, pre-specified workflows. In such settings, ef…
Sparse Reward Subsystem in Large Language Models
Guowei Xu, Mert Yuksekgonul, James Zou
Recent studies show that LLM hidden states encode reward-related information, such as answer correctness and model confidence. However, existing approaches typically fit black-box…
ACT: Agentic Classification Tree
Vincent Grari, Tim Arni, Thibault Laugel +3
When used in high-stakes settings, AI systems are expected to produce decisions that are transparent, interpretable and auditable, a requirement increasingly expected by regulation…
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
Qizheng Zhang, Changran Hu, Shubhangi Upasani +10
Large language model (LLM) applications such as agents and domain-specific reasoning increasingly rely on context adaptation: modifying inputs with instructions, strategies, or evi…