4 papers
PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research
Tingjia Miao, Wenkai Jin, Muhua Zhang +19
The paradigm of agentic science requires AI systems to conduct robust reasoning and engage in long-horizon, autonomous exploration. However, current scientific benchmarks remain co…
Automated Extraction of Collins-Soper Kernel from Lattice QCD using An Autonomous AI Physicist System
Jin-Xin Tan, Ting-Jia Miao, Mu-Hua Zhang +5
We employ {PhysMaster}, an autonomous agentic AI system integrating theoretical reasoning, numerical computation, and exploitation strategies towards ultra-long horizon automation,…
Bohrium + SciMaster: Building the Infrastructure and Ecosystem for Agentic Science at Scale
Linfeng Zhang, Siheng Chen, Yuzhu Cai +46
AI agents are emerging as a practical way to run multi-step scientific workflows that interleave reasoning with tool use and verification, pointing to a shift from isolated AI-assi…
Inverse Knowledge Search over Verifiable Reasoning: Synthesizing a Scientific Encyclopedia from a Long Chains-of-Thought Knowledge Base
Yu Li, Yuan Huang, Tao Wang +19
Most scientific materials compress reasoning, presenting conclusions while omitting the derivational chains that justify them. This compression hinders verification by lacking expl…