3 papers
cs.LG2026
PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research
Tingjia Miao, Wenkai Jin, Muhua Zhang +19
The paradigm of agentic science requires AI systems to conduct robust reasoning and engage in long-horizon, autonomous exploration. However, current scientific benchmarks remain co…
cs.AI2025
PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research
Tingjia Miao, Jiawen Dai, Jingkun Liu +23
Advances in LLM reasoning and tool use have enabled agentic science, yet frontier theoretical and computational physics remains challenging because research requires deep domain ex…
cs.HC2025
LLMartini: Seamless and Interactive Leveraging of Multiple LLMs through Comparison and Composition
Yingtian Shi, Jinda Yang, Yuhan Wang +4
The growing diversity of large language models (LLMs) means users often need to compare and combine outputs from different models to obtain higher-quality or more comprehensive res…