5 papers
Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints
Xinge Liu, Terry Jingchen Zhang, Bernhard Schölkopf +2
The rise of autonomous AI agents suggests that dynamic benchmark environments with built-in feedback on scientifically grounded tasks are needed to evaluate the capabilities of the…
Meridional circulation molecular-weighted
Deepayan Banik, Kristen Menou, Evan H. Anders
Meridional circulation in stratified stellar/planetary interiors in the presence of stable molecular weight gradients remains poorly understood, thereby affecting angular momentum…
Gravity-Bench-v1: A Benchmark on Gravitational Physics Discovery for Agents
Nolan Koblischke, Hyunseok Jang, Kristen Menou +1
Modern science emerged from reasoning over repeatedly-observed planetary motions. We present Gravity-Bench-v1, an environment-based benchmark that challenges AI agents on tasks tha…
Meridional Circulation Streamlined
Deepayan Banik, Kristen Menou
Time-dependent meridional circulation and differential rotation in radiative zones are central open issues in stellar evolution theory. We streamline this challenging problem using…
Physics simulation capabilities of LLMs
Mohamad Ali-Dib, Kristen Menou
[Abridged abstract] Large Language Models (LLMs) can solve some undergraduate-level to graduate-level physics textbook problems and are proficient at coding. Combining these two ca…