3 papers
cs.LG2026
Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints
Xinge Liu, Terry Jingchen Zhang, Bernhard Schölkopf +2
The rise of autonomous AI agents suggests that dynamic benchmark environments with built-in feedback on scientifically grounded tasks are needed to evaluate the capabilities of the…
astro-ph.SR2025
Meridional circulation molecular-weighted
Deepayan Banik, Kristen Menou, Evan H. Anders
Meridional circulation in stratified stellar/planetary interiors in the presence of stable molecular weight gradients remains poorly understood, thereby affecting angular momentum…
cs.AI2025
Gravity-Bench-v1: A Benchmark on Gravitational Physics Discovery for Agents
Nolan Koblischke, Hyunseok Jang, Kristen Menou +1
Modern science emerged from reasoning over repeatedly-observed planetary motions. We present Gravity-Bench-v1, an environment-based benchmark that challenges AI agents on tasks tha…