Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Gravity-Bench-v1: A Benchmark on Gravitational Physics Discovery for Agents
Nolan Koblischke, Hyunseok Jang, Kristen Menou +1
Modern science emerged from reasoning over repeatedly-observed planetary motions. We present Gravity-Bench-v1, an environment-based benchmark that challenges AI agents on tasks tha…
cs.AI2024
Physics simulation capabilities of LLMs
Mohamad Ali-Dib, Kristen Menou
[Abridged abstract] Large Language Models (LLMs) can solve some undergraduate-level to graduate-level physics textbook problems and are proficient at coding. Combining these two ca…