1 paper · 1 filter
Saeed Mohammadzadeh, Erfan Hamdi, Joel Shor +1
As LLMs advance their reasoning capabilities about the physical world, the absence of rigorous benchmarks for evaluating their ability to generate scientifically valid physical mod…