1 paper
Joseph Chan, Utkarsh Jha, Xiyin Yang +6
Large language models (LLMs) perform strongly on static science benchmarks, yet their ability to reason about the physical world through active experimentation remains poorly under…