1 paper
Noah van der Vleuten, Anthony Flores, Shray Mathur +4
Evaluating large language models (LLMs) for instrument control requires methods that go beyond standard, stateless algorithmic benchmarks, since the behavior of physical systems ca…