1 paper
Callum Sharrock, Lukas Petersson, Hanna Petersson +4
We present Butter-Bench, a benchmark evaluating large language model (LLM) controlled robots for practical intelligence, defined as the ability to navigate the messiness of the phy…