1 paper
Yueru Yan, Tuc Nguyen, Bo Su +2
By evaluating Large Language Models (LLMs) through uniform, text-only interfaces, current academic benchmarks obscure how the unique designs and affordances of distinct commercial…