1 paper
Darvin Yi, Teng Liu, Mattie Terzolo +4
As large language model (LLM) agents increasingly undertake digital work, reliable frameworks are needed to evaluate their real-world competence, adaptability, and capacity for hum…