1 paper
Hengle Jiang, Qijun Cai, Ziying Luo +1
As LLM-based agents continue to advance, their evaluation has become increasingly multifaceted: a capable agent must not only achieve high task completion accuracy but also perform…