1 paper
Sirui Liang, Bohan Yu, Peiyu Wang +8
Large language models are increasingly used to power personal agents for everyday applications, but evaluating these agents remains a challenge. Existing benchmarks still rely on s…