1 paper
Xuexiong Yin, Zechuan Chen, Yongsen Zheng +5
Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personal…