3 papers
cs.CL2025
StepTool: Enhancing Multi-Step Tool Usage in LLMs via Step-Grained Reinforcement Learning
Yuanqing Yu, Zhefan Wang, Weizhi Ma +4
Despite their powerful text generation capabilities, large language models (LLMs) still struggle to effectively utilize external tools to solve complex tasks, a challenge known as…
cs.IR2024
RecSys Arena: Pair-wise Recommender System Evaluation with Large Language Models
Zhuo Wu, Qinglin Jia, Chuhan Wu +4
Evaluating the quality of recommender systems is critical for algorithm design and optimization. Most evaluation methods are computed based on offline metrics for quick algorithm e…
cs.IR2024
Beyond Utility: Evaluating LLM as Recommender
Chumeng Jiang, Jiayin Wang, Weizhi Ma +4
With the rapid development of Large Language Models (LLMs), recent studies employed LLMs as recommenders to provide personalized information services for distinct users. Despite ef…