2 papers
cs.IR2024
RecSys Arena: Pair-wise Recommender System Evaluation with Large Language Models
Zhuo Wu, Qinglin Jia, Chuhan Wu +4
Evaluating the quality of recommender systems is critical for algorithm design and optimization. Most evaluation methods are computed based on offline metrics for quick algorithm e…
cs.CL2024
StepTool: Enhancing Multi-Step Tool Usage in LLMs via Step-Grained Reinforcement Learning
Yuanqing Yu, Zhefan Wang, Weizhi Ma +4
Despite their powerful text generation capabilities, large language models (LLMs) still struggle to effectively utilize external tools to solve complex tasks, a challenge known as…