9 papers
Mitigating Context Interference for Reliable and Efficient Search Agents
Boyang Xue, Bin Wu, Shuofei Qiao +8
Recent research empowers Large Language Models (LLMs) as multi-turn search agents to iteratively retrieve and generate outputs until complex tasks are solved. However, the contexts…
InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem
Shuofei Qiao, Yunxiang Wei, Xuehai Wang +10
The rapid evolution of Large Language Models has catalyzed a surge in scientific idea production, yet this leap has not been accompanied by a matching advance in idea evaluation. T…
AgentSearchBench: A Benchmark for AI Agent Search in the Wild
Bin Wu, Arastun Mammadli, Xiaoyu Zhang +1
The rapid growth of AI agent ecosystems is transforming how complex tasks are delegated and executed, creating a new challenge of identifying suitable agents for a given task. Unli…
Collaborative User Prompt for Personalized Generative Recommendation
Jerome Ramos, Bin Wu, Aldo Lipani
Large Language Models (LLMs) have become powerful foundations for generative recommender systems, framing recommendation tasks as text generation tasks. However, existing generativ…
A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
Jinyuan Fang, Yanwen Peng, Xi Zhang +12
Recent advances in large language models have sparked growing interest in AI agents capable of solving complex, real-world tasks. However, most existing agent systems rely on manua…
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
Xiao Fu, Hossein A. Rahmani, Bin Wu +3
Personalised text generation is essential for user-centric information systems, yet most evaluation methods overlook the individuality of users. We introduce \textbf{PREF}, a \text…