activity
20212026
most citedExpertPrompting: Instructing Large Language Models to be Distinguished Experts

56 citations · 97 across the 26 of their papers we have counts for

collaborators
Showing 2025 · cs.CLShow all

6 papers · 2 filters

cs.CL2025

MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools

Zikang Guo, Benfeng Xu, Chiwei Zhu +3

The Model Context Protocol (MCP) is rapidly emerging as a pivotal open standard, designed to enhance agent-tool integration and interoperability, and is positioned to unlock a new…

cs.CL2025★ 2 cited

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

Mingxuan Du, Benfeng Xu, Chiwei Zhu +2

Deep Research Agents are a prominent category of LLM-based agents. By autonomously orchestrating multistep web exploration, targeted retrieval, and higher-order synthesis, they tra…

cs.CL2025

From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding

Chiwei Zhu, Benfeng Xu, Xiaorui Wang +1

The pursuit of diverse, complex, and large-scale instruction data is crucial for automatically aligning large language models (LLMs). While there are methods capable of generating…

cs.CL2025

Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability

Chiwei Zhu, Benfeng Xu, An Yang +4

Training language models with rationales augmentation has been shown to be beneficial in many existing works. In this paper, we identify that such a prevailing view does not hold c…

cs.CL2025

Training LLM-Based Agents with Synthetic Self-Reflected Trajectories and Partial Masking

Yihan Chen, Benfeng Xu, Xiaorui Wang +2

Autonomous agents, which perceive environments and take actions to achieve goals, have become increasingly feasible with the advancements in large language models (LLMs). However,…

cs.CL2025

Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach

Ruizhe Li, Chiwei Zhu, Benfeng Xu +2

Creative writing is a key capability of Large Language Models (LLMs), with potential applications in literature, storytelling, and various creative domains. However, evaluating the…