Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use
Zijian Wu, Xiangyan Liu, Xinyuan Zhang +12
MCP standardizes how LLMs interact with external systems, forming the foundation for general agents. However, existing MCP benchmarks remain narrow in scope: they focus on read-hea…
cs.CL2024
Is Self-knowledge and Action Consistent or Not: Investigating Large Language Model's Personality
Yiming Ai, Zhiwei He, Ziyin Zhang +5
In this study, we delve into the validity of conventional personality questionnaires in capturing the human-like personality traits of Large Language Models (LLMs). Our objective i…