1 citations · 1 across the 1 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2024
PromptBench: A Unified Library for Evaluation of Large Language Models
Kaijie Zhu, Qinlin Zhao, Hao Chen +2
The evaluation of large language models (LLMs) is crucial to assess their performance and mitigate potential security risks. In this paper, we introduce PromptBench, a unified libr…
cs.AI2024
CompeteAI: Understanding the Competition Dynamics in Large Language Model-based Agents
Qinlin Zhao, Jindong Wang, Yixuan Zhang +4
Large language models (LLMs) have been widely used as agents to complete different tasks, such as personal assistance or event planning. While most of the work has focused on coope…