3 papers
cs.CL2025
CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
Kung-Hsiang Huang, Akshara Prabhakar, Onkar Thorat +6
While AI agents hold transformative potential in business, effective performance benchmarking is hindered by the scarcity of public, realistic business data on widely used platform…
cs.CL2025
BingoGuard: LLM Content Moderation Tools with Risk Levels
Fan Yin, Philippe Laban, Xiangyu Peng +7
Malicious content generated by large language models (LLMs) can pose varying degrees of harm. Although existing LLM-based moderators can detect harmful content, they struggle to as…
cs.CL2025
Evaluating Cultural and Social Awareness of LLM Web Agents
Haoyi Qiu, Alexander R. Fabbri, Divyansh Agarwal +4
As large language models (LLMs) expand into performing as agents for real-world applications beyond traditional NLP tasks, evaluating their robustness becomes increasingly importan…