2 papers
cs.AI2025
Establishing Best Practices for Building Rigorous Agentic Benchmarks
Yuxuan Zhu, Tengjun Jin, Yada Pruksachatkun +22
Benchmarks are essential for quantitatively tracking progress in AI. As AI agents become increasingly capable, researchers and practitioners have introduced agentic benchmarks to e…
cs.CL2025
A Framework to Assess the Persuasion Risks Large Language Model Chatbots Pose to Democratic Societies
Zhongren Chen, Joshua Kalla, Quan Le +3
In recent years, significant concern has emerged regarding the potential threat that Large Language Models (LLMs) pose to democratic societies through their persuasive capabilities…