2 citations · 3 across the 2 of their papers we have counts for
3 papers
cs.AI2025★ 1 cited
Establishing Best Practices for Building Rigorous Agentic Benchmarks
Yuxuan Zhu, Tengjun Jin, Yada Pruksachatkun +22
Benchmarks are essential for quantitatively tracking progress in AI. As AI agents become increasingly capable, researchers and practitioners have introduced agentic benchmarks to e…
cs.CL2025★ 2 cited
A Framework to Assess the Persuasion Risks Large Language Model Chatbots Pose to Democratic Societies
Zhongren Chen, Joshua Kalla, Quan Le +3
In recent years, significant concern has emerged regarding the potential threat that Large Language Models (LLMs) pose to democratic societies through their persuasive capabilities…
cs.HC2023
Assessing the Usability of GutGPT: A Simulation Study of an AI Clinical Decision Support System for Gastrointestinal Bleeding Risk
Colleen Chan, Kisung You, Sunny Chung +10
Applications of large language models (LLMs) like ChatGPT have potential to enhance clinical decision support through conversational interfaces. However, challenges of human-algori…