3 citations · 5 across the 2 of their papers we have counts for
3 papers
cs.AI2026
Game Arena: Strategic LLM Evaluation in Competitive Environments
Bovard Doerschuk-Tiberi, Yao Yan, Justin Chiu +59
We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena…
cs.CL2025★ 2 cited
RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models
Bang An, Shiyue Zhang, Mark Dredze
Efforts to ensure the safety of large language models (LLMs) include safety fine-tuning, evaluation, and red teaming. However, despite the widespread use of the Retrieval-Augmented…
cs.CL2025★ 3 cited
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Alon Jacovi, Andrew Wang, Chris Alberti +23
We introduce FACTS Grounding, an online leaderboard and associated benchmark that evaluates language models' ability to generate text that is factually accurate with respect to giv…