2 papers
cs.AI2025
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
Sayash Kapoor, Benedikt Stroebl, Peter Kirgis +28
AI agents have been developed for complex real-world tasks from coding to customer service. But AI agent evaluations suffer from many challenges that undermine our understanding of…
cs.CL2025
Localized Cultural Knowledge is Conserved and Controllable in Large Language Models
Veniamin Veselovsky, Berke Argin, Benedikt Stroebl +5
Just as humans display language patterns influenced by their native tongue when speaking new languages, LLMs often default to English-centric responses even when generating in othe…