1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.LG2025
Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
Yangsibo Huang, Milad Nasr, Anastasios Angelopoulos +10
It is now common to evaluate Large Language Models (LLMs) by having humans manually vote to evaluate model outputs, in contrast to typical benchmarks that evaluate knowledge or ski…
cs.LG2024★ 1 cited
Human-aligned Chess with a Bit of Search
Yiming Zhang, Athul Paul Jacob, Vivian Lai +2
Chess has long been a testbed for AI's quest to match human intelligence, and in recent years, chess AI systems have surpassed the strongest humans at the game. However, these syst…