Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Mubashara Akhtar, Anka Reuel, Prajna Soni +36
Artificial intelligence benchmarks are an important mechanism to measure model progress and guide deployment decisions. However, benchmarks quickly "saturate", making it difficult…
cs.AI2025
GraphicBench: A Planning Benchmark for Graphic Design with Language Agents
Dayeon Ki, Tianyi Zhou, Marine Carpuat +3
Large Language Model (LLM)-powered agents have unlocked new possibilities for automating human tasks. While prior work has focused on well-defined tasks with specified goals, the c…