Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents
Dhaval C. Patel, Kaoutar El Maghraoui, Shuxin Lin +58
Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated d…
cs.AI2024
AI Can Enhance Creativity in Social Networks
Raiyan Abdul Baten, Ali Sarosh Bangash, Krish Veera +2
Can peer recommendation engines elevate people's creative performances in self-organizing social networks? Answering this question requires resolving challenges in data collection…