1 citations · 1 across the 1 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
GameDevBench: Evaluating Agentic Capabilities Through Game Development
Wayne Chi, Yixiong Fang, Arnav Yayavaram +8
Despite rapid progress on coding agents, progress on their multimodal counterparts has lagged behind. A key challenge is the scarcity of evaluation testbeds that combine the comple…
cs.AI2025★ 1 cited
ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning
Zexi Liu, Yuzhu Cai, Xinyu Zhu +6
As AI capabilities advance toward and potentially beyond human-level performance, a natural transition emerges where AI-driven development becomes more efficient than human-centric…