2 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.CL2026
ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues
Shanda Li, Qiuhong Anna Wei, Jingwu Tang +5
Reproducing research results from papers and released code is central to scientific progress. Existing works have introduced benchmarks to evaluate whether LLM agents can assist wi…
cs.AI2026
GameDevBench: Evaluating Agentic Capabilities Through Game Development
Wayne Chi, Yixiong Fang, Arnav Yayavaram +8
Despite rapid progress on coding agents, progress on their multimodal counterparts has lagged behind. A key challenge is the scarcity of evaluation testbeds that combine the comple…
cs.CV2024★ 2 cited
Open-Universe Indoor Scene Generation using LLM Program Synthesis and Uncurated Object Databases
Rio Aguina-Kang, Maxim Gumin, Do Heon Han +7
We present a system for generating indoor scenes in response to text prompts. The prompts are not limited to a fixed vocabulary of scene descriptions, and the objects in generated…