13 papers
OdinEval: A Reproducible Benchmark for LLM-Based Program Repair in the Odin Programming Language
Bang Xie, Hao Liu, Zhiyuan Peng +8
Repository-level repair benchmarks still center on a few mainstream languages, leaving systems languages such as Odin largely untested. We present OdinEval, a reproducible benchmar…
AppEval: A Unified Benchmark for LLM-Based Mobile Application Repair in ArkTS, Swift, and Kotlin
Bang Xie, Hao Liu, Zhenyu Shi +11
Repository-level LLM agents are typically evaluated on projects whose tests run on the build host. It remains unclear whether their repairs survive the mobile build-install-launch-…
EvoClawBench: Can Agents Learn Reusable Skills from Their Own Runs?
Zhiyuan Peng, Xin Yin, Chenhao Ying +5
Existing agent benchmarks primarily test task completion, tool use, or skill utility, but do not isolate whether a runtime can convert evidence from its own runs into reusable skil…
UGround: Towards Unified Visual Grounding with Unrolled Transformers
Rui Qian, Xin Yin, Chuanhang Deng +4
We present UGround, a \textbf{U}nified visual \textbf{Ground}ing paradigm that dynamically selects intermediate layers across \textbf{U}nrolled transformers as ``mask as prompt,''…
PlayCoder: Making LLM-Generated GUI Code Playable
Zhiyuan Peng, Wei Tao, Xin Yin +3
Large language models (LLMs) have achieved strong results in code generation, but their ability to generate GUI applications, especially games, remains insufficiently studied. Exis…
RepoGenesis: Benchmarking End-to-End Microservice Generation from Readme to Repository
Zhiyuan Peng, Xin Yin, Pu Zhao +7
Large language models and agents have achieved remarkable progress in code generation. However, existing benchmarks focus on isolated function/class-level generation (e.g., ClassEv…