most citedSolAgent: A Specialized Multi-Agent Framework for Solidity Code Generation

1 citations · 3 across the 15 of their papers we have counts for

collaborators

15 papers

cs.LG2026

From Reasoning Strings to Partial Orders: Verifier-Certified Rule Transport through Quotient Policy Optimization

Bang Xie, Hao Liu, Zhiyuan Peng +5

Many computations admit several valid execution orders because independent subgoals or disjoint state updates can commute. Reinforcement learning with verifiable rewards usually tr…

cs.SE2026

PonyEval: Evaluating LLM-Based Program Repair for Capability-Safe and Actor-Oriented Pony Software

Bang Xie, Hao Liu, Zhenyu Shi +8

Repository-level issue-resolution benchmarks have made executable evaluation central to software-engineering agents, but their language coverage remains concentrated in mainstream…

cs.SE2026

OdinEval: A Reproducible Benchmark for LLM-Based Program Repair in the Odin Programming Language

Bang Xie, Hao Liu, Zhiyuan Peng +8

Repository-level repair benchmarks still center on a few mainstream languages, leaving systems languages such as Odin largely untested. We present OdinEval, a reproducible benchmar…

cs.SE2026

AppEval: A Unified Benchmark for LLM-Based Mobile Application Repair in ArkTS, Swift, and Kotlin

Bang Xie, Hao Liu, Zhenyu Shi +11

Repository-level LLM agents are typically evaluated on projects whose tests run on the build host. It remains unclear whether their repairs survive the mobile build-install-launch-…

cs.LG2026

EvoClawBench: Can Agents Learn Reusable Skills from Their Own Runs?

Zhiyuan Peng, Xin Yin, Chenhao Ying +5

Existing agent benchmarks primarily test task completion, tool use, or skill utility, but do not isolate whether a runtime can convert evidence from its own runs into reusable skil…

cs.SE2026

PlayCoder: Making LLM-Generated GUI Code Playable

Zhiyuan Peng, Wei Tao, Xin Yin +3

Large language models (LLMs) have achieved strong results in code generation, but their ability to generate GUI applications, especially games, remains insufficiently studied. Exis…