3 papers
cs.LG2026
Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents
Jiawei Huang, Qingping Yang, Renjie Zheng +1
Despite recent progress in Large Language Model (LLM) Agents for Software Engineering (SWE) tasks, end-to-end fine-tuning typically relies on verifiable terminal rewards such as wh…
cs.SE2025
Alibaba LingmaAgent: Improving Automated Issue Resolution via Comprehensive Repository Exploration
Yingwei Ma, Qingping Yang, Rongyu Cao +3
This paper presents Alibaba LingmaAgent, a novel Automated Software Engineering method designed to comprehensively understand and utilize whole software repositories for issue reso…
cs.CL2025
UTMath: Math Evaluation with Unit Test via Reasoning-to-Coding Thoughts
Bo Yang, Qingping Yang, Yingwei Ma +1
The evaluation of mathematical reasoning capabilities is essential for advancing Artificial General Intelligence (AGI). While Large Language Models (LLMs) have shown impressive per…