activity
20242026
most citedAgentic Software Engineering: Foundational Pillars and a Research Roadmap

5 citations · 5 across the 3 of their papers we have counts for

collaborators

19 papers

cs.SE2026

LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation

Han Li, Zhemin Fang, Rili Feng +8

Coding agent infrastructure is shifting from harness engineering toward loop engineering as coding agents are deployed for sustained long-horizon software development. Existing ben…

cs.SE2026

Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment

Kirill Vasilevski, Ximing Dong, Benjamin Rombaut +8

LLMs have substantially improved software engineering yet real-world development requires architectural understanding. Such understanding is prohibitively expensive to label manual…

cs.SE20265 cited

Agentic Software Engineering: Foundational Pillars and a Research Roadmap

Ahmed E. Hassan, Hao Li, Dayi Lin +4

Agentic Software Engineering (SE 3.0) represents a new era where intelligent agents are tasked not with simple code generation, but with achieving complex, goal-oriented SE objecti…

cs.SE2026

From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap

Gopi Krishnan Rajbahadur, Gustavo A. Oliva, Dayi Lin +2

The rapid expansion of foundation models (FMs), such as large language models (LLMs), has given rise to FMware, software systems that integrate FM(s) as core components. While buil…

cs.SE2026

Towards Reliable Generation of Executable Workflows by Foundation Models

Sogol Masoumzadeh, Keheliya Gallaba, Dayi Lin +1

Recent advancements in Foundation Models (FMs) have demonstrated significant progress in processing complex natural language to perform intricate tasks. Successfully executing thes…

cs.SE2026

When Elo Lies: Hidden Biases in Codeforces-Based Evaluation of Large Language Models

Shenyu Zheng, Ximing Dong, Xiaoshuang Liu +6

As Large Language Models (LLMs) achieve breakthroughs in complex reasoning, Codeforces-based Elo ratings have emerged as a prominent metric for evaluating competitive programming c…