benchmark validation 1large language models 1multi-agent systems 1pr-issue alignment 1software engineering 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.SE2026
PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks
Manyi Wang, Junjielong Xu, Pinjia He
The paper introduces PAIChecker, a multi‑agent system that automatically detects misalignments between pull requests and their linked issues in SWE‑bench‑style benchmarks, improvin…
cs.SE2025
Rethinking the Evaluation of Microservice RCA with a Fault Propagation-Aware Benchmark
Aoyang Fang, Songhan Zhang, Yifan Yang +7
While cloud-native microservice architectures have revolutionized software development, their inherent operational complexity makes failure Root Cause Analysis (RCA) a critical yet…