3 citations · 3 across the 6 of their papers we have counts for
6 papers
SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
Boxi Yu, Yang Cao, Yuzhong Zhang +9
The SWE-Bench Verified leaderboard is approaching saturation, with the top system achieving 78.80%. However, we show that this performance is inflated. Our re-evaluation reveals th…
BackportBench: A Multilingual Benchmark for Automated Backporting of Patches
Zhiqing Zhong, Jiaming Huang, Pinjia He
Many modern software projects evolve rapidly to incorporate new features and security patches. It is important for users to update their dependencies to safer versions, but many st…
SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints
Zhiyu Fan, Kirill Vasilevski, Dayi Lin +6
The advancement of large language models (LLMs) and code agents has demonstrated significant potential to assist software engineering (SWE) tasks, such as autonomous issue resoluti…
An Empirical Study on Package-Level Deprecation in Python Ecosystem
Zhiqing Zhong, Shilin He, Haoxuan Wang +3
Open-source software (OSS) plays a crucial role in modern software development. Utilizing OSS code can greatly accelerate software development, reduce redundancy, and enhance relia…
ROME: Testing Image Captioning Systems via Recursive Object Melting
Boxi Yu, Zhiqing Zhong, Jiaqi Li +3
Image captioning (IC) systems aim to generate a text description of the salient objects in an image. In recent years, IC systems have been increasingly integrated into our daily li…
Automated Testing of Image Captioning Systems
Boxi Yu, Zhiqing Zhong, Xinran Qin +3
Image captioning (IC) systems, which automatically generate a text description of the salient objects in an image (real or synthetic), have seen great progress over the past few ye…