most citedSkeleton-Guided-Translation: A Benchmarking Framework for Code Repository Translation with Fine-Grained Quality Evaluation

1 citations · 1 across the 3 of their papers we have counts for

collaborators

5 papers

cs.SE2025

Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation

Spandan Garg, Benjamin Steenhoek, Yufan Huang

Current benchmarks for evaluating software engineering agents, such as SWE-Bench Verified, are predominantly derived from GitHub issues and fail to accurately reflect how developer…

cs.AI2025

Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs

Hankun Dai, Maoquan Wang, Mengnan Qi +6

Large language models (LLMs) are increasingly being applied to programming tasks, ranging from single-turn code completion to autonomous agents. Current code agent designs frequent…

cs.SE2025

SWE-bench Goes Live!

Linghao Zhang, Shilin He, Chaoyun Zhang +12

The issue-resolving task, where a model generates patches to fix real-world bugs, has emerged as a critical benchmark for evaluating the capabilities of large language models (LLMs…

cs.SE20251 cited

Skeleton-Guided-Translation: A Benchmarking Framework for Code Repository Translation with Fine-Grained Quality Evaluation

Xing Zhang, Jiaheng Wen, Fangkai Yang +11

The advancement of large language models has intensified the need to modernize enterprise applications and migrate legacy systems to secure, versatile languages. However, existing…

cs.CL2025

DI-BENCH: Benchmarking Large Language Models on Dependency Inference with Testable Repositories at Scale

Linghao Zhang, Junhao Wang, Shilin He +13

Large Language Models have advanced automated software development, however, it remains a challenge to correctly infer dependencies, namely, identifying the internal components and…