7 citations · 12 across the 3 of their papers we have counts for
3 papers
cs.CL2024★ 2 cited
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
Jia Li, Ge Li, Xuanming Zhang +6
How to evaluate Large Language Models (LLMs) in code generation remains an open question. Existing benchmarks have two limitations - data leakage and lack of domain-specific evalua…
cs.CL2024★ 7 cited
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
Jia Li, Ge Li, Xuanming Zhang +2
How to evaluate Large Language Models (LLMs) in code generation is an open question. Existing benchmarks demonstrate poor alignment with real-world code repositories and are insuff…
cs.SE2024★ 3 cited
DevEval: Evaluating Code Generation in Practical Software Projects
Jia Li, Ge Li, Yunfei Zhao +14
How to evaluate Large Language Models (LLMs) in code generation is an open question. Many benchmarks have been proposed but are inconsistent with practical software projects, e.g.,…