23 citations · 25 across the 8 of their papers we have counts for
Showing 2024Show all
2 papers · 1 filter
cs.SE2024
Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark
Dewu Zheng, Yanlin Wang, Ensheng Shi +4
With the rapid advancement of large language models (LLMs), extensive research has been conducted to investigate the code generation capabilities of LLMs. However, existing efforts…
cs.SE2024★ 2 cited
HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation
Dewu Zheng, Yanlin Wang, Ensheng Shi +4
To evaluate the repository-level code generation capabilities of Large Language Models (LLMs) in complex real-world software development scenarios, many evaluation methods have bee…