1 paper · 1 filter
Forough Mehralian, Ryan Shar, James R. Rae +1
As large language models become increasingly capable of generating code, evaluating their performance remains a complex and evolving challenge. Existing benchmarks primarily focus…