1 paper
Ze Sheng, Aleksandar Kezic, Zhicheng Chen +1
Evaluating the ability of large language models (LLMs) to discover software bugs is increasingly important. Existing benchmarks typically evaluate this capability by asking the mod…