65 citations · 205 across the 35 of their papers we have counts for
1 paper · 1 filter
Shahin Honarvar, Mark van der Wilk, Alastair Donaldson
We present a method for systematically evaluating the correctness and robustness of instruction-tuned large language models (LLMs) for code generation via a new benchmark, Turbulen…