1 paper
Zhenghan Song, Yulong Liu, Cheng Wan +4
Execution-based evaluation of LLM-generated code implicitly treats successful execution as a proxy for correctness. In scientific simulation, this proxy is insufficient: a generate…