1 paper
Qi Hu, Yifeng Tang, Qinghua Wang +7
Large language models are increasingly deployed as coding agents, shifting safety from individual responses to action sequences. Existing benchmarks, however, primarily assess whet…