From the 1 of 1 linked paper with an AI index.
1 paper
Ananda Dhakal, Krish Neupane, Aarjan Chaudhary
The paper conducts a controlled evaluation of plain coding agents versus specialized penetration‑testing architectures on the XBOW benchmark, showing that baseline agents already s…