1 paper
Justin Chavarria, Rohan Raizada, Justin White +1
We introduce SOCK, a benchmark command line interface (CLI) that measures large language models' (LLMs) ability to self-replicate without human intervention. In this benchmark, sel…