2 papers
cs.CR2026
VirtualCrime: Evaluating the Criminal Potential and Agentic Behaviors of Large Language Models via Sandbox Simulation
Yilin Tang, Yu Wang, Lanlan Qiu +4
Large language models (LLMs) have shown strong capabilities in multi-step decision-making, and are increasingly integrated into real-world agentic applications. While existing safe…
cs.CL2026
Can LLMs Generate Reliable Test Case Generators? A Study on Competition-Level Programming Problems
Yuhan Cao, Zian Chen, Kun Quan +17
Large Language Models (LLMs) have demonstrated remarkable capabilities in code generation, capable of tackling complex tasks during inference. However, the extent to which LLMs can…