2 papers
cs.CR2026
PatchBench: Evaluating AI Agents for Vulnerability Patching
Chihao Shen, Jiacheng Li, Aastha Mahajan +3
AI agents have recently demonstrated strong performance in automated vulnerability patching. However, existing evaluations often validate a patch only by testing whether the provid…
cs.SE2025
Benchmarking Correctness and Security in Multi-Turn Code Generation
Ruchit Rawal, Jeffrey Yang Fan Chiang, Chihao Shen +4
AI coding assistants powered by large language models (LLMs) have transformed software development, significantly boosting productivity. While existing benchmarks evaluate the corr…