2 papers
cs.CR2026
From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts
Bo Chen
Security research artifacts---repositories, PoC exploits, and validation pipelines---are increasingly produced by LLM/agent-driven vulnerability workflows, yet the gap between \emp…
cs.AI2026
WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader
Boshui Chen, Huiping Liu, Shaolei Zhang
Large language models increasingly generate complete websites from natural-language descriptions, and reinforcement learning has become a central approach to closing their remainin…