2 papers
cs.CR2026
VICBench: A Multi-Language Benchmark for Code Vulnerability Detection
Jin Lu, Xuening Han, Yang Zhong +4
Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VICs) - the commits that first introduce vulnerabilities into cod…
cs.PL2026
Learning Context-Free Grammars for Grammar-Constrained Decoding via Declarative Agentic Programming with Guarantees
Kevin Cheang, Geoff Hulette, Rahul Kumar +5
Language models (LMs) are increasingly used to interact with external services via programs written in domain-specific languages (DSLs). Unfortunately, since DSLs are often low-res…