2 papers
cs.SE2026
VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation
Zichen Xie, Mrigank Pawagi, Yuxin Liu +5
Large language models can generate useful code from natural language, but their outputs come without correctness guarantees. Verifiable code generation offers a path beyond testing…
cs.LG2026
Property-Driven Evaluation of GNN Expressiveness at Scale: Datasets, Framework, and Study
Sicong Che, Jiayi Yang, Sarfraz Khurshid +1
Advancing trustworthy AI requires principled software engineering approaches to model evaluation. Graph Neural Networks (GNNs) have achieved remarkable success in processing graph-…