2 papers
cs.SE2026
The Illusion of Safety: Multi-Tier Verification of AI vs. Human C++ Code
Saif Mahmud, Fadul Sikder, Yuede Ji +3
As large language models (LLMs) are increasingly deployed for systems programming, their ability to generate secure C++ code, where a single memory-safety failure creates an exploi…
cs.SE2026
EGSS: Entropy-guided Stepwise Scaling for Reliable Software Engineering
Chenhui Mao, Yuanting Lei, Zhixiang Wei +6
Agentic Test-Time Scaling (TTS) has delivered state-of-the-art (SOTA) performance on complex software engineering tasks such as code generation and bug fixing. However, its practic…