3 papers
cs.CR2026
SFCoT: Safer Chain-of-Thought via Active Safety Evaluation and Calibration
Yu Pan, Wenlong Yu, Tiejun Wu +4
Large language models (LLMs) have demonstrated remarkable capabilities in complex reasoning tasks. However, they remain highly susceptible to jailbreak attacks that undermine their…
cs.SE2025
VulnRepairEval: An Exploit-Based Evaluation Framework for Assessing Large Language Model Vulnerability Repair Capabilities
Weizhe Wang, Wei Ma, Qiang Hu +6
The adoption of Large Language Models (LLMs) for automated software vulnerability patching has shown promising outcomes on carefully curated evaluation sets. Nevertheless, existing…
cs.CR2025
Unlocking User-oriented Pages: Intention-driven Black-box Scanner for Real-world Web Applications
Weizhe Wang, Yao Zhang, Kaitai Liang +5
Black-box scanners have played a significant role in detecting vulnerabilities for web applications. A key focus in current black-box scanning is increasing test coverage (i.e., ac…