collaborators

5 papers

cs.SE2026

PenForge: On-the-Fly Expert Agent Construction for Automated Penetration Testing

Huihui Huang, Jieke Shi, Junkai Chen +6

Penetration testing is essential for identifying vulnerabilities in web applications before real adversaries can exploit them. Recent work has explored automating this process with…

cs.SE2025

Open the Oyster: Empirical Evaluation and Improvement of Code Reasoning Confidence in LLMs

Shufan Wang, Xing Hu, Junkai Chen +2

With the widespread application of large language models (LLMs) in the field of code intelligence, increasing attention has been paid to the reliability and controllability of thei…

cs.CR2025

Out of Distribution, Out of Luck: How Well Can LLMs Trained on Vulnerability Datasets Detect Top 25 CWE Weaknesses?

Yikun Li, Ngoc Tan Bui, Ting Zhang +16

Automated vulnerability detection research has made substantial progress, yet its real-world impact remains limited. Prior work found that current vulnerability datasets suffer fro…

cs.SE2025

Assessing and Advancing Benchmarks for Evaluating Large Language Models in Software Engineering Tasks

Xing Hu, Feifei Niu, Junkai Chen +5

Large language models (LLMs) are gaining increasing popularity in software engineering (SE) due to their unprecedented performance across various applications. These models are inc…

cs.SE2025

R2Vul: Learning to Reason about Software Vulnerabilities with Reinforcement Learning and Structured Reasoning Distillation

Martin Weyssow, Chengran Yang, Junkai Chen +12

Large language models (LLMs) have shown promising performance in software vulnerability detection, yet their reasoning capabilities remain unreliable. We propose R2Vul, a method th…