collaborators

6 papers

cs.CR2026

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems

Jiaqi Luo, Jiarun Dai, Zhile Chen +8

Nowadays, the autonomous execution of cyberattacks capable of causing substantial real-world harm is widely regarded as one of the critical red lines that frontier AI systems must…

cs.CR2026

AgentGuard: An Attribute-Based Access Control Framework for Tool-Use LLM-Based Agent

Jiaqi Luo, Songyang Peng, Jiarun Dai +6

LLM-based agents have recently attracted significant attention due to their ability to autonomously invoke relevant tools to accomplish complex tasks. However, recent studies have…

cs.CR2025

Shell or Nothing: Real-World Benchmarks and Memory-Activated Agents for Automated Penetration Testing

Wuyuao Mai, Geng Hong, Qi Liu +5

Penetration testing is critical for identifying and mitigating security vulnerabilities, yet traditional approaches remain expensive, time-consuming, and dependent on expert human…

cs.SE2025

CXXCrafter: An LLM-Based Agent for Automated C/C++ Open Source Software Building

Zhengmin Yu, Yuan Zhang, Ming Wen +3

Project building is pivotal to support various program analysis tasks, such as generating intermediate rep- resentation code for static analysis and preparing binary code for vulne…

cs.CR2025

You Can't Eat Your Cake and Have It Too: The Performance Degradation of LLMs with Jailbreak Defense

Wuyuao Mai, Geng Hong, Pei Chen +5

With the rise of generative large language models (LLMs) like LLaMA and ChatGPT, these models have significantly transformed daily life and work by providing advanced insights. How…

cs.CR2025

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity

Zhengmin Yu, Jiutian Zeng, Siyi Chen +7

Over the past year, there has been a notable rise in the use of large language models (LLMs) for academic research and industrial practices within the cybersecurity field. However,…