6 papers
OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajectories
Yibing Liu, Yangze Liu, Xiaolong Yin +4
Task success can hide process anomalies in real-world agent executions. An agent may pass the final task oracle while still accumulating unresolved ambiguity, unsafe external write…
Multi-Agent Honeypot-Based Request-Response Context Dataset for Improved SQL Injection Detection Performance
Hao Yu, Hui Li, FengYuan Shi +4
SQL injection remains a major threat to web applications, as existing defenses often fail against obfuscation and evolving attacks because of neglecting the request-response contex…
Argus: A Multi-Agent Sensitive Information Leakage Detection Framework Based on Hierarchical Reference Relationships
Bin Wang, Hui Li, Liyang Zhang +5
Sensitive information leakage in code repositories has emerged as a critical security challenge. Traditional detection methods that rely on regular expressions, fingerprint feature…
RefleXGen:The unexamined code is not worth using
Bin Wang, Hui Li, AoFan Liu +7
Security in code generation remains a pivotal challenge when applying large language models (LLMs). This paper introduces RefleXGen, an innovative method that significantly enhance…
RA-Gen: A Controllable Code Generation Framework Using ReAct for Multi-Agent Task Execution
Aofan Liu, Haoxuan Li, Bin Wang +2
Code generation models based on large language models (LLMs) have gained wide adoption, but challenges remain in ensuring safety, accuracy, and controllability, especially for comp…
PiCo: Jailbreaking Multimodal Large Language Models via Pictorial Code Contextualization
Aofan Liu, Lulu Tang, Ting Pan +3
Multimodal Large Language Models (MLLMs), which integrate vision and other modalities into Large Language Models (LLMs), significantly enhance AI capabilities but also introduce ne…