collaborators

5 papers

cs.CR2026

Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection

Zonghao Ying, Xiangfan Wu, Huiyu Wu +4

We assess indirect prompt injection in DeepSeek Harness (DSH), using AI-Infra-Guard (A.I.G) to construct tests, deliver controlled taint, execute DSH, collect traces, and judge out…

cs.CR2026

Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs

Xiangfan Wu, Zonghao Ying, Huiyu Wu +4

As large language models become increasingly widespread, third-party providers that deploy open-weight models have become an important part of the ecosystem. Auditing the quality o…

cs.CR2026

SkillJack: Persistent Skill Backdoors in Self-Evolving Agents

Zonghao Ying, Xiangfan Wu, Huiyu Wu +4

Self-evolving agents increasingly convert interaction histories into reusable skills that persist beyond individual tasks. While prior work studies memory and retrieval poisoning,…

cs.CR2026

Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming

Yong Yang, Xing Zheng, Huiyu Wu +7

The fast growth of open-source AI infrastructure, from model serving engines and agent platforms to the Model Context Protocol (MCP) ecosystem and the language models themselves, h…

cs.CR2025

SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity

Pengfei Jing, Mengyun Tang, Xiaorong Shi +5

Evaluating Large Language Models (LLMs) is crucial for understanding their capabilities and limitations across various applications, including natural language processing and code…