From the 1 of 25 linked papers with an AI index.
25 papers
Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection
Zonghao Ying, Xiangfan Wu, Huiyu Wu +4
We assess indirect prompt injection in DeepSeek Harness (DSH), using AI-Infra-Guard (A.I.G) to construct tests, deliver controlled taint, execute DSH, collect traces, and judge out…
Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs
Xiangfan Wu, Zonghao Ying, Huiyu Wu +4
As large language models become increasingly widespread, third-party providers that deploy open-weight models have become an important part of the ecosystem. Auditing the quality o…
Mitigating Bias in Large Vision-Language Models via Counterfactual Ensemble Decoding
Yisong Xiao, Aishan Liu, Yongxin Huang +6
Large Vision-Language Models (LVLMs) have achieved remarkable performance across a wide range of tasks; however, they often inherit social biases from their training data, resultin…
SkillJack: Persistent Skill Backdoors in Self-Evolving Agents
Zonghao Ying, Xiangfan Wu, Huiyu Wu +4
Self-evolving agents increasingly convert interaction histories into reusable skills that persist beyond individual tasks. While prior work studies memory and retrieval poisoning,…
SkillSentry: Adaptive Honey Worlds for Dynamic Safety Testing of Agent Skills
Nizhang Li, Zonghao Ying, Xiangfan Wu +7
External skills extend the capabilities of large language model agents, but also introduce an execution-time attack surface: a skill that appears benign under inspection may reveal…
SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems
Haowen Dai, Zonghao Ying, Wenfeng Li +10
Multi-agent systems improve capability through task decomposition and role specialization, but these same mechanisms introduce an important safety blind spot: a harmful objective c…