activity
20232026
most citedSafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image Models

19 citations · 48 across the 48 of their papers we have counts for

collaborators
Showing cs.CRShow all

15 papers · 1 filter

cs.CR2026

Beyond Over-Refusal: Defending Indirect Prompt Injection via Latent Instruction Manifolds

Jiahao Chen, Rui Yin, Xinfeng Li +6

Large Language Models (LLMs) have been integrated into complex ecosystems (e.g., Code Agents), while Indirect Prompt Injection (IPI) attacks have emerged as critical barriers to th…

cs.CR2026

Understanding Implicit Trust Errors in Core Carrier Networks through Multi-Agent Flaw Discovery and Analysis

Ziyu Lin, Ziting Wang, Xinfeng Li +2

Cellular core networks (CNs) are critical infrastructure, yet their internal security model has historically relied on physical isolation: interfaces between core components often…

cs.CR2026

Customization under Fire: Plugin Poisoning in Text-to-Image Ecosystem

Jiahao Chen, Xing He, Yong Yang +6

The prosperity of text-to-image (T2I) models has fostered a vibrant share-and-play ecosystem centered on Low-Rank Adaptation (LoRA) plugins, which allow users to customize and shar…

cs.CR2026

You Told Me to Do It: Measuring Instructional Text-induced Private Data Leakage in LLM Agents

Ching-Yu Kao, Xinfeng Li, Shenyu Dai +4

High-privilege LLM agents that autonomously process external documentation are increasingly trusted to automate tasks by reading and executing project instructions, yet they are gr…

cs.CR2026

The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis

Peiran Wang, Xinfeng Li, Chong Xiang +5

The evolution of Large Language Models (LLMs) has resulted in a paradigm shift towards autonomous agents, necessitating robust security against Prompt Injection (PI) vulnerabilitie…

cs.CR2026

DP-MGTD: Privacy-Preserving Machine-Generated Text Detection via Adaptive Differentially Private Entity Sanitization

Lionel Z. Wang, Yusheng Zhao, Jiabin Luo +6

The deployment of Machine-Generated Text (MGT) detection systems necessitates processing sensitive user data, creating a fundamental conflict between authorship verification and pr…