Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
IterInject: Indirect Prompt Injection Against LLM Agents via Feedback-Guided Iterative Optimization
Zixuan Chen, Jiaxiang Chen, Li Luo +4
LLM-based agents are increasingly deployed for complex tasks requiring planning, tool use, and interaction with external services. Their reliance on untrusted external content expo…
cs.LG2026
Dynamic Optimization and Safety Indicator Injection for Jailbreaking Text-to-Image Models with Multimodal Safety Filters
Zixuan Chen, Hao Lin, Ke Xu +2
Text-to-image (T2I) models can generate not-safe-for-work (NSFW) content, motivating multi-stage safety pipelines with both text and image filters. Newer LLM-based filters detect l…