Showing cs.CRShow all
2 papers · 1 filter
cs.CR2025
Black-Box Guardrail Reverse-engineering Attack
Hongwei Yao, Yun Xia, Shuo Shao +3
Large language models (LLMs) increasingly employ guardrails to enforce ethical, legal, and application-specific constraints on their outputs. While effective at mitigating harmful…
cs.CR2025
ControlNET: A Firewall for RAG-based LLM System
Hongwei Yao, Haoran Shi, Yidou Chen +3
Retrieval-Augmented Generation (RAG) has significantly enhanced the factual accuracy and domain adaptability of Large Language Models (LLMs). This advancement has enabled their wid…