3 papers
cs.CR2026
Document-Authored Control-Signal Impersonation: A Low-Cost Indirect Prompt Attack on RAG Safety Boundaries
Jianguo Zhu
Retrieval-augmented generation (RAG) systems often serialize user queries, retrieved documents, metadata, system labels, and task instructions into one natural-language prompt. We…
cs.CL2026
Discourse-Role Labels as Presentation-Time Variables for Context Use in Language Models
Jianguo Zhu, Xiangmei Li, Wenjie Liu
Context-augmented language model systems often wrap supplied content with labels such as Reference:, Evidence:, Instruction:, Note:, or Example:, but the effect of these labels on…
cs.AI2025
Unbiased Evaluation of Large Language Models from a Causal Perspective
Meilin Chen, Jian Tian, Liang Ma +3
Benchmark contamination has become a significant concern in the LLM evaluation community. Previous Agents-as-an-Evaluator address this issue by involving agents in the generation o…