Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
Jingwei Yi, Yueqi Xie, Bin Zhu +4
The integration of large language models with external content has enabled applications such as Microsoft Copilot but also introduced vulnerabilities to indirect prompt injection a…
cs.CL2024
Elephant in the Room: Unveiling the Impact of Reward Model Quality in Alignment
Yan Liu, Xiaoyuan Yi, Xiaokang Chen +6
The demand for regulating potentially risky behaviors of large language models (LLMs) has ignited research on alignment methods. Since LLM alignment heavily relies on reward models…