5 citations · 5 across the 2 of their papers we have counts for
2 papers
cs.CR2024★ 5 cited
Defending Against Indirect Prompt Injection Attacks With Spotlighting
Keegan Hines, Gary Lopez, Matthew Hall +3
Large Language Models (LLMs), while powerful, are built and trained to process a single text input. In common applications, multiple inputs can be processed by concatenating them t…
cs.LG2023
Reckoning with the Disagreement Problem: Explanation Consensus as a Training Objective
Avi Schwarzschild, Max Cembalest, Karthik Rao +2
As neural networks increasingly make critical decisions in high-stakes settings, monitoring and explaining their behavior in an understandable and trustworthy manner is a necessity…