2 papers
cs.CR2024
Soft Begging: Modular and Efficient Shielding of LLMs against Prompt Injection and Jailbreaking based on Prompt Tuning
Simon Ostermann, Kevin Baum, Christoph Endres +2
Prompt injection (both direct and indirect) and jailbreaking are now recognized as significant issues for large language models (LLMs), particularly due to their potential for harm…
cs.CL2023
Investigating the Encoding of Words in BERT's Neurons using Feature Textualization
Tanja Baeumel, Soniya Vijayakumar, Josef van Genabith +2
Pretrained language models (PLMs) form the basis of most state-of-the-art NLP technologies. Nevertheless, they are essentially black boxes: Humans do not have a clear understanding…