1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Jiawen Wang, Pritha Gupta, Ivan Habernal +3
Recent studies demonstrate that Large Language Models (LLMs) are vulnerable to attacks that generate harmful or sensitive outputs. As open-source LLMs are increasingly adopted in h…