3 citations · 5 across the 4 of their papers we have counts for
4 papers · 1 filter
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
Samuele Poppi, Zheng-Xin Yong, Yifei He +4
Recent advancements in Large Language Models (LLMs) have sparked widespread concerns about their safety. Recent work demonstrates that safety alignment of LLMs can be easily remove…
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi +8
We introduce Llama Guard, an LLM-based input-output safeguard model geared towards Human-AI conversation use cases. Our model incorporates a safety risk taxonomy, a valuable tool f…
Intent Classification and Slot Filling for Privacy Policies
Wasi Uddin Ahmad, Jianfeng Chi, Tu Le +3
Understanding privacy policies is crucial for users as it empowers them to learn about the information that matters to them. Sentences written in a privacy policy document explain…
PolicyQA: A Reading Comprehension Dataset for Privacy Policies
Wasi Uddin Ahmad, Jianfeng Chi, Yuan Tian +1
Privacy policy documents are long and verbose. A question answering (QA) system can assist users in finding the information that is relevant and important to them. Prior studies in…