From the 1 of 7 linked papers with an AI index.
7 papers
RAFP: Identifying LLM Lineages via Rare-Region Fingerprints
Yun-Yun Tsai, Jia Hao Liang, Chuan Guo +2
The paper proposes RAFP, a non‑invasive method that creates fingerprints from rare prompt‑response regions to reliably identify the lineage of large language models even after vari…
Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks
Sizhe Chen, Arman Zharmagambetov, David Wagner +1
Prompt injection attacks, where untrusted data contains an injected prompt to manipulate the system, have been listed as the top security threat to LLM-integrated applications. Mod…
AdvPrefix: An Objective for Nuanced LLM Jailbreaks
Sicheng Zhu, Brandon Amos, Yuandong Tian +2
Many jailbreak attacks on large language models (LLMs) rely on a common objective: making the model respond with the prefix ``Sure, here is (harmful request)''. While straightforwa…
The Imitation Game: Using Large Language Models as Chatbots to Combat Chat-Based Cybercrimes
Yifan Yao, Baojuan Wang, Jinhao Duan +4
Chat-based cybercrime has emerged as a pervasive threat, with attackers leveraging real-time messaging platforms to conduct scams that rely on trust-building, deception, and psycho…
SecAlign: Defending Against Prompt Injection with Preference Optimization
Sizhe Chen, Arman Zharmagambetov, Saeed Mahloujifar +3
Large language models (LLMs) are becoming increasingly prevalent in modern software systems, interfacing between the user and the Internet to assist with tasks that require advance…
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
Anselm Paulus, Arman Zharmagambetov, Chuan Guo +2
Large Language Models (LLMs) are vulnerable to jailbreaking attacks that lead to generation of inappropriate or harmful content. Manual red-teaming requires a time-consuming search…