adversarial auditing 1benchmark provenance 1inference-time defenses 1multimodal large language models 1safety evaluation 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CR2026
When the Defense Writes the Refusal: Auditing Keyword-Scored Evaluation of Inference-Time Defenses for Multimodal Large Language Models
Bulat Nutfullin, Vladimir Evgrafov, Dmitry Namiot
The paper audits the evaluation of inference-time safety defenses for multimodal large language models, revealing provenance and implementation errors in benchmark pipelines and pr…
cs.CL2024
Prompt Injection Attacks in Defended Systems
Daniil Khomsky, Narek Maloyan, Bulat Nutfullin
Large language models play a crucial role in modern natural language processing technologies. However, their extensive use also introduces potential security risks, such as the pos…
cs.CL2024
Trojan Detection in Large Language Models: Insights from The Trojan Detection Challenge
Narek Maloyan, Ekansh Verma, Bulat Nutfullin +1
Large Language Models (LLMs) have demonstrated remarkable capabilities in various domains, but their vulnerability to trojan or backdoor attacks poses significant security risks. T…