889 citations · 1.7k across the 40 of their papers we have counts for
14 papers · 1 filter
Out-of-Distribution Detection for Continual Learning: Design Principles and Benchmarking
Srishti Gupta, Riccardo Balia, Daniele Angioni +7
Recent years have witnessed significant progress in the development of machine learning models across a wide range of fields, fueled by increased computational resources, large-sca…
Counterfeit Answers: Adversarial Forgery against OCR-Free Document Visual Question Answering
Marco Pintore, Maura Pintor, Dimosthenis Karatzas +1
Document Visual Question Answering (DocVQA) enables end-to-end reasoning grounded on information present in a document input. While recent models have shown impressive capabilities…
SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models
Giorgio Piras, Raffaele Mura, Fabio Brau +3
Refusal refers to the functional behavior enabling safety-aligned language models to reject harmful or unethical prompts. Following the growing scientific interest in mechanistic i…
LatentBreak: Jailbreaking Large Language Models through Latent Space Feedback
Raffaele Mura, Giorgio Piras, Kamilė Lukošiūtė +3
Jailbreaks are adversarial attacks designed to bypass the built-in safety mechanisms of large language models. Automated jailbreaks typically optimize an adversarial suffix or adap…
S2AP: Score-space Sharpness Minimization for Adversarial Pruning
Giorgio Piras, Qi Zhao, Fabio Brau +3
Adversarial pruning methods have emerged as a powerful tool for compressing neural networks while preserving robustness against adversarial attacks. These methods typically follow…
Evaluating Line-level Localization Ability of Learning-based Code Vulnerability Detection Models
Marco Pintore, Giorgio Piras, Angelo Sotgiu +2
To address the extremely concerning problem of software vulnerability, system security is often entrusted to Machine Learning (ML) algorithms. Despite their now established detecti…