7 papers
Adversarial Frontiers: Minimum-Norm Attack Ensembles for Robustness Evaluation
Luca Scionis, Luca Melis, Maura Pintor +5
Adversarial robustness is commonly evaluated with predefined attack ensembles, such as AutoAttack, at a single perturbation budget and on a selective choice of pertur…
Latent-space Attacks for Refusal Evasion in Language Models
Giorgio Piras, Raffaele Mura, Fabio Brau +4
Safety-aligned language models are trained to refuse harmful requests, yet refusal behavior can be suppressed by steering their internal representations. Existing methods do so by…
SAGE-5GC: Security-Aware Guidelines for Evaluating Anomaly Detection in the 5G Core Network
Cristian Manca, Christian Scano, Giorgio Piras +3
Machine learning-based anomaly detection systems are increasingly being adopted in 5G Core networks to monitor complex, high-volume traffic. However, most existing approaches are e…
BlackCATT: Black-box Collusion Aware Traitor Tracing in Federated Learning
Elena RodrÃguez-Lois, Fabio Brau, Maura Pintor +2
Federated Learning has been popularized in recent years for applications involving personal or sensitive data, as it allows the collaborative training of machine learning models th…
Out-of-Distribution Detection for Continual Learning: Design Principles and Benchmarking
Srishti Gupta, Riccardo Balia, Daniele Angioni +7
Recent years have witnessed significant progress in the development of machine learning models across a wide range of fields, fueled by increased computational resources, large-sca…
SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models
Giorgio Piras, Raffaele Mura, Fabio Brau +3
Refusal refers to the functional behavior enabling safety-aligned language models to reject harmful or unethical prompts. Following the growing scientific interest in mechanistic i…