3 papers
cs.AI2025
SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models
Giorgio Piras, Raffaele Mura, Fabio Brau +3
Refusal refers to the functional behavior enabling safety-aligned language models to reject harmful or unethical prompts. Following the growing scientific interest in mechanistic i…
cs.CR2025
Demystifying the Role of Rule-based Detection in AI Systems for Windows Malware Detection
Andrea Ponte, Luca Demetrio, Luca Oneto +3
Malware detection increasingly relies on AI systems that integrate signature-based detection with machine learning. However, these components are typically developed and combined i…
cs.CR2025
Empirical Quantification of Spurious Correlations in Malware Detection
Bianca Perasso, Ludovico Lozza, Andrea Ponte +3
End-to-end deep learning exhibits unmatched performance for detecting malware, but such an achievement is reached by exploiting spurious correlations -- features with high relevanc…