5 papers
Towards Reliable Local Security Agents: Verifiable Post-Training for Linux Privilege Escalation
Philipp Normann, Andreas Happe, Jürgen Cito +1
LLM agents are becoming increasingly important in the security domain, but leading systems are often closed-source, cloud-based, hard to reproduce or use with sensitive code. This…
Beyond the TESSERACT:Trustworthy Dataset Curation for Sound Evaluations of Android Malware Classifiers
Theo Chow, Mario D'Onghia, Lorenz Linhardt +4
The reliability of machine learning critically depends on dataset quality. While machine learning applied to computer vision and natural language processing benefits from high-qual…
Chasing Shadows: Pitfalls in LLM Security Research
Jonathan Evertz, Niklas Risse, Nicolai Neuer +12
Large language models (LLMs) are increasingly prevalent in security research. Their unique characteristics, however, introduce challenges that undermine established paradigms of re…
On the Effectiveness of Adversarial Training on Malware Classifiers
Hamid Bostani, Jacopo Cortellazzi, Daniel Arp +3
Adversarial Training (AT) is a key defense against Machine Learning evasion attacks, but its effectiveness for real-world malware detection remains poorly understood. This uncertai…
TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and Time (Extended Version)
Zeliang Kan, Shae McFadden, Daniel Arp +5
Machine learning (ML) plays a pivotal role in detecting malicious software. Despite the high F1-scores reported in numerous studies reaching upwards of 0.99, the issue is not compl…