Comprehensive evaluation of Mal-API-2019 dataset by machine learning in malware detection
arXiv:2403.02232 · doi:10.62051/ijcsit.v2n1.01
Abstract
This study conducts a thorough examination of malware detection using machine learning techniques, focusing on the evaluation of various classification models using the Mal-API-2019 dataset. The aim is to advance cybersecurity capabilities by identifying and mitigating threats more effectively. Both ensemble and non-ensemble machine learning methods, such as Random Forest, XGBoost, K Nearest Neighbor (KNN), and Neural Networks, are explored. Special emphasis is placed on the importance of data pre-processing techniques, particularly TF-IDF representation and Principal Component Analysis, in improving model performance. Results indicate that ensemble methods, particularly Random Forest and XGBoost, exhibit superior accuracy, precision, and recall compared to others, highlighting their effectiveness in malware detection. The paper also discusses limitations and potential future directions, emphasizing the need for continuous adaptation to address the evolving nature of malware. This research contributes to ongoing discussions in cybersecurity and provides practical insights for developing more robust malware detection systems in the digital era.
References in corpus (8)
- LESSON: Multi-Label Adversarial False Data Injection Attack for Deep Learning Locational Detection
- Large Language Models for Forecasting and Anomaly Detection: A Systematic Literature Review
- BotShape: A Novel Social Bots Detection Approach via Behavioral Patterns
- Robust Data Preprocessing for Machine-Learning-Based Disk Failure Prediction in Cloud Production Environments
- Utilizing Deep Learning for Enhancing Network Resilience in Finance
- Attention Hijacking in Trojan Transformers
- BotTriNet: A Unified and Efficient Embedding for Social Bots Detection via Metric Learning
- Non-Exhaustive Learning Using Gaussian Mixture Generative Adversarial Networks