Automatic Labeling for Entity Extraction in Cyber Security
arXiv:1308.4941
Abstract
Timely analysis of cyber-security information necessitates automated information extraction from unstructured text. While state-of-the-art extraction methods produce extremely accurate results, they require ample training data, which is generally unavailable for specialized applications, such as detecting security related entities; moreover, manual annotation of corpora is very costly and often not a viable solution. In response, we develop a very precise method to automatically label text from several data sources by leveraging related, domain-specific, structured data and provide public access to a corpus annotated with cyber-security entities. Next, we implement a Maximum Entropy Model trained with the average perceptron on a portion of our corpus (750,000 words) and achieve near perfect precision, recall, and accuracy, with training times under 17 seconds.
10 pages
Cited by in corpus (7)
- Towards a relation extraction framework for cyber-security concepts
- What are the attackers doing now? Automating cyber threat intelligence extraction from text on pace with the changing threat landscape: A survey
- An Automated, End-to-End Framework for Modeling Attacks From Vulnerability Descriptions
- Few-Sample Named Entity Recognition for Security Vulnerability Reports by Fine-Tuning Pre-Trained Language Models
- Deep Learning Approach for Intelligent Named Entity Recognition of Cyber Security
- PACE: Pattern Accurate Computationally Efficient Bootstrapping for Timely Discovery of Cyber-Security Concepts
- Detecting Cybersecurity Events from Noisy Short Text