Evaluation of LLM Chatbots for OSINT-based Cyber Threat Awareness
arXiv:2401.15127 · doi:10.1016/j.eswa.2024.125509
Abstract
Knowledge sharing about emerging threats is crucial in the rapidly advancing field of cybersecurity and forms the foundation of Cyber Threat Intelligence (CTI). In this context, Large Language Models are becoming increasingly significant in the field of cybersecurity, presenting a wide range of opportunities. This study surveys the performance of ChatGPT, GPT4all, Dolly, Stanford Alpaca, Alpaca-LoRA, Falcon, and Vicuna chatbots in binary classification and Named Entity Recognition (NER) tasks performed using Open Source INTelligence (OSINT). We utilize well-established data collected in previous research from Twitter to assess the competitiveness of these chatbots when compared to specialized models trained for those tasks. In binary classification experiments, Chatbot GPT-4 as a commercial model achieved an acceptable F1 score of 0.94, and the open-source GPT4all model achieved an F1 score of 0.90. However, concerning cybersecurity entity recognition, all evaluated chatbots have limitations and are less effective. This study demonstrates the capability of chatbots for OSINT binary classification and shows that they require further improvement in NER to effectively replace specially trained models. Our results shed light on the limitations of the LLM chatbots when compared to specialized models, and can help researchers improve chatbots technology with the objective to reduce the required effort to integrate machine learning in OSINT-based CTI tools.
References in corpus (16)
- LLaMA: Open and Efficient Foundation Language Models
- LoRA: Low-Rank Adaptation of Large Language Models
- ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks
- ChatGPT: Jack of all trades, master of none
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications
- How Generative AI models such as ChatGPT can be (Mis)Used in SPC Practice, Education, and Research? An Exploratory Study
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- Evaluation of ChatGPT Model for Vulnerability Detection
- Chatbots to ChatGPT in a Cybersecurity Space: Evolution, Vulnerabilities, Attacks, Challenges, and Future Recommendations
- Chatbots in a Honeypot World
- Representational Strengths and Limitations of Transformers
- Pushing the Limits of ChatGPT on NLP Tasks
- Chatbots As Fluent Polyglots: Revisiting Breakthrough Code Snippets
- MLCopilot: Unleashing the Power of Large Language Models in Solving Machine Learning Tasks
- Are LLMs the Master of All Trades? : Exploring Domain-Agnostic Reasoning Skills of LLMs