Safety Analysis in the Era of Large Language Models: A Case Study of STPA using ChatGPT
arXiv:2304.01246 · doi:10.1016/j.mlwa.2025.100622
Abstract
Can safety analysis make use of Large Language Models (LLMs)? A case study explores Systems Theoretic Process Analysis (STPA) applied to Automatic Emergency Brake (AEB) and Electricity Demand Side Management (DSM) systems using ChatGPT. We investigate how collaboration schemes, input semantic complexity, and prompt guidelines influence STPA results. Comparative results show that using ChatGPT without human intervention may be inadequate due to reliability related issues, but with careful design, it may outperform human experts. No statistically significant differences are found when varying the input semantic complexity or using common prompt guidelines, which suggests the necessity for developing domain-specific prompt engineering. We also highlight future challenges, including concerns about LLM trustworthiness and the necessity for standardisation and regulation in this domain.
Under Review
References in corpus (6)
- A Survey of Large Language Models
- Scaling Laws for Neural Language Models
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Recent Advances in Natural Language Processing via Large Pre-Trained Language Models: A Survey
- On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective
- Welcome Your New AI Teammate: On Safety Analysis by Leashing Large Language Models