This peer-reviewed study evaluates whether four leading LLMs (GPT-4o, GPT-4o-mini, LLaMA, and Gemini) can autonomously perform a complete HAZOP analysis from a P&ID without human intervention. Each model generated full HAZOP worksheets that were benchmarked against an expert-prepared reference. Results showed that only 19–37% of AI-generated scenarios were technically valid, and safeguards suggested were predominantly procedural rather than engineering-based. The authors conclude that LLMs can serve as a useful starting point for HAZOP but cannot replace expert-led studies. This paper is directly relevant to PSM practitioners evaluating AI augmentation of PHA workflows, providing empirical evidence of both the promise and current limitations of LLM-assisted HAZOP, a cornerstone of RBPS Element 7 (Hazard Identification & Risk Analysis).
AUTHORS
J. Lee, S. Park, S. Oh, B. Ma
CITATIONS
J. Lee, S. Park, S. Oh, and B. Ma, "Can large language models automate the HAZOP process without human intervention?" Safety Science, vol. 194, p. 107039, 2026. https://doi.org/10.1016/j.ssci.2025.107039