CERESResearch Repository

Large language models powered system safety assessment: applying STPA and FRAM

Loading...
Thumbnail Image

Date published

Free to read from

2025-08-20

Supervisor/s

Industry supervisor/s

Journal Title

Journal ISSN

Volume Title

Publisher

Department

Course name

ISSN

0925-7535

Format

Citation

Kaya GK, Bovell D, Sujan M, Braithwaite G. (2025) Large language models powered system safety assessment: applying STPA and FRAM. Safety Science, Volume 191, November 2025, Article number 106960

Abstract

The advancement of large language models (LLMs) shows immense promise in many domains. However, their reliability is still questionable. This study aims to comparatively examine the performance of ChatGPT and Gemini in conducting a stand-alone systems-based risk assessment using System-Theoretic Process Analysis (STPA) and the Functional Resonance Analysis Method (FRAM). Our findings revealed that both LLMs demonstrated weaknesses in their analyses, with ChatGPT generally outperforming Gemini regarding response comprehensiveness and adhering to the prompted format. Specifically, LLMs failed to use systems thinking in their stand-alone applications and failed to follow up on previous prompt outputs. While LLMs can provide substantial amounts of information quickly, the effectiveness of LLMs in system safety assessment is contingent on addressing their limitations and implementing strategies to improve their capabilities.

Description

Software description

Software language

Git repository

Keywords

40 Engineering, 52 Psychology, Human Factors, 42 Health sciences, LLM, AI chatbot, ChatGPT, FRAM, STPA

DOI

Rights

Attribution 4.0 International

Funder/s

Grant number

Relationships

Relationships

Resources