Large language models powered system safety assessment: applying STPA and FRAM
| dc.contributor.author | Kaya, Gulsum Kubra | |
| dc.contributor.author | Bovell, Dominique | |
| dc.contributor.author | Sujan, Mark | |
| dc.contributor.author | Braithwaite, Graham R. | |
| dc.date.accessioned | 2025-08-20T15:02:49Z | |
| dc.date.available | 2025-08-20T15:02:49Z | |
| dc.date.freetoread | 2025-08-20 | |
| dc.date.issued | 2025-11 | |
| dc.date.pubOnline | 2025-08-06 | |
| dc.description.abstract | The advancement of large language models (LLMs) shows immense promise in many domains. However, their reliability is still questionable. This study aims to comparatively examine the performance of ChatGPT and Gemini in conducting a stand-alone systems-based risk assessment using System-Theoretic Process Analysis (STPA) and the Functional Resonance Analysis Method (FRAM). Our findings revealed that both LLMs demonstrated weaknesses in their analyses, with ChatGPT generally outperforming Gemini regarding response comprehensiveness and adhering to the prompted format. Specifically, LLMs failed to use systems thinking in their stand-alone applications and failed to follow up on previous prompt outputs. While LLMs can provide substantial amounts of information quickly, the effectiveness of LLMs in system safety assessment is contingent on addressing their limitations and implementing strategies to improve their capabilities. | |
| dc.description.journalName | Safety Science | |
| dc.identifier.citation | Kaya GK, Bovell D, Sujan M, Braithwaite G. (2025) Large language models powered system safety assessment: applying STPA and FRAM. Safety Science, Volume 191, November 2025, Article number 106960 | en_UK |
| dc.identifier.eissn | 1879-1042 | |
| dc.identifier.elementsID | 848837 | |
| dc.identifier.issn | 0925-7535 | |
| dc.identifier.paperNo | 106960 | |
| dc.identifier.uri | https://doi.org/10.1016/j.ssci.2025.106960 | |
| dc.identifier.uri | https://dspace.lib.cranfield.ac.uk/handle/1826/24312 | |
| dc.identifier.volumeNo | 191 | |
| dc.language | English | |
| dc.language.iso | en | |
| dc.publisher | Elsevier | en_UK |
| dc.publisher.uri | https://www.sciencedirect.com/science/article/pii/S0925753525001857?via%3Dihub | |
| dc.rights | Attribution 4.0 International | en |
| dc.rights.uri | http://creativecommons.org/licenses/by/4.0/ | |
| dc.subject | 40 Engineering | en_UK |
| dc.subject | 52 Psychology | en_UK |
| dc.subject | Human Factors | en_UK |
| dc.subject | 42 Health sciences | en_UK |
| dc.subject | LLM | en_UK |
| dc.subject | AI chatbot | en_UK |
| dc.subject | ChatGPT | en_UK |
| dc.subject | FRAM | en_UK |
| dc.subject | STPA | en_UK |
| dc.title | Large language models powered system safety assessment: applying STPA and FRAM | en_UK |
| dc.type | Article | |
| dc.type.subtype | Journal Article | |
| dcterms.dateAccepted | 2025-07-29 |
