CERESResearch Repository

Interpreting the observed behavior of a class of autonomous linear systems using explainable inverse reinforcement learning

dc.contributor.authorPerrusquía, Adolfo
dc.contributor.authorZou, Mengbang
dc.contributor.authorGuo, Weisi
dc.date.accessioned2026-06-24T08:53:49Z
dc.date.available2026-06-24T08:53:49Z
dc.date.freetoread2026-06-24
dc.date.issued2026-12-31
dc.date.pubOnline2026-06-08
dc.description.abstractOne of the main challenges faced by society is how to verify the safety of autonomous systems. As the level of autonomy grows, it becomes critical to understand why an autonomous system exhibits a particular behavior and what we need to do to ensure it does not pose risks to people or the surrounding environment. To this end, we need to infer, from observational data, the necessary evidence or decision-making factors to effectively interpret the autonomous system's behavior. Previous approaches in the literature have used explainable models to provide simple input-output mappings to interpret the observer behavior. However, autonomous systems do not pose simple input-output relationships, which compromises the veracity of the interpretations. To alleviate this issue, this article proposes a novel explainable inverse reinforcement learning (EXIRL) that provides the evidence to interpret the behavior of a class of autonomous linear systems. The approach uses a combined model-free and model-based mechanism that improves the learning phase of traditional inverse learning algorithms to accurately infer the evidence from the data. The inferred evidence is used to design counterfactual explanations to interpret the observed behavior and provide the mechanisms to modify it into a desired one. Simulation studies using a DJI Phantom 4 and a quadrotor model are provided to show the benefits and challenges of the proposed work.
dc.description.journalNameIEEE Transactions on Neural Networks and Learning Systems
dc.format.mediumPrint-Electronic
dc.identifier.citationPerrusquía A, Zou M, Guo W. (2026) Interpreting the observed behavior of a class of autonomous linear systems using explainable inverse reinforcement learning. IEEE Transactions on Neural Networks and Learning Systems, Available online 8 June 2026en_UK
dc.identifier.eissn2162-2388
dc.identifier.elementsID871177
dc.identifier.issn2162-237X
dc.identifier.urihttps://doi.org/10.1109/tnnls.2026.3699708
dc.identifier.urihttps://dspace.lib.cranfield.ac.uk/handle/1826/25360
dc.languageeng
dc.language.isoen
dc.publisherInstitute of Electrical and Electronics Engineers (IEEE)en_UK
dc.publisher.urihttps://ieeexplore.ieee.org/document/11554117
dc.rightsAttribution 4.0 Internationalen
dc.rights.urihttp://creativecommons.org/licenses/by/4.0/
dc.subject46 Information and Computing Sciencesen_UK
dc.subject4611 Machine Learningen_UK
dc.subjectBasic Behavioral and Social Scienceen_UK
dc.subjectBioengineeringen_UK
dc.subjectBehavioral and Social Scienceen_UK
dc.subjectGeneric health relevanceen_UK
dc.subjectAutonomous systemsen_UK
dc.subjectcounterfactual explanationsen_UK
dc.subjectevidence spaceen_UK
dc.subjectexplainable inverse reinforcement learning (EXIRL)en_UK
dc.subjectinterpretationsen_UK
dc.titleInterpreting the observed behavior of a class of autonomous linear systems using explainable inverse reinforcement learningen_UK
dc.typeArticle
dc.type.subtypeJournal Article
dcterms.dateAccepted2026-05-31

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Interpreting_the_observed_behavior-2026.pdf
Size:
3.95 MB
Format:
Adobe Portable Document Format
Description:
Accepted version

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.63 KB
Format:
Plain Text
Description: