CERESResearch Repository

Interpreting the observed behavior of a class of autonomous linear systems using explainable inverse reinforcement learning

Loading...
Thumbnail Image

Date published

Free to read from

2026-06-24

Supervisor/s

Industry supervisor/s

Journal Title

Journal ISSN

Volume Title

Department

Course name

ISSN

2162-237X

Format

Citation

Perrusquía A, Zou M, Guo W. (2026) Interpreting the observed behavior of a class of autonomous linear systems using explainable inverse reinforcement learning. IEEE Transactions on Neural Networks and Learning Systems, Available online 8 June 2026

Abstract

One of the main challenges faced by society is how to verify the safety of autonomous systems. As the level of autonomy grows, it becomes critical to understand why an autonomous system exhibits a particular behavior and what we need to do to ensure it does not pose risks to people or the surrounding environment. To this end, we need to infer, from observational data, the necessary evidence or decision-making factors to effectively interpret the autonomous system's behavior. Previous approaches in the literature have used explainable models to provide simple input-output mappings to interpret the observer behavior. However, autonomous systems do not pose simple input-output relationships, which compromises the veracity of the interpretations. To alleviate this issue, this article proposes a novel explainable inverse reinforcement learning (EXIRL) that provides the evidence to interpret the behavior of a class of autonomous linear systems. The approach uses a combined model-free and model-based mechanism that improves the learning phase of traditional inverse learning algorithms to accurately infer the evidence from the data. The inferred evidence is used to design counterfactual explanations to interpret the observed behavior and provide the mechanisms to modify it into a desired one. Simulation studies using a DJI Phantom 4 and a quadrotor model are provided to show the benefits and challenges of the proposed work.

Description

Software description

Software language

Git repository

Keywords

46 Information and Computing Sciences, 4611 Machine Learning, Basic Behavioral and Social Science, Bioengineering, Behavioral and Social Science, Generic health relevance, Autonomous systems, counterfactual explanations, evidence space, explainable inverse reinforcement learning (EXIRL), interpretations

DOI

Rights

Attribution 4.0 International

Funder/s

Grant number

Relationships

Relationships

Resources