CERESResearch Repository

Learning what matters now: a dual-critic context-aware RL framework for priority-driven information gain

Loading...
Thumbnail Image

Date published

Free to read from

2026-03-20

Supervisor/s

Industry supervisor/s

Journal Title

Journal ISSN

Volume Title

Department

Course name

ISSN

Format

Citation

Panagopoulos D, Perrusquía A, Guo W. (2025) Learning what matters now: a dual-critic context-aware RL framework for priority-driven information gain. In: Proceedings of the 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC), 5-8 Oct 2025, Vienna, Austria, pp. 5655-5660

Abstract

Autonomous systems operating in high-stakes search-and-rescue (SAR) missions must continuously gather mission-critical information while flexibly adapting to shifting operational priorities. We propose CA-MIQ (Context-Aware Max-Information Q-learning), a lightweight dual-critic reinforcement learning (RL) framework that dynamically adjusts its exploration strategy whenever mission priorities change. CA-MIQ pairs a standard extrinsic critic for task reward with an intrinsic critic that fuses state-novelty, information-location awareness, and real-time priority alignment. A built-in shift detector triggers transient exploration boosts and selective critic resets, allowing the agent to re-focus after a priority revision. In a simulated SAR grid-world, where experiments specifically test adaptation to changes in the priority order of information types the agent is expected to focus on, CA-MIQ achieves nearly four times higher mission-success rates than baselines after a single priority shift and more than three times better performance in multiple-shift scenarios, achieving 100% recovery while baseline methods fail to adapt. These results highlight CA-MIQ’s effectiveness in any discrete environment with piecewise-stationary information-value distributions.

Description

Software description

Software language

Git repository

Keywords

46 Information and Computing Sciences, 4602 Artificial Intelligence, 4611 Machine Learning, Behavioral and Social Science, Basic Behavioral and Social Science, Information gain, intrinsic motivation, priority shift, reinforcement learning

DOI

Rights

Attribution 4.0 International

Funder/s

This work is funded by EPSRC iCASE with Thales UK (EP/X52475X/1)

Grant number

Relationships

Relationships

Resources