Temporal-augmented observation for navigation of unmanned aerial vehicles: a recurrent reinforcement learning architecture
Date published
Free to read from
Supervisor/s
Industry supervisor/s
Journal Title
Journal ISSN
Volume Title
Department
Course name
Type
ISSN
Format
Citation
Abstract
In aerial robotics, data-driven Reinforcement Learning (RL) approaches have proven highly effective for obstacle avoidance and goal-directed navigation, especially when operating on high-dimensional sensor data that provide only partial, local information about the environment. Such limited observability, combined with irregularly shaped obstacles, poses significant challenges for reactive control policies that rely solely on instantaneous observations. To address these issues, this paper introduces a Twin Deep Deterministic Policy Gradient (TD3)-based algorithm that leverages explicit Temporal Augmentation of the Observation space (TAO-TD3). The proposed method preserves the simplicity of the original TD3 framework by augmenting the observation with a short history of past states and incorporating a lightweight recurrent network, without requiring changes to the TD3 training paradigm. Extensive simulations across diverse environmental topographies and irregular obstacle shapes demonstrate that the proposed approach nearly halves the collision rate and improves overall navigation success compared to feedforward RL-based architectures.
