Tactical planning interception enhancement using expert learning - twin delayed deep deterministic policy gradient
Date published
Free to read from
Supervisor/s
Industry supervisor/s
Journal Title
Journal ISSN
Volume Title
Department
Course name
Type
ISSN
Format
Citation
Abstract
The accurate interception of adversarial unmanned aerial vehicles (UAVs) is paramount for the protection of people and national facilities. Urban cities pose several challenges for target interception algorithms due to the presence of buildings and flying constraints that limit the manoeuvrability of UAVs for target interception. Deep Reinforcement Learning (DRL) algorithms have been deployed to solve the task effectively. However, the design of its inner elements such as the reward function and action distribution limits its generalisation to different environments. To solve this issue, this paper proposes a novel twin-delayed deep deterministic policy gradient (TD3) based expert learning algorithm that combines previous expert experiences with on-line learning to regularise and improve the policy learning effectively. This is done by following an action distribution algorithm that allows a learner agent to mix its own actions with expert ones for learning improvement and fast convergence. Extensive simulation studies are carried out under diverse urban cities configurations to show the robustness and high-accuracy of the proposed approach compared with traditional DRL baseline algorithms.
