A multi-robot path-planning algorithm for autonomous navigation using meta-reinforcement learning based on transfer learning

被引：35

作者：

Wen, Shuhuan ^{[1
,2
]}

Wen, Zeteng ^{[1
,2
]}

Zhang, Di ^{[1
,2
]}

Zhang, Hong ^{[3
]}

Wang, Tao ^{[1
,2
]}

机构：

[1] Yanshan Univ, Engn Res Ctr, Minist Educ Intelligent Control Syst & Intelligen, Qinhuangdao 066004, Hebei, Peoples R China

[2] Yanshan Univ, Key Lab Ind Comp Control Engn Hebei Prov, Qinhuangdao 066004, Hebei, Peoples R China

[3] Southern Univ Sci & Technol, Dept Elect & Elect Engn, Shenzhen 518000, Peoples R China

来源：

APPLIED SOFT COMPUTING | 2021年 / 110卷 / 110期

关键词：

Multi-robot system; Path planning; Deep reinforcement learning; Meta learning; Transfer learning;

D O I：

10.1016/j.asoc.2021.107605

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

The adaptability of multi-robot systems in complex environments is a hot topic. Aiming at static and dynamic obstacles in complex environments, this paper presents dynamic proximal meta policy optimization with covariance matrix adaptation evolutionary strategies (dynamic-PMPO-CMA) to avoid obstacles and realize autonomous navigation. Firstly, we propose dynamic proximal policy optimization with covariance matrix adaptation evolutionary strategies (dynamic-PPO-CMA) based on original proximal policy optimization (PPO) to obtain a valid policy of obstacles avoidance. The simulation results show that the proposed dynamic-PPO-CMA can avoid obstacles and reach the designated target position successfully. Secondly, in order to improve the adaptability of multi-robot systems in different environments, we integrate meta-learning with dynamic-PPO-CMA to form the dynamic-PMPO-CMA algorithm. In training process, we use the proposed dynamic-PMPO-CMA to train robots to learn multi-task policy. Finally, in testing process, transfer learning is introduced to the proposed dynamic-PMPO-CMA algorithm. The trained parameters of meta policy are transferred to new environments and regarded as the initial parameters. The simulation results show that the proposed algorithm can have faster convergence rate and arrive the destination more quickly than PPO, PMPO and dynamic-PPO-CMA. (C) 2021 Elsevier B.V. All rights reserved.

引用

页数：15

共 50 条

[31] Meta-Reinforcement Learning Algorithm Based on Reward and Dynamic Inference
Chen, Jinhao
Zhang, Chunhong
Hu, Zheng
[J]. ADVANCES IN KNOWLEDGE DISCOVERY AND DATA MINING, PT III, PAKDD 2024, 2024, 14647 : 223 - 234
[32] Path planning of autonomous UAVs using reinforcement learning
Chronis, Christos
Anagnostopoulos, Georgios
Politi, Elena
Garyfallou, Antonios
Varlamis, Iraklis
Dimitrakopoulos, George
[J]. 12TH EASN INTERNATIONAL CONFERENCE ON "INNOVATION IN AVIATION & SPACE FOR OPENING NEW HORIZONS", 2023, 2526
[33] Connectivity Guaranteed Multi-robot Navigation via Deep Reinforcement Learning
Lin, Juntong
Yang, Xuyun
Zheng, Peiwei
Cheng, Hui
[J]. CONFERENCE ON ROBOT LEARNING, VOL 100, 2019, 100
[34] Cooperative Multi-Robot Navigation in Dynamic Environment with Deep Reinforcement Learning
Han, Ruihua
Chen, Shengduo
Hao, Qi
[J]. 2020 IEEE INTERNATIONAL CONFERENCE ON ROBOTICS AND AUTOMATION (ICRA), 2020, : 448 - 454
[35] Path-planning for an autonomous robot using a simulating annealing
Chiu, Min-Chie
[J]. JOURNAL OF INFORMATION & OPTIMIZATION SCIENCES, 2011, 32 (02): : 297 - 314
[36] Fuzzy Reinforcement Learning and Curriculum Transfer Learning for Micromanagement in Multi-Robot Confrontation
Hu, Chunyang
Xu, Meng
[J]. INFORMATION, 2019, 10 (11)
[37] Indoor Mobile Robot Path Planning and Navigation System Based on Deep Reinforcement Learning
Pai, Neng-Sheng
Tsai, Xiang-Yan
Chen, Pi-Yun
Lin, Hsu -Yung
[J]. SENSORS AND MATERIALS, 2024, 36 (05) : 1959 - 1982
[38] Sequencing of multi-robot behaviors using reinforcement learning
Pietro Pierpaoli
Thinh T. Doan
Justin Romberg
Magnus Egerstedt
[J]. Control Theory and Technology, 2021, 19 : 529 - 537
[39] Coordinated Multi-Robot Exploration using Reinforcement Learning
Mete, Atharva
Mouhoub, Malek
Farid, Ali Moltajaei
[J]. 2023 INTERNATIONAL CONFERENCE ON UNMANNED AIRCRAFT SYSTEMS, ICUAS, 2023, : 265 - 272
[40] Sequencing of multi-robot behaviors using reinforcement learning
Pierpaoli, Pietro
Doan, Thinh T.
Romberg, Justin
Egerstedt, Magnus
[J]. CONTROL THEORY AND TECHNOLOGY, 2021, 19 (04) : 529 - 537

← 1 2 3 4 5 →