Action Candidate Driven Clipped Double Q-Learning for Discrete and Continuous Action Tasks

被引：5

作者：

Jiang, Haobo ^{[1
,2
,3
]}

Li, Guangyu ^{[1
,2
,3
]}

Xie, Jin ^{[1
,2
,3
]}

Yang, Jian ^{[1
,2
,3
]}

机构：

[1] Nanjing Univ Sci & Technol, PCA Lab, Sch Comp Sci & Engn, Nanjing 210094, Peoples R China

[2] Nanjing Univ Sci & Technol, Sch Comp Sci & Engn, Key Lab Intelligent Percept & Syst High Dimens In, Minist Educ, Nanjing 210094, Peoples R China

[3] Nanjing Univ Sci & Technol, Jiangsu Key Lab Image & Video Understanding Socia, Sch Comp Sci & Engn, Nanjing 210094, Peoples R China

来源：

IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS | 2024年 / 35卷 / 04期

关键词：

Q-learning; Task analysis; Benchmark testing; Approximation algorithms; Toy manufacturing industry; Markov processes; Learning systems; Clipped double Q-learning; overestimation bias; reinforcement learning; underestimation bias; REINFORCEMENT; BIAS;

D O I：

10.1109/TNNLS.2022.3203024

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Double Q-learning is a popular reinforcement learning algorithm in Markov decision process (MDP) problems. Clipped double Q-learning, as an effective variant of double Q-learning, employs the clipped double estimator to approximate the maximum expected action value. Due to the underestimation bias of the clipped double estimator, the performance of clipped Double Q-learning may be degraded in some stochastic environments. In this article, in order to reduce the underestimation bias, we propose an action candidate-based clipped double estimator (AC-CDE) for Double Q-learning. Specifically, we first select a set of elite action candidates with high action values from one set of estimators. Then, among these candidates, we choose the highest valued action from the other set of estimators. Finally, we use the maximum value in the second set of estimators to clip the action value of the chosen action in the first set of estimators and the clipped value is used for approximating the maximum expected action value. Theoretically, the underestimation bias in our clipped Double Q-learning decays monotonically as the number of action candidates decreases. Moreover, the number of action candidates controls the tradeoff between the overestimation and underestimation biases. In addition, we also extend our clipped Double Q-learning to continuous action tasks via approximating the elite continuous action candidates. We empirically verify that our algorithm can more accurately estimate the maximum expected action value on some toy environments and yield good performance on several benchmark problems. Code is available at https://github.com/Jiang-HB/ac_CDQ.

引用

页码：5269 / 5279

页数：11

共 50 条

[1] Action Candidate Based Clipped Double Q-learning for Discrete and Continuous Action Tasks
Jiang, Haobo
Xie, Jin
Yang, Jian
[J]. THIRTY-FIFTH AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE, THIRTY-THIRD CONFERENCE ON INNOVATIVE APPLICATIONS OF ARTIFICIAL INTELLIGENCE AND THE ELEVENTH SYMPOSIUM ON EDUCATIONAL ADVANCES IN ARTIFICIAL INTELLIGENCE, 2021, 35 : 7979 - 7986
[2] Continuous-action Q-learning
Millán, JDR
Posenato, D
Dedieu, E
[J]. MACHINE LEARNING, 2002, 49 (2-3) : 247 - 265
[3] Continuous-Action Q-Learning
José del R. Millán
Daniele Posenato
Eric Dedieu
[J]. Machine Learning, 2002, 49 : 247 - 265
[4] Q-learning in continuous state and action spaces
Gaskett, C
Wettergreen, D
Zelinsky, A
[J]. ADVANCED TOPICS IN ARTIFICIAL INTELLIGENCE, 1999, 1747 : 417 - 428
[5] Fuzzy Q-learning in continuous state and action space
Xu, Ming-Liang
Xu, Wen-Bo
[J]. Journal of China Universities of Posts and Telecommunications, 2010, 17 (04): : 100 - 109
[6] Fuzzy Q-learning in continuous state and action space
XU Ming-liang1
[J]. The Journal of China Universities of Posts and Telecommunications, 2010, 17 (04) : 100 - 109
[7] CONTINUOUS ACTION GENERATION OF Q-LEARNING IN MULTI-AGENT COOPERATION
Hwang, Kao-Shing
Chen, Yu-Jen
Jiang, Wei-Cheng
Lin, Tzung-Feng
[J]. ASIAN JOURNAL OF CONTROL, 2013, 15 (04) : 1011 - 1020
[8] Double action Q-learning for obstacle avoidance in a dynamically changing environment
Ngai, DCK
Yung, NHC
[J]. 2005 IEEE Intelligent Vehicles Symposium Proceedings, 2005, : 211 - 216
[9] Q-Learning with probability based action policy
Ugurlu, Ekin Su
Biricik, Goksel
[J]. 2006 IEEE 14TH SIGNAL PROCESSING AND COMMUNICATIONS APPLICATIONS, VOLS 1 AND 2, 2006, : 210 - +
[10] Performance evaluation of Double Action Q-learning in moving obstacle avoidance problem
Ngai, DCK
Yung, NHC
[J]. INTERNATIONAL CONFERENCE ON SYSTEMS, MAN AND CYBERNETICS, VOL 1-4, PROCEEDINGS, 2005, : 865 - 870

← 1 2 3 4 5 →