Self-scaled conjugate gradient training algorithms

被引：19

作者：

Kostopoulos, A. E. ^{[1
]}

Grapsa, T. N. ^{[1
]}

机构：

[1] Univ Patras, Dept Math, GR-26504 Patras, Greece

来源：

NEUROCOMPUTING | 2009年 / 72卷 / 13-15期

关键词：

Neural network; Training; Self-scaled conjugate gradient; Perry's method; Line search; LEARNING ALGORITHMS; RESTART PROCEDURES; CONVERGENCE;

D O I：

10.1016/j.neucom.2009.04.006

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

This article presents some efficient training algorithms, based on conjugate gradient optimization methods. In addition to the existing conjugate gradient training algorithms, we introduce Perry's conjugate gradient method as a training algorithm [A. Perry, A modified conjugate gradient algorithm, Operations Research 26 (1978) 26-43]. Perry's method has been proven to be a very efficient method in the context of unconstrained optimization, but it has never been used in MLP training. Furthermore, a new class of conjugate gradient (CG) methods is proposed, called self-scaled CG methods, which are derived from the principles of Hestenes-Stiefel, Fletcher-Reeves, Polak-Ribiere and Perry's method. This class is based on the spectral scaling parameter introduced in [J. Barzilai, J.M. Borwein, Two point step size gradient methods, IMA journal of Numerical Analysis 8 (1988) 141-148]. The spectral scaling parameter contains second order information without estimating the Hessian matrix. Furthermore, we incorporate to the CG training algorithms an efficient line search technique based on the Wolfe conditions and on safeguarded cubic interpolation [D.F. Shanno, K.H. Phua, Minimization of unconstrained multivariate functions, ACM Transactions on Mathematical Software 2 (1976) 87-94]. In addition, the initial learning rate parameter, fed to the line search technique, was automatically adapted at each iteration by a closed formula proposed in [D.F. Shanno, K.H. Phua, Minimization of unconstrained multivariate functions, ACM Transactions on Mathematical Software 2 (1976) 87-94; D.G. Sotiropoulos, A.E. Kostopoulos, T.N. Grapsa, A spectral version of Perry's conjugate gradient method for neural network training, in: D.T. Tsahalis (Ed.), Fourth GRACM Congress on Computational Mechanics, vol. 1, 2002, pp. 172-179]. Finally, an efficient restarting procedure was employed in order to further improve the effectiveness of the CG training algorithms. Experimental results show that, in general, the new class of methods can perform better with a much lower computational cost and better success performance. (C) 2009 Elsevier B.V. All rights reserved.

引用

页码：3000 / 3019

页数：20

共 50 条

[21] Two modified scaled nonlinear conjugate gradient methods
Babaie-Kafaki, Saman
[J]. JOURNAL OF COMPUTATIONAL AND APPLIED MATHEMATICS, 2014, 261 : 172 - 182
[22] Temporal differences learning with the scaled conjugate gradient algorithm
Falas, T
Stafyopatis, A
[J]. ICONIP'02: PROCEEDINGS OF THE 9TH INTERNATIONAL CONFERENCE ON NEURAL INFORMATION PROCESSING: COMPUTATIONAL INTELLIGENCE FOR THE E-AGE, 2002, : 2625 - 2629
[23] A Scaled Conjugate Gradient Backpropagation Algorithm for Keyword Extraction
Aich, Ankit
Dutta, Amit
Chakraborty, Aruna
[J]. INFORMATION SYSTEMS DESIGN AND INTELLIGENT APPLICATIONS, INDIA 2017, 2018, 672 : 674 - 684
[24] A scaled conjugate gradient method for nonlinear unconstrained optimization
Fatemi, Masoud
[J]. OPTIMIZATION METHODS & SOFTWARE, 2017, 32 (05): : 1095 - 1112
[25] A scaled nonlinear conjugate gradient algorithm for unconstrained optimization
Andrei, Neculai
[J]. OPTIMIZATION, 2008, 57 (04) : 549 - 570
[26] Speeding up the scaled conjugate gradient algorithm and its application in neuro-fuzzy classifier training
Cetisli, Bayram
Barkana, Atalay
[J]. SOFT COMPUTING, 2010, 14 (04) : 365 - 378
[27] Speeding up the scaled conjugate gradient algorithm and its application in neuro-fuzzy classifier training
Bayram Cetişli
Atalay Barkana
[J]. Soft Computing, 2010, 14 : 365 - 378
[28] ON THE CONVERGENCE OF CONJUGATE-GRADIENT ALGORITHMS
PYTLAK, R
[J]. IMA JOURNAL OF NUMERICAL ANALYSIS, 1994, 14 (03) : 443 - 460
[29] Oscillations make a self-scaled model for honeybees' visual odometer reliable regardless of flight trajectory
Bergantin, Lucia
Harbaoui, Nesrine
Raharijaona, Thibaut
Ruffier, Franck
[J]. JOURNAL OF THE ROYAL SOCIETY INTERFACE, 2021, 18 (182)
[30] Self-scaled bounds for atomic cone ranks: applications to nonnegative rank and cp-rank
Fawzi, Hamza
Parrilo, Pablo A.
[J]. MATHEMATICAL PROGRAMMING, 2016, 158 (1-2) : 417 - 465

← 1 2 3 4 5 →