Evaluating the extrapolation potential of random forest digital soil mapping

被引:2
|
作者
Hateffard, Fatemeh [1 ]
Steinbuch, Luc [2 ]
Heuvelink, Gerard B. M. [2 ,3 ]
机构
[1] Univ Debrecen, Dept Landscape Protect & Environm Geog, Egyet Ter 1, H-4032 Debrecen, Hungary
[2] Wageningen Univ & Res, Soil Geog & Landscape Grp, Wageningen, Netherlands
[3] ISRIC World Soil Informat, Wageningen, Netherlands
关键词
Spatial soil information; Extrapolation effects; Prediction accuracy; Similarities; REGRESSION; UNCERTAINTY; INFORMATION; PREDICTION; SUPPORT;
D O I
10.1016/j.geoderma.2023.116740
中图分类号
S15 [土壤学];
学科分类号
0903 ; 090301 ;
摘要
Spatial soil information is essential for informed decision-making in a wide range of fields. Digital soil mapping (DSM) using machine learning algorithms has become a popular approach for generating soil maps. DSM capitalises on the relation between environmental variables (i.e., features) and a soil property of interest. It typically needs a training dataset that covers the feature space well. Mapping in areas where there are no training data is challenging, because extrapolation in geographic space often induces extrapolation in feature space and can seriously deteriorate prediction accuracy. The objective of this study was to analyse the extrapolation effects of random forest DSM models by predicting topsoil properties (OC, clay, and pH) in four African countries using soil data from the ISRIC Africa Soil Profiles database. The study was conducted in eight experiments whereby soil data from one or three countries were used to predict in the other countries. We calculated similarities between donor and recipient areas using four measures, including soil type similarity, homosoil, dissimilarity index by area of applicability (AOA), and quantile regression forest (QRF) prediction interval width. The aim was to determine the level of agreement between these four measures and identify the method that had the strongest agreement with common validation metrics. The results indicated a positive correlation between soil type similarity, homosoil and dissimilarity index by AOA. Surprisingly, we observed a negative correlation between dissimilarity index by AOA and QRF prediction interval width. Although the cross-validation results for the trained models were acceptable, the extrapolation results were unsatisfactory, highlighting the risk of extrapolation. Using soil data from three countries instead of one increased the similarities for all measures, but it had a limited effect on improving extrapolation. Also, none of the measures had a strong correlation with the validation metrics. This was particularly disappointing for AOA and QRF, which we had expected to be strong indicators of extrapolation prediction performance. Results showed that homosoil and soil type methods had the strongest correlation with validation metrics. The results for this case study revealed limitations of using AOA and QRF as measures of extrapolation effects, highlighting the importance of not relying on these methods blindly. Further research and more case studies are needed to address the effects of extrapolation of DSM models.
引用
收藏
页数:12
相关论文
共 50 条
  • [1] Multivariate random forest for digital soil mapping
    van der Westhuizen, Stephan
    Heuvelink, Gerard B. M.
    Hofmeyr, David P.
    GEODERMA, 2023, 431
  • [2] Comparing and evaluating digital soil mapping methods in a Hungarian forest reserve
    Illes, Gabor
    Kovacs, Gabor
    Heil, Balint
    CANADIAN JOURNAL OF SOIL SCIENCE, 2011, 91 (04) : 615 - 626
  • [3] Digital mapping of soil texture classes using Random Forest classification algorithm
    Dharumarajan, Subramanian
    Hegde, Rajendra
    SOIL USE AND MANAGEMENT, 2022, 38 (01) : 135 - 149
  • [4] Extrapolation of a structural equation model for digital soil mapping
    Angelini, M. E.
    Kempen, B.
    Heuvelink, G. B. M.
    Temme, A. J. A. M.
    Ransom, M. D.
    GEODERMA, 2020, 367
  • [5] Provincial-scale digital soil mapping using a random forest approach for British Columbia
    Heung, Brandon
    Bulmer, Chuck E.
    Schmidt, Margaret G.
    Zhang, Jin
    CANADIAN JOURNAL OF SOIL SCIENCE, 2022,
  • [6] Digital mapping of soil quality index to evaluate orchard fields using random forest models
    Barikloo, Ali
    Alamdari, Parisa
    Rezapour, Salar
    Taghizadeh-Mehrjardi, Ruhollah
    MODELING EARTH SYSTEMS AND ENVIRONMENT, 2024, : 6787 - 6803
  • [7] Multinomial Logistic Regression and Random Forest Classifiers in Digital Mapping of Soil Classes in Western Haiti
    Jeune, Wesly
    Francelino, Marcio Rocha
    de Souza, Eliana
    Fernandes Filho, Elpidio Inacio
    Rocha, Genelicio Crusoe
    REVISTA BRASILEIRA DE CIENCIA DO SOLO, 2018, 42
  • [8] Provincial-scale digital soil mapping using a random forest approach for British Columbia
    Heung, Brandon
    Bulmer, Chuck E.
    Schmidt, Margaret G.
    Zhang, Jin
    CANADIAN JOURNAL OF SOIL SCIENCE, 2022, 102 (03) : 597 - 620
  • [9] Digital mapping of soil classes using spatial extrapolation with imbalanced data
    Neyestani, Mehrnaz
    Sarmadian, Fereydoon
    Jafari, Azam
    Keshavarzi, Ali
    Sharififar, Amin
    GEODERMA REGIONAL, 2021, 26
  • [10] Sampling design optimization for soil mapping with random forest
    Wadoux, Alexandre M. J-C.
    Brus, Dick J.
    Heuvelink, Gerard B. M.
    GEODERMA, 2019, 355