Explainable housing price prediction with determinant analysis

被引:7
|
作者
Teoh, Ean Zou [1 ]
Yau, Wei-Chuen [1 ]
Ong, Thian Song [2 ]
Connie, Tee [2 ]
机构
[1] Xiamen Univ Malaysia, Sch Comp & Data Sci, Sepang, Malaysia
[2] Multimedia Univ, Fac Informat Sci & Technol, Melaka, Malaysia
关键词
House price prediction; Regression model; SHAP analysis; Multinomial logistic regression; Determinant analysis; Housing market analysis;
D O I
10.1108/IJHMA-02-2022-0025
中图分类号
TU98 [区域规划、城乡规划];
学科分类号
0814 ; 082803 ; 0833 ;
摘要
Purpose This study aims to develop a regression-based machine learning model to predict housing price, determine and interpret factors that contribute to housing prices using different data sets available publicly. The significant determinants that affect housing prices will be first identified by using multinomial logistics regression (MLR) based on the level of relative importance. A comprehensive study is then conducted by using SHapley Additive exPlanations (SHAP) analysis to examine the features that cause the major changes in housing prices. Design/methodology/approach Predictive analytics is an effective way to deal with uncertainties in process modelling and improve decision-making for housing price prediction. The focus of this paper is two-fold; the authors first apply regression analysis to investigate how well the housing independent variables contribute to the housing price prediction. Two data sets are used for this study, namely, Ames Housing dataset and Melbourne Housing dataset. For both the data sets, random forest regression performs the best by achieving an average R-2 of 86% for the Ames dataset and 85% for the Melbourne dataset, respectively. Second, multinomial logistic regression is adopted to investigate and identify the factor determinants of housing sales price. For the Ames dataset, the authors find that the top three most significant factor variables to determine the housing price is the general living area, basement size and age of remodelling. As for the Melbourne dataset, properties having more rooms/bathrooms, larger land size and closer distance to central business district (CBD) are higher priced. This is followed by a comprehensive analysis on how these determinants contribute to the predictability of the selected regression model by using explainable SHAP values. These prominent factors can be used to determine the optimal price range of a property which are useful for decision-making for both buyers and sellers. Findings By using the combination of MLR and SHAP analysis, it is noticeable that general living area, basement size and age of remodelling are the top three most important variables in determining the house's price in the Ames dataset, while properties with more rooms/bathrooms, larger land area and closer proximity to the CBD or to the South of Melbourne are more expensive in the Melbourne dataset. These important factors can be used to estimate the best price range for a housing property for better decision-making. Research limitations/implications A limitation of this study is that the distribution of the housing prices is highly skewed. Although it is normal that the properties' price is normally cluttered at the lower side and only a few houses are highly price. As mentioned before, MLR can effectively help in evaluating the likelihood ratio of each variable towards these categories. However, housing price is originally continuous, and there is a need to convert the price to categorical type. Nonetheless, the most effective method to categorize the data is still questionable. Originality/value The key point of this paper is the use of explainable machine learning approach to identify the prominent factors of housing price determination, which could be used to determine the optimal price range of a property which are useful for decision-making for both the buyers and sellers.
引用
收藏
页码:1021 / 1045
页数:25
相关论文
共 50 条
  • [1] Housing Price Prediction by Divided Regression Analysis
    Goh, Yann Ling
    Goh, Yeh Huann
    Yip, Chun-Chieh
    Ng, Kooi Huat
    CHIANG MAI JOURNAL OF SCIENCE, 2022, 49 (06): : 1669 - 1682
  • [2] Explainable, Multi-Region Price Prediction
    Ghatnekar, Atharva
    Shanbhag, Aakash Dhananjay
    INTERNATIONAL CONFERENCE ON ELECTRICAL, COMPUTER AND ENERGY TECHNOLOGIES (ICECET 2021), 2021, : 930 - 936
  • [3] Housing Price Prediction Based on CNN
    Piao, Yong
    Chen, Ansheng
    Shang, Zhendong
    2019 9TH INTERNATIONAL CONFERENCE ON INFORMATION SCIENCE AND TECHNOLOGY (ICIST2019), 2019, : 491 - 495
  • [4] Analysis of the housing price
    Yanbing, Liang
    Yongsheng, Ma
    Shuo, Zhao
    Journal of Chemical and Pharmaceutical Research, 2014, 6 (07) : 1168 - 1172
  • [5] Housing Price Prediction Using Neural Networks
    Lim, Wan Teng
    Wang, Lipo
    Wang, Yaoli
    Chang, Qing
    2016 12TH INTERNATIONAL CONFERENCE ON NATURAL COMPUTATION, FUZZY SYSTEMS AND KNOWLEDGE DISCOVERY (ICNC-FSKD), 2016, : 518 - 522
  • [6] Understanding the effects of socioeconomic factors on housing price appreciation using explainable AI
    Jin, Shengxiang
    Zheng, Huixin
    Marantz, Nicholas
    Roy, Avipsa
    APPLIED GEOGRAPHY, 2024, 169
  • [7] Several determinant factors of the secondhand housing price: an application of the hedonic methodology
    Garcia Pozo, Alejandro
    REVISTA DE ESTUDIOS REGIONALES, 2008, (82) : 135 - 158
  • [8] Capturing the distance decay effect of amenities on housing price using explainable artificial intelligence
    Lee, Hojun
    Han, Hoon
    Pettit, Chris
    APPLIED GEOGRAPHY, 2025, 174
  • [9] An optimized and interpretable carbon price prediction: Explainable deep learning model
    Sayed, Gehad Ismail
    El-Latif, Eman I. Abd
    Darwish, Ashraf
    Snasel, Vaclav
    Hassanien, Aboul Ella
    CHAOS SOLITONS & FRACTALS, 2024, 188
  • [10] Housing Price Prediction Based on Multiple Linear Regression
    Zhang, Qingqi
    SCIENTIFIC PROGRAMMING, 2021, 2021