Are more data always better? - Machine learning forecasting of algae based on long-term observations

被引:0
|
作者
Beckmann, D. Atton [1 ]
Werther, M. [2 ]
Mackay, E. B. [3 ]
Spyrakos, E. [1 ]
Hunter, P. [1 ,4 ]
Jones, I. D. [1 ]
机构
[1] Univ Stirling, Sch Nat Sci, Biol & Environm Sci, Stirling, Scotland
[2] Swiss Fed Inst Aquat Sci & Technol, Dept Surface Waters Res & Management, Dubendorf, Switzerland
[3] UK Ctr Ecol & Hydrol, Lancaster Environm Ctr, Lancaster LA1 4AP, England
[4] Univ Stirling, Sch Nat Sci, Scotlands Int Environm Ctr, Stirling, Scotland
基金
英国自然环境研究理事会;
关键词
Algal blooms; Cyanobacteria; Forecasting; Freshwater; Early warning; Machine learning; ARTIFICIAL NEURAL-NETWORK; CLIMATE-CHANGE; CYANOBACTERIAL BLOOMS; WATER-QUALITY; FRESH-WATER; ENVIRONMENTAL-FACTORS; GLOBAL EXPANSION; CHLOROPHYLL-A; LAKE; PREDICTION;
D O I
10.1016/j.jenvman.2024.123478
中图分类号
X [环境科学、安全科学];
学科分类号
08 ; 0830 ;
摘要
Bloom-forming algae present a unique challenge to water managers as they can significantly impair provision of important ecosystem services and cause health risks to humans and animals. Consequently, effective short-term algae forecasts are important as they provide early warnings and enable implementation of mitigation strategies. In this context, machine learning (ML) emerges as a promising forecasting tool. However, the performance of ML models is heavily dependent on the availability of appropriate training data. Consequently, it is essential to determine the volume of data necessary to develop reliable ML forecasts. Understanding this will guide future monitoring strategies, optimize resource allocation, and set realistic expectations for management outcomes. In this study, we used 30 years of fortnightly measurements of 13 different parameters from a lake in the English Lake District (UK) to examine the impact of training data duration on the performance of ML models for forecasting chlorophyll-a two weeks in advance. Once training data availability exceeded four years, a Random Forest model was found to consistently outperform naive benchmarks (mean absolute percentage error 16.4 % lower than the best-performing benchmark). With more than 5 years of training data, model performance generally continued to improve, but with diminishing returns. Furthermore, it was found that equivalent and, in some cases, better performance could be achieved by only using a subset of the most important input features. Additionally, it was found that reducing the sampling frequency had negative impacts on performance, both due to the reduced number of training observations available, and increased forecast horizon. Our findings demonstrate that for lakes ecologically similar to the study site, a consistent and regular sampling programme focused on monitoring a limited number of key parameters can provide sufficient observations for generating short-term algae forecasts after approximately five years of data collection. Importantly, this result provides justification for the initiation of new monitoring programmes for sites where algal blooms are a concern, and suggests that there are likely many pre-existing monitoring datasets which would be suitable for training algae forecast models.
引用
收藏
页数:13
相关论文
共 50 条
  • [21] Long Term Forecasting using Machine Learning Methods
    Sangrody, Hossein
    Zhou, Ning
    Tutun, Salih
    Khorramdel, Benyamin
    Motalleb, Mahdi
    Sarailoo, Morteza
    2018 IEEE POWER AND ENERGY CONFERENCE AT ILLINOIS (PECI), 2018,
  • [22] Review on the Application of Photovoltaic Forecasting Using Machine Learning for Very Short- to Long-Term Forecasting
    Radzi, Putri Nor Liyana Mohamad
    Akhter, Muhammad Naveed
    Mekhilef, Saad
    Shah, Noraisyah Mohamed
    SUSTAINABILITY, 2023, 15 (04)
  • [23] Machine Learning Based Univariate Models For Long Term Wind Speed Forecasting
    Akash, R.
    Rangaraj, A. G.
    Meenal, R.
    Lydia, M.
    PROCEEDINGS OF THE 5TH INTERNATIONAL CONFERENCE ON INVENTIVE COMPUTATION TECHNOLOGIES (ICICT-2020), 2020, : 779 - 784
  • [24] Long-term natural gas peak demand forecasting in Tunisia Using machine learning
    Ben Brahim, Sami
    Slimane, Mohamed
    PROCEEDINGS OF THE 2022 5TH INTERNATIONAL CONFERENCE ON ADVANCED SYSTEMS AND EMERGENT TECHNOLOGIES IC_ASET'2022), 2022, : 222 - 227
  • [25] A Machine Learning Model for Long-Term Power Generation Forecasting at Bidding Zone Level
    Moschella, Michela
    Tucci, Mauro
    Crisostomi, Emanuele
    Betti, Alessandro
    PROCEEDINGS OF 2019 IEEE PES INNOVATIVE SMART GRID TECHNOLOGIES EUROPE (ISGT-EUROPE), 2019,
  • [26] An Intelligent SARIMAX-Based Machine Learning Framework for Long-Term Solar Irradiance Forecasting at Muscat, Oman
    Baloch, Mazhar
    Honnurvali, Mohamed Shaik
    Kabbani, Adnan
    Jumani, Touqeer Ahmed
    Chauhdary, Sohaib Tahir
    ENERGIES, 2024, 17 (23)
  • [27] Deterioration forecasting of joint members based on long-term monitoring data
    Kobayashi, Kiyoshi
    Kaito, Kiyoyuki
    Kazumi, Kosuke
    EURO JOURNAL ON TRANSPORTATION AND LOGISTICS, 2015, 4 (01) : 5 - 30
  • [28] Machine learning modeling structures and framework for short-term forecasting and long-term projection of Streamflow
    Trung Duc Tran
    Jongho Kim
    Stochastic Environmental Research and Risk Assessment, 2024, 38 : 793 - 813
  • [29] Machine learning modeling structures and framework for short-term forecasting and long-term projection of Streamflow
    Tran, Trung Duc
    Kim, Jongho
    STOCHASTIC ENVIRONMENTAL RESEARCH AND RISK ASSESSMENT, 2024, 38 (02) : 793 - 813
  • [30] Is deeper always better? Evaluating deep learning models for yield forecasting with small data
    Filip Sabo
    Michele Meroni
    François Waldner
    Felix Rembold
    Environmental Monitoring and Assessment, 2023, 195