A

Can machine learning algorithms improve upon classical palaeoenvironmental reconstruction models?

Abstract

Abstract. Classical palaeoenvironmental reconstruction models often incorporate biological ideas and commonly assume that the taxa comprising a fossil assemblage exhibit unimodal response functions of the environmental variable of interest. In contrast, machine learning approaches do not rely upon any biological assumptions, but instead need training with large data-sets to extract some understanding of the relationships between biological assemblages and their environment. We have developed a two-layered machine learning reconstruction model MEMLM (Multi Ensemble Machine Learning Model). The first layer applies three different ensemble machine learning models of random forests, extra random trees and lightGBM, trained on the modern taxon assemblage and associated environmental data to make reconstructions based on the three different models, while the second layer uses multiple linear regression to integrate these three reconstructions into a consensus reconstruction. We consider three versions of the model: 1) A standard version of MEMLM, which uses only taxon abundance data, 2) MEMLMe, which uses embedded assemblage information, using a natural language processing model (GLOVE) to detect associations between taxa across the training data-set and 3) MEMLMc which incorporates both taxon abundance and assemblage data. We train these MEMLM model variants with three high quality diatom and pollen training sets and compare their reconstruction performance with three weighted averaging (WA) approaches of WA-Cla (classical deshrinking), WA-Inv (inverse deshrinking) and WA-PLS (partial least squares). In general, the MEMLM approaches, even when trained on only embedded assemblage data, perform substantially better than the WA approaches under cross-validation in the larger data-sets. However, when applied to fossil data, MEMLM and WA approaches sometimes generate qualitatively different palaeoenvironmental reconstructions. We applied a statistical significance test to all the reconstructions. This successfully identified each incidence where the reconstruction is not robust with respect to the model choice. We find that machine learning approaches can outperform classical approaches, but can sometimes catastrophically fail, despite showing high performance under cross-validation, likely indicating problems when extrapolation occurs. We find that the classical approaches are generally more robust, although they can also generate reconstructions which have modest statistical significance, and therefore may be unreliable. We conclude that cross-validation is not a sufficient measure of transfer-function performance, and we recommend that the results of statistical significance tests are provided alongside the down-core reconstructions based on fossil assemblages.

Authors 3

  1. Leiden University

    Affiliation as printed

    Institute of Environmental Sciences (CML), Leiden University, 2333 CC Leiden, the Netherlands

  2. Philip B. Holden corresponding

    The Open University

    Affiliation as printed

    Environment, Earth and Ecosystem Sciences, The Open University, Walton Hall, Milton Keynes, MK7 6AA, UK

  3. H. J. B. Birks corresponding

    University College London · University of Bergen · Bjerknes Centre for Climate Research

    Affiliation as printed

    Department of Biological Sciences and Bjerknes Centre for Climate Research, University of Bergen, P.O. Box 7803, Bergen 5020, Norway

    Environmental Change Research Centre, University College London, London WC1 6BT, UK

Cited by 1 stored of 2

1 result

No patents citing this paper on Lens.org (checked 2026-10-11).

References 52