Ontology pre-training improves machine learning-based predictions for metabolites
bioRxiv (Cold Spring Harbor Laboratory)
Abstract
Abstract Recent advances in the field of machine learning have shown that integration of expert knowledge improves performances, in particular for complex domains such as biology. Bio-ontologies offer a rich source of curated biological knowledge that can be harnessed to this end. Here, we describe an intuitive and generalisable approach to embed the knowledge contained in a classification hierarchy derived from a bio-ontology into a machine learning model as an intermediate training step between general-purpose pre-training and task-specific fine-tuning in a process that we call ‘ontology pre-training’. We show that this approach leads to an improvement in predictive performance and a reduction in training time for a broad range of prediction tasks relevant to understanding metabolite functions in living systems, using a range of datasets derived from MoleculeNet. We see the biggest improvement for regression tasks, e.g. prediction of lipophilicity and aqueous solubility of molecules, and a robust improvement for most classification tasks. Our approach can be adapted for a wide range of knowledge sources, models and prediction tasks.
Authors 7
-
Affiliation as printed
University of Zurich;
-
Otto-von-Guericke-Universität Magdeburg
Affiliation as printed
Otto-von-Guericke University Magdeburg;
-
Affiliation as printed
Institute of Computer Science , Osnabrück University , Neuer Graben/ Schloss , 49074 Osnabrück , Germany,
-
Martin F. Larralde Aachen
Leiden University Medical Center
Affiliation as printed
Leiden University Medical Center
-
Otto-von-Guericke-Universität Magdeburg
Affiliation as printed
Otto-von-Guericke University Magdeburg;
-
Affiliation as printed
Osnabrueck University;
-
Janna Hastings corresponding
Affiliation as printed
University of Zurich;
Cited by 0 stored of 0
No patents citing this paper on Lens.org (checked 2026-10-11).