Machine Learning approaches for the detection of disease-causing variants in whole-genome data need to address the expression of functional genes

Scritto il 28/08/2026
da Camilla Mapstone

PLoS One. 2026 Aug 28;21(8):e0355557. doi: 10.1371/journal.pone.0355557. eCollection 2026.

ABSTRACT

Gene-dosage combinations have been recognised as leading factors of disease. Given that those combinations may include dozens of genes, it is hypothesised that machine learning (ML) approaches may be useful in the classification of cases and controls and the identification of causative genes. We aimed to assess the validity of this hypothesis. Here, we have constructed a benchmark that includes real data (with ground truth knowledge) and synthetic data with known generating mechanisms and various dataset sizes and levels of noise. We trained standard statistical learning/ ML models on these datasets to classify disease phenotype. We present an analysis of how model performance varies across different synthetic genetic scenarios, and how it is impacted by dataset size. The logistic regression model was found to be the most reliable at causative gene identification across the synthetic datasets, despite not always performing the best in terms of classification performance and, in some cases, having a relatively low ROC AUC score. When our training attempts on the UK Biobank datasets failed, we performed an analysis into model performance vs dataset richness. Our results show that it is necessary to take into account the expression of functional genes in order to successfully predict disease.

PMID:42664201 | DOI:10.1371/journal.pone.0355557