Close Menu
Health  JustFineHealth  JustFine

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Suns’ Dillon Brooks says youth camp focuses on mental wellness, hoops

    September 13, 2026

    Red Bank Catholic High School and Conscientia Health Open New Wellness Room, Marking Expanded Commitment to Student Mental and Physical Health

    September 13, 2026

    Mental wellness: NIMHANS to be nodal body for BRICS Centres of Excellence

    September 13, 2026
    Facebook X (Twitter) Instagram
    Health  JustFineHealth  JustFine
    Facebook X (Twitter) Instagram
    • Home
    • General Health News
    • Sleep Health
    • Mental Wellness
    • Fitness & Recovery
    • Health Tech & Wearables
    • More
      • Longevity & Anti-Aging
      • Women’s Hormone Health
      • Gut Health & Microbiome
      • Metabolic Health & Blood Sugar
      • Nutrition & Anti-Inflammatory Foods
    Health  JustFineHealth  JustFine
    Home»Gut Health & Microbiome»Frontiers | A data-driven universal gut microbiome health assessment: a machine learning framework trained on large metagenomic data
    Gut Health & Microbiome

    Frontiers | A data-driven universal gut microbiome health assessment: a machine learning framework trained on large metagenomic data

    HealthJustfine TeamBy HealthJustfine TeamSeptember 13, 2026No Comments80 Mins Read
    Facebook Twitter Pinterest LinkedIn Tumblr Reddit WhatsApp Email
    Frontiers | A data-driven universal gut microbiome health assessment: a machine learning framework trained on large metagenomic data
    Share
    Facebook Twitter LinkedIn Pinterest WhatsApp Email

    A data-driven universal gut microbiome health assessment: a machine learning framework trained on large metagenomic data

    • 1. Department of Oncology and Hematology-Oncology, Università degli Studi di Milano, Milan, Italy

    • 2. Department of Biosciences, Biotechnology and Environment, University of Bari A. Moro, Bari, Italy

    • 3. National Research Council, Institute of Biomembranes, Bioenergetics and Molecular Biotechnologies, Bari, Italy

    Article metrics

    View details

    Abstract

    The gut microbiota is essential to maintain host physiology, and its disruption (dysbiosis) is associated with a wide range of diseases. Machine learning (ML) offers a powerful tool to model species-level microbiome profiles, but classifiers that reliably separate healthy from diseased individuals across independent cohorts are still lacking. In this study, we developed a ML classifiers trained on 7,452 publicly available stool metagenomes spanning 32 studies and 12 diseases, designed to distinguish healthy individuals (absence of a clinically diagnosed disease) from non-healthy individuals (presence of a clinically diagnosed disease) based on species-level gut microbiome profiles. We trained 16 supervised models combining four algorithms (RF, SVM-LIN, SVM-RBF, and LR-ElasticNet) combined with all feature sets and three feature-selection algorithms. Performance was assessed by F1 score and ROC–AUC on held-out test data and externally validated on 642 samples from six independent cohorts, including previously unseen diseases. On the test set, all models achieved F1 scores of 78–86% and ROC–AUC values of 89–95%. An SVM-RBF model using permutation-based feature selection performed best (F1 = 86.6%, ROC–AUC = 95.5%; healthy F1 = 86.6%, non-healthy F1 = 88.7%). Importantly, external validation confirmed the generalizability of the full-feature SVM-RBF model (overall F1 = 70.6%; ROC–AUC = 84.7%), including unseen disease types such as Clostridioides difficile infection (F1 = 90.3%) and type 2 diabetes (F1 = 77.4%). Feature-importance and multivariate analyses revealed both shared and disease-specific microbial signatures, suggesting that the model captures biologically meaningful patterns rather than cohort-specific artifacts. Disease-associated taxa included Klebsiella pneumoniae, Raoultella ornithinolytica, Sutterella wadsworthensis, Gemmiger formicilis, and Lactobacillus crispatus. In contrast, healthy status was consistently associated with commensal species such as Extibacter hylemonae and Ruthenibacterium lactatiformans. These results show that models trained on pooled metagenomes predict gut health status accurately and transfer to independent cohorts, providing a scalable, non-invasive framework and a set of candidate microbial biomarkers for further clinical evaluation.

    1 Introduction

    The gut microbiota is integral to host physiology, contributing to maintain intestinal homeostasis and host wellbeing (Cho and Blaser, 2012; Fan and Pedersen, 2021; He et al., 2026). The dysbiosis is an alteration of the resident community ecology, and has been associated with several pathological conditions (Fan and Pedersen, 2021; Hou et al., 2022; Van Hul et al., 2024; Joos et al., 2025). Consistent gut microbiota changes across different cohorts support supervised classifiers for identifying disease-specific microbial signatures (Pasolli et al., 2016; Giliberti et al., 2022; Li et al., 2023; Jin et al., 2024; Kumar et al., 2024; Li et al., 2025; Porcari et al., 2025). However, despite a large application of classification models in microbiome field, relying on both machine learning (ML) or Artificial Intelligence (AI) approaches, there is still a gap in training a model recapping a common dysbiosis signature that cuts across multiple conditions.

    Shotgun metagenomic allows to sequence total community DNA, enabling species- and strain-level profiling alongside functional characterization, including for uncultured taxa (Quince et al., 2017; Bars-Cortina et al., 2024; Kumar et al., 2024; Visci et al., 2026). This resolution yields the high-dimensional, species-level abundance data required to train models that predict host health status from the whole-community composition rather than a small panel of preselected taxa (Asnicar et al., 2023).

    The mechanisms by which microbial communities influence gut homeostasis remain largely uncharacterized, with few experimentally demonstrated causal pathways (Lloyd-Price et al., 2016; de Vos et al., 2022). Indeed, despite consistent microbial patterns recur across different populations, discerning correlation from causation still remain challenging (Wade and Hall, 2019; Basic et al., 2022; Corander et al., 2022; Fonseca et al., 2024).

    Several technical factors make metagenomics difficult for conventional statistical methods. Species typically outnumber samples, abundance matrices are zeros inflated, and relative abundances sum to a constant (i.e., data compositionality) (Gloor et al., 2017; Chan and Li, 2024). Further, cohort-specific differences in sampling, DNA extraction, library preparation, and sequencing depth introduce technical variation that mask biological signals when cases and controls are unevenly distributed across studies (Costea et al., 2017; Wirbel et al., 2021). These properties inflate spurious associations (Gloor et al., 2017; Chan and Li, 2024), violate the independence assumptions underlying classical models, and limit the capacity of standard statistical frameworks to capture the nonlinear, higher-order interactions that characterize microbial ecosystems (Sczyrba et al., 2017; Walsh et al., 2023).

    In this scenario, the lack of a widely accepted definition of a “healthy” gut microbiome is, in itself, part of the problem, because community composition is shaped by host and environmental factors, that varies both between people and within the same person over time (Hou et al., 2022; Van Hul et al., 2024; Joos et al., 2025; Wang et al., 2026; Zeng et al., 2026). Microbial communities act as complex adaptive ecosystems in which ecological interactions among taxa generate emergent community-level properties (Chang et al., 2023). Functional redundancy is a key: distinct taxa share metabolic niches and co-occur, so very different taxonomic configurations can support similar functions, while similar profiles can yield different functional outcomes (Widder et al., 2016; Tian et al., 2020). Indeed, the metabolic pathways that defining a dysbiotic signature remain poorly characterized (Vieira-Silva et al., 2016; Tian et al., 2020; Zielińska et al., 2025) and the boundary between normal variability in a healthy microbiome and pathological deviation has never been defined (McBurney et al., 2019; Shanahan et al., 2021; Ghosh et al., 2022; Van Hul et al., 2024). Health-associated states are therefore better captured as distributed community structures than as single markers (Zhu et al., 2026), underlining substantial interpersonal variability (Falony et al., 2016). Finally, it remains to be tested whether it is possible to guess host’s health status (i.e., classification) from a single faecal metagenome in the absence of a diagnosis, and whether this inference is generalizable (Gupta et al., 2020; Chang et al., 2024). While functional profiling and longitudinal sampling might address the definition of a eubiotic/dysbiotic signature; the classification challenge can be addressed by leveraging and pooling existing data, such as we propose in the present study.

    Machine-learning (ML)-based classifiers are particularly suited to address classification tasks, because they integrate information across many taxa rather than depending on the reproducibility of any individual feature (Goodswen et al., 2021; Casimiro-Soriguer et al., 2022; Asnicar et al., 2023; Bekhet et al., 2025; Karwowska et al., 2025). ML methods can identify such latent community-level signals by modeling nonlinear relationships and higher-order interactions within high-dimensional, sparse, and compositional microbiome data without relying on assumptions of linearity or feature independence (Camacho et al., 2018; Ghannam and Techtmann, 2021; Zhou and Zhao, 2025; Yu et al., 2026).

    As larger and more diverse microbiome datasets become available, these approaches can increasingly distinguish robust biological signals from cohort-specific variation (Topçuoğlu et al., 2020; Lee and Lee, 2024; Romano et al., 2025; Pekel et al., 2026). Thus, the objective shifts from identifying individual taxa that characterize health or disease toward predicting host health status from distributed community-level microbial signatures (Gupta et al., 2020; Chang et al., 2024).

    Despite the prediction of healthy versus diseased individuals from gut microbiome profiles has been successfully established for different pathologies by using machine-learning models trained on thousands of publicly available microbiome datasets (Thomas et al., 2019; Kraszewski et al., 2021; Chanda and De, 2024; Wu et al., 2024; Zheng et al., 2024; Piccinno et al., 2025; Romano et al., 2025; Yan R. et al., 2025), the reported performance and discriminative taxa substantially differ between these studies. This reflects limited sample sizes, cohort-specific confounding, and methodological heterogeneity (Tierney et al., 2022; Teixeira et al., 2024), affecting models generalizability beyond the study cohort (Li et al., 2025; Barcan et al., 2026).

    Pooling cohorts across multiple diseases addresses challenges that cannot be solved even by large cohorts focused on a single disease. (i) It markedly increases the overall sample size, which boosts statistical power and supports more dependable machine (Gupta et al., 2020; Kumar et al., 2024); (ii) it incorporates both biological and technical variability from independent cohorts, enabling the model to detect stable consistent microbial signatures (Li et al., 2025) and (iii) it diminishes the impact of study-specific biases by pushing the model to learn reproducible biological signals (Duvallet et al., 2017); (iv) it provides a broader and more generalizable definition of the healthy population by including individuals from different countries, age groups, and populations (Gupta et al., 2020; Wilmanski et al., 2021; Chang et al., 2024; Huang et al., 2026); (v) it helps reduce diagnosis-related confounding factors, making it easier to distinguish disease-specific effects from changes genuinely associated with dysbiosis (Vujkovic-Cvijin et al., 2020). For instance, Pasolli et al. (2016) demonstrated that cross-study pooling improves phenotype prediction from taxonomic profiles, though on small cohorts. Pooled ML prediction has also been extended to non-intestinal conditions (Giliberti et al., 2022; Pietrucci et al., 2022; Bao et al., 2024; Jin et al., 2024; Lee and Lee, 2024) and recently extend to fungal mycobiome (Defazio et al., 2026). Collectively, these studies show that pooling across cohorts improves the robustness and generalizability of microbiome-based prediction, even if the focus on individual diseases, with only one exception (Defazio et al., 2026).

    Taken together, we hypothesize that a supervised ML model trained on species-level relative abundance profiles across multiple disease cohorts’ studies, can classify individuals as healthy (absence of diagnosed disease) or non-healthy (presence of diagnosed disease). Our objective is accurate classification and external validation, not causal inference; feature importance is examined only as a secondary, exploratory analysis to inform future mechanistic work. To test this, we trained a single supervised binary classifier on 7,452 stool shotgun metagenomes from 32 independent studies spanning 12 disease conditions, rather than optimizing for one disease in one cohort. Model selection and evaluation prioritize cross-cohort robustness and external validation over maximal within-dataset performance, directly addressing the generalization failures documented in single-cohort studies. The result is a scalable framework for microbiome-based health prediction with potential relevance to disease screening, independent of any single disease’s specific microbial etiology.

    2 Materials and methods

    2.1 Stool-based shotgun metagenomic data collection

    We searched PubMed and Google Scholar for the peer-reviewed papers published from inception to 31/08/2024 using the keywords “gut microbiome,” “shotgun sequencing,” “humans,” “stool/faecal,” and “WGS/metagenomes.” We retained studies that primarily focused on stool-based shotgun metagenomic sequencing and include metadata specific to healthy (absence of a diagnosed disease) and non-healthy (presence of a diagnosed disease) individuals, as per the original study. After that, to reduce potential unfairness and to ensure consistency in a large cohort data set, we applied inclusion and exclusion criteria. Inclusion Criteria were (i) human stool samples sequenced using shotgun metagenomic sequencing methods on an Illumina platform; (ii) each sample should include information on the health status (i.e., if the individual is healthy or diagnosed with a disease); (iii) at least 10 metagenomes per study. Exclusion criteria: (i) samples with lacking information on health status; (ii) data sets generated by non-Illumina platforms, such as Oxford Nanopore, 454 GS FLX Titanium, Ion Torrent PGM, and Ion Torrent Proton. The associated metadata were manually curated from either the Supplementary material or retrieved directly from the Sequence Read Archive (SRA) when not available. For each study, we identified the corresponding “BioProject ID” accession and retrieved all associated Run Accession IDs. These Run Accessions were then used to download the raw DNA sequencing reads (FASTQ files) from the NCBI SRA using the SRA Toolkit (fastq-dump).1

    2.2 Preprocessing of shotgun metagenomic sequences

    The quality of raw sequencing reads was assessed using FastQC (v0.12.1).2 Then, poor quality and adapter regions were trimmed using Fastp (v0.24.0) (Chen et al., 2018) with the following parameters: –detect_adapter_for_pe (to enable paired-end adapter detection), −q 20 (maintain quality score threshold), −u 20 (maximum unqualified base percentage), and -l 45 (minimum read length). After quality filtering, the remaining reads were referred to as ‘cleaned reads.’

    The cleaned reads were then aligned to the human reference genome (GRCh38)3 using Bowtie2 (v2.5.4) with default parameters to remove potential human derived DNA reads contaminants. The –un-con option was used to retain reads that failed to align to the human reference genome, while aligned reads (considered as potentially human DNA reads) were discarded. Only the resulting non-host (unaligned) reads were retained for downstream taxonomic classification.

    Taxonomic classification of the retained reads at the species level was performed using Kraken2 (v2.1.3) with default parameters (Wood et al., 2019) against the Unified Human Gastrointestinal Genome (UHGG) v2.0.2 reference database.4 This reference database is optimized for prokaryotes and contains 286,997 genomes, representing 4,644 bacteria and archaea species from the human gut (Almeida et al., 2021). To ensure high-quality metagenomes, only samples with >100,000 reads assigned at the species level by kraken2 were retained for downstream analyses.

    Next, Bracken (v2.9) was used to correct the relative species abundances (i.e., the proportion of reads assigned to each species) per sample. For Bracken, the Kraken2 report was used as an input with the following parameters: read length (r) = 100 bp, taxonomic level (l) = species (S), and classification threshold (t) = 10 reads. Then, using a Python script,5 we combined the Bracken result into a single consolidated report to generating a species-by-sample abundance table.

    2.3 Microbiome diversity analysis

    To assess the alpha diversity within groups, the Chao1 estimator and the Shannon index were calculated using the iNEXT R package (v2.0.20) (Hsieh et al., 2016). This package provides standardized estimates of diversity between samples with different sequencing depths through a Hill-number framework based on rarefaction and extrapolation. The Mann–Whitney U test was used to compare alpha-diversity between healthy and non-healthy groups, and between each disease phenotype and the healthy group. The effect size of these differences was assessed using Cliff’s delta (Cliff, 1993). Because diversity estimates obtained from iNEXT did not follow a normal distribution, the non-parametric Mann–Whitney U test was used for all group-level comparisons, consistent with standard practice in large-scale microbiome studies (Wirbel et al., 2019; Bastiaanssen et al., 2023).

    To evaluate differences in microbial composition between healthy and non-healthy groups, the Aitchison distance was used, calculated as the Euclidean distance based on centered log-ratio (CLR) transformed species-level abundance data. The analysis was performed using the Python package scikit-bio (v0.5.6). To visualize the compositional similarity and dissimilarity between samples, Principal Coordinate Analysis (PCoA) was performed using the skbio.stats.ordination.pcoa function. Additionally, Permutational Multivariate Analysis of Variance (PERMANOVA) was used to test for statistically significant differences in microbial community structure between health status groups and different disease phenotypes. This analysis was performed using the skbio.stats.distance.permanova function with 999 permutations.

    2.4 Batch correction across studies

    To correct batch effects arising due to studies cohorts (refers as BioProject), we used the MMUPHin (v1.16.0) package in R, especially designed for zero-inflated microbiome data (Ma et al., 2022). The adjust_batch function was applied to a species-level read count tables (contained number of reads for each species) to adjust for batch effects, specifying “BioProjectID” as the batch variable. To measure how much of the variation in microbial composition was explained by BioProject before and after batch correction, Aitchison distance (Euclidean distance) matrices were calculated using the vegdist function from the R package vegan (v2.6.10). Permutational multivariate analysis of variance (PERMANOVA) was then performed with the adonis2 function (999 permutations, by = “terms”) using the formula read_abundance_matrix ~ BioProject. The explained variance (R2 values) by BioProject, obtained before and after batch correction, were compared to evaluate the reduction of batch effects.

    2.5 Predictive model development for disease diagnosis

    We developed an ML-based predictive model to predict individuals as healthy or non-healthy using species-level gut microbiome profiles from stool metagenomic sequencing. Model development proceeded in four stages: (1) data preprocessing and splitting, (2) baseline training of four supervised algorithms, (3) hyperparameter optimization, (4) feature selection, and (5) evaluation on held out and external data (Section 2.6).

    2.5.1 Data preprocessing and splitting

    Before model training, the batch-corrected species count matrix was processed as follows:

    • To reduce sparsity and the influence of rare species, we applied prevalence-based filtering to the batch-corrected species count matrix. Prevalence was calculated as the percentage of samples in which a given species was detected, where detection was defined as having at least one non-zero read count (Gupta et al., 2020; Cao et al., 2021). To determine the impact of species prevalence filtering on model performance and to choose the optimal filtering threshold, six species-level feature sets were generated using prevalence thresholds from 0 to 25%, evaluated at 5% intervals across the pooled dataset, and each set was benchmarked across all ML models. We retained the 20% threshold (4,022 species), which gave the highest and most stable F1 scores across all models while excluding low-prevalence, sample-specific taxa.

    • Following prevalence filtering, the retained species count matrix was CLR-transformed in Python (v3.11.0) using the skbio.stats.composition.clr function from the scikit-bio package, to account for the compositional structure of microbiome data (Lee and Lee, 2024; Karwowska et al., 2025; Li et al., 2025). Before transformation, a pseudo-count of 1 was added to all feature counts to handle zero counts in the matrix. This value is the most commonly used choice and ensures that zeros in the input remain zeros after log transformation (Badsha et al., 2020). Each species abundance was then divided by the geometric mean of all species abundances within the same sample, and the natural logarithm of the resulting ratio was taken. Importantly, CLR is a sample-wise transformation that does not use metadata labels and therefore cannot, on its own, introduce information leakage during model training (Austin and Korem, 2025).

    • After CLR transformation, the full dataset of 7,452 samples was split in an 80:20 ratio into training (n = 5,961; 3,189 non-healthy and 2,772 healthy) and test (n = 1,491; 798 non-healthy and 693 healthy) sets. ML models were fitted on the training data, and their predictive performance was evaluated on the held-out test set

    2.5.2 Model algorithms

    We evaluated four supervised algorithms: SVM with linear kernel (SVM-Linear), SVM with RBF kernel (SVM-RBF), logistic regression with ElasticNet regularization (LogReg-ElasticNet), and random forest (RF), spanning both linear and non-linear model classes. These algorithms have been commonly applied in microbiome research for modeling disease prediction models and follow the established best practices for microbiome machine learning (Knights et al., 2011; Statnikov et al., 2013; Pasolli et al., 2016; Camacho et al., 2018; Zhou and Gallins, 2019; Topçuoğlu et al., 2020; Goodswen et al., 2021; Greener et al., 2022; Su et al., 2022; Asnicar et al., 2023; Lee and Lee, 2024; Wu et al., 2024).

    • Random Forest (RF): RF is an ensemble-based learning approach that builds multiple decision trees during training, each tree was trained on a randomly sampled subset of the input dataset. The final prediction is obtained by aggregating the output of individual trees through majority voting. Taking advantage of a collection of weak learners, RF improves predictive accuracy and reduces the risk of overfitting compared to a single decision tree. In this study, each tree was trained on species-level features to distinguish microbiome profiles and to identify microbial taxa associated with specific health phenotypes.

    • Support Vector Machine (SVM): The SVM algorithm classifies data by finding the optimal hyperplane that maximizes the margin of separation between target classes in a multidimensional feature space. Previous microbiome-based studies have demonstrated the effectiveness of SVM-based models for analyzing high-dimensional metagenomic datasets, where the number of features exceeds the number of samples (Knights et al., 2011; Lee and Lee, 2024). To determine the most effective Kernel separation configuration for our metagenomic dataset, we evaluated two kernel functions, that is, linear and RBF kernels, with several parameter settings. The RBF kernel captures complex, non-linear relationships by mapping features into a higher-dimensional space, allowing for a linear separator to be defined. On the other hand, the linear kernel identifies the hyperplane that best separates two classes while minimizing classification errors.

    • Logistic Regression with ElasticNet Regularization (LogReg-ElasticNet): Logistic Regression is a widely used supervised learning algorithm for binary classification, modeling the probability of a categorical outcome based on input features. In this study, ElasticNet regularization was applied to take advantage of the strength of both L1 (Lasso) and L2 (Ridge) penalties, thereby enhancing model generalizability. This approach is well-suited for metagenomic data, which often exhibits multicollinearity and sparsity. The L1 component accounts for sparsity by reducing less informative feature coefficients to zero, while the L2 component stabilizes the model by penalizing large coefficients and minimizing the risk of overfitting.

    For all four algorithms, the prediction task was identically: each stool metagenome was assigned one binary label, healthy (absence of clinically diagnosed disease) or non-healthy (presence of clinically diagnosed disease). The ML input features were CLR-transformed species-level relative read abundances obtained from UHGG v2.0.2-based taxonomic profiling. Each algorithm was given the same feature matrix and label (healthy/non-healthy), so the models differed only in the learning algorithm, and feature-selection strategy used.

    2.5.3 Model hyperparameter tuning

    Hyperparameters were tuned on the training set (80% of the data) using a two-step procedure, random search followed by grid search, applied identically to every algorithm (LaPierre et al., 2019; Lo and Marculescu, 2019; van der Sommen et al., 2020)

    In the first step, RandomizedSearchCV was used to explore 50 random hyperparameter combinations within predefined search spaces. In the second step, the most promising parameter combinations identified during the random search step were subsequently refined using GridSearchCV to determine the final optimized model configuration. To prevent overfitting, all hyperparameter tuning was performed using 10-fold stratified cross-validation, with the F1-score as the optimization metric. Tested ranges and final optimized values for each model are summarized (Supplementary Table S3).

    2.5.4 Evaluation of the model on the test and external validation dataset

    The model predictive performances were assessed using three metrics: (1) F1 score (the harmonic mean of precision and recall) (2) the area under the receiver operating characteristic curve (ROC-AUC) (which measures performance across all classification thresholds) as a performance matrix, (3) Class-wise performance was further assessed using the classification report, which provides precision, recall, F1 score, and support for each class (healthy and non-healthy), enabling detailed evaluation of how well the model predict both class. It is important to note that F1 scoring matrices were chosen as our primary metrics in the light of imbalanced datasets, consistent with established practice (Wu et al., 2018; LaPierre et al., 2019; Asnicar et al., 2023).

    All models were evaluated on both the held-out test set (20% of the entire dataset) and an independent external validation dataset (642 samples from six independent studies). The external validation dataset was solely used to evaluate model generalizability and was not involved in any phase of model development, including training, testing, and feature selection

    2.5.5 Feature selection algorithm

    To identify robust and discriminative microbial biomarkers associated with gut health, three complementary feature selection strategies were applied: (i) Least Absolute Shrinkage and Selection Operator (LASSO), (ii) Permutation Feature Importance (PFI), and (iii) Recursive Feature Elimination with Cross-Validation (RFECV)

    • Least Absolute Shrinkage and Selection Operator (LASSO): Feature selection by LASSO was performed using a logistic regression model with L1 regularization (sklearn.linear_model. LogisticRegression with L1 penalty). The L1 penalty constrains regression coefficients, shrinking less informative features toward zero and thereby enforcing sparsity in the selected feature set. The regularization parameter (α) was optimized over a logarithmic range (‘alpha’: np.logspace(−5, 2, 100)) using both random and grid search within internal cross-validation splits. This approach enables us to select features that contribute most to predictive performance while reducing overfitting (Susin et al., 2020; Ranalli et al., 2023; Garach Vélez et al., 2025).

    • Permutation-based feature importance (PFI) was applied as a model-agnostic methods to quantify the individual contribution of each microbial species to target predictions. This approach assesses the relevance of features by randomly permuting the values of each predictor (species) while maintaining the underlying data structure and then quantifying the resulting change in classification performance. Features that cause a significant decrease in model performance are interpreted as having high predictive importance, while those that have minimal impact on performance are deemed weakly informative or redundant. In this analysis, PFI was implemented using an RF classifier as the base estimator, enabling the estimation of feature importance in a non-parametric and ensemble-based learning context. This configuration enables the capture of non-linear relationships and higher-order interactions among microbial taxa, while also maintaining the model-agnostic interpretability of the resulting importance scores.

    • Recursive Feature Elimination with Cross-Validation (RFECV): RFECV was used to identify an optimal subset of features for predicting the health status via iterative feature elimination. This approach is based on repeatedly fitting a base estimator and systematically removing the least informative feature(s) at each iteration (Sanz et al., 2018). Here, a trained linear kernel of SVM (SVM-Linear) acts as the base estimator. The linear kernel was chosen to keep the model complexity low, which is advantageous when the sample size is limited but the feature space is high-dimensional. Feature importance was derived from the magnitude of the model coefficients (weight vectors), and at each iteration the feature with the lowest absolute contribution to the decision function was eliminated. This recursive process was continued until an optimal feature subset was identified. Optimality was operationally defined as the feature set that maximized the F1 score across 100 bootstrap iterations of internal train–test resampling, thereby ensuring robustness, stability, and reproducibility of the selected feature panel.

    2.6 Multivariate analysis of factors contributing to machine learning models utilizing MaAsLin2

    Multivariable association analyses were performed using the MaAsLin2 (v 1.16.0) package in R6 to assess biological relevance of ML derived features (Mallick et al., 2021). The analysis included the top 50 microbial species ranked by their importance scores in the ML models. Default MaAsLin2 parameters were used, except that healthy individuals were defined as the reference group, following recent recommendations (Nearing et al., 2022; Yang and Chen, 2023). MaAsLin2 applies a linear model to the transformed abundance of each feature across sample groups, estimates significance using a Wald test, and adjusts p-values for multiple testing using the Benjamini–Hochberg false discovery rate (FDR) method. The resulting output “significant_results.tsv” file was used to selectively retain statistically significant species for visualization, thus improving heatmap clarity and interpretability. The filtered data were imported into R, and heat maps were generated using the pheatmap package (v1.0.12).

    2.7 SHAPley values inference

    To interpret the SVM-RBF model, SHAP (SHapley Additive exPlanations) values were computed by using the Python library SHAP (v0.51.0) (Lundberg and Lee, 2017). To address the complexity of non-linear kernel and cope with computational requests the algorithm permutation was employed in the KernelExplainer class and shap values inference was parallelized among samples by using 100 chunks on 100 CPU each accessing to 100 Gb memory. Globally feature importance was measured as the mean absolute SHAP value across all samples, while the average signed SHAP value was used to assess the direction of the contribution toward each class. Features with positive mean SHAP values were interpreted as promoting the Healthy class, whereas negative values indicated an association with the Non-Healthy class. Only features with a mean SHAP values greater than zero were retained.

    3 Result

    3.1 Overview of the data collection and reprocessing using bioinformatics pipelines

    To train a machine learning model able to predict host health status based on gut microbiome profiles, we collected a total of 11,492 publicly available stool-derived shotgun metagenomes from 41 studies. Following application of predefined inclusion criteria (Methods and Figures 1A,C), 7,452 metagenomes from 32 studies, encompassing 12 distinct disease conditions, were retained for downstream analysis. Criteria for metagenome exclusion are detailed (Supplementary Section 1.2; Supplementary Table S1). For each included study, associated metadata such as sample size, disease classification, and geographical origin were manually curated and are summarized in Supplementary Table S2.

    Based on the health status reported in the original studies, metagenomes were stratified into two classes: non-healthy (n = 3,987; 53.6%) and healthy (n = 3,465; 46.4%) (defined as the presence and absence of clinically diagnosed specific disease). Within the non-healthy group, major disease categories included colorectal cancer (CRC; n = 1,132; 15.1%), Parkinson’s disease (PD; n = 690; 9.3%), obesity (OB; n = 672; 9.0%), Crohn’s disease (CD; n = 402; 5.4%), ulcerative colitis (UC; n = 293; 3.2%), and type 2 diabetes (T2D; n = 164; 1.8%), with additional disease categories summarized in Figure 1B. We acknowledge that the individuals classified as healthy were defined solely by the absence of a diagnosed disease according to the original studies. This definition may not account for undiagnosed comorbidities, microbiome-modulating medications (e.g., antibiotics, proton pump inhibitors), or demographic and dietary heterogeneity across cohorts, which may introduce label noise into the healthy class.

    We downloaded raw sequencing reads (FASTQ format) from the SRA database by using the “run accession IDs” for all selected studies. These raw reads were reprocessed and taxonomically profiled at the species level using a unified bioinformatics pipeline (Figure 1C), ensuring consistency across all datasets and minimizing technical variability introduced by differences in original processing workflows. Across all metagenomic datasets, we identified a total of 4,602 bacterial species, with an average of 2,620 species per sample, highlighting the diverse nature of gut microbial communities (Supplementary Section 1.3; Supplementary Figures S1A,B; Supplementary Table S3). We focused exclusively on bacterial species, as bacteria constitute the major component of the gut microbiome (Staudacher and Loughman, 2021; de Vos et al., 2022) and species-level resolution provides the most accurate and complete representation for discriminating disease-associated microbial signatures compared to broader taxonomic ranks (Giliberti et al., 2022; Lee and Lee, 2024; Van Hul et al., 2024).

    To characterize baseline differences between healthy and non-healthy individuals, we assessed alpha diversity (Shannon and Chao1 indices) and beta diversity (Aitchison distance) across 7,452 samples (3,465 healthy, 3,987 non-healthy) (Supplementary Section 1.4). Both Shannon and Chao1 were significantly higher in healthy individuals (p = 2.07 × 10−14 and p = 4.16 × 10−5, respectively, Supplementary Figures S2A,C), but effect sizes were small (Cliff’s Δ = 0.103 and 0.055), and the two indices did not always move in the same direction across the 12 disease phenotypes—for example, Shannon diversity decreased in Crohn’s disease but increased in colorectal cancer. In contrast, Chao1 richness showed the opposite pattern in several conditions (Supplementary Figures S2B,D; Supplementary Table S4), consistent with disease-specific rather than uniform shifts in community diversity. Similarly in beta diversity, PERMANOVA test on Aitchison distance showed a statistically significant but modest separation between healthy and non-healthy groups (R2 = 0.56%, p = 0.001, Supplementary Figure S3A), whereas phenotype-level classification across all 12 diseases explained substantially more variance (R2 = 4.04%, ~7-fold higher, Supplementary Figure S3B), with Crohn’s disease, liver disease, and colorectal cancer showing the largest separation (Supplementary Table S5)—indicating that disease-specific compositional signal exists but is diluted when diverse conditions are pooled into a single non-healthy category. In sum, both alpha and beta diversity analysis results indicate that conventional diversity metrics capture only modest, disease-heterogeneous signal (Williams et al., 2024).

    Furthermore, because the pooled dataset comprised samples from multiple independent studies, microbial profiles varied with study cohort, laboratory protocol, sequencing platform, and geographic origin. Such heterogeneity can adversely affect model performance (Ma et al., 2022; Kumar et al., 2024). We therefore applied MMUPHin to correct batch effects, using BioProject ID as the batch variable. We assessed the effect of this correction using PERMANOVA based on Bray–Curtis dissimilarities before and after correction, evaluating both study of origin and health status as explanatory variables (Supplementary Section 1.5; Supplementary Figures S4A,B). Before correction, study of origin explained 13.0% of microbial community variation in the training set and 10.4% in the external validation set (PERMANOVA R2 = 0.130 and 0.104, respectively; p = 0.001 for both), whereas health status explained only 0.10% (R2 = 0.00101, p = 0.001) and 0.51% (R2 = 0.00514, p = 0.001) of the variation, respectively. After correction, study-driven variance fell by approximately half, to 6.8 and 5.8% (R2 = 0.068 and 0.058), while the variance explained by health status remained stable or slightly increased (training 0.09%, R2 = 0.00092; validation 0.61%, R2 = 0.00608; p = 0.001), confirming preservation of disease-associated microbial signatures. Throughout, “training datasets” denotes the pooled cohorts split 80:20 into training and test subsets, and “external validation set” denotes the independent cohorts, held out entirely from model development, on which final performance was assessed.

    3.2 Development of binary classification machine learning models leveraging stool-derived microbiome data

    Given the growing potential of the gut microbiome for disease prediction and the development of machine learning-based diagnostic frameworks (Lee and Lee, 2024; Lee et al., 2024; Liu et al., 2024; Li et al., 2025; Tegegne and Savidge, 2025), we trained and evaluated four machine learning models — SVM-RBF, SVM-Linear, Logistic Regression with ElasticNet regularization (LogReg-ElasticNet), and Random Forest (RF) to develop a predictive model to predict healthy (absence of diagnosed disease) and non-healthy (presence of diagnosed disease) using microbial species profiles obtained from stool-based shotgun metagenomic datasets (Figure 1D).

    Before model training, we applied prevalence-based species filtering (referred to as the presence/absence of taxa) to reduce noise from low-prevalence and sample-specific taxa in overall metagenomic datasets. Species were retained based on the presence/absence prevalence thresholds ranging from 0 to 25%, evaluated systematically to identify the optimal filtering threshold for downstream classification (Supplementary Figure S6A). Following centered log-ratio (CLR) transformation on filtered species and their read counts, the modeling datasets was then randomly split to form a training (80%) and hold-out test (20%) sets. All four models were trained on training datasets (healthy n = 2,772; non-healthy n = 3,189), and their predictive performance was evaluated on test datasets (healthy n = 693; non-healthy n = 798) (Figure 1D).

    Benchmarking model performance across prevalence thresholds revealed that retaining species present in ≥20% of samples (n = 4,022) produced the most stable and reproducible predictive performance across all classifiers, with F1 scores consistently exceeding 79.4% (Supplementary Section 1.7; Supplementary Figure S6B). This threshold was therefore selected for all downstream analyses, and the resulting feature set is hereafter referred to as the “all-feature set.”

    Next, using the trained model on the all-feature set, we observed that the SVM-RBF model outperformed all other evaluated models. In terms of overall performance, on the test dataset, it achieved an F1 score of 86.4% (Figure 2A) and a ROC–AUC of 94.6% (Figure 2B). Class-specific performance was similarly best, with F1 scores of 88.2 and 86.4% for the non-healthy and healthy classes, respectively (Figure 2C). Additionally, other models demonstrate competitive performance but consistently had lower F1 score than the SVM-RBF model. The LogReg-ElasticNet model achieved an F1 score of 81.8% and ROC–AUC of 91.4% (Figures 2A,B). Similarly, the SVM-Linear model with an F1 score of 81.7% and an ROC–AUC of 90.8%, while the RF model achieved an F1 score of 79.4% and an ROC–AUC of 90.0% (Figures 2A,B). Class-wise performances across healthy and non-healthy groups are summarized in Figure 2C, demonstrating comparatively lower predictive performance of all evaluated models compared to the SVM-RBF classifier. A comprehensive summary of all the other classification metrics, such as balanced accuracy, overall accuracy, precision, and recall, for all models, is provided in Supplementary Table S7.

    Overall, the SVM-RBF model demonstrated the best and most consistent predictive performance across both healthy and non-healthy classes, outperforming all other evaluated classifiers in discriminating gut health status from species-level microbial profiles. We speculated that this model not only accurately captures disease-associated microbial changes but also effectively identifies stable and consistent microbial signatures present in healthy individuals. Consequently, all subsequent analyses were conducted using the SVM-RBF model to facilitate further comparisons.

    To assess whether predictive performance could be enhanced by reducing the dimensionality of the original subset of features, the SVM–RBF model was retrained on the feature subset of the top 50 microbial species (Supplementary Table S8). Focusing on this reduced set was intended to prioritize the most informative and putatively influential predictors of health status. These species were selected directly from the original SVM–RBF model, ranked by their feature importance scores. However, we found that the retrained model exhibited reduced performance, the overall F1 score declined to 77.6%, and the ROC–AUC decreased to 88.2% on the test set (Supplementary Figure S8). We speculated that the decrease in performance is due to the exclusion of additional informative species. This indicates that feature ranking based solely on marginal importance scores is insufficient to capture the full complexity of health- and disease-associated microbial signatures and inter-species co-dependencies within the gut microbiome. The collective predictive signal is distributed across a broader feature space than any single ranking metric can identify (Layeghifard et al., 2017; Coyte and Rakoff-Nahoum, 2019; D’Elia et al., 2023). To overcome this limitation and improve predictive accuracy, we next sought to apply machine learning–based feature selection strategies to more effectively identify the most informative gut microbial species, followed by retraining all models on each optimized feature subset. This approach also offers an effective way to select microbial biomarkers and address the intrinsic sparsity of microbiome data, as previously recommended (Asnicar et al., 2023).

    3.3 SVM-RBF with PFI based feature selection methods improved predictions of human gut health

    To enhance overall predictive performance and reduce the dimensionality of the metagenomic feature space, we applied three feature selection techniques: RFECV, PFI, and LASSO (see Method 2.5.5). Each method systematically reduces the dimensionality of the metagenomic feature space by retaining the most informative variables while discarding less discriminative ones, thereby improving model efficiency, interpretability, and robustness (Garach Vélez et al., 2025). Utilizing all three feature-selection approaches, RFECV retained 3,472 species, PFI retained 3,127 species, and LASSO selected 828 species. All four base classifiers were then retrained on each of the three feature-selected subsets, resulting in twelve further model configurations: RF-PFI, RF-RFECV, RF-LASSO, SVM-RBF-PFI, SVM-RBF-RFECV, SVM-RBF-LASSO, SVM-LIN-PFI, SVM-LIN-RFECV, SVM-LIN-LASSO, LogReg-ElasticNet-PFI, LogReg-ElasticNet-RFECV, and LogReg-ElasticNet-LASSO. During retraining, we applied the identical procedure as used for the original “all-feature” configuration, thereby maintaining methodological consistency and enabling fair and robust comparison across all model configurations.

    Among all retrained models, the SVM-RBF model trained on PFI-based features (SVM-RBF-PFI, n = 3,127 species) demonstrated the best overall performance on the hold-out test dataset, achieving the highest F1 score of 86.6% (Figure 2A) and a ROC–AUC of 95.5% (Figure 2B). This performance was marginally superior to the SVM-RBF model trained on the full feature set (F1 = 86.4%), indicating that PFI-based feature selection preserved and slightly improved the discriminative signal of the complete microbial profile while reducing feature dimensionality. Class-wise evaluation further demonstrates that consistent improvement in predictive performance across both classes, with an F1 score of 86.6% for the healthy group and 88.7% for the non-healthy group (Figure 2C). These observations suggest that PFI effectively identifies a compact yet highly informative subset of microbial features that captures the most discriminative biological signals associated with host health status.

    For other feature selection methods, the SVM-RBF-RFECV model achieved an F1 score of 85.7% and a ROC-AUC of 94.5%, slightly lower than the PFI-based model. The SVM-RBF-LASSO model achieved an F1 score of 85.5% and a ROC-AUC score of 93.2% (Figures 2A,B). While LASSO achieved comparable performance with substantially fewer features, the broader PFI feature set may better capture complex microbial interactions and community-level dynamics relevant to health status, consistent with the known ecological redundancy of the gut microbiome (Coyte and Rakoff-Nahoum, 2019). For LogReg-ElasticNet, both PFI and RFECV configurations showed identical performance, each achieving an F1 score of 81.6% and a ROC-AUC score of 91% (Figures 2A,B). The LogReg-ElasticNet-LASSO model underperformed slightly, with an F1 score of 80.2% and an ROC-AUC score of 89% (Figures 2A,B). Similarly, SVM-Linear models followed the same trend. The SVM-LIN-RFECV model attained an F1 score of 81.6% and a ROC-AUC score of 88%, followed by SVM-LIN-PFI with an F1 score of 80.5% and a ROC-AUC score of 89%, and SVM-LIN-LASSO with an F1 score of 79.6% and a ROC-AUC score of 88.6% (Figures 2A,B). RF models performed comparatively lower compared to all other models. The RF-LASSO model achieved an F1 score of 79.8% and a ROC-AUC score of 90.3%, followed by RF-PFI with an F1 score of 79.1% and a ROC-AUC score of 89.3%, and RF-RFECV with an F1 score of 78.9% and a ROC-AUC score of 89.8%. Class-wise performance of all models is provided (Figure 2C).

    In sum, across all sixteen models, including those trained on the full feature set, the SVM-RBF model with PFI-based feature selection (i.e., SVM-RBF-PFI) exhibited the best predictive performance on the test dataset for classifying gut health status. PFI-based feature selection outperformed both LASSO and RFECV; it is likely that PFI effectively identified most discriminative microbial taxa, reducing the dimensionality of the datasets while improving model accuracy. It also captures the complex relationships between gut microbiome composition and host health status, without excessive dimensionality reduction (Papoutsoglou et al., 2023; Garach Vélez et al., 2025). Consequently, PFI-based selection outperformed alternative feature selection strategies in capturing microbial signals relevant to health status prediction.

    3.4 Identification of key gut microbiome species and the use of MaAsLin2 to enhance biological interpretability in disease phenotypes

    To identify the microbial predictors driving model performance and evaluate their biological relevance, we used the best-performing model, SVM–RBF–PFI, to rank gut microbial species in the overall test dataset according to their contribution to health-status prediction. We focused subsequent analyses on the top 50 species that were consistently ranked across the overall test dataset (Figure 3A), as a result yielding a prioritized set of predictive taxa.

    Because PFI ranking reflects predictive importance but does not directly estimate the direction or strength of association with health status, we conducted complementary multivariate association testing using MaAsLin2 (v.1.26.0). MaAsLin2 applies generalized linear models to identify statistically significant associations between microbial taxa and host phenotypes while controlling for covariates and false discovery, an approach previously applied in large-scale metagenomic studies to validate machine learning-derived microbial signatures (Su et al., 2022; Bao et al., 2024; Lee and Lee, 2024; Lee et al., 2026). This integrated strategy, combining model-based importance ranking with MaAsLin2-derived association testing, ensures that the identified taxa are not only key contributors to predictive performance but also exhibit statistically robust and biologically relevant associations with health and disease phenotypes.

    Interestingly, we also found that a subset of microbial species was associated with both healthy and disease phenotypes (LV, ACVD, GC), including Gemella sp002871655, Finegoldia sp900766215, Lactobacillus crispatus, MGYG000003698, and Peptoniphilus_A harei_A

    In addition, several taxa were significantly associated with healthy individuals, including MGYG000003061, Peptoniphilus_A phoceensis, Massilistercora sp902406105, Extibacter hylemonae, and Ruthenibacterium lactatiformans, suggesting further validation required for the candidate biomarkers of healthy gut status. Another species, Parvimonas micra, was identified among taxa associated with healthy individuals in our dataset.

    Overall, our findings emphasize the biological relevance of the SVM-RBF-PFI model and demonstrate its robustness in identifying gut microbial signatures associated with both physiological and pathological states. These signatures may serve as reliable predictive biomarkers, highlighting the need for validation before any identified taxa are proposed as clinical biomarkers

    3.5 Evaluation of SVM-RBF model generalization capability on external validation cohorts

    To this end, we found that our SVM–RBF model, trained on all features and incorporating feature selection, generally outperformed all other models within the test dataset. Next, to evaluate how well they generalize to previously unseen data, we assessed their performance on independent metagenomic datasets, hereafter referred to as the validation dataset. In the validation datasets, we obtained a total of 642 stool-derived metagenomes from six published studies (associated metadata are provided in Supplementary Table S10). These datasets included samples from six disease conditions: type 2 diabetes (T2D, n = 127), Parkinson’s disease (PD, n = 96), Colorectal Cancer (CRC, n = 70), Obesity (OB, n = 61), Clostridioides Difficile Infection (CDI, n = 22), and Atherosclerotic Cardiovascular Disease (ACVD, n = 12) (Figure 4A). Notably, the CDI and T2D disease datasets represent a unique case, as neither this disease type nor the corresponding studies were represented in the training dataset. All validation metagenomes were reprocessed through the same unified bioinformatic pipeline applied to the training data to minimize methodological biases (Methods Section 2.2). Importantly, all validation data were excluded from model training, internal cross-validation, and feature selection. This design ensured cohort-level independence and prevented data leakage, thereby enabling an unbiased evaluation of the model’s generalization performance.

    On the external validation datasets, the SVM–RBF model trained on all sets of features demonstrated the highest overall performance, achieving an F1 score of 70.64% and an ROC–AUC of 84.7% (Figures 4B,D; Supplementary Table 11). Class-wise evaluation revealed significant performance for both healthy individuals (F1 = 70.6%) and non-healthy individuals (F1 = 83%) (Figure 4C). In comparison, the SVM–RBF–RFECV model showed slightly lower performance, with an overall F1 score of 68.72% and an ROC–AUC of 84.1%, and class-specific for healthy individuals (F1 = 68.7%) and for non-healthy individuals (F1 = 82.9%) (Figure 4C). The SVM–RBF–PFI model demonstrated lower performance, with an overall F1 score of 60.43% and an ROC–AUC of 82.5% (Figures 4A,D). While this model showed marginally improved performance of non-healthy samples, achieving an F1 score of 83.7%. However, its performance on healthy samples declined substantially (F1 = 60.4%). Similarly, the SVM–RBF–LASSO model achieved an overall F1 score of 60.91% and an ROC–AUC of 79.1%. Class-wise performance for other models is provided in Figure 4C. Such performance from an independent validation cohort further confirmed the robustness and generalizability of our SVM-RBF with all features model across different disease datasets.

    To further characterize model performance, we evaluated all SVM-RBF configurations on each disease cohort within the external validation dataset. Our findings indicate that the F1 scores for all models of SVM-RBF ranged from 49 to 94% (Figure 5A), with ROC-AUC scores ranging from 75.9 to 93.5% (Figure 5B) across the validation cohorts. Notably, the SVM-RBF (all features) outperformed the other models, achieving F1 scores between 60 and 90.3% across all disease datasets. Specifically, this model showed best performance on ACVD (F1 score of 94.1% and ROC-AUC score of 93.5%), CDI (F1 score of 90.3% and ROC-AUC score of 92.2%) cohorts, and for T2D (F1 score of 77.4% and ROC-AUC score of 83.2%).

    In terms of microbial predictors in the external validation cohort, we examined feature importance profiles across individual disease datasets. The SVM-RBF model identified both shared and cohort-specific microbial taxa contributing to health status predictions (Supplementary Table S12). Our model identified several taxa that were shared across multiple disease datasets, including Victivallis sp002998355, MGYG000001921, Desulfovibrio sp900319575, Akkermansia muciniphila_B, CAG-269 sp000437215, and the Ruminococcus genus. Among these, Akkermansia species is well-established as a health-associated commensal whose depletion is linked to T2D and CDI (Panzetta and Valdivia, 2024; Zhao et al., 2024). Similarly, Desulfovibrio species have been consistently reported at elevated abundance across multiple disease states, including PD, OB, and T2D, where they contribute to gut barrier disruption through hydrogen sulphide production (Murros et al., 2021; Nie et al., 2023). The Ruminococcus genus encompasses both health-associated fiber-degrading commensals and disease-linked species such as Ruminococcus gnavus, which has been implicated in CD (Henke et al., 2019), and metabolic dysregulation (Meadows et al., 2025), consistent with its cross-disease predictive relevance observed here. In addition, we also found cohort-specific taxa (Supplementary Table S12), Thelissonella pneumosintes and Dialister hominis were prevalent in the CRC dataset, UCA11452 sp003526375 and Prevotella sp000431975 were specific to ACVD, Dialister invisus and Acidaminococcus massiliensis were associated with T2D, and Bulleidia moorei, Odoribacter splanchnicus, and Clostridium species were specific to CDI. Overall, the SVM-RBF-all features model demonstrated superior performance in the external validation cohorts, suggesting that our ML models more effectively and robustly capture health- and disease-associated microbial patterns when applied to independent datasets.

    3.6 Evaluation of SVM-RBF features relevance through XAI

    Explainable artificial intelligence (XAI) methods can help identify the features driving model predictions. To determine the features contributing most strongly to the predictions of the SVM-RBF model, we calculated SHapley Additive exPlanations (SHAP) values using the Python SHAP library. We considered both the mean absolute SHAP value, which quantifies the overall importance of each feature, and the mean signed SHAP value, which indicates the average direction of its contribution to model predictions (Supplementary Table S15).

    A total of 2,525 features with mean absolute SHAP values greater than zero were retained. Figure 6 shows the 20 features with the highest mean absolute SHAP values. The direction of each bar indicates the average direction of the feature’s contribution: negative values indicate a contribution toward classification as non-healthy, whereas positive values indicate a contribution toward classification as healthy. The bar color represents the magnitude of the mean absolute SHAP value, with darker shades indicating features with greater overall influence on model predictions.

    Accordingly, Sutterella wadsworthensis A contributed, on average, toward classification as non-healthy, whereas Fusobacterium A sp900555845 contributed toward classification as healthy. Among these 20 features, only Fusobacterium A sp900543175 was also selected by the PFI procedure. Overall, 29 of the 2,525 features retained in the SHAP analysis were also identified through PFI (Supplementary Table S15).

    4 Discussion

    The human gut harbors a dense and taxonomically diverse microbial community whose composition is closely associated to host health and disease (Asnicar et al., 2023; Rosenberg, 2024). In this study, we developed a ML framework to predict gut health status—defined as the presence or absence of a clinically diagnosed disease from species-level stool metagenomic profiles. Classifiers trained on a single cohort tend to learn study-specific technical and population variation rather than microbial signals shared across diseases, which limits their performance when applied to independent cohorts (Li et al., 2023). We therefore pooled 7,452 publicly available stool metagenomes from 32 independent studies covering 12 disease conditions, and trained ML models on this combined dataset to identify dysbiosis signatures that recur across both cohorts and disease types, extending earlier applications of this pooling strategy (Li et al., 2023; Bao et al., 2024; Jin et al., 2024; Lee and Lee, 2024; Defazio et al., 2026).

    Our diversity and compositional analyses showed that alpha and beta diversity explain only a small fraction of the disease-associated variation in the gut microbiome (Supplementary Section 1.3). The small effect sizes and inconsistent diversity patterns across diseases indicate that no single diversity metric provides a universal marker of gut health, consistent with previous studies showing that microbiome diversity changes vary across diseases and cohorts (Sun et al., 2024; Williams et al., 2024; Corral López et al., 2026). These findings motivate supervised machine learning on species-level profiles, which can detect multivariate compositional shifts that alpha- and beta-diversity summaries cannot resolve (Kumar et al., 2024; Zhou and Zhao, 2025; Yu et al., 2026).

    Pooling large volumes of metagenomic data can introduce study-level batch effects that degrade model performance (Yu et al., 2023). We therefore corrected batch effects using MMUPHin, with BioProject ID as the batch label. Because technical metadata were largely unavailable (Supplementary Section 1.3), BioProject ID served as a proxy for methodological differences in library preparation and sequencing, as samples from the same project often share uniform protocols (Barrett et al., 2012). This approach was intended to reduce study-associated technical variation that could compromise cross-cohort generalizability. We acknowledge, however, that BioProject ID may be confounded with disease status, population characteristics, and other biological factors. Consequently, correction based on study identifiers may attenuate genuine biological signals while leaving some residual technical variation. Moreover, disease-associated microbial signatures may themselves be partly attributable to measured or unmeasured confounders (Li et al., 2025). Our analysis therefore focused on reducing inter-study heterogeneity, but it cannot completely separate technical variation from cohort-specific biological variation.

    In total, we trained and benchmarked 16 supervised models, combining four algorithms with the full feature set and three feature-selection strategies. In terms of classification performance for predicting health status (healthy vs. non-healthy), the non-linear SVM-RBF model consistently outperformed linear classifiers (SVM-Linear, LogReg-ElasticNet) and RF, both on the held-out test set and in external validation. On the test dataset (n = 1,491), the SVM-RBF-PFI model achieved the best performance compared with the full-feature model (F1 = 86.6%, ROC-AUC = 95.5%), with comparably high F1-scores for non-healthy (88.7%) and healthy (86.6%) individuals. Gut metagenomic profiles are high-dimensional, and many features represent rare, low-prevalence taxa detected in only a small proportion of samples. These features may introduce sparse noise rather than reproducible disease-associated information (Papoutsoglou et al., 2023). By retaining only, the species that contributed most strongly to classification, the PFI-based approach may have reduced irrelevant variation and improved the signal-to-noise ratio, thereby enhancing the internal predictive performance of the SVM-RBF-PFI model. Moreover, the SVM-RBF model trained on the full feature set exhibited superior generalization to independent validation cohorts (F1 = 70.6%, ROC-AUC = 84.7%), maintaining robust performance for both healthy (70.6%) and non-healthy (83.0%) individuals, including disease types completely unseen in the training set, such as Clostridioides difficile infection (F1 = 90.3%) and type 2 diabetes (F1 = 77.4%).

    We speculate that this trade-off may reflect a broader ecological phenomenon: different bacterial species can perform similar functions in the gut, and a species carrying a disease-associated signal in one cohort may be replaced by a functionally similar species in another (Moya and Ferrer, 2016; Rosenberg, 2024). Such ecological variation has been reported across disease cohorts at scale (Jiang et al., 2025). Restricting the model to species that were most discriminative in the training cohorts may therefore have removed alternative or partially redundant taxa that retained predictive information in independent cohorts and previously unseen diseases. By retaining the complete species set, the full-feature model may have preserved more of these signals. This interpretation is consistent with reports that microbiome biomarkers are often shared across diseases rather than being strictly disease-specific (Duvallet et al., 2017). An ecological perspective has also been used to extend conventional enterotype models. For example, Lai et al. (2023) applied the concept of microbial ecological factors (MEFs) to identify microbial guilds whose collective dynamics were associated with host health (Zhu et al., 2023). Validating models on diseases absent from the training data, as performed here for CDI and T2D, provides a more stringent assessment of cross-disease generalizability than within-cohort cross-validation alone (Steyerberg and Harrell, 2016; Li et al., 2025). he superior performance of the nonlinear RBF kernel relative to the linear classifiers may similarly reflect the complexity of microbial communities, in which taxa interact through processes such as cross-feeding and competition rather than contributing independently (Culp and Goodman, 2023). Consistent with prior work, our results indicate that non-linear SVM-RBF decision boundaries between healthy and non-healthy individuals are more generalizable than linearly constrained boundaries, underscoring its effectiveness in capturing non-linear relationships between microbial community composition and host health status (Kubinski et al., 2022).

    Beyond evaluating predictive performance, we identified taxa associated with disease phenotypes and examined the features contributing to model predictions across diverse conditions. Feature-importance rankings, considered alongside MaAsLin2 multivariable association analyses, showed that several influential model features were also statistically associated with host phenotypes. This convergence supports the biological relevance of some of the predictive signals, although it does not exclude residual confounding or other pipeline-related effects. Among the highest-ranked taxa, Klebsiella pneumoniae was linked to nine disease states, while Salinicola tamaricis, Raoultella ornithinolytica, and Actinomyces oris were linked to multiple conditions each—a shared dysbiosis signature rather than disease-specific markers (Supplementary Table S8). Furthermore, taxa such as Gemella sp002871655, Finegoldia sp900766215, Lactobacillus crispatus, MGYG000003698, and Peptoniphilus harei_A displayed context-dependent associations, being enriched in healthy samples in some comparisons but showing altered abundance or associations in disease in others. These patterns may reflect context-dependent ecological roles in which taxa behave as commensals under some conditions but undergo abundance or functional changes in altered gut environments. They also underscore the dynamic nature of host–microbe and microbe–microbe interactions. Consistent with previous studies, taxa may be associated with both health and multiple disease states rather than being disease-specific (Duvallet et al., 2017; Wilkins et al., 2019; Abbas-Egbariya et al., 2022; Priya et al., 2022; Corral López et al., 2026; da Silva et al., 2026). Such patterns may also arise from strain-level variation, as strains within a species differ in accessory gene content, metabolic capacity, and virulence factors, yielding context-dependent effects on host–microbiome interactions (van der Sommen et al., 2020; Doran et al., 2025). These observations highlight the limitations of species-level interpretation and the need for strain-resolved analyses.

    In the case of Lactobacillus crispatus, our findings are concordant with previous reports describing its context-dependent role in health and disease (Pan et al., 2020). The model also identified several taxa predominantly associated with health, including Parvimonas micra, MGYG000003061, Peptoniphilus_A phoceensis, Massilistercora sp902406105, Extibacter hylemonae, and Ruthenibacterium lactatiformans. Notably, although P. micra is a normal commensal of the gastrointestinal tract, it is also a well-documented pathobiont enriched in colorectal cancer, periodontitis, and various systemic infections (Xu et al., 2020; Zhao et al., 2022; Higashi et al., 2023). Its association with healthy status in our analysis may reflect its role as a low-abundance commensal organism under healthy conditions. Our training cohort may also not have captured the disease contexts that promote its pathogenic behavior. This is consistent with its role as an opportunistic pathogen that remains harmless in a healthy gut but may contribute to dysbiosis and disease when the gut environment is disturbed (Chow et al., 2011; Kamada et al., 2013; Thomas et al., 2019; Xu et al., 2020; Higashi et al., 2023). Indeed, P. micra has been described as an inflammophilic pathobiont whose overgrowth is driven by inflammation and inflammation-damaged tissue (Bergsten et al., 2023). Two main P. micra phylotypes (A and B) have recently been described, with phylotype A predominantly associated with colorectal cancer. Variation in P. micra abundance was reported at early and late tumor stages but not in adenomas, suggesting that its association with disease may depend on disease stage and local inflammatory conditions (Bergsten et al., 2023).

    In external validation datasets comprising entirely independent studies, our model identified both microbial species shared across diseases and taxa specific to particular cohorts or phenotypes (Supplementary File 2). These results suggest that the model did not rely exclusively on patterns associated with the diseases represented in the training data and could recover biologically plausible microbial signatures in previously unseen disease types. Dialister invisus and Acidaminococcus massiliensis were associated with T2D, while Bulleidia moorei, Odoribacter splanchnicus, and Clostridium species were specific to CDI. Several of these associations are supported by prior evidence. For example, Dialister invisus has been linked to increased intestinal permeability and metabolic dysregulation in diabetic patients (Maffeis et al., 2016; Loaiza et al., 2025), underscoring its relevance to metabolic disease. Similarly, Odoribacter splanchnicus, detected specifically in the CDI cohort, has been reported to inhibit Clostridioides difficile toxin production through competitive inhibition in clinical and in vitro settings, with its abundance negatively correlated with CDI severity (Wang et al., 2026). Collectively, the biological plausibility of these species-level associations across two entirely unseen disease types suggests that the model has learned shared features of gut microbial dysbiosis rather than memorized disease-specific patterns.

    Finally, considering its generalization ability we inferred SHAP values on the SVM-RBF model. The returned features partially reprise what we found by combining the SVM-RBF-PFI model with MaAsLin2 analysis. For example, Sutterella wadsworthensis A was among the features contributing most strongly toward classification as non-healthy. Other members of this genus were also associated with several disease phenotypes in our analyses. Although S. wadsworthensis is commonly present in the gut microbiome and has not itself been characterized as directly pro-inflammatory (Mukhopadhya et al., 2011), recent studies have reported genes encoding proteases capable of targeting IgA1 and IgA2, suggesting a potential mechanism through which it could influence host–microbiome interactions and community composition (Dupraz et al., 2026; Majzoub et al., 2026).

    Another taxon contributing toward classification as non-healthy was Collinsella aerofaciens A. This observation is consistent with the findings of Chen et al. (2016), who reported enrichment of the genus Collinsella, particularly C. aerofaciens, in patients with rheumatoid arthritis. Its abundance was correlated with production of the pro-inflammatory cytokine IL-17A, and experiments in an arthritis model supported a role in increasing intestinal permeability. These findings provide biological context for its contribution toward the non-healthy class in our model, although they do not establish the same mechanism across the diverse diseases included here. Interestingly, two Fusobacterium species – Fusobacterium A sp900555845 and Fusobacterium A sp900543175 – contributed toward classification as healthy. The taxonomy of this genus has recently been extensively revised, revealing substantial genetic heterogeneity among its members; consequently, disease associations reported for some Fusobacterium taxa should not be generalized to all species within the genus (Molteni et al., 2024). Conversely, P. micra contributed toward classification as non-healthy, although its mean signed SHAP value was very small (8.9 × 10−6). This finding illustrates how model-interpretation methods can reveal subtle predictive contributions, but the small effect magnitude warrants cautious biological interpretation. From a translational perspective, this framework provides a foundation for developing microbiome-based classifiers designed to generalize across multiple disease types rather than being restricted to a single disease. The shared microbial signatures identified here may help prioritize candidates for mechanistic investigation and microbiome-directed interventions and may inform participant stratification in future studies. However, the model classifies samples according to the reported presence or absence of clinically diagnosed disease; it is not a quantitative health index or a substitute for clinical diagnosis. Prospective validation in well-characterized populations, together with standardized demographic, clinical, dietary, medication, technical, functional, and strain-level data, will be necessary before any clinical application can be considered.

    4.1 Limitations of the studies and future directions

    Although the models showed strong predictive performance, the limitations described below should be considered when interpreting these results

    • The model captures statistical associations between gut microbiome composition and disease-based health classification; it does not establish causality. It should not be used as a clinical diagnostic tool or interpreted as a quantitative index of gut health because it was developed for binary classification rather than individualized health scoring. Its intended purpose is to identify candidate microbial biomarkers and prioritize species for further biological and clinical investigation. Unmeasured host factors—including medication use, diet, age, and BMI—also influence microbiome composition. If these factors are unevenly distributed between the healthy and non-healthy groups, they may confound microbiota–disease associations and generate spurious signals (Vujkovic-Cvijin et al., 2020; Zhu et al., 2026). Our labels follow operational definitions previously used in microbiome health-status frameworks (Gupta et al., 2020; Chang et al., 2024). Because they are based primarily on the reported presence or absence of clinically diagnosed disease, undiagnosed comorbidities, microbiome-altering medication use, and between-cohort biological and technical heterogeneity may introduce label noise, particularly within the healthy group.

    • A direct comparison with frameworks such as MetaML (Pasolli et al., 2016) was not feasible because the models were developed using different taxonomic reference databases and, consequently, different feature spaces and relative-abundance estimates. A valid head-to-head comparison would require reprocessing the same raw sequencing data using harmonized taxonomic profiling, preprocessing, and validation procedures. Without such harmonization, observed performance differences could reflect database and pipeline variation rather than differences between modeling frameworks. To our knowledge, no universally accepted shotgun-metagenomic benchmark dataset currently exists for binary classification of healthy and non-healthy samples. Developing standardized benchmark datasets and evaluation protocols should therefore be a priority for the field.

    • Our model was trained exclusively on bacterial species because bacteria are the best-characterized component of the gut microbiome and the reference database used in this study, UHGG v2.0.2, is designed primarily for prokaryotic profiling (Almeida et al., 2021). The virome and mycobiome were therefore beyond the scope of the present analysis. Future studies incorporating virus- and fungi-specific reference databases could provide a more comprehensive representation of the gut microbial ecosystem.

    • Clinical and demographic variables (e.g., age, sex, BMI) were not used during model training because they were missing for most studies (Supplementary Table S10). Systematic extraction and harmonization of such metadata from published datasets remains an open challenge in the field (Kumar et al., 2024). The same metadata gap meant our batch correction relied on study identity alone, which cannot fully separate technical from biological variation. This limits our ability to fully characterize model behavior, given the influence of host clinical factors on gut microbiota composition (Bajinka et al., 2020) and the potential for unmatched host variables to confound microbiota–disease associations (Vujkovic-Cvijin et al., 2020. Future studies integrating harmonized clinical and demographic metadata with microbial profiles should help disentangle host- and microbiome-driven effects and improve model robustness and interpretability (Wang et al., 2023).

    • Model generalizability was assessed using a limited number of external validation datasets. Although the model retained predictive performance in independent cohorts and disease types not represented during training, validation in larger, more diverse, and prospectively collected clinical populations is required before its translational relevance can be established. Future work should also evaluate a broader range of diseases because dysbiosis signatures may differ among conditions with distinct pathophysiological mechanisms (Kim et al., 2024).

    • The model used species-level abundance profiles and therefore did not capture strain-level variation or functional potential at the gene or pathway level. Strains belonging to the same species can differ in gene content, metabolic capacity, and virulence factors and may consequently exhibit distinct associations with disease despite sharing the same species-level classification. Numerous strain–phenotype and strain–geography associations have been reported across global gut microbiome datasets (Faith et al., 2015; Zhang and Zhao, 2016; Van Rossum et al., 2020; Andreu-Sánchez et al., 2025). Functional changes may also occur without corresponding taxonomic shifts, while pathway profiling can reveal metabolic and ecological relationships that cannot be inferred from taxonomic composition alone (Fuggle et al., 2025; Zielińska et al., 2025). Future work could generate HUMAnN-derived gene-family and pathway profiles as additional model inputs. StrainGE (van Dijk et al., 2022) or StrainScan (Liao et al., 2023) could also be applied to sufficiently prevalent and well-covered biomarker species to determine whether functional or strain-level resolution improves disease-based sample classification and cross-cohort generalization.

    One potential extension of our binary classification framework is normative modeling, which quantifies the continuous deviation of an individual’s microbiome from an appropriately defined reference population rather than assigning a binary label (Zhu et al., 2026). Such an approach could potentially improve the characterization of early, intermediate, or subclinical dysbiosis. Two recent frameworks illustrate related strategies. GMWI2 provides a continuous microbiome-based wellness score (Chang et al., 2024), whereas Wiredancer represents the microbiome through multiple continuous ecological factors rather than a single score. The latter framework aims to characterize dysbiotic, protective, and intermediate microbial states and may therefore position samples more precisely along a multidimensional health–disease continuum (Zhu et al., 2026). A complementary direction would be to integrate taxonomic profiles with strain-level and functional information, including pathway and gene-family abundances generated using HUMAnN 3. These microbial data layers could subsequently be combined with other omics data, such as metatranscriptomics and metabolomics, as well as clinical and physiological variables. In disease contexts in which they are relevant, neuroimaging and electrophysiological measurements such as electroencephalography could also be incorporated. Recent large-scale brain–gut cohorts have demonstrated the potential value of combining gut microbiome profiles with neuroimaging, electrophysiological data, peripheral biomarkers, and clinical phenotypes to characterize microbiome–host relationships (Wu et al., 2026). Deep-learning approaches—including Transformers, multilayer perceptrons, and convolutional neural networks—may capture complex microbial patterns as larger and more comprehensively annotated datasets become available (Oh and Zhang, 2020; Su et al., 2022; Bao et al., 2024; Yan B. et al., 2025; Liu et al., 2026). In the present study, however, we deliberately restricted the benchmark to established classical ML algorithms that have been extensively applied to high-dimensional microbiome data (Goodswen et al., 2021; Asnicar et al., 2023). Future studies should compare these methods with appropriately regularized deep-learning architectures using the same cohort-stratified external-validation framework. This would help determine whether more complex architectures provide reproducible improvements in cross-cohort generalization rather than gains confined to the training distribution.

    Our study demonstrates that the large-scale integration of publicly available shotgun-metagenomic datasets, combined with standardized processing and explicit management of study-associated variation, can support the development of ML models for gut microbiome-based classification of samples from individuals with and without reported clinically diagnosed disease. The framework is scalable and showed generalization across independent cohorts and disease types absent from the training data. Nevertheless, it is intended primarily as a research framework for cross-disease biomarker discovery and prioritization, not as a clinical diagnostic tool or population-screening instrument. Prospective validation in deeply phenotyped populations, accompanied by harmonized clinical, demographic, technical, functional, and strain-level data, will be required before potential clinical applications can be evaluated.

    Statements

    Data availability statement

    Raw sequencing data were retrieved from the NCBI Sequence Read Archive (SRA, https://www.ncbi.nlm.nih.gov/sra). The corresponding BioProject accession numbers and associated metadata are provided in Supplementary Tables S2 and S10. All scripts and features table used for data processing, analysis and models training are available at: https://github.com/bablukr056/ML4GutHealth

    Funding

    The author(s) declared that financial support was received for this work and/or its publication. This work was supported by projects DigitAl Lifelong PRevEntion (DARE) (PNC0000002 -CUP: B53C22006420001), Life Science Hub Regione Puglia (LSH-Puglia, T4-AN-01 H93C22000560003), INNOVA -Italian network of excellence for advanced diagnosis (PNC-EJ-2022-23683266 PNC-HLS-DA) and by ELIXIR-IT through the empowering project ELIXIRNextGenIT (Grant No. IR0000010).

    Acknowledgments

    BK is a PhD student within the European School of Molecular Medicine (SEMM)

    Conflict of interest

    The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest

    The author(s) BF and GP declared that they were an editorial board member of Frontiers, at the time of submission. This had no impact on the peer review process and the final decision

    Generative AI statement

    The author(s) declared that Generative AI was used in the creation of this manuscript. Generative AI was used to improve English writing and grammar

    Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us

    Publisher’s note

    All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher

    Supplementary material

    The Supplementary material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fmicb.2026.1925500/full#supplementary-material

    Footnotes

    1.^https://github.com/ncbi/sra-tools

    2.^https://github.com/s-andrews/fastqc

    3.^Indexed downloaded from: https://genome-idx.s3.amazonaws.com/bt/GRCh38_noalt_as.zip

    4.^http://ftp.ebi.ac.uk/pub/databases/metagenomics/mgnify_genomes/human-gut/v2.0.2/kraken2_db_uhgg_v2.0.2/

    5.^https://github.com/jenniferlu717/KrakenTools/blob/master/combine_kreports.py

    6.^https://huttenhower.sph.harvard.edu/maaslin/

    References

    • 1

      Abbas-EgbariyaH.HabermanY.BraunT.HadarR.DensonL.Gal-MorO.et al. (2022). Meta-analysis defines predominant shared microbial responses in various diseases and a specific inflammatory bowel disease signal. Genome Biol.23:61. doi: 10.1186/s13059-022-02637-7

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 2

      AlmeidaA.NayfachS.BolandM.StrozziF.BeracocheaM.ShiZ. J.et al. (2021). A unified catalog of 204,938 reference genomes from the human gut microbiome. Nat. Biotechnol.39, 105–114. doi: 10.1038/s41587-020-0603-3

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 3

      Andreu-SánchezS.Blanco-MíguezA.WangD.GolzatoD.ManghiP.HeidrichV.et al. (2025). Global genetic diversity of human gut microbiome species is related to geographic location and host health. Cell188, 3942–3959.e9. doi: 10.1016/j.cell.2025.04.014

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 4

      AsnicarF.ThomasA. M.PasseriniA.WaldronL.SegataN. (2023). Machine learning for microbiologists. Nat. Rev. Microbiol.22, 191–205. doi: 10.1038/s41579-023-00984-1

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 5

      AustinG. I.KoremT. (2025). Compositional transformations can reasonably introduce phenotype-associated values into sparse features. mSystems10, e0002125–e0002125. doi: 10.1128/msystems.00021-25

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 6

      BadshaM. B.LiR.LiuB.LiY. I.XianM.BanovichN. E.et al. (2020). Imputation of single-cell gene expression with an autoencoder neural network. Quant Biol8, 78–94. doi: 10.1007/s40484-019-0192-7

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 7

      BajinkaO.TanY.AbdelhalimK. A.ÖzdemirG.QiuX. (2020). Extrinsic factors influencing gut microbes, the immediate consequences and restoring eubiosis. AMB Express10:130. doi: 10.1186/s13568-020-01066-8

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 8

      BaoZ.YangZ.SunR.ChenG.MengR.WuW.et al. (2024). Predicting host health status through an integrated machine learning framework: insights from healthy gut microbiome aging trajectory. Sci. Rep.14:31143. doi: 10.1038/s41598-024-82418-3

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 9

      BarcanR. A.CarradoriS.SamsingF.NguyenN.-L.HeL.WangY.et al. (2026). Machine learning in applied microbiology, from data quality to model validation and implementation. Microbiol. Res.311:128588. doi: 10.1016/j.micres.2026.128588

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 10

      BarrettT.ClarkK.GevorgyanR.GorelenkovV.GribovE.Karsch-MizrachiI.et al. (2012). BioProject and BioSample databases at NCBI: facilitating capture and organization of metadata. Nucleic Acids Res.40, D57–D63. doi: 10.1093/nar/gkr1163

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 11

      Bars-CortinaD.RamonE.Rius-SansalvadorB.GuinóE.Garcia-SerranoA.MachN.et al. (2024). Comparison between 16S rRNA and shotgun sequencing in colorectal cancer, advanced colorectal lesions, and healthy human gut microbiota. BMC Genomics25:730. doi: 10.1186/s12864-024-10621-7

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 12

      BasicM.DardevetD.AbujaP. M.BolsegaS.BornesS.CaesarR.et al. (2022). Approaches to discern if microbiome associations reflect causation in metabolic and immune disorders. Gut Microbes14:2107386. doi: 10.1080/19490976.2022.2107386

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 13

      BastiaanssenT. F. S.QuinnT. P.LoughmanA. (2023). Bugs as features (part 1): concepts and foundations for the compositional data analysis of the microbiome–gut–brain axis. Nat. Ment. Health1, 930–938. doi: 10.1038/s44220-023-00148-3

      • CrossRef
      • Google Scholar
    • 14

      BekhetS.AlsherefF. K.AbdelAzizA. M.AlshazlyH. (2025). Revolutionizing microbial classification: leveraging machine learning for enhanced classification with feature-based data. Multimed. Tools Appl.84, 45041–45060. doi: 10.1007/s11042-025-20933-9

      • CrossRef
      • Google Scholar
    • 15

      BergstenE.MestivierD.DonnadieuF.PedronT.BarauC.MedaL. T.et al. (2023). Parvimonas micra, an oral pathobiont associated with colorectal cancer, epigenetically reprograms human colonocytes. Gut Microbes15:2265138. doi: 10.1080/19490976.2023.2265138

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 16

      CamachoD. M.CollinsK. M.PowersR. K.CostelloJ. C.CollinsJ. J. (2018). Next-generation machine learning for biological networks. Cell173, 1581–1592. doi: 10.1016/j.cell.2018.05.015

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 17

      CaoQ.SunX.RajeshK.ChalasaniN.GelowK.KatzB.et al. (2021). Effects of rare microbiome taxa filtering on statistical analysis. Front. Microbiol.11:607325. doi: 10.3389/fmicb.2020.607325

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 18

      Casimiro-SoriguerC. S.LouceraC.Peña-ChiletM.DopazoJ. (2022). Towards a metagenomics machine learning interpretable model for understanding the transition from adenoma to colorectal cancer. Sci. Rep.12:450. doi: 10.1038/s41598-021-04182-y

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 19

      ChanL. S.LiG. (2024). Zero is not absence: censoring-based differential abundance analysis for microbiome data. Bioinformatics40:btae071. doi: 10.1093/bioinformatics/btae071

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 20

      ChandaD.DeD. (2024). Meta-analysis reveals obesity associated gut microbial alteration patterns and reproducible contributors of functional shift. Gut Microbes16:2304900. doi: 10.1080/19490976.2024.2304900

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 21

      ChangC.-Y.BajićD.VilaJ. C. C.EstrelaS.SanchezA. (2023). Emergent coexistence in multispecies microbial communities. Science381, 343–348. doi: 10.1126/science.adg0727

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 22

      ChangD.GuptaV. K.HurB.Cobo-LópezS.CunninghamK. Y.HanN. S.et al. (2024). Gut microbiome wellness index 2 enhances health status prediction from gut microbiome taxonomic profiles. Nat. Commun.15:7447. doi: 10.1038/s41467-024-51651-9

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 23

      ChenJ.WrightK.DavisJ. M.JeraldoP.MariettaE. V.MurrayJ.et al. (2016). An expansion of rare lineage intestinal microbes characterizes rheumatoid arthritis. Genome Med.8:43. doi: 10.1186/s13073-016-0299-7

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 24

      ChenS.ZhouY.ChenY.GuJ. (2018). Fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics34, i884–i890. doi: 10.1093/bioinformatics/bty560

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 25

      ChoI.BlaserM. J. (2012). The human microbiome: at the interface of health and disease. Nat. Rev. Genet.13, 260–270. doi: 10.1038/nrg3182

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 26

      ChowJ.TangH.MazmanianS. K. (2011). Pathobionts of the gastrointestinal microbiota and inflammatory disease. Curr. Opin. Immunol.23, 473–480. doi: 10.1016/j.coi.2011.07.010

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 27

      CliffN. (1993). Dominance statistics: ordinal analyses to answer ordinal questions. Psychol. Bull.114, 494–509. doi: 10.1037/0033-2909.114.3.494

      • CrossRef
      • Google Scholar
    • 28

      CoranderJ.HanageW. P.PensarJ. (2022). Causal discovery for the microbiome. Lancet Microbe3, e881–e887. doi: 10.1016/S2666-5247(22)00186-0

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 29

      Corral LópezR.BonachelaJ. A.Dominguez-BelloM. G.ManhartM.LevinS. A.BlaserM. J.et al. (2026). Imbalance in gut microbial interactions as a marker of health and disease. Science391, 890–895. doi: 10.1126/science.ady1729

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 30

      CosteaP. I.ZellerG.SunagawaS.PelletierE.AlbertiA.LevenezF.et al. (2017). Towards standards for human fecal sample processing in metagenomic studies. Nat. Biotechnol.35, 1069–1076. doi: 10.1038/nbt.3960

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 31

      CoyteK. Z.Rakoff-NahoumS. (2019). Understanding competition and cooperation within the mammalian gut microbiome. Curr. Biol.29, R538–R544. doi: 10.1016/j.cub.2019.04.017

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 32

      CulpE. J.GoodmanA. L. (2023). Cross-feeding in the gut microbiome: ecology and mechanisms. Cell Host Microbe31, 485–499. doi: 10.1016/j.chom.2023.03.016

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 33

      D’EliaD.TruuJ.LahtiL.BerlandM.PapoutsoglouG.CeciM.et al. (2023). Advancing microbiome research with machine learning: key findings from the ML4Microbiome COST action. Front. Microbiol.14:1257002. doi: 10.3389/fmicb.2023.1257002

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 34

      da SilvaA. C.LapkinJ.YinQ.MullerE.AlmeidaA. (2026). Meta-analysis of the uncultured gut microbiome across 11,115 global metagenomes reveals a candidate signature of health. Cell Host Microbe34, 379–392.e5. doi: 10.1016/j.chom.2026.01.013

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 35

      de VosW. M.TilgH.HulM. V.CaniP. D. (2022). Gut microbiome and health: mechanistic insights. Gut71, 1020–1032. doi: 10.1136/gutjnl-2021-326789

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 36

      DefazioG.LorussoE.De RobertisM.MelloT.GalliA.PesoleG.et al. (2026). Machine learning-based assessment of the healthy human gut mycobiota landscape using ITS1 DNA metabarcoding data. BioData Min.19:35. doi: 10.1186/s13040-026-00532-6

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 37

      DoranB. A.ChenR. Y.GibaH.BeheraV.BaratB.SundararajanA.et al. (2025). Subspecies phylogeny in the human gut revealed by co-evolutionary constraints across the bacterial kingdom. Cell Syst.16:101167. doi: 10.1016/j.cels.2024.12.008

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 38

      DuprazL.OrianneG.Da CostaG.GauthierR.BlondeauA.BoinetM.et al. (2026). Inhibition of aryl hydrocarbon receptor interleukin-22 signaling and worsening of intestinal inflammation by Sutterella species. Gut Microbes18:2690688. doi: 10.1080/19490976.2026.2690688

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 39

      DuvalletC.GibbonsS. M.GurryT.IrizarryR. A.AlmE. J. (2017). Meta-analysis of gut microbiome studies identifies disease-specific and shared responses. Nat. Commun.8:1784. doi: 10.1038/s41467-017-01973-8

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 40

      FaithJ. J.ColombelJ.-F.GordonJ. I. (2015). Identifying strains that contribute to complex diseases through the study of microbial inheritance. Proc. Natl. Acad. Sci. USA112, 633–640. doi: 10.1073/pnas.1418781112

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 41

      FalonyG.JoossensM.Vieira-SilvaS.WangJ.DarziY.FaustK.et al. (2016). Population-level analysis of gut microbiome variation. Science352, 560–564. doi: 10.1126/science.aad3503

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 42

      FanY.PedersenO. (2021). Gut microbiota in human metabolic health and disease. Nat. Rev. Microbiol.19, 55–71. doi: 10.1038/s41579-020-0433-9

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 43

      FonsecaD. C.Marques Gomes da RochaI.Depieri BalmantB.CalladoL.Aguiar PrudêncioA. P.Tepedino Martins AlvesJ.et al. (2024). Evaluation of gut microbiota predictive potential associated with phenotypic characteristics to identify multifactorial diseases. Gut Microbes16:2297815. doi: 10.1080/19490976.2023.2297815

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 44

      FuggleR.MatiasM. G.Mayer-PintoM.MarzinelliE. M. (2025). Multiple stressors affect function rather than taxonomic structure of freshwater microbial communities. NPJ Biofilms Microbiomes11:60. doi: 10.1038/s41522-025-00700-2

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 45

      Garach VélezI.Ortuño GuzmánF. M.Rojas RuizI.Herrera MaldonadoL. J. (2025). Exploring the role of normalization and feature selection in microbiome disease classification pipelines. Gigascience14:giaf096. doi: 10.1093/gigascience/giaf096

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 46

      GhannamR. B.TechtmannS. M. (2021). Machine learning applications in microbial ecology, human microbiome studies, and environmental monitoring. Comput. Struct. Biotechnol. J.19, 1092–1107. doi: 10.1016/j.csbj.2021.01.028

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 47

      GhoshT. S.ShanahanF.O’TooleP. W. (2022). Toward an improved definition of a healthy microbiome for healthy aging. Nat. Aging2, 1054–1069. doi: 10.1038/s43587-022-00306-9

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 48

      GilibertiR.CavaliereS.MaurielloI. E.ErcoliniD.PasolliE. (2022). Host phenotype classification from human microbiome data is mainly driven by the presence of microbial taxa. PLoS Comput. Biol.18:e1010066. doi: 10.1371/journal.pcbi.1010066

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 49

      GloorG. B.MacklaimJ. M.Pawlowsky-GlahnV.EgozcueJ. J. (2017). Microbiome datasets are compositional: and this is not optional. Front. Microbiol.8:2224. doi: 10.3389/fmicb.2017.02224

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 50

      GoodswenS. J.BarrattJ. L. N.KennedyP. J.KauferA.CalarcoL.EllisJ. T. (2021). Machine learning and applications in microbiology. FEMS Microbiol. Rev.45:fuab015. doi: 10.1093/femsre/fuab015

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 51

      GreenerJ. G.KandathilS. M.MoffatL.JonesD. T. (2022). A guide to machine learning for biologists. Nat. Rev. Mol. Cell Biol.23, 40–55. doi: 10.1038/s41580-021-00407-0

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 52

      GuptaV. K.KimM.BakshiU.CunninghamK. Y.DavisJ. M.LazaridisK. N.et al. (2020). A predictive index for health status using species-level gut microbiome profiling. Nat. Commun.11:4635. doi: 10.1038/s41467-020-18476-8

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 53

      HeY.ShaoyongW.ChenY.LiM.GanY.SunL.et al. (2026). The functions of gut microbiota-mediated bile acid metabolism in intestinal immunity. J. Adv. Res.80, 351–370. doi: 10.1016/j.jare.2025.05.015

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 54

      HenkeM. T.KennyD. J.CassillyC. D.VlamakisH.XavierR. J.ClardyJ. (2019). Ruminococcus gnavus, a member of the human gut microbiome associated with Crohn’s disease, produces an inflammatory polysaccharide. Proc. Natl. Acad. Sci. USA116, 12672–12677. doi: 10.1073/pnas.1904099116

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 55

      HigashiD. L.KriegerM. C.QinH.ZouZ.PalmerE. A.KrethJ.et al. (2023). Who is in the driver’s seat? Parvimonas micra: An understudied pathobiont at the crossroads of dysbiotic disease and cancer. Environ. Microbiol. Rep.15, 254–264. doi: 10.1111/1758-2229.13153

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 56

      HouK.WuZ.-X.ChenX.-Y.WangJ.-Q.ZhangD.XiaoC.et al. (2022). Microbiota in health and diseases. Sig Transduct Target Ther7, 135–128. doi: 10.1038/s41392-022-00974-4

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 57

      HsiehT. C.MaK. H.ChaoA.McInernyG. (2016). iNEXT: an R package for rarefaction and extrapolation of species diversity (Hill numbers). Methods Ecol. Evol.7, 1451–1456. doi: 10.1111/2041-210X.12613

      • CrossRef
      • Google Scholar
    • 58

      HuangS.ChaudhariD. S.ShuklaR.KananiP.ZeidanR. S.LinY.et al. (2026). Global microbiome: Core and unique signatures across diverse populations. Int. J. Mol. Sci.27:1776. doi: 10.3390/ijms27041776

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 59

      JiangY.CheL.LiS. C. (2025). Deciphering the personalized functional redundancy hierarchy in the gut microbiome. Microbiome14:17. doi: 10.1186/s40168-025-02273-w

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 60

      JinD.-M.MortonJ. T.BonneauR. (2024). Meta-analysis of the human gut microbiome uncovers shared and distinct microbial signatures between diseases. mSystems9, e0029524–e0029524. doi: 10.1128/msystems.00295-24

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 61

      JoosR.BoucherK.LavelleA.ArumugamM.BlaserM. J.ClaessonM. J.et al. (2025). Examining the healthy human microbiome concept. Nat. Rev. Microbiol.23, 192–205. doi: 10.1038/s41579-024-01107-0

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 62

      KamadaN.SeoS.-U.ChenG. Y.NúñezG. (2013). Role of the gut microbiota in immunity and inflammatory disease. Nat. Rev. Immunol.13, 321–335. doi: 10.1038/nri3430

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 63

      KarwowskaZ.AasmetsO.Estonian Biobank research teamKosciolekT.OrgE. (2025). Effects of data transformation and model selection on feature importance in microbiome classification data. Microbiome13:2. doi: 10.1186/s40168-024-01996-6

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 64

      KimH. S.OhS. J.KimB. K.KimJ. E.KimB.-H.ParkY.-K.et al. (2024). Dysbiotic signatures and diagnostic potential of gut microbial markers for inflammatory bowel disease in Korean population. Sci. Rep.14:23701. doi: 10.1038/s41598-024-74002-6

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 65

      KnightsD.CostelloE. K.KnightR. (2011). Supervised classification of human microbiota. FEMS Microbiol. Rev.35, 343–359. doi: 10.1111/j.1574-6976.2010.00251.x

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 66

      KraszewskiS.SzczurekW.SzymczakJ.RegułaM.NeubauerK. (2021). Machine learning prediction model for inflammatory bowel disease based on laboratory markers. Working model in a discovery cohort study. J. Clin. Med.10:4745. doi: 10.3390/jcm10204745

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 67

      KubinskiR.Djamen-KepaouJ.-Y.ZhanabaevT.Hernandez-GarciaA.BauerS.HildebrandF.et al. (2022). Benchmark of data processing methods and machine learning models for gut microbiome-based diagnosis of inflammatory bowel disease. Front. Genet.13:784397. doi: 10.3389/fgene.2022.784397

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 68

      KumarB.LorussoE.FossoB.PesoleG. (2024). A comprehensive overview of microbiome data in the light of machine learning applications: categorization, accessibility, and future directions. Front. Microbiol.15:1343572. doi: 10.3389/fmicb.2024.1343572

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 69

      LaiS.YanY.PuY.LinS.QiuJ.-G.JiangB.-H.et al. (2023). Enterotypes of the human gut mycobiome. Microbiome11:179. doi: 10.1186/s40168-023-01586-y

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 70

      LaPierreN.JuC. J.-T.ZhouG.WangW. (2019). MetaPheno: a critical evaluation of deep learning and machine learning in metagenome-based disease prediction. Methods166, 74–82. doi: 10.1016/j.ymeth.2019.03.003

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 71

      LayeghifardM.HwangD. M.GuttmanD. S. (2017). Disentangling interactions in the microbiome: a network perspective. Trends Microbiol.25, 217–228. doi: 10.1016/j.tim.2016.11.008

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 72

      LeeS.LeeI. (2024). Comprehensive assessment of machine learning methods for diagnosing gastrointestinal diseases through whole metagenome sequencing data. Gut Microbes16:2375679. doi: 10.1080/19490976.2024.2375679

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 73

      LeeS.PortlockT.Le ChatelierE.Garcia-GuevaraF.ClasenF.OñateF. P.et al. (2024). Global compositional and functional states of the human gut microbiome in health and disease. Genome Res.34, 967–978. doi: 10.1101/gr.278637.123

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 74

      LeeJ.-Y.YooJ.-H.KimJ. E.BaeJ.-W.LeeC. K. (2026). Translating gut microbiota into diagnostics: a multidimensional approach for the diagnosis of inflammatory bowel disease. Gut and Liver20, 199–212. doi: 10.5009/gnl250360

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 75

      LiP.LiM.ChenW.-H. (2025). Best practices for developing microbiome-based disease diagnostic classifiers through machine learning. Gut Microbes17:2489074. doi: 10.1080/19490976.2025.2489074

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 76

      LiM.LiuJ.ZhuJ.WangH.SunC.GaoN. L.et al. (2023). Performance of gut microbiome as an independent diagnostic tool for 20 diseases: cross-cohort validation of machine-learning classifiers. Gut Microbes15:2205386. doi: 10.1080/19490976.2023.2205386

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 77

      LiaoH.JiY.SunY. (2023). High-resolution strain-level microbiome composition analysis from short reads. Microbiome11:183. doi: 10.1186/s40168-023-01615-w

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 78

      LiuY.FachrulM.InouyeM.MéricG. (2024). Harnessing human microbiomes for disease prediction. Trends Microbiol.32, 707–719. doi: 10.1016/j.tim.2023.12.004

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 79

      LiuG.ZhangX.LaiX.ZhuX.SuL.WangJ. (2026). A comparative study of machine learning models for microbiome-based diagnosis and multi-class staging of colorectal cancer. Sci. Rep. doi: 10.1038/s41598-026-53441-3

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 80

      Lloyd-PriceJ.Abu-AliG.HuttenhowerC. (2016). The healthy human microbiome. Genome Med.8:51. doi: 10.1186/s13073-016-0307-y

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 81

      LoC.MarculescuR. (2019). MetaNN: accurate classification of host phenotypes from metagenomic data using neural networks. BMC Bioinformatics20:314. doi: 10.1186/s12859-019-2833-2

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 82

      LoaizaL. I. S.Fernández-EdreiraD.Liñares-BlancoJ.CepedaA.Cardelle-CobasA.Fernandez-LozanoC. (2025). Fecal microbiome analysis in patients with metabolic syndrome and type 2 diabetes. PeerJ13:e19108. doi: 10.7717/peerj.19108

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 83

      LundbergS.LeeS.-I. (2017). A unified approach to interpreting model predictions. arXiv.org. Available online at: https://arxiv.org/abs/1705.07874v2 (Accessed October 4, 2024)

      • Google Scholar
    • 84

      MaS.ShunginD.MallickH.SchirmerM.NguyenL. H.KoldeR.et al. (2022). Population structure discovery in meta-analyzed microbial communities and inflammatory bowel disease using MMUPHin. Genome Biol.23:208. doi: 10.1186/s13059-022-02753-4

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 85

      MaffeisC.MartinaA.CorradiM.QuarellaS.NoriN.TorrianiS.et al. (2016). Association between intestinal permeability and faecal microbiota composition in Italian children with beta cell autoimmunity at risk for type 1 diabetes. Diabetes Metab. Res. Rev.32, 700–709. doi: 10.1002/dmrr.2790

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 86

      MajzoubM. E.SantiagoF. S.RaichS. S.SirigeriP.SimovicI.TedlaN.et al. (2026). Immunoglobulin a protease from Sutterella wadsworthensis modifies outcome of infection with campylobacter jejuni and is associated with microbiome diversity. Gut Microbes18:2611543. doi: 10.1080/19490976.2025.2611543

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 87

      MallickH.RahnavardA.McIverL. J.MaS.ZhangY.NguyenL. H.et al. (2021). Multivariable association discovery in population-scale meta-omics studies. PLoS Comput. Biol.17:e1009442. doi: 10.1371/journal.pcbi.1009442

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 88

      McBurneyM. I.DavisC.FraserC. M.SchneemanB. O.HuttenhowerC.VerbekeK.et al. (2019). Establishing what constitutes a healthy human gut microbiome: state of the science, regulatory considerations, and future directions. J. Nutr.149, 1882–1895. doi: 10.1093/jn/nxz154

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 89

      MeadowsV.AntonioJ. M.FerrarisR. P.GaoN. (2025). Ruminococcus gnavus in the gut: driver, contributor, or innocent bystander in steatotic liver disease?FEBS J.292, 1252–1264. doi: 10.1111/febs.17327

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 90

      MolteniC.ForniD.CaglianiR.SironiM. (2024). Comparative genomics reveal a novel phylotaxonomic order in the genus Fusobacterium. Commun. Biol.7:1102. doi: 10.1038/s42003-024-06825-y

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 91

      MoyaA.FerrerM. (2016). Functional redundancy-induced stability of gut microbiota subjected to disturbance. Trends Microbiol.24, 402–413. doi: 10.1016/j.tim.2016.02.002

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 92

      MukhopadhyaI.HansenR.NichollC. E.AlhaidanY. A.ThomsonJ. M.BerryS. H.et al. (2011). A comprehensive evaluation of colonic mucosal isolates of Sutterella wadsworthensis from inflammatory bowel disease. PLoS One6:e27076. doi: 10.1371/journal.pone.0027076

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 93

      MurrosK. E.HuynhV. A.TakalaT. M.SarisP. E. J. (2021). Desulfovibrio Bacteria are associated With Parkinson’s disease. Front. Cell. Infect. Microbiol.11:652617. doi: 10.3389/fcimb.2021.652617

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 94

      NearingJ. T.DouglasG. M.HayesM. G.MacDonaldJ.DesaiD. K.AllwardN.et al. (2022). Microbiome differential abundance methods produce different results across 38 datasets. Nat. Commun.13:342. doi: 10.1038/s41467-022-28034-z

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 95

      NieS.JingZ.WangJ.DengY.ZhangY.YeZ.et al. (2023). The link between increased Desulfovibrio and disease severity in Parkinson’s disease. Appl. Microbiol. Biotechnol.107, 3033–3045. doi: 10.1007/s00253-023-12489-1

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 96

      OhM.ZhangL. (2020). DeepMicro: deep representation learning for disease prediction based on microbiome data. Sci. Rep.10:6026. doi: 10.1038/s41598-020-63159-5

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 97

      PanM.Hidalgo-CantabranaC.BarrangouR. (2020). Host and body site-specific adaptation of Lactobacillus crispatus genomes. NAR Genom Bioinform2:lqaa001. doi: 10.1093/nargab/lqaa001

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 98

      PanzettaM. E.Valdi. Akkermansia in the gastrointestinal tract as a modifier of human health. Gut Microbes16:2406379. doi: 10.1080/19490976.2024.2406379

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 99

      PapoutsoglouG.TarazonaS.LopesM. B.KlammsteinerT.IbrahimiE.EckenbergerJ.et al. (2023). Machine learning approaches in microbiome research: challenges and best practices. Front. Microbiol.14:1261889. doi: 10.3389/fmicb.2023.1261889

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 100

      PasolliE.TruongD. T.MalikF.WaldronL.SegataN. (2016). Machine learning Meta-analysis of large metagenomic datasets: tools and biological insights. PLoS Comput. Biol.12:e1004977. doi: 10.1371/journal.pcbi.1004977

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 101

      PekelS.KarcherN.EssexM.SpringerF.RomanoS.DucarmonQ. R.et al. (2026). Meta-analysis reveals microbiome signatures for colorectal cancer that are universal across age groups and sequencing methods. Cell Host Microbe34, 1462–1476.e5. doi: 10.1016/j.chom.2026.05.030

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 102

      PiccinnoG.ThompsonK. N.ManghiP.GhaziA. R.ThomasA. M.Blanco-MíguezA.et al. (2025). Pooled analysis of 3,741 stool metagenomes from 18 cohorts for cross-stage and strain-level reproducible microbial biomarkers of colorectal cancer. Nat. Med.31, 2416–2429. doi: 10.1038/s41591-025-03693-9

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 103

      PietrucciD.TeofaniA.MilanesiM.FossoB.PutignaniL.MessinaF.et al. (2022). Machine learning data analysis highlights the role of Parasutterella and Alloprevotella in autism Spectrum disorders. Biomedicine10:2028. doi: 10.3390/biomedicines10082028

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 104

      PorcariS.MullishB. H.AsnicarF.NgS. C.ZhaoL.HansenR.et al. (2025). International consensus statement on microbiome testing in clinical practice. Lancet Gastroenterol. Hepatol.10, 154–167. doi: 10.1016/S2468-1253(24)00311-X

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 105

      PriyaS.BurnsM. B.WardT.MarsR. A. T.AdamowiczB.LockE. F.et al. (2022). Identification of shared and disease-specific host gene–microbiome associations across human diseases using multi-omic integration. Nat. Microbiol.7, 780–795. doi: 10.1038/s41564-022-01121-z

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 106

      QuinceC.WalkerA. W.SimpsonJ. T.LomanN. J.SegataN. (2017). Shotgun metagenomics, from sampling to analysis. Nat. Biotechnol.35, 833–844. doi: 10.1038/nbt.3935

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 107

      RanalliM. G.SalvatiN.PetrellaL.PantaloneF. (2023). M-quantile regression shrinkage and selection y and traffic on air quality. Biom. J.65:e2100355. doi: 10.1002/bimj.202100355

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 108

      RomanoS.WirbelJ.AnsorgeR.SchudomaC.DucarmonQ. R.NarbadA.et al. (2025). Machine learning-based meta-analysis reveals gut microbiome alterations associated with Parkinson’s disease. Nat. Commun.16:4227. doi: 10.1038/s41467-025-56829-3

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 109

      RosenbergE. (2024). Diversity of bacteria within the human gut and its contribution to the functional unity of holobionts. NPJ Biofilms Microbiomes10:134. doi: 10.1038/s41522-024-00580-y

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 110

      SanzH.ValimC.VegasE.OllerJ. M.ReverterF. (2018). SVM-RFE: selection and visualization of the most relevant features through non-linear kernels. BMC Bioinformatics19:432. doi: 10.1186/s12859-018-2451-4

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 111

      SczyrbaA.HofmannP.BelmannP.KoslickiD.JanssenS.DrögeJ.et al. (2017). Critical assessment of metagenome interpretation-a benchmark of metagenomics software. Nat. Methods14, 1063–1071. doi: 10.1038/nmeth.4458

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 112

      ShanahanF.GhoshT. S.O’TooleP. W. (2021). The healthy microbiome-what is the definition of a healthy gut microbiome?Gastroenterology160, 483–494. doi: 10.1053/j.gastro.2020.09.057

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 113

      StatnikovA.HenaffM.NarendraV.KongantiK.LiZ.YangL.et al. (2013). A comprehensive evaluation of multicategory classification methods for microbiomic data. Microbiome1:11. doi: 10.1186/2049-2618-1-11

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 114

      StaudacherH. M.LoughmanA. (2021). Gut health: definitions and determinants. Lancet Gastroenterol. Hepatol.6:269. doi: 10.1016/S2468-1253(21)00071-6

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 115

      SteyerbergE. W.HarrellF. E. (2016). Prediction models need appropriate internal, internal-external, and external validation. J. Clin. Epidemiol.69, 245–247. doi: 10.1016/j.jclinepi.2015.04.005

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 116

      SuQ.LiuQ.LauR. I.ZhangJ.XuZ.YeohY. K.et al. (2022). Faecal microbiome-based machine learning for multi-class disease diagnosis. Nat. Commun.13:6818. doi: 10.1038/s41467-022-34405-3

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 117

      SunW.ZhangY.GuoR.ShaS.ChenC.UllahH.et al. (2024). A population-scale analysis of 36 gut microbiome studies reveals universal species signatures for common diseases. NPJ Biofilms Microbiomes10:96. doi: 10.1038/s41522-024-00567-9

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 118

      SusinA.WangY.Lê CaoK.-A.CalleM. L. (2020). Variable selection in microbiome compositional data analysis. NAR Genom Bioinform2:lqaa029. doi: 10.1093/nargab/lqaa029

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 119

      TegegneH. A.SavidgeT. C. (2025). Gut microbiome metagenomics in clinical practice: bridging the gap between research and precision medicine. Gut Microbes17:2569739. doi: 10.1080/19490976.2025.2569739

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 120

      TeixeiraM.SilvaF.FerreiraR. M.PereiraT.FigueiredoC.OliveiraH. P. (2024). A review of machine learning methods for cancer characterization from microbiome data. npj Precis. Onc.8, 1–16. doi: 10.1038/s41698-024-00617-7

      • CrossRef
      • Google Scholar
    • 121

      ThomasA. M.ManghiP.AsnicarF.PasolliE.ArmaniniF.ZolfoM.et al. (2019). Metagenomic analysis of colorectal cancer datasets identifies cross-cohort microbial diagnostic signatures and a link with choline degradation. Nat. Med.25, 667–678. doi: 10.1038/s41591-019-0405-7

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 122

      TianL.WangX.-W.WuA.-K.FanY.FriedmanJ.DahlinA.et al. (2020). Deciphering functional redundancy in the human microbiome. Nat. Commun.11:6217. doi: 10.1038/s41467-020-19940-1

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 123

      TierneyB. T.TanY.YangZ.ShuiB.WalkerM. J.KentB. M.et al. (2022). Systematically assessing microbiome-disease associations identifies drivers of inconsistency in metagenomic research. PLoS Biol.20:e3001556. doi: 10.1371/journal.pbio.3001556

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 124

      TopçuoğluB. D.LesniakN. A.RuffinM. T.WiensJ.SchlossP. D. (2020). A framework for effective application of machine learning to microbiome-based classification problems. MBio11:e00434-20. doi: 10.1128/mBio.00434-20

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 125

      van der SommenF.de GroofJ.StruyvenbergM.van der PuttenJ.BoersT.FockensK.et al. (2020). Machine learning in GI endoscopy: practical guidance in how to interpret a novel field. Gut69, 2035–2045. doi: 10.1136/gutjnl-2019-320466

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 126

      van DijkL. R.WalkerB. J.StraubT. J.WorbyC. J.GroteA.SchreiberH. L.et al. (2022). StrainGE: a toolkit to track and characterize low-abundance strains in complex microbial communities. Genome Biol.23:74. doi: 10.1186/s13059-022-02630-0

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 127

      Van HulM.CaniP. D.PetifilsC.De VosW. M.TilgH.El OmarE. M. (2024). What defines a healthy gut microbiome?Gut:gutjnl-2024-333378. doi: 10.1136/gutjnl-2024-333378

      • CrossRef
      • Google Scholar
    • 128

      Van RossumT.FerrettiP.MaistrenkoO. M.BorkP. (2020). Diversity within species: interpreting strains in microbiomes. Nat. Rev. Microbiol.18, 491–506. doi: 10.1038/s41579-020-0368-1

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 129

      Vieira-SilvaS.FalonyG.DarziY.Lima-MendezG.Garcia YuntaR.OkudaS.et al. (2016). Species–function relationships shape ecological properties of the human gut microbiome. Nat. Microbiol.1:16088. doi: 10.1038/nmicrobiol.2016.88

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 130

      VisciG.NotarioE.DefazioG.CaratozzoloM. F.CoxS. N.FossoB.et al. (2026). Benchmarking short- and long-read sequencing technologies for metagenomic profiling of microbiomes. Sci. Rep.16:22610. doi: 10.1038/s41598-026-49725-3

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 131

      Vujkovic-CvijinI.SklarJ.JiangL.NatarajanL.KnightR.BelkaidY. (2020). Host variables confound gut microbiota studies of human disease. Nature587, 448–454. doi: 10.1038/s41586-020-2881-9

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 132

      WadeK. H.HallL. J. (2019). Improving causality in microbiome research: can human genetic epidemiology help?Wellcome Open Res4:199. doi: 10.12688/wellcomeopenres.15628.3

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 133

      WalshC.Stallard-OliveraE.FiererN. (2023). Nine (not so simple) steps: a practical guide to using machine learning in microbial ecology. MBio15, e0205023–e0205023. doi: 10.1128/mbio.02050-23

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 134

      WangS.AlmeidaA.MuD.ZengS. (2026). Spatially resolved architecture of the human gut microbiome and its health implications. The Lancet Microbe7:101389. doi: 10.1016/j.lanmic.2026.101389

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 135

      WangS.MuL.YuC.HeY.HuX.JiaoY.et al. (2023). Microbial collaborations and conflicts: unraveling interactions in the gut ecosystem. Gut Microbes16:2296603. doi: 10.1080/19490976.2023.2296603

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 136

      WidderS.AllenR. J.PfeifferT.CurtisT. P.WiufC.SloanW. T.et al. (2016). Challenges in microbial ecology: building predictive understanding of community function and dynamics. ISME J.10, 2557–2568. doi: 10.1038/ismej.2016.45

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 137

      WilkinsL. J.MongaM.MillerA. W. (2019). Defining Dysbiosis for a cluster of chronic diseases. Sci. Rep.9:12918. doi: 10.1038/s41598-019-49452-y

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 138

      WilliamsC. E.HammerT. J.WilliamsC. L. (2024). Diversity alone does not reliably indicate the healthiness of an animal microbiome. ISME J.18:wrae133. doi: 10.1093/ismejo/wrae133

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 139

      WilmanskiT.DienerC.RappaportN.PatwardhanS.WiedrickJ.LapidusJ.et al. (2021). Gut microbiome pattern reflects healthy ageing and predicts survival in humans. Nat. Metab.3, 274–286. doi: 10.1038/s42255-021-00348-0

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 140

      WirbelJ.PylP. T.KartalE.ZychK.KashaniA.MilaneseA.et al. (2019). Meta-analysis of fecal metagenomes reveals global microbial signatures that are specific for colorectal cancer. Nat. Med.25, 679–689. doi: 10.1038/s41591-019-0406-6

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 141

      WirbelJ.ZychK.EssexM.KarcherN.KartalE.SalazarG.et al. (2021). Microbiome meta-analysis and cross-disease comparison enabled by the SIAMCAT machine learning toolbox. Genome Biol.22:93. doi: 10.1186/s13059-021-02306-1

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 142

      WoodD. E.LuJ.LangmeadB. (2019). Improved metagenomic analysis with kraken 2. Genome Biol.20:257. doi: 10.1186/s13059-019-1891-0

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 143

      WuH.CaiL.LiD.WangX.ZhaoS.ZouF.et al. (2018). Metagenomics biomarkers selected for prediction of three different diseases in Chinese population. Biomed. Res. Int.2018, 1–7. doi: 10.1155/2018/2936257

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 144

      WuG.XuT.ZhaoN.LamY. Y.DingX.WeiD.et al. (2024). A core microbiome signature as an indicator of health. Cell187, 6550–6565.e11. doi: 10.1016/j.cell.2024.09.019

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 145

      WuF.ZhuB.FengS.LiH.ZhouJ.NingY.et al. (2026). The brain–gut health initiative (BIGHI): a prospective cohort on psychiatric disorders in China. Research9:1142. doi: 10.34133/research.1142

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 146

      XuJ.YangM.WangD.ZhangS.YanS.ZhuY.et al. (2020). Alteration of the abundance of Parvimonas micra in the gut along the adenoma-carcinoma sequence. Oncol. Lett.20:106. doi: 10.3892/ol.2020.11967

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 147

      YanB.NamY.LiL.DeekR. A.LiH.MaS. (2025). Recent advances in deep learning and language models for studying the microbiome. Front. Genet.15:1494474. doi: 10.3389/fgene.2024.1494474

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 148

      YanR.ZhengR.HanY.SongG.HuoB.SunH. (2025). Meta-analysis of gut microbiome reveals patterns of dysbiosis in colorectal cancer patients. J. Med. Microbiol.74:002042. doi: 10.1099/jmm.0.002042

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 149

      YangL.ChenJ. (2023). Benchmarking differential abundance analysis methods for correlated microbiome sequencing data. Brief. Bioinform.24:bbac607. doi: 10.1093/bib/bbac607

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 150

      YuW.QuH.WangS.ShiM.LuY.XiaS.et al. (2026). From microbiomes to predictive ecosystems: challenges and opportunities in artificial intelligence-based approaches. The Lancet Microbe0:101428. doi: 10.1016/j.lanmic.2026.101428

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 151

      YuY.ZhangN.MaiY.RenL.ChenQ.CaoZ.et al. (2023). Correcting batch effects in large-scale multiomics studies using a reference-material-based ratio method. Genome Biol.24:201. doi: 10.1186/s13059-023-03047-z

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 152

      ZengS.AlmeidaA.MuD.WangS. (2026). Temporal variations of the gut microbiome in human health. The Lancet Microbe7:101388. doi: 10.1016/j.lanmic.2026.101388

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 153

      ZhangC.ZhaoL. (2016). Strain-level dissection of the contribution of the gut microbiome to human metabolic disease. Genome Med.8:41. doi: 10.1186/s13073-016-0304-1

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 154

      ZhaoY.YangH.WuP.YangS.XueW.XuB.et al. (2024). Akkermansia muciniphila: a promising probiotic against inflammation and metabolic disorders. Virulence15:2375555. doi: 10.1080/21505594.2024.2375555

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 155

      ZhaoL.ZhangX.ZhouY.FuK.LauH. C.-H.ChunT. W.-Y.et al. (2022). Parvimonas micra promotes colorectal tumorigenesis and is associated with prognosis of colorectal cancer patients. Oncogene41, 4200–4210. doi: 10.1038/s41388-022-02395-7

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 156

      ZhengJ.SunQ.ZhangM.LiuC.SuQ.ZhangL.et al. (2024). Noninvasive, microbiome-based diagnosis of inflammatory bowel disease. Nat. Med.30, 3555–3567. doi: 10.1038/s41591-024-03280-4

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 157

      ZhouY.-H.GallinsP. (2019). A review and tutorial of machine learning methods for microbiome host trait prediction. Front. Genet.10:579. doi: 10.3389/fgene.2019.00579

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 158

      ZhouT.ZhaoF. (2025). AI-empowered human microbiome research. Gut75, 1432–1446. doi: 10.1136/gutjnl-2025-335946

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 159

      ZhuB.ChenS.DiaoY.WangW.HuangY.LiangL.et al. (2026). Dissecting the ecological structure of health and disease in the global gut microbiome. Adv. Sci.13:e17087. doi: 10.1002/advs.202517087

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 160

      ZhuJ.XieH.YangZ.ChenJ.YinJ.TianP.et al. (2023). Statistical modeling of gut microbiota for personalized health status monitoring. Microbiome11:184. doi: 10.1186/s40168-023-01614-x

      • Pubmed Abstract
      • CrossRef
      • Google Scholar
    • 161

      ZielińskaK.UdekwuK. I.RudnickiW.FrolovaA.ŁabajP. P. (2025). Healthy microbiome—moving towards functional interpretation. Gigascience14:giaf015. doi: 10.1093/gigascience/giaf015

      • Pubmed Abstract
      • CrossRef
      • Google Scholar

    Summary

    Keywords

    dysbiosis, gut microbiome, health-disease classification, machine learning, microbial biomarkers, microbiome-based prediction, shotgun metagenomics

    Citation

    Kumar B, Lorusso E, Fosso B and Pesole G (2026) A data-driven universal gut microbiome health assessment: a machine learning framework trained on large metagenomic data. Front. Microbiol. 17:1925500. doi: 10.3389/fmicb.2026.1925500

    Received

    01 July 2026

    Revised

    22 August 2026

    Accepted

    24 August 2026

    Published

    10 September 2026

    Volume

    17 – 2026

    Edited by

    Edoardo Pasolli, University of Naples Federico II, Italy

    Reviewed by

    Guang Liu, Xi’an Jiaotong University, China

    Baoyuan Zhu, South China University of Technology, China

    Updates

    Copyright

    © 2026 Kumar, Lorusso, Fosso and Pesole

    This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.

    Disclaimer

    All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher

    datadriven Frontiers Health microbiome Universal
    HealthJustfine Team
    • Website
    • Facebook

    Related Posts

    Why More Women Are Prioritizing Preventive Health at Biograph – Grit Daily News

    September 12, 2026

    Beyond Digestion: The Hidden Power Of The Gut In Ageing, Immunity And Cancer Care

    September 12, 2026

    Mental Health And Wellness Block Party Coming To Elmhurst

    September 12, 2026
    Leave A Reply Cancel Reply

    Don't Miss
    Mental Wellness

    Suns’ Dillon Brooks says youth camp focuses on mental wellness, hoops

    By HealthJustfine TeamSeptember 13, 20260

    SUNS Suns’ Dillon Brooks says youth camp focuses on mental wellness, hoops Jorge I. GuajardoArizona Republic Sept. 12, 2026, 6:04 p.m. MT

    Red Bank Catholic High School and Conscientia Health Open New Wellness Room, Marking Expanded Commitment to Student Mental and Physical Health

    September 13, 2026

    Mental wellness: NIMHANS to be nodal body for BRICS Centres of Excellence

    September 13, 2026

    Why ‘anti-ageing’ products are disappearing from shelves, and what’s taking their place

    September 13, 2026
    Stay In Touch
    • Facebook
    • Twitter
    • Pinterest
    • Instagram
    • YouTube
    • Vimeo
    Our Picks

    Expert shares 6 tips to recover faster and stronger after intense workout sessions- Moneycontrol.com

    June 28, 2026

    These Viral Fitness & Wellness Recovery Products Are Taking Over TikTok Ahead of Prime Day

    June 28, 2026

    Life Time Has Created a Fitness and Recovery Paradise – Muscle & Fitness

    June 28, 2026

    The Movement Twenty Four: New 24-Hour Fitness and Recovery Hub Opens Down South

    June 28, 2026

    Subscribe to Updates

    Get the latest creative news from SmartMag about art & design.

    About Us

    Welcome to HealthJustFine.com, your trusted destination for reliable health news, wellness insights, and evidence-based information that empowers you to live a healthier life.
    Our mission is to make quality health information accessible, easy to understand, and relevant for everyone. We believe that staying informed is the first step toward making better decisions about your health, nutrition, fitness, and overall well-being. That’s why we deliver timely updates on the latest medical research, healthy living trends, preventive care, and wellness innovations from around the world.

    Our Picks

    Suns’ Dillon Brooks says youth camp focuses on mental wellness, hoops

    September 13, 2026

    Red Bank Catholic High School and Conscientia Health Open New Wellness Room, Marking Expanded Commitment to Student Mental and Physical Health

    September 13, 2026

    Mental wellness: NIMHANS to be nodal body for BRICS Centres of Excellence

    September 13, 2026
    Latest Posts

    Expert shares 6 tips to recover faster and stronger after intense workout sessions- Moneycontrol.com

    June 28, 2026

    These Viral Fitness & Wellness Recovery Products Are Taking Over TikTok Ahead of Prime Day

    June 28, 2026

    Life Time Has Created a Fitness and Recovery Paradise – Muscle & Fitness

    June 28, 2026
    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Contact Us
    • Terms & Conditions
    • Privacy Policy
    • Disclaimer

    © 2026 healthjustfine.com. All rights reserved. Designed by DD.

    Type above and press Enter to search. Press Esc to cancel.