# A Machine-Learning Framework for Predicting Bacteriophage–Host Interactions Using Genomic Features
## Introduction
Bacteriophages — viruses that infect bacteria — represent one of the most promising frontiers in antimicrobial therapy, biotechnology, and microbial ecology. The ability to predict which phages can infect which bacterial hosts is critical for designing effective phage cocktails, understanding co-evolutionary dynamics, and deploying phage-based interventions in clinical and agricultural settings. Yet predicting phage–host interactions remains a formidable challenge, given the enormous diversity of both bacterial and viral genomes and the complex molecular mechanisms governing infection specificity.
Recent advances in machine learning offer powerful tools for uncovering patterns in large biological datasets. In the work described here, researchers developed and systematically evaluated a machine-learning framework capable of predicting phage–host interactions from genomic and protein-level features. By leveraging multiple datasets spanning diverse bacterial genera and phage families, the study aimed to identify generalizable predictors of infection and to provide a practical toolkit for phage discovery and engineering.
This article provides a comprehensive overview of the methodology, key findings, and practical implications of this predictive framework, including answers to frequently asked questions and a discussion of broader significance.
—
## Dataset Compilation and Curation
### Data Sources and Composition
The analysis drew on six published phage–host interaction datasets, each containing binary phenotype data — indicating whether a given phage could infect a given bacterial strain — alongside the corresponding complete genomes. Four of these datasets were designated for model development, while the remaining two were held out exclusively for validation purposes.
The development datasets covered a wide range of biological systems. One dataset focused on *Escherichia coli*, comprising over 16,000 interactions across nearly 200 strains and approximately 90 distinct phages. A second dataset pooled multiple *Klebsiella* species, yielding roughly 3,600 interactions from about 60 bacterial strains and 60 phages. A third dataset was restricted to *Klebsiella pneumoniae* specifically, containing over 6,300 interactions from 138 strains and 46 phages. Finally, a dataset centered on *Pseudomonas* species included several hundred interactions across two dozen strains and nineteen phages. Together, these development datasets encompassed more than 27,000 individual interactions, of which only a small fraction — roughly 2 to 4 percent in some cases and over 36 percent in others — represented confirmed positive infections.
The validation datasets were equally diverse. They included a large mixed Vibrionaceae dataset with over 63,000 interactions involving more than 250 strains and 248 phages, an independent *Klebsiella* dataset with several thousand interactions, and a complete *E. coli* dataset with nearly 38,000 interactions. By testing on these independent datasets, the researchers could assess how well the trained models would generalize to entirely new bacterial–phage pairs.
### Phenotype Definition and Assay Conditions
A notable distinction among the source datasets concerns how infection was measured. The *E. coli* dataset was generated using solid culture medium, where bacterial lawn clearance by phage dilutions served as the readout for infection. In that dataset, any observable clearance was classified as a positive interaction. In contrast, all other datasets were generated using liquid culture, where the area under an optical density curve over time quantified bacterial growth inhibition. Each of these datasets applied its own author-defined threshold for distinguishing positive from negative outcomes, introducing some variability in how “infection” was operationalized across the aggregated data.
—
## Phylogenetic Characterization of Strains and Phages
### Bacterial Phylogenomics
To understand how evolutionary relationships among bacterial hosts influence infection patterns, the researchers constructed phylogenetic trees for each dataset. They began by identifying the core genome — the set of genes shared by every strain in a given dataset — using a pangenome analysis tool. After extracting these conserved genes and aligning their protein sequences, they built maximum-likelihood phylogenetic trees using a standard tree-building algorithm. Pairwise distances between all strains were then calculated from the resulting trees, producing a comprehensive distance matrix for each bacterial group.
### Phage Evolutionary Distances
For phages, evolutionary relationships were assessed using protein-family similarity scores that quantify how much gene content two phage genomes share. These similarity scores were converted into distance metrics so that phages with highly overlapping gene catalogs were considered closely related, while those with divergent gene content were placed farther apart. Additionally, phage gene-sharing networks were constructed to visualize clusters of phages with shared proteins, helping to identify groups of related viral agents across datasets.
### Phylogenetic Isolation Metric
To quantify how evolutionarily distinct each bacterial strain or phage was from its closest relatives, the researchers computed a “phylogenetic isolation” score for every entity in each dataset. This score represents the average distance from a given strain or phage to its five nearest neighbors on the phylogenetic tree. By excluding self-distance and focusing on local rather than global tree position, this metric captures the immediate evolutionary context of each organism. The team then tested whether phylogenetic isolation correlated with prediction accuracy, finding that more isolated entities did not necessarily pose greater challenges for the model.
—
## Feature Engineering from Genomic Data
### Protein Family Construction
The first major category of features was derived from protein families. All proteins encoded in the genomes under study were clustered based on sequence similarity using a high-performance sequence comparison tool. The clustering process used a sequence identity threshold of 40 percent and a coverage requirement of 80 percent, ensuring that only proteins sharing substantial similarity were grouped together. Singleton clusters — those present in only a single genome — were removed so that the resulting features would reflect shared, rather than unique, biological content.
This workflow produced a binary matrix where each row represented a genome (bacterial strain or phage) and each column represented a protein family, with entries indicating whether that family was present or absent in a given genome. Redundant features with identical presence–absence patterns across all genomes were consolidated to reduce dimensionality.
A systematic optimization process tested multiple sequence identity thresholds to determine the best clustering stringency. Importantly, the results showed that the choice of clustering threshold had minimal impact on model performance, providing confidence that the feature set was robust to parameter variation. An alternative approach using a different homology search tool followed by graph-based clustering was also evaluated and produced comparable results, but the primary method was retained for its superior computational efficiency.
### K-mer Based Features
The second category of features was built from short amino acid subsequences called *k*-mers. These are sliding-window extracts of length *k* from every protein sequence in a genome. The researchers tested *k*-mer lengths ranging from 3 to 15 amino acids, finding that *k* = 6 offered the best balance of predictive power and consistency across training runs. However, when working with complete proteomes rather than protein families, a shorter *k* value of 4 was preferred, as longer *k* values would have produced feature tables too large for practical computation during iterative model training.
*K*-mer features were filtered to remove those present in only a single gene, and identical patterns were consolidated. For new genomes encountered during prediction, *k*-mers were assigned to features if they appeared in at least a small fraction of the proteins belonging to a given protein family, ensuring that feature assignment remained biologically meaningful even when sequences were not perfectly identical to reference data.
—
## Model Selection and Training
### Algorithms Compared
Eight distinct machine-learning algorithms were benchmarked across all development datasets. These included gradient-boosted decision trees (specifically CatBoost), random forest, k-nearest neighbors, logistic regression, multilayer perceptrons (a type of neural network), naïve Bayes, support vector machines, and standard gradient boosting. Each algorithm was subjected to exhaustive grid search over algorithm-specific hyperparameter ranges, ensuring that every model was tuned for optimal performance.
### Handling Class Imbalance
A central challenge in this work is that positive phage–host interactions are rare relative to negatives, with positive rates ranging from roughly 2 percent to over 36 percent depending on the dataset. To address this imbalance, the researchers implemented phage-specific class weighting. For each phage, a weight was calculated based on the ratio of negative to positive interactions it participates in, so that phages with few known infections would be given greater importance during training. Two weighting schemes were tested — one based on logarithmic scaling and another based on inverse frequency — with the inverse-frequency approach being incorporated into the final workflow.
### Ensemble Strategy and Feature Selection
Rather than relying on a single model, the final predictions were generated by aggregating outputs across multiple training iterations, using the median confidence score to improve stability and reduce the risk of overfitting. Six different feature selection methods were systematically tested, and recursive feature elimination — which iteratively removes the least important features — emerged as the best-performing approach in the majority of datasets, while also being computationally efficient.
The optimal feature selection process used the best-performing algorithm as the underlying estimator and tracked which features were retained across many rounds of selection. Features that appeared consistently across iterations were considered stable predictors, while those selected only in specific data splits were treated with caution.
### Dataset Balancing Strategy
To further guard against overfitting and ensure generalizability, the researchers compared fourteen different strategies for handling dataset imbalance and strain heterogeneity. The winning combination involved hierarchical clustering of bacterial strains into twenty groups, applying the inverse-frequency class weights per phage, and requiring that predictive features be present in at least two different strain clusters. This ensured that features unique to a single group of bacteria were filtered out, leaving only those with broader relevance.
—
## Evaluation Framework
### Cross-Validation Design
Model generalization was assessed using three distinct experimental configurations, each designed to probe a different aspect of biological transferability. In the first configuration, the model was trained on 90 percent of bacterial strains (with all available phages) and tested on the remaining 10 percent of strains, simulating the challenge of predicting infection for a bacterium that has never been encountered before. In the second configuration, the model was trained on all strains but only 90 percent of phages, testing whether it could predict the host range of a previously unseen phage. The most demanding configuration trained the model on interactions between 90 percent of strains and 90 percent of phages (covering roughly 81 percent of all possible interactions), then validated it on the interactions between the held-out 10 percent of strains and the held-out 10 percent of phages, representing the most realistic scenario of novel discovery.
### Performance Metrics
Because of the class imbalance inherent in phage–host interaction data, the primary performance metric was the Matthews Correlation Coefficient (MCC), which provides a balanced assessment that takes into account all four categories of prediction outcome — true positives, true negatives, false positives, and false negatives. The MCC ranges from −1 (perfect disagreement) to +1 (perfect agreement), with zero indicating performance no better than random chance.
In addition to MCC, the researchers reported several other metrics, including accuracy, precision, recall, F1-score, and the area under the receiver operating characteristic curve (AUROC), which evaluates the model’s ability to rank positive interactions above negative ones regardless of the classification threshold. For imbalanced datasets, the area under the precision–recall curve (AUPR) was also computed and normalized against a random baseline, enabling fair comparison across datasets with different positive rates.
### Calibration Assessment
To evaluate how well the model’s predicted probabilities matched observed frequencies, calibration analysis was performed. Predictions were binned into ten intervals, and the average predicted probability within each bin was compared to the actual fraction of positive interactions. A perfectly calibrated model would have all points lying on the diagonal line where predicted probability equals observed frequency. The Brier score — which measures the mean squared difference between predicted probabilities and actual outcomes — was also calculated, with lower values indicating better overall calibration and accuracy.
—
## Key Findings
### Best-Performing Model
Across all development datasets, CatBoost gradient-boosted decision trees consistently achieved the highest or near-highest MCC scores while requiring fewer computational resources than several other algorithms tested. This combination of strong predictive performance and efficiency made CatBoost the algorithm of choice for the final models.
### Feature Importance and Biological Interpretability
By computing Shapley values — a method for attributing each prediction to its individual input features — the researchers were able to identify which protein families and *k*-mer patterns most strongly influenced infection predictions. Features were classified as promoting infection (positive impact) or reducing it (negative impact) based on their average contribution across all predictions.
The predictive features identified by the model showed significant overlap with genes previously validated as mediators of phage–host interactions through high-throughput genetic screens such as transposon mutagenesis and CRISPR interference experiments. This convergence between computational predictions and experimental validation provides strong evidence that the model is capturing genuine biological determinants of phage susceptibility, including genes involved in surface structures, defense systems, and metabolic pathways relevant to viral entry and replication.
### Error Analysis
A comprehensive error analysis framework was employed to identify the types of predictions the model gets wrong and why. The analysis included examination of high-confidence misclassifications (cases where the model was very certain but incorrect), systematic biases (phages or strains that were consistently over- or under-predicted), and cases where predicted interaction patterns diverged sharply from experimentally observed patterns. These analyses revealed that the model’s most challenging cases often involved phages or strains with highly divergent evolutionary histories, suggesting that phylogenetic novelty can complicate prediction — though not insurmountably.
### Phage Cocktail Design
Beyond predicting individual interactions, the framework was applied to the practical problem of designing multi-phage cocktails. The researchers evaluated ten different cocktail design strategies that varied in how phages were ranked (by model confidence or by observed promiscuity in training data), how diversity was enforced (no clustering, hierarchical clustering, or density-based clustering), and how phages were represented (by their interaction profiles or by genomic features). Cocktail performance was measured by the fraction of target strains for which at least one phage in the cocktail was predicted to be effective. The results highlighted the importance of diversity-aware selection, showing that clustering-based strategies often outperformed simple rank-based selection, especially at larger cocktail sizes.
—
## Experimental Validation
### Laboratory Testing of Predictions
To confirm that the model’s predictions held up in real biological experiments, the researchers selected 52 phages from the Basel Bacteriophage Selection (BASEL) collection and tested them against 25 *E. coli* strains from the ECOR reference collection. Phage suspensions were prepared at standardized titers and spotted onto bacterial lawns using an automated liquid-handling system. After overnight incubation, plaque formation was scored manually and compared against model predictions.
A complementary experiment assessed whether the observed clearance phenotypes resulted from productive phage infection or from non-productive killing mechanisms such as mechanical lysis or endolysin activity. By measuring the efficiency of plating — the ratio of phage titer on a test strain relative to the amplification host — the researchers could distinguish productive from non-productive interactions and better understand the biological basis of the model’s binary predictions.
### Genetic Validation Using Transposon Mutagenesis
In a separate line of experimental validation, the team used transposon insertion sequencing (RB-TnSeq) in *E. coli* ECOR27 to identify genes essential for phage infection or resistance. A barcoded mutant library covering over 4,000 genes was exposed to a panel of phages, and the survival of each mutant was quantified by sequencing. Genes showing strong enrichment or depletion in the final pool were considered critical for phage susceptibility or resistance, respectively. The overlap between these experimentally validated genes and the model’s computationally identified predictive features confirmed that the model was capturing biologically meaningful signals.
—
## Frequently Asked Questions (FAQ)
**Q1: What is a phage–host interaction prediction model, and why is it useful?**
A: A phage–host interaction prediction model is a computational tool that uses genomic and protein-level features of bacteria and bacteriophages to forecast whether a given phage can infect a given bacterial strain. Such models are valuable for accelerating phage discovery, guiding the design of phage-based therapeutics, and reducing the need for labor-intensive experimental screening of every possible phage–host combination.
**Q2: What kinds of genomic features does the model use?**
A: The model primarily relies on two types of features: protein family presence–absence profiles and amino acid *k*-mer patterns. Protein families are groups of evolutionarily related proteins identified by clustering sequences across all genomes in the dataset. *K*-mers are short subsequences of amino acids extracted from proteins, which capture local sequence patterns that may be relevant to infection.
**Q3: How does the model handle the fact that most phages do not infect most bacteria?**
A: The model addresses class imbalance through phage-specific class weighting, which gives greater importance to rare positive interactions during training. Additionally, performance is evaluated using metrics such as the Matthews Correlation Coefficient (MCC) that are designed to be informative even when one class is much more common than the other.
**Q4: Can this model predict interactions for completely novel bacteria or phages that were not part of the training data?**
A: Yes. The model is designed to assign features to new genomes by searching for similarity to the protein clusters and *k*-mer patterns established during training. This allows features to be transferred to novel organisms even when they share only partial sequence identity with the reference data, simulating realistic inference conditions.
**Q5: How accurate are the predictions?**
A: Prediction accuracy varies across datasets depending on factors such as the proportion of positive interactions, the genetic diversity of the organisms involved, and the degree of phylogenetic novelty. Across the datasets tested, the CatBoost-based models consistently achieved strong MCC scores, and the framework demonstrated robust generalization to independent validation datasets.
**Q6: What are the limitations of the approach?**
A: The model predicts binary infection outcomes based on genomic features alone and does not capture every nuance of phage–host biology, such as environmental conditions, multi-phage dynamics, or the mechanistic details of infection. Additionally, the binary phenotype used in training may conflate productive infection with non-productive lysis, which can have different biological implications.
**Q7: How can this framework be applied to phage therapy?**
A: The framework can be used to pre-screen large numbers of phages against a target pathogen, identify cocktails of phages with complementary host ranges, and prioritize phages for experimental testing. The cocktail design strategies evaluated in the study specifically address the practical challenge of assembling effective multi-phage treatments that maximize bacterial coverage.
**Q8: Is the software and methodology publicly available?**
A: The machine-learning package, analysis workflows, and downstream processing tools described in this work are available through public repositories, enabling other researchers to apply, modify, and extend the framework for their own phage–host prediction needs.
—
## Conclusion
The development of machine-learning models for predicting phage–host interactions represents a significant step forward in the field of phage biology and its practical applications. By integrating genomic features derived from both protein families and short sequence patterns, and by systematically evaluating model performance across phylogenetically diverse datasets, this framework demonstrates that computational prediction can reliably identify infection outcomes across a wide range of bacterial–phage systems.
The convergence of computational predictions with independent experimental validation — including laboratory phage susceptibility assays and transposon-based genetic screens — strengthens confidence that the identified features capture genuine biological determinants of infection. Moreover, the extension of the framework to phage cocktail design illustrates the practical utility of these models in therapeutic development, where selecting the right combination of phages is critical for treatment success.
While challenges remain — including the need to incorporate additional biological variables, to refine predictions at the mechanistic level, and to continuously update models as new genomes and interaction data become available — the approach presented here provides a robust, publicly accessible foundation for accelerating phage discovery and application. As antibiotic resistance continues to pose a growing threat to global health, predictive tools for phage–host interactions will play an increasingly important role in the development of phage-based interventions.
Thank you for reading.



