How artificial intelligence is revealing thousands of previously unknown bacterial antiviral defense systems, and why they could become the next generation of biotechnology tools
TL;DR: Two recently published science studies demonstrate how protein and genomic language models can identify thousands of previously unknown bacterial antiviral defense systems that traditional homology searches miss. Beyond expanding our understanding of microbial immunity, these discoveries could uncover entirely new biotechnology tools. just as restriction enzymes and CRISPR revolutionized molecular biology.
From Bacterial–Phage Arms Races to Biotechnology Breakthroughs
Bacteria are in a constant evolutionary arms race with viruses (bacteriophages), driving the evolution of remarkably diverse bacterial immunity and bacterial defense systems against bacteriophage infection. Previous discoveries from this battle, including restriction enzymes and CRISPR, have transformed biotechnology. Now, two studies published in Science demonstrate how artificial intelligence, machine learning, protein language models, and genomic language models can identify previously unknown antiphage defense systems that traditional homology searches miss. These discoveries not only expand our understanding of innate immunity in bacteria but may also provide the next generation of biotechnology tools.
Until recently, studies describing bacterial antiviral defense systems rarely attracted widespread attention, despite the discovery of transformative technologies such as restriction enzymes and, later, CRISPR/Cas systems. Today, advances in the field have become major scientific news because they continue to reveal new biology with enormous potential for biotechnology (see Gene Editing Revamped with CRISPR Prime Editor – Result Is A High Efficiency, More Precise Gene Editor With Versatile Editing Capabilities And Lower Off-Target Effects). The two independent studies published in Science by DeWeirdt et al. (2026) and Mordret et al., (2026) combine protein and genomic language models with machine learning to dramatically expand the known universe of bacterial antiviral defense systems.
“Bacterial antiviral defenses represent a largely unexplored source of new molecular biology tools.”
Bacteria are in a constant battle with bacteriophages; in some environments, phage action can result in a significant daily turnover of 10% to 25% (DeWeirdt et al. 2026). Consequently, it is estimated that approximately 0.5% of the bacterial pangenome encodes defense systems against these bacteriophages (Mordret et al., 2026). Historically, searches for defense proteins or systems have relied on one-off directed studies or by sequence homology screens . More recently, it has been noted that many prokaryotic defense proteins and multi-gene operons reside in so-called “defense islands” within the bacterial genome, as well as in mobile genetic elements (MGEs) and integrons.
Because many components of eukaryotic innate immune systems are analogous or even directly homologous to components of prokaryotic defense systems, advances in this area could offer valuable insights into the detailed functions of our own immune system. Furthermore, restriction enzymes and CRISPR may only be the tip of the iceberg for new biotechnology tools. Two new studies have now begun leveraging large language models (LLMs) to search for these dense proteins wherever they reside in bacterial genomes and associated elements.
Defense Predictor: Using Protein Language Models to Discover Bacterial Immune Systems
To train the model, the researchers analyzed approximately 17,000 prokaryotic reference genomes (DeWeirdt et al., 2026). They identified around 244,000 homologs belonging to defense systems for the positive dataset, and approximately 14 million genes with non-defensive annotations for the negative dataset. After clustering and data reduction, the final dataset yielded 15,000 defense proteins and 186,000 control proteins.
Using the Evolutionary Scale Model 2 (ESM2) protein language model (PLM), (Lin et al. 2023), the team represented each protein alongside its four nearest neighbor genes. This approach captured both individual protein features and genomic context, resulting in over 3,000 dimensions in the final DefensePredictor model. In terms of performance, DefensePredictor achieved an average precision of 0.86 in predicting defense genes. For comparison, a random baseline scores 0.1, a structure-based homology approach scored 0.38, and ESM2 alone scored 0.51.
When applied to the genomes of 69 diverse Escherichia coli isolates, Defense Predictor identified 624 high-confidence defense protein clusters. Notably:
- 512 of these 624 clusters were identified by DefensePredictor but missed by DefenseFinder (the previous standard).
- Slightly more than 100 clusters (of the 624) shared no homology to known defense proteins, meaning they would be missed by simpler homology-only search approaches.
- Nearly half of the identified proteins were found outside of traditional “hotspot” locations (like defense islands, prophages, and plasmids), demonstrating this approach’s superiority to “guilt by location” methods alone.
To validate these predictions, 94 of the predicted defense systems were cloned into E. coli and challenged with a panel of phages. Of these, 42 provided protection against at least one of the 24 phages tested. Within the 42 validated systems, 15 protein domains previously unassociated with phage defense were identified. Several of these proteins were examined in greater detail (See Figure1):
- DS-8: A single protein containing a metalloproteinase domain which is homologous to the human protein SMPDL3A (a phosphodiesterase involved in cGAS-STING immunity signaling).
- DS-11: A system containing both CBS and HEPN domains. CBS domains are prevalent in bacteria, archaea and eukaryotes, but were not previously known in defense proteins, while HEPN domains are related to Cas13 nuclease domains.
- DS-6A: A two-protein system featuring a predicted S/T-kinase related to the toxin HipA. Mutation studies suggest DS-6A functions as a toxin-antitoxin system analogous to HipAB.
Figure 1: Genomic and domain architecture of Defense Predictor newly validated defense systems plus protein structures of a key subset. Image credit: DeWeirdt et al. (2026)
Even this small sample of detailed analyses highlights how much these newly uncovered defense systems can teach us about novel immune mechanisms (DeWeirdt et al., 2026).
Comparing Protein and Genomic Language Models for Antiphage Defense Discovery
In contrast to other methods, Mordret and team (2026) took an incremental approach to search for new prokaryotic defense proteins.. Their first model, a natural language protocol called ALBERTDF, A Lite Bidirectional Encoder Representations from Transformers, was trained to assert defensive function using only local genomic context. The training set consisted of 10,796 genomes from the Actinomycetota phylum. In this type of model, each protein family is treated as a “word” and each stretch of surrounding genes is a “sentence,” with the model attempting to understand the “grammar” of the genome. While natural language models typically use tens of thousands of words, genomic context is orders of magnitude larger, which represents a limitation. To address this, the authors limited their input data to one phylum and restricted the vocabulary to approximately 500,000 of the most common protein families. Using ALBERTDF, the investigators successfully predict 1,930 protein families as antiphage candidates. Ten of these were selected for experimental validation, cloned into a Streptomycin albus model strain, and challenged with a phage panel. Although two products were toxic under the test conditions, six of the remaining eight successfully protected against at least one phage. Ultimately, this demonstrates that despite some scale limitations, this genomic-context-based language model can effectively identify antiphage candidates, even those that lack obvious homology to known systems.
Mordret and team (2026) developed an ESM2 refinement to leverage this protein language model without needing to predefine protein families or apply “grammar” restrictions. They used approximately 500,000 proteins annotated by DefenseFinder, an algorithm that identifies defense proteins, as positive training data, and 1.4 million core genes with asserted nondefensive functions as negatives. The top-performing model from this family, ESM-650MDF, was used to predict candidate defensive proteins from previously uncharacterized proteins. Out of ten candidates selected for experimental validation, six were successful.
To combine the benefits of both sequence- and context-based approaches, Mordret et al. (2026) created GeneCLRDF (contrast learning of visual representation), which integrates features from both ALBERTDF and their ESM mode (See Figure 2 for a summary)l. When analyzed against benchmarks, GeneCLRDF was their most accurate and precise model, achieving 99% precision and 92% recall. They then applied GeneCLRDF to a collection of more than 32,000 bacterial genomes. Based on this, they estimate that approximately 1.5% of the genes in a typical genome are devoted to antiphage defense – a significant increase from older estimates. Interestingly, the fraction varied widely from 6.5% in Helicobacter pylori to 0% Chlamydia trachomatis. In this run, the model predicts 2.39 million antiphage proteins, with 85% having no prior link to immune function. A substantial number of these new defense candidates occur as single gene/protein operons with no previously asserted defense function. These findings indicate a prokaryotic antiviral immunity universe that is both deeper and more diverse than previously thought.

Figure 2: This image depicts a scientific workflow and results for an integrated genomic and protein machine learning model named GeneCLRDF, designed to predict bacterial antiphage defense systems. Includes the diagram of model construction, phylogenetic trees distribution of antiphage domains, and examples of genomic regions annotated with different methods (Image credit: Mordret et al. (2026))
Future Applications for Biotechnology, CRISPR, and Antimicrobial Research
Based on our discussion of the recent works by DeWeirdt et al. (2026) and Mordret et al. (2026), it is clear that there is an enormous, largely uncharacterized, genomic storehouse of novel prokaryotic proteins and systems evolved for defense against phage infections. Investigating these systems will teach us a great deal about the regulation and evolution of microbiomes. Furthermore, because both groups only sampled currently available genomes, the ongoing accumulation of the microbial pangenome sequence will undoubtedly reveal even more novel systems for discovery. Ultimately, understanding these novel mechanisms will both expand our knowledge of the microbial world and uncover new facets of eukaryotic innate immunity.
Since this field is still expanding rapidly, few companies are currently known to be developing products related to bacterial defense systems. However, as noted in our 2025 blog post on the fight against rising antimicrobial resistance (see Recent Advances in Overcoming Antimicrobial Resistance), several companies are working on phage-related treatments for resistant infections. Gaining a better understanding of antiphage defense systems will be critical in designing more effective phage therapies moving forward.
Just as restriction enzymes and CRISPR emerged from studying bacterial defenses against viruses, today’s AI-driven exploration of bacterial immunity may uncover the next generation of molecular biology tools. As protein and genomic language models continue to improve and more microbial genomes become available, our understanding of bacterial immune systems is likely still in its earliest stages.
“The hidden universe of bacterial immunity may hold both the next generation of biotechnology tools and new strategies for combating antibiotic-resistant infections.”
References
Lin et al., Evolutionary-scale prediction of atomic-level protein structure with a language model. (2023) Science, Mar 17;379(6637):1123-1130. doi: 10.1126/science.ade2574. Epub 2023 Mar 16.
Mordret et al., Protein and genomic language models uncover the unexplored diversity of bacterial immunity. (2026) Science, Apr 2;392(6793):eadv8275. doi: 10.1126/science.adv8275. Epub 2026 Apr 2.
DeWeirdt et al., DefensePredictor: A machine learning model to discover prokaryotic immune systems. (2026) Science, Apr 2;392(6793):eadv7924. doi: 10.1126/science.adv7924. Epub 2026 Apr 2.






