# De Novo Design of NRPS Thiolation Domains Using Generative AI: A Platform for Modular Peptide Biosynthesis
## Introduction
Nonribosomal peptide synthetases (NRPSs) are massive enzymatic machines that build structurally complex peptides without relying on ribosomes. Each NRPS module typically contains an adenylation (A) domain responsible for selecting amino acid substrates, a thiolation (T) domain that carries the activated amino acid intermediate, and a condensation (C) domain that catalyzes bond formation. The T-domain, also known as the peptidyl carrier protein (PCP), is often considered a central hub of interdomain communication — it must interact with upstream and downstream catalytic partners in a highly specific, state-dependent manner throughout the assembly cycle.
Despite decades of structural and mechanistic characterization, engineering T-domains de novo has remained a formidable challenge. The functional constraints of these domains extend beyond simple substrate recognition to include precise geometric and chemical complementarity with flanking A- and C-domains across multiple catalytic states. Traditional mutagenesis approaches, while informative, explore sequence space in a localized and often low-throughput manner.
Recent advances in generative artificial intelligence have opened new possibilities for protein design. Large language models trained on protein sequences, diffusion-based generative frameworks, and energy-based protein design algorithms now offer complementary strategies for proposing novel protein sequences tailored to specific structural contexts. However, the extent to which AI-generated sequences can recapitulate or improve the function of complex biosynthetic enzymes in vivo had not been systematically demonstrated.
Here, we describe an integrated design-build-test-learn (DBTL) platform that leverages generative AI to create de novo T-domains for NRPS assembly lines, followed by high-throughput in vivo validation in a split, bipartite system. Our approach demonstrates that AI-designed thiolation domains can support peptide biosynthesis across minimal, full-length, repositioned, and chimeric NRPS architectures — and that their functional performance is shaped by the identity of downstream partner domains in a manner that mirrors natural evolutionary patterns.
## The Type S GxpS Bipartite Platform
To enable rapid, quantitative evaluation of T-domain variants in a native NRPS context, we developed a modified version of the GameXPeptide synthetase (GxpS) from *Photorhabdus luminescens* TTO1, a bacterium associated with entomopathogenic nematodes. The native GxpS assembly line produces a family of cyclic nonribosomal peptides known as GameXPeptides, including the well-characterized pentapeptide cyclo(vLfLL).
Our platform, termed Type S, partitions the native GxpS assembly line across two independently expressed protein subunits that reconstitute in vivo through engineered synthetic zipper (SZ) interactions. These zippers are built from high-affinity coiled-coil pairs inspired by leucine zipper motifs, enabling controlled and efficient subunit association. The system provides a two-tier functional readout:
**Tier 1 — SU2-only prescreen.** When only the downstream subunit (SU2, comprising the A3–TE5 region) is expressed, the system produces two dominant linear tripeptides. The promiscuous A3-domain incorporates either phenylalanine or leucine, generating a mixture of products including di-D-L and di-L-D configured tripeptides. This configuration serves as a high-throughput first-pass screen for T-domain function, since product formation requires successful thiolation, adenylation, and interdomain transfer within a simplified module context.
**Tier 2 — Full assembly line validation.** Coexpression of SU2 with the upstream subunit (SU1, comprising A1–C3) reconstitutes the complete five-module assembly line. The primary readout is the macrocyclic pentapeptide cyclo(vLfLL), whose production demands functional interactions across all modules and all catalytic states, including thiolation, condensation donor, and condensation acceptor interfaces.
T-domain variants are introduced through a streamlined Golden Gate cloning workflow that replaces a BsaI-flanked counterselection cassette in the SU2 construct with candidate T-domain sequences. This design enables rapid, modular swapping of T-domains at the T3 position and facilitates iterative rounds of design and testing.
## Generative Design Strategies
Candidate T-domain sequences were generated using three complementary generative approaches, each offering distinct strengths:
– **ESM3**, a large-scale evolutionary model, generates sequences using structure tokens — encoded representations of protein structural information — allowing it to condition generation on three-dimensional context.
– **EvoDiff**, a diffusion-based generative framework, provided a sequence-only baseline that explored the local sequence neighborhood without explicit structural input.
– **ProteinMPNN**, an inverse folding model, generated sequences conditioned on backbone coordinates, enabling direct control over the three-dimensional fold of the T-domain.
In all cases, generation was guided by an explicit local-context constraint: the flanking A and C domain sequences (and structural coordinates where available) were provided as conditioning information, while the T-domain segment itself was masked during sequence generation. This ensured that proposed designs were tailored to the specific junction environment of interest.
As the campaign progressed, we incorporated progressive candidate prioritization using lightweight in silico filters and, in later rounds, data-driven surrogate machine-learning models trained on experimentally measured sequence–activity relationships. These filters considered motif conservation, sequence identity to wild-type sequences, perplexity scores from language models, and predicted functional likelihood.
## Iterative Design Rounds and Functional Validation
### Round 1: Proof of Concept
The first design round tested 11 AI-generated T-domain variants at the T3 position of the Type S GxpS scaffold. Two generative models were used: ESM3 and ProteinMPNN. Notably, ESM3-derived designs uniformly retained the canonical phosphopantetheinylation motif (FFxxGGHS), which is essential for post-translational activation of the T-domain through attachment of the 4′-phosphopantetheine cofactor. In contrast, ProteinMPNN designs exhibited variation at the residue immediately preceding the catalytic serine.
Four of the four ESM3-derived variants supported detectable production of the prescreen tripeptides, with one variant exceeding wild-type output by approximately 15%. Among the seven ProteinMPNN-derived variants, only two were functional. Analysis of sequence features revealed that functional ProteinMPNN designs retained naturally occurring residues at the critical position preceding the phosphopantetheinylation serine, consistent with the known dependence of this position on the downstream C-domain identity.
### Round 2: Expanded Model Coverage
The second round expanded the generative toolkit to include EvoDiff and introduced simple sequence-based filtering criteria informed by round 1 outcomes. Candidate sequences were required to retain the conserved FFxxGGxS motif, maintain a minimum of 50% sequence identity to wild-type GxpS_T3, and exhibit low perplexity scores as assessed by the ESM2 650M language model. These filters increased the fraction of functional designs from the initial generative pool.
### Round 3: Surrogate-Guided Prioritization
The third round incorporated surrogate models trained on sequence–activity data from rounds 1 and 2, along with an expanded set of error-prone polymerase chain reaction (epPCR)-derived variants. This data-driven approach enabled the ranking and prioritization of newly generated sequences based on their predicted likelihood of functional performance. The steady increase in the proportion of moderately and highly active variants across rounds demonstrated the value of iterative learning within the DBTL cycle.
## Benchmarking with Error-Prone Mutagenesis
After round 2, we used epPCR to benchmark classical local sequence exploration of the wild-type T3-domain in parallel with diversification of three AI-designed scaffolds. This comparison revealed strong scaffold dependence in epPCR outcomes. The wild-type T3 library showed modest improvement, with 24% of variants exceeding wild-type performance and 16% yielding no detectable product.
The AI-2 scaffold — an early design that had exceeded wild-type activity in the prescreen — proved especially amenable to epPCR refinement, with 60% of its derived library variants outperforming wild-type and one variant reaching an extraordinary 285% of wild-type output. By contrast, AI-12 and AI-15 scaffolds showed more limited headroom for diversification, with lower fractions of improved variants and higher rates of functional loss.
These results demonstrate that AI-designed scaffolds can exhibit different mutational landscapes than wild-type sequences, with some designs offering substantially greater opportunity for local optimization. The combination of generative design and directed evolution thus emerges as a powerful strategy for exploring sequence space and maximizing functional output.
## Functionality in the Full-Length Assembly Line
A critical question in NRPS engineering is whether variants that perform well in simplified minimal systems will translate to the full-length context, where the T-domain must engage with multiple interdomain partners across different catalytic states. To address this, we selected the eight highest-performing AI-designed T-domains across all three rounds and model classes and installed them at the T3 position of the reconstituted full-length Type S GxpS assembly line.
All eight variants produced detectable pentapeptide, confirming that de novo T-domain sequences can maintain compatibility with both upstream and downstream interfaces in the complete module context. Two variants showed substantial improvements over wild-type titers: one reached 257% of wild-type output (7.4 mg/L), and another achieved 144% of wild-type (4.1 mg/L). Interestingly, some variants that had shown modest gains in the prescreen increased their relative performance in the full assembly line, while others did not, underscoring the importance of testing in the complete architectural context.
We also examined how cultivation conditions influenced titer measurements. The choice of culture medium had a dramatic effect: wild-type production increased approximately 17-fold when switching from succinate XPP medium to LB medium at the same culture volume. Across different culture volumes, AI-designed variants showed pattern-dependent responses, with some performing best at small scale and others at larger scales. These findings highlight the need for standardized cultivation conditions to ensure internal consistency across DBTL rounds and suggest that absolute titers should be interpreted as condition-dependent functional labels rather than intrinsic measures of sequence quality.
## Biochemical Characterization of AI-2
AI-2 was prioritized for detailed biochemical characterization based on its consistent performance across multiple assay conditions and its status as an early high-performing design. Comparative analysis of wild-type GxpS_T3 and AI-2 across several domain contexts — including isolated T-domains and multi-domain constructs — revealed that AI-2 consistently accumulated at higher levels in *E. coli* and was readily recovered in the soluble fraction, whereas the wild-type domain was predominantly insoluble and required denaturing purification.
Thermal denaturation experiments showed that AI-2 melted at approximately 52°C, compared with roughly 40°C for refolded wild-type GxpS_T3, indicating enhanced thermal robustness. Size-exclusion chromatography after denaturation and refolding further demonstrated that AI-2 recovered a native-like conformation efficiently, whereas the wild-type domain largely aggregated.
However, the improved biochemical tractability of AI-2 did not fully account for its functional performance in the complete assembly line, as it did not maximize titers across all cultivation conditions and architectural contexts. This separation between intrinsic developability (solubility, stability, expression level) and context-dependent productivity (peptide output in the full system) highlights the multifaceted nature of T-domain function in NRPS assembly lines.
## State-Dependent Interdomain Dynamics from Molecular Simulations
To understand the molecular basis of T-domain function and the differences between wild-type and AI-designed variants, we performed 100-nanosecond atomistic molecular dynamics (MD) simulations of both wild-type GxpS_T3 and AI-2 in three catalytic states: thiolation (T-domain interacting with the A3 domain), condensation donor (T-domain positioned to deliver the tethered intermediate to the downstream C/E4 domain), and condensation acceptor (T-domain engaging the upstream C3 domain).
Despite differences in specific interdomain contacts, both variants maintained comparable global structural stability, with the T-domain core exhibiting a root-mean-square deviation of approximately 1.5–1.8 Å across all states. The increased apparent motion observed in some simulations arose primarily from flexible linker regions rather than destabilization of folded domains.
Comparative contact analysis revealed state-dependent differences in how the two variants organized their interdomain interactions. In the thiolation state, wild-type GxpS_T3 engaged A3 through a distributed interface involving multiple helical and loop regions, whereas AI-2 adopted a more focused interaction dominated by helix 2. In the condensation donor state, AI-2 showed a greater contribution of polar and hydrogen-bond-mediated contacts, including interactions involving arginine residues on helices 2 and 3. In the condensation acceptor state, wild-type maintained a higher number of persistent contacts with the upstream C3 and A3 partners.
These observations are consistent with the emerging view that NRPS carrier domains engage their catalytic partners through networks of weak, transient interactions that balance specificity with dynamic exchange. They also help explain why T-domain compatibility is context-dependent: different partner domains and catalytic states select for different intercontact features, and a variant optimized for one set of interactions may not perform equally well in another.
## Positional Dependence Within the Assembly Line
Given the state-dependent interface differences observed in simulations, we next asked whether T-domain variants exhibit positional preferences when installed at different module locations within the GxpS assembly line. We selected three T-domain designs spanning different design rounds and functional profiles — AI-2 (an early robust scaffold), AI-12 (WT-like in the prescreen but underperforming in the full system), and AI-38 (a top-performing variant in both configurations) — and installed each at positions T1, T2, T4, and T5.
The results revealed pronounced and position-specific effects. At T1, AI-2 and AI-12 abolished detectable peptide production, while AI-38 retained approximately 24% of the native T1 control activity. At T2, AI-2 supported roughly 42% of native output, whereas AI-38 was non-functional and AI-12 performed poorly at approximately 16%. In contrast, T4 was broadly more permissive: AI-2 reached roughly 43% of native T4 levels, AI-12 exceeded the native control by approximately 37%, and AI-38 produced roughly 61% of native output. T5 was the most permissive position overall, with AI-2 and AI-12 exceeding the native T5 control by more than twofold, while AI-38 again failed to produce detectable product.
This pronounced positional dependence is consistent with the T-domain encoding partner- and state-specific interface features, as suggested by the MD simulations. It also reflects the identity of the immediate downstream partner domain at each position: at T3 and T4, the downstream partner is a dual C/E-type condensation domain, whereas at T2, the downstream partner is an L-type condensation domain with distinct stereochemical and mechanistic requirements. Our T-domain designs, which were generated in the T3 neighborhood where the downstream partner is C/E-type, showed reduced compatibility when transplanted to positions where the downstream partner is of the L-type, consistent with the known co-variation of T-domain features with downstream partner identity in natural NRPS systems.
## Engineering Chimeric NRPSs with AI-Designed T-Domains
A key question for practical NRPS engineering is whether AI-designed T-domains can serve as modular parts for constructing hybrid assembly lines from non-cognate biosynthetic gene clusters. To test this, we generated chimeric systems by recombining modules from GxpS with those from two distinct bacterial natural product pathways: the xenotetrapeptide synthetase (XtpS) from *Xenorhabdus nematophila* HGB081 and the szentiamide synthetase (SzeS) from *Xenorhabdus szentirmaii*.
These hybrid systems impose non-cognate A–T–C interfaces and therefore provide a stringent test of functional generalization. In the first GxpS–XtpS hybrid (GxhS-1), constructs carrying native GxpS_T3 produced only minimal peptide output. In contrast, constructs carrying AI-2, AI-27, and AI-31 yielded dramatic increases in product formation — improvements of more than three orders of magnitude in LC-MS signal — while constructs carrying AI-12 and AI-32 produced no detectable peptide.
In the GxpS–SzeS hybrid (GshS), the native SzeS T4 domain supported production of the pentapeptide szentiamide, and replacing it with AI-2 or AI-12 increased output by approximately 25% and 65%, respectively. Notably, AI-31 performed comparably to the native domain, while AI-27 and AI-32 reduced product formation, further illustrating the context-dependent nature of T-domain compatibility.
To test context-conditioned design more directly, we generated a second set of AI-designed T-domains in an L-type C-domain context and evaluated them in a GxpS–XtpS hybrid (GxhS-2) where the engineered carrier was positioned immediately before an L-type condensation domain. In this setting, only five of the 76 AI-designed variants remained functional, whereas constructs carrying native GxpS_T2 produced detectable output. The vast majority of designs generated in a C/E-domain context failed entirely in the L-type context.
Sequence-distance-based phylogenetic analysis of T-domains annotated by downstream partner identity revealed that AI-designed variants clustered predominantly with wild-type T-domains associated with their respective downstream partner class. This separation mirrors the functional patterns observed in the positional and hybrid experiments and suggests that generative models capture downstream-partner-specific features embedded in natural T-domain sequences — features that may not be fully apparent from sequence alone.
## Discussion
The results presented here establish a generalizable framework for de novo design of NRPS thiolation domains using generative artificial intelligence, validated through high-throughput in vivo functional assays across multiple NRPS architectures. Several key insights emerge from this work.
First, AI-generated T-domains are frequently functional, with the fraction of active designs increasing substantially across iterative DBTL rounds when generative proposals are combined with progressive candidate prioritization using in silico filters and data-driven surrogate models. This finding challenges the conventional assumption that de novo design of such complex, multifunctional protein domains is infeasible and demonstrates that modern generative models can navigate the sequence–function landscape of NRPS T-domains with meaningful success.
Second, T-domain function is strongly context-dependent, influenced by the identity of immediate downstream partner domains, the specific catalytic state under consideration, and the position of the T-domain within the assembly line. This context dependence is not a limitation but rather a feature of NRPS biology — it reflects the evolved modularity and specificity of interdomain communication in these megasynthetases. The fact that AI-designed variants capture this context dependence, with designs generated in one partner context performing poorly when transplanted to a non-cognate context, suggests that generative models learn biologically meaningful sequence–structure–function relationships even when trained on sequence alone.
Third, the combination of AI design with directed evolution (epPCR) can yield substantial functional gains, with some AI-designed scaffolds providing a higher baseline and greater headroom for optimization than wild-type sequences. This synergy between generative design and classical protein engineering represents a promising direction for optimizing not only T-domains but entire NRPS assembly lines.
Fourth, the enhanced expression, solubility, and thermal stability of certain AI-designed T-domains — as exemplified by AI-2 — indicate that generative design can simultaneously improve both the developability and the functional performance of biosynthetic enzymes. This dual benefit is particularly valuable for practical applications in combinatorial biosynthesis and synthetic biology, where soluble, stable, and highly expressed proteins dramatically simplify downstream processing and pathway optimization.
## Frequently Asked Questions (FAQ)
**Q1: What is an NRPS thiolation domain, and why is it important?**
A1: The thiolation (T) domain, also known as the peptidyl carrier protein (PCP), is a central component of nonribosomal peptide synthetases. It covalently attaches amino acid intermediates via a phosphopantetheine cofactor and delivers them to condensation domains for peptide bond formation. Because the T-domain must interact productively with upstream adenylation domains and downstream condensation domains across multiple catalytic states, it plays a critical role in determining the specificity and efficiency of peptide assembly.
**Q2: How does the Type S bipartite system enable high-throughput testing?**
A2: The Type S system splits the native GxpS assembly line into two independently expressed subunits that assemble in vivo through engineered synthetic zipper interactions. This architecture allows researchers to first test T-domain variants in a simplified, minimal context (SU2-only, producing linear tripeptides) before validating the best performers in the full, reconstituted assembly line (producing macrocyclic pentapeptides). The modular design and streamlined Golden Gate cloning workflow make it possible to rapidly swap T-domain sequences and quantify outputs.
**Q3: What generative AI models were used, and how did they differ?**
A3: Three complementary models were employed. ESM3 is an evolutionary model that uses structure tokens to condition sequence generation on three-dimensional context. EvoDiff is a diffusion-based generative framework that operates on sequence alone, providing a baseline for exploration. ProteinMPNN is an inverse folding model that generates sequences conditioned on specified backbone coordinates, enabling direct control over the T-domain fold. Each model brought different strengths, and their combined use increased the diversity and coverage of the explored sequence space.
**Q4: Can AI-designed T-domains outperform wild-type sequences?**
A4: Yes. Multiple AI-designed variants exceeded wild-type performance in both the minimal prescreen and the full-length assembly line. One variant reached 257% of wild-type output in the full system, and an epPCR-diversified library derived from AI-2 produced a variant at 285% of wild-type. Importantly, improved performance in the AI-designed scaffold was not solely attributable to higher expression or solubility, as variants showed position-dependent and context-dependent gains that reflected genuine functional optimization.
**Q5: Why do AI-designed T-domains show positional dependence?**
A5: Each position in an NRPS assembly line features a specific downstream partner domain with distinct structural and catalytic properties. T-domains from natural NRPSs have evolved to match their downstream partners through co-variation of interface residues. When AI-designed T-domains, which were generated in the context of a C/E-type downstream partner, are installed at positions where the downstream partner is of an L-type, the interface complementarity is often disrupted. This positional dependence mirrors natural evolutionary patterns and underscores the importance of matching T-domain design to the intended downstream partner.
**Q6: How does epPCR complement AI-based design?**
A6: Error-prone polymerase chain reaction (epPCR) introduces random mutations across the T-domain sequence, enabling classical directed evolution to explore local sequence neighborhoods. This approach is complementary to generative design: AI models propose novel sequences informed by broad sequence and structural knowledge, while epPCR enables fine-tuning and optimization of those designs or wild-type sequences. In our study, epPCR proved particularly effective on the AI-2 scaffold, yielding a high fraction of improved variants and the most active T-domain identified in the entire study.
**Q7: What are the practical implications for synthetic biology and combinatorial biosynthesis?**
A7: The ability to design functional T-domains de novo opens the door to building hybrid NRPS assembly lines that combine modules from different natural product pathways, enabling the synthesis of novel peptides with non-natural amino acid compositions or unusual structural features. The demonstrated dependence on downstream partner identity provides a design rule for matching T-domains to their intended C-domain partners, while the enhanced biochemical properties of certain AI-designed variants improve the tractability of heterologous expression in production hosts like *E. coli*.
**Q8: How were peptide products identified and quantified?**
A8: All peptide products were confirmed by tandem mass spectrometry (MS/MS). Quantification was performed using targeted liquid chromatography–mass spectrometry in selected reaction monitoring (LC-MS/MS SRM) mode with external calibration standards. In high-throughput screening campaigns, a RapidFire LC-MS system enabled semi-quantitative assessment of peptide production across large numbers of constructs.
## Conclusion
This work establishes that generative artificial intelligence can be used to design functional thiolation domains for nonribosomal peptide synthetases from scratch, with in vivo validation across minimal, full-length, repositioned, and chimeric assembly line architectures. The integrated DBTL workflow — coupling conditional generative design with high-throughput in vivo screening and iterative machine learning — demonstrates that AI-designed biosynthetic enzymes can match or exceed wild-type performance while offering improved biochemical tractability. The pronounced context dependence of T-domain function, governed primarily by the identity of downstream partner domains, provides both a design constraint and a design principle: successful T-domain engineering must account for the specific catalytic and interaction environment in which the domain will operate. These findings expand the toolkit available for combinatorial biosynthesis and synthetic biology and highlight the growing role of AI-driven protein design in unlocking the biosynthetic potential of modular enzyme assembly lines. Thank you for reading.



