# How Pharmaceutical Companies Are Training the Next Generation of AI Protein-Folding Models
### A growing shortage of public data is pushing drugmakers to share their most closely guarded molecular secrets—and the results are already paying off.
**Artificial intelligence tools that predict how proteins fold have become indispensable in modern drug discovery, but they face a critical limitation: they were built on data that simply doesn’t capture how proteins interact with drug-like molecules.** Now, a landmark collaboration among pharmaceutical companies is tackling this gap head-on by pooling thousands of proprietary molecular structures into a single, shared training set—and the early results are remarkable.
—
## The Data Bottleneck Holding Back AI Drug Discovery
AI models like AlphaFold made headlines in 2024 when their creator, DeepMind, was awarded the Nobel Prize in Chemistry for achieving near-perfect accuracy in predicting how proteins fold into their three-dimensional shapes. These breakthroughs were built on the Protein Data Bank, a massive public repository containing more than 200,000 experimentally determined protein structures obtained through techniques like X-ray crystallography and cryo-electron microscopy.
However, the Protein Data Bank has a blind spot. While it is rich in structures of proteins on their own, it contains relatively few examples of proteins interacting with small-molecule drug candidates—known as ligands. Industry estimates suggest there are fewer than 10,000 such structures available publicly. For AI systems tasked with predicting how potential drugs will bind to their protein targets, this is a serious shortfall.
Research has shown that the predictive accuracy of co-folding models—AI systems designed to simulate how a protein and a small molecule interact—degrades sharply when presented with molecular pairs that look very different from those in the training data. In practical terms, this means a model trained mostly on publicly available data may struggle to predict how an experimental drug compound will behave when it encounters a novel protein target.
The structures pharmaceutical companies generate during their own internal research programs are precisely the data that fill this gap. Yet most of these structures have never been deposited in public databases, as they are tied to proprietary drug-development pipelines. Some experts believe the total volume of these private structures may actually exceed what is held in the public Protein Data Bank.
—
## The Consortium Approach: Pooling Private Data Without Giving It Away
Recognizing the potential of shared data to accelerate progress, a group of pharmaceutical companies recently formed a collaborative initiative called the AI Structural Biology Network. The goal was ambitious but carefully designed: use proprietary molecular structures to retrain an open-source protein-folding model without ever exposing any company’s confidential data.
The collaboration selected OpenFold, an open-source replica of a leading commercial protein-folding system, as its base model. Previously, this model had been trained exclusively on publicly available structures from the Protein Data Bank. The team then introduced an additional 20,000+ proprietary structures to the training process—each one depicting a protein in complex with a potential drug molecule. These structures were sourced from five participating companies and were integrated into the model in a way that preserved the privacy of each firm’s intellectual property.
After retraining, the new model was put to the test. The team held back 1,056 protein-ligand structures that the model had never seen before and used them as a benchmark. The results were striking: the retrained model accurately predicted the structure of more than half of these held-out examples at a high confidence level. By comparison, the original publicly available version of the same model achieved comparable accuracy on only about one-third of the test cases, while another popular open-source tool managed around 40%.
Perhaps equally important, the consortium model outperformed versions that had been trained on each company’s proprietary data alone. This finding underscores a broader truth about AI development: shared, diverse datasets tend to produce more robust and generalizable models than siloed ones.
## Why This Matters for the Future of Drug Discovery
The implications of these results extend well beyond a single benchmark test. They provide strong evidence that pharmaceutical companies’ private vaults of molecular data represent an enormous, largely untapped resource for improving AI tools used across the industry. By demonstrating that such data can be safely shared in a privacy-preserving way, the consortium has opened the door to a new paradigm in collaborative AI development—one where competitors work together to build better tools without compromising their competitive advantages.
Scientists involved in the effort are now calling for the creation of more publicly accessible datasets modeled on this approach. One such initiative, supported by millions in government funding, has already begun releasing hundreds of newly generated protein structures to the public, with thousands more in development.
—
## Frequently Asked Questions
**What is protein folding, and why is it important for drug discovery?**
Proteins are long chains of molecules that fold into specific three-dimensional shapes, and those shapes determine how proteins function in the body. Many diseases arise when proteins misfold or interact with other molecules in harmful ways. By predicting how a protein folds and how it might interact with a drug candidate, AI models can help researchers design more effective therapies faster and at lower cost.
**Why is the public Protein Data Bank not enough for training these AI models?**
The Protein Data Bank is rich in structures of proteins by themselves, but it contains relatively few examples of proteins bound to small drug-like molecules. This limits the ability of AI models to learn how drugs interact with proteins—an essential capability for designing new treatments.
**How do pharmaceutical companies share proprietary data without giving it away?**
The consortium used a privacy-preserving approach in which proprietary structures were provided to the shared AI model during training but were never stored or exposed in a way that could reveal a company’s confidential research. The resulting model learns from all the data collectively without retaining access to any individual firm’s proprietary information.
**What is OpenFold, and how is it different from AlphaFold?**
OpenFold is an open-source replication of a leading protein-folding AI system. While AlphaFold was developed by DeepMind and is proprietary, OpenFold allows researchers to inspect, modify, and retrain the model—a critical feature for collaborative projects like the one described here.
**Are the results of this study peer-reviewed?**
At the time of reporting, the findings had been shared as a blog post and had not yet undergone peer review. The team behind the collaboration has stated that it plans to submit a formal paper to a peer-reviewed journal for publication.
**What is OpenBind, and how does it relate to this work?**
OpenBind is a publicly funded initiative aimed at generating and sharing new protein structures that capture drug-protein interactions. It represents a complementary approach to the consortium’s work—building a public resource that could be used to train future generations of AI models without relying on private data.
—
## Conclusion
The collaboration among pharmaceutical companies to retrain protein-folding AI models on proprietary data represents a significant step forward for computational drug discovery. It demonstrates that shared, diverse datasets can dramatically improve model performance, that private data can be used without compromising intellectual property, and that the path to more powerful AI tools lies not in hoarding information but in finding smart ways to collaborate. As initiatives like OpenBind continue to expand the pool of publicly available protein-ligand structures, the future looks promising for an era in which AI and human expertise work hand in hand to bring new medicines to patients faster than ever before.
Thank you for reading



