Artificial-intelligence-based models of protein structure could be improved by incorporating data from pharmaceutical companies.Credit: Miyako Nakamura/Getty For drug discovery, protein-folding models such as AlphaFold have a data problem: there isn’t enough of it in public databases. Improving the performance of these artificial-intelligence-based tools will require extra data that provide examples of how proteins and drugs

Artificial-intelligence-based models of protein structure could be improved by incorporating data from pharmaceutical companies.Credit: Miyako Nakamura/Getty
For drug discovery, protein-folding models such as AlphaFold have a data problem: there isn’t enough of it in public databases. Improving the performance of these artificial-intelligence-based tools will require extra data that provide examples of how proteins and drugs interact, some scientists argue.
Protein structures – locked away by the thousands in drug company vaults – offer one promising source. Today, a consortium of pharmaceutical companies reports that using such data to train AI models of protein folding improves model performance markedly.
The group used OpenFold3 — an open-source replication of AlphaFold 3 — to develop a new model trained on more than 20,000 proprietary protein structures. The system outperformed both comparable ones trained on public data alone and those trained on the siloed datasets of individual firms. The study, described in a blog post, has not been peer-reviewed, and the model is not publicly available.
“You add all this data, and you get a pretty big bump in performance,” says Mohammed AlQuraishi, a computational biologist at Columbia University in New York City, who was part of the effort.

AlphaFold is running out of data — so drug firms are building their own version
The findings, he says, strengthen the case for generating similar publicly available datasets to supercharge protein-folding AIs. One such project, called OpenBind and supported by up to £8 million (US$10.8 million) in UK government funding, released hundreds of new protein structures last month, with thousands more in the works.
An untapped vein
The Protein Data Bank (PDB), an open repository of more than 200,000 experimentally determined protein structures, was the bedrock of AlphaFold 2’s training data. It enabled the tool to predict protein structures with startling accuracy — a breakthrough recognized with the 2024 Nobel Prize in Chemistry.
The model’s successors, including AlphaFold 3, added the ability to predict how proteins will interact with other molecules, including potential drugs. But the PDB has relatively few examples of experimentally determined structures interacting with drug-like molecules — maybe just 10,000, says Paul Mortenson, vice-president for computational chemistry and informatics at Astex Pharmaceuticals in Cambridge, UK.
That lack of data is a problem for drug-discovery efforts. Research has suggested that the accuracy of AlphaFold 3 and other ‘co-folding’ models — which predict the structure of proteins interacting with each other — drops off a cliff when the models are challenged to predict interactions between molecules highly dissimilar to those on which they were trained1.
To make such tools more useful in drug discovery, researchers say, they need access to more data — which is why they have turned to vaults of molecular structures from pharmaceutical companies.

What’s next for AlphaFold and the AI protein-folding revolution
These data are generated during drug-discovery programmes, using techniques such as X-ray crystallography and cryo-electron microscopy. Many of the protein structures have never been deposited in public databases because they relate to proprietary drug-development efforts. The total size of these vaults is unknown, but some have estimated that they could contain more data than the PDB does.
“The data that’s missing from the PDB is exactly the data that’s present in our internal data,” John Karanicolas, head of computational drug discovery at the pharma company AbbVie in Chicago, Illinois, told Nature last year.
Better predictions
To test whether their data could be useful for protein-folding models, AbbVie, Astex and several other drug companies last year formed a collaboration called the AI Structural Biology (AISB) Network.
It involved ‘fine-tuning’ OpenFold3 – previously trained only with PDB data – on an additional 20,167 structures capturing proteins bound to potential drugs, or ligands. The structures came from five companies and were provided to the model in such a way that proprietary data remained private.
The AISB study found that the extra data enhanced predictions. When tested on 1,056 protein–ligand structures that were set aside from the training data, the AISB model predicted more than half of them to a high level of accuracy. By contrast, the publicly available version of OpenFold3 achieved the same performance on just one-third of the structures, and a competing open-source model called Boltz-2 achieved around 40%. The team plans to submit a paper describing the work to a peer-reviewed journal.
The fact that the AISB model also outperformed co-folding tools that were trained only on each company’s individual data highlights the benefits of pooling information, says Karanicolas.
Keep following us for the latest insights.

















