Drug firms’ secret data supercharge AI protein models

September 15, 2026

A protein model

(Nature) – An AI system trained on more than 20,000 protein structures from pharmaceutical companies outperforms AlphaFold-like models that use only public data.

For drug discovery, protein-folding models such as AlphaFold have a problem: there aren’t enough data in public databases. Improving the performance of these artificial-intelligence-based tools will require extra data that provide examples of how proteins and drugs interact, some scientists argue.

Protein structures — locked away by the thousands in drug company vaults — offer one promising source. Today, a consortium of pharmaceutical companies reports that using such data to train AI models of protein folding improves model performance markedly.

The group used OpenFold3 — an open-source replication of AlphaFold 3 — to develop a new model trained on more than 20,000 proprietary protein structures. (Read More)