Drug firms’ secret data supercharge AI protein models
September 15, 2026

(Nature) – An AI system trained on more than 20,000 protein structures from pharmaceutical companies outperforms AlphaFold-like models that use only public data.
For drug discovery, protein-folding models such as AlphaFold have a problem: there aren’t enough data in public databases. Improving the performance of these artificial-intelligence-based tools will require extra data that provide examples of how proteins and drugs interact, some scientists argue.
Protein structures — locked away by the thousands in drug company vaults — offer one promising source. Today, a consortium of pharmaceutical companies reports that using such data to train AI models of protein folding improves model performance markedly.
The group used OpenFold3 — an open-source replication of AlphaFold 3 — to develop a new model trained on more than 20,000 proprietary protein structures. (Read More)