Pharma Giants Pool Data for Antibody AI
Apheris and Ginkgo have launched a consortium allowing major pharmaceutical companies to train AI on shared antibody data without exposing raw proprietary sequences.

Apheris and Ginkgo are creating a shared antibody dataset with major pharmaceutical companies while trying to preserve each participant's proprietary data.
What happened
The companies launched the Antibody Developability Consortium with AbbVie, argenx, Lundbeck and Takeda as founding members.
The project aims to build a standardised dataset covering 10,000 antibodies.
Participating companies can contribute data and train or benchmark models without directly exposing their raw proprietary sequences to one another.
Why it matters
Biological AI is constrained by data quality as much as model quality.
Pharmaceutical companies possess valuable internal datasets, but competitive and intellectual-property concerns make them reluctant to pool that information openly.
Federated or privacy-preserving infrastructure offers a way to create larger training datasets without centralising all the raw data.
That can improve model quality while reducing the commercial risk of collaboration.
The bigger picture
Some of the most valuable AI datasets may never become fully open.
Instead, industries such as pharmaceuticals, banking and manufacturing may build shared training environments where data stays under controlled access.
The antibody consortium is a useful example of that model.
If it works, the competitive advantage may shift toward infrastructure companies that can let rivals collaborate on model training without forcing them to surrender the underlying proprietary information.
