AI and Biobank Consent: Why Historical Consent Forms Don’t Cover AI Training

Research institutions are sitting on extraordinary resources: biobanks containing millions of stored tissue samples and biological specimens, accumulated over decades of clinical care and research enrollment, linked to decades of electronic health record data. For AI-powered drug discovery, this is precisely the training material that companies need — large, longitudinal, richly annotated datasets that no single institution could assemble prospectively.

The problem is that the people whose tissue and data fill those biobanks gave consent under documents that were never designed to cover AI. They consented to biobank storage. They consented to use in future research. Some gave broad consent for unspecified future research purposes. But consent to AI training — specifically to using biological samples and linked clinical data as training material for commercial machine learning models that may be licensed, monetized, and deployed in ways that did not exist when consent was obtained — is a fundamentally different thing. And most historical consent forms do not cover it.

This consent gap is not a technicality. It exposes research institutions to regulatory liability under the Common Rule and HIPAA, creates barriers to commercialization of AI-derived discoveries, and — if it becomes a reputational issue — threatens the trust that the research enterprise depends on for future enrollment.

The Common Rule (45 CFR Part 46) governs federally funded human subjects research in the United States. For biobank operations, the most relevant provisions are those governing informed consent for the storage, maintenance, and future use of identifiable private information and identifiable biospecimens.

The 2018 revisions to the Common Rule, which took effect in 2019, introduced the concept of “broad consent” — a mechanism that allows institutions to obtain a single consent for future unspecified research uses of stored biospecimens and identifiable private information, rather than requiring new consent for each future study. This was a significant practical accommodation for biobank research.

But even broad consent has limits. Under 45 CFR § 46.116(d), broad consent must include a description of the types of research that may be conducted with the stored biospecimens and data. It must include a statement about whether information derived from the research will be shared, whether it may be used for commercial profit, and whether subjects will share in that profit. And it must address the possibility that the research might include whole genome sequencing.

None of those disclosures, as written into most pre-2019 consent forms, contemplate AI training in any meaningful sense. More fundamentally, the concept of using stored specimens and linked clinical records to train a machine learning model — a process in which the “research” is not studying the subject but using the subject’s data to build a tool that will be used on entirely different patients — is a novel use that most institutional consent frameworks have not yet addressed.

The 2018 Common Rule explicitly leaves open whether protections are adequate for secondary research uses of data to develop AI applications. Policymakers, including HHS’s Office for Human Research Protections (OHRP), have signaled awareness of this gap but have not yet issued comprehensive guidance addressing it.

Why “De-Identified” Data Doesn’t Solve the Problem

Institutions often point to de-identification as the solution to the consent problem: if data is properly de-identified under HIPAA’s Safe Harbor or Expert Determination standards, it falls outside HIPAA’s requirements and may fall outside the Common Rule’s protections as well.

This is technically accurate in many circumstances, but it does not fully resolve the AI-specific consent gap.

First, re-identification risk is a genuine and growing concern in the AI context. Modern machine learning models trained on de-identified biological data — particularly genomic data, rare disease phenotype data, or highly specific clinical histories — can be used in ways that allow re-identification of individuals. A biobank that de-identifies genomic data before providing it for AI training cannot fully guarantee that the AI model itself will not inadvertently enable re-identification through the patterns it learns.

Second, even for data that poses minimal re-identification risk, the ethical dimension of consent does not entirely disappear with technical de-identification. If participants were told their samples would be used for medical research to advance understanding of their disease, and instead their data is used to train a commercial AI model that has no direct relationship to their disease or their care, there is a reasonable argument that the actual use exceeds the reasonable expectations created by the consent, even if technically permissible under HIPAA and the Common Rule.

Research institutions that dismiss this as a non-legal concern misunderstand how consent controversies develop. The UK Biobank has faced scrutiny over exactly these questions. US institutions that become the subject of investigative journalism or advocacy campaigns about their commercial AI partnerships — particularly if those partnerships involve large financial transfers to the institution and no benefit to participants — face exactly the kind of public trust damage that harms future enrollment.

Third, HIPAA’s de-identification standards were not designed with AI training in mind. The 18-item Safe Harbor standard strips specific identifiers from records but does not account for the contextual re-identification possibilities that arise when records are combined with other datasets or processed through models capable of inferring identity from patterns. Expert determination may provide more robust protection, but it requires affirmative analysis by a qualified statistician — something many AI data sharing arrangements do not include.

For research institutions that are not just sharing data but entering commercial partnerships — providing biobank data to pharmaceutical companies or AI startups in exchange for licensing fees, milestone payments, or equity — the consent gap takes on an additional dimension.

Under 45 CFR § 46.116(c)(7), broad consent must include “a statement about whether information derived from the research will be shared with researchers outside this institution, and, if so, a description of the types of researchers with whom it will be shared.” It must also include, at subsection (c)(8), “a statement about whether commercial profit may be generated from research using the subject’s specimens or information and whether the subject will share in that profit.”

Most pre-2019 consent forms contain none of this language. Many post-2019 forms address it only generically. A consent form that tells participants their data may be used for research that could lead to commercial treatments — and that they will not share in any resulting profits — is different from a consent form that specifically tells participants their biological samples will be incorporated into a commercial AI training dataset licensed to a pharmaceutical company for drug discovery.

The gap matters for IP reasons as well. If a pharmaceutical company licenses biobank data, trains an AI model on it, and that model generates a lead compound that becomes a marketed drug, the chain of title for the IP runs through the biobank data. If the data was used outside the scope of participant consent, the institution and the company have a cloud over that title. In litigation — whether from a participant advocacy organization, a class action, or a disgruntled institutional partner — the consent validity question could be used to challenge the IP derived from the data.

This is not a hypothetical risk. It is a structural vulnerability that sophisticated IP counsel in the pharmaceutical industry are already flagging in due diligence for AI-related licensing transactions.

What should research institutions do when they have identified a gap between historical consent and current or planned AI use? Several remediation approaches exist, each with different practical and legal implications.

Re-consent campaigns. The most straightforward remediation is to contact participants and obtain new consent specifically addressing AI training uses. This is the approach most clearly consistent with research ethics principles. It is also operationally demanding — large biobanks may have millions of participants, many of whom are no longer reachable, and re-consent campaigns take time and resources.

Re-consent is most appropriate when: the gap between historical consent and planned use is significant; the planned use involves sensitive data (genomic data, mental health records, data about minors); or the commercial dimension of the use is substantial and would reasonably surprise participants.

IRB waiver of consent. Under the Common Rule, an IRB may waive the requirement for informed consent (or modify consent requirements) if the research meets specified criteria — including that the research could not practicably be carried out without the waiver and that the waiver poses no more than minimal risk to subjects.

For large-scale biobank data uses, IRB waiver has historically been used. But the AI training context poses specific challenges to the minimal risk determination. Using biospecimen data to train a commercial AI model that will be licensed to a pharmaceutical company is not obviously “minimal risk” in the traditional sense — the risks may be different in kind rather than in magnitude, including reputational harm, privacy harm from potential re-identification, and loss of autonomy over use of one’s biological materials for commercial purposes.

Institutions seeking IRB waiver for AI training uses should build a waiver analysis that specifically addresses these AI-specific risk dimensions — not simply apply the waiver analysis framework developed for traditional research secondary use.

Transition to broad consent for future enrollment. For new participants enrolling in biobanks after 2019, institutions should be using updated consent forms that specifically address AI training uses, commercial partnerships, and data sharing with technology companies. If your institution has not updated its biobank consent forms since 2019, start there.

The updated consent form should: specifically describe AI training as a potential future use; address commercial use and whether participants will share in any resulting profits; identify the categories of external parties with whom data may be shared (including pharmaceutical companies and AI companies); address genomic data specifically if applicable; and explain the de-identification approach and its limitations.

Tiered consent. An increasingly favored approach is tiered consent — a consent structure that gives participants specific choices about what uses they agree to, rather than a binary all-or-nothing broad consent. Under a tiered approach, participants might consent to use in academic non-profit research while declining commercial use, or consent to use for their specific disease while declining other uses.

Tiered consent is more consistent with genuine participant autonomy, but it requires data management infrastructure that can track and enforce participant preferences across a large biobank with diverse consent status records. Institutions that have not built this infrastructure will find tiered consent difficult to implement retroactively.

What Vendor Contracts Must Address

When institutions are contracting with pharmaceutical companies or AI companies to provide biobank data for AI training, contracts need to address the consent gap directly. Ignoring it in the contract does not make the liability go away — it just creates a dispute about who bears it when it materializes.

Consent representation and warranty. The institution providing data should represent that the data was collected under consent adequate for the intended AI training use — or, if that representation cannot be made, clearly disclose the nature of the consent obtained and the gap. A data provider that warrants adequacy of consent and is later found to have had inadequate consent has significant contract liability. A data provider that specifically discloses the consent limitation and transfers risk appropriately has a defensible position.

Regulatory compliance warranty. The contract should address which party is responsible for ensuring that the use of the biobank data complies with the Common Rule, HIPAA, and applicable state laws (several states have adopted biobank-specific statutes). This allocation should not default to vague mutual representations — it should specify which party bears what compliance obligation.

Participant rights management. Participants whose data is used in AI training may have rights to access information about that use, to withdraw consent, or — under state laws adopting elements of the GDPR framework — to request deletion. Contracts must address how participant rights requests will be handled, including what the AI company’s obligations are if a participant exercises a right that affects data already incorporated into a trained model.

Indemnification for consent-related claims. Given the potential for participant claims, class action risk, and regulatory enforcement arising from consent gaps, the allocation of indemnification risk in biobank AI contracts is not a boilerplate issue. Both parties need to understand who indemnifies whom if a participant or regulatory agency challenges the data use on consent grounds.

Chain of custody documentation. The AI company’s model training may incorporate data from multiple institutional sources. Contracts should require the company to maintain documentation of the consent status for each data source, so that if a challenge arises, the company can demonstrate the provenance and consent status of the data used to train the model.

State Law Adds Complexity

Beyond the federal Common Rule and HIPAA framework, several states have enacted biobank-specific legislation that creates additional consent requirements. California, Texas, Florida, and New York all have statutes or regulations addressing consent for biospecimen use. Some of these statutes explicitly require consent for commercial use. Others impose requirements for destruction of specimens on participant request.

For institutions and companies operating nationally, this patchwork of state requirements creates compliance complexity. The most conservative approach — building consent forms that satisfy the most demanding applicable state standard — is also the most protective against litigation risk.

The Bottom Line

The intersection of historical biobank data and AI drug discovery is one of the most commercially promising areas of pharmaceutical research. It is also one of the areas with the most significant unresolved consent law questions.

Research institutions and their commercial partners that acknowledge the consent gap — assess it carefully, remediate where possible, and build robust consent practices for future data collection — are positioned to participate in this space in a way that is legally defensible and consistent with the ethical foundations of human subjects research.

Institutions and companies that treat historical consent forms as covering AI uses they were never designed to address, and that proceed with commercial AI partnerships without addressing the consent gap, are building on a foundation that is both legally vulnerable and ethically precarious.

The investment in getting consent right is modest compared to the cost of a regulatory action, a class action, or a reputational crisis. And in the current environment, where research participants are increasingly sophisticated about data rights and commercial uses of biological materials, the trust dimension is not a soft consideration — it is a business necessity.


This post is for general informational purposes only and does not constitute legal advice. Reading this post does not create an attorney-client relationship. If you have questions about your specific situation, consult a qualified attorney.



Leave a Reply