Risk-based does not mean lightly controlled
It means the depth of evidence and oversight is deliberately concentrated where an AI failure could materially affect a patient, a product, a regulated record, or a regulatory decision.
A defensible AI quality framework links intended use, patient and product risk, evidence, human accountability, and lifecycle control, without forcing every use case into the same validation model.
Artificial intelligence (AI) is moving rapidly from innovation pilots into routine life-sciences operations. Organizations are evaluating AI-enabled tools for clinical development, manufacturing, regulatory intelligence, quality systems, medical information, and pharmacovigilance.
The potential benefits are substantial: faster review of complex data, more consistent processing, improved prioritization, and greater capacity for qualified personnel to focus on higher-value decisions. In a regulated environment, however, organizations cannot treat AI as merely another productivity tool.
AI and machine-learning (ML) systems can be probabilistic, data-dependent, difficult to interpret, and subject to change through model updates, altered source data, new integrations, or user configuration. These characteristics introduce risks that are familiar to quality assurance (QA), including inaccurate records, uncontrolled change, poor supplier oversight, and weak accountability.
But they can also create AI-specific failure modes such as model drift, bias, hallucinated content, opaque reasoning, and automation bias. The practical question is therefore not whether an organization can use AI.
It is whether the organization can show that the system is fit for its intended use, that controls are proportional to the potential GxP impact, and that performance remains acceptable throughout the system lifecycle. This requires QA involvement from concept through retirement, rather than a validation exercise performed after a tool has already been selected or deployed.
Regulators are not yet operating from a single harmonized rulebook for enterprise AI used in GxP activities. The direction of travel is nevertheless increasingly consistent.
FDA’s 2021 AI/ML Software as a Medical Device Action Plan, although developed for medical-device software, emphasized total product lifecycle oversight, good ML practices, transparency, planned control of model changes, and real-world performance monitoring.1
FDA’s January 2025 draft guidance for AI used to support regulatory decisions for drugs and biological products moved the discussion toward a risk-based credibility assessment tied to a clearly defined context of use. In January 2026, FDA and the European Medicines Agency (EMA) jointly identified 10 principles for good AI practice in drug development, including human-centric design, risk-based approaches, clear context of use, data governance, risk-based performance assessment, lifecycle management, and clear information for users and decision-makers.2,3
EMA’s final 2024 reflection paper addresses AI across the medicinal-product lifecycle, including discovery, clinical development, manufacturing, regulatory submission activities, and post-authorization safety surveillance. The paper reinforces a human-centric approach and the need to consider the relevance and consequences of AI outputs within the applicable legal and regulatory framework.4
The United Kingdom’s public position is also becoming more operational. The Medicines and Healthcare products Regulatory Agency’s (MHRA) established GxP data-integrity guidance remains directly relevant to AI-supported records and decisions.
In June 2026, the MHRA Inspectorate further stated that AI-assisted inspection responses must be factually accurate and verifiable, technically reviewed by appropriately experienced personnel, approved by an accountable individual, supported by evidence, and appropriate to the regulatory context.
The emphasis is not on prohibiting AI; it is on the effectiveness of the quality system governing its use.5,6 Health Canada’s work on generative AI and ML-enabled devices has highlighted concerns that apply more broadly to regulated AI use: third-party model updates, lack of transparency regarding training data, hallucinated information, privacy and cybersecurity risks, model cards, human review, fail-safe mechanisms, controlled change, and ongoing performance monitoring.
Health Canada has also stated its intention to integrate AI into scientific and regulatory processes and has previously identified pharmacovigilance as an area for AI exploration and future guidance harmonization.7-10 Together, these perspectives do not point to one universal compliance package.
They point to a defensible relationship between intended use, risk, evidence, transparency, human oversight, change control, and continued performance.
The following implementation model adapts established GxP quality principles to the characteristics of AI/ML systems. It is intended to help QA and compliance leaders translate high-level governance commitments into operational controls.
Every AI implementation should begin with a concise intended-use statement. The statement should define what the system will do, who will use it, what data it will process, what output it will produce, where the output enters the workflow, and whether that output may influence a GxP decision.
It should also describe what the system is not authorized to do. The risk assessment should focus on consequence rather than technology labels.
A large language model used to reformat an approved document may be lower risk than a conventional algorithm that determines which safety cases receive expedited review. Conversely, a generative model becomes higher risk when its output directly affects patient safety, product quality, data integrity, regulatory reporting, or safety-signal detection.
Useful risk factors include the criticality of the process, the directness of the decision's impact, the degree of autonomy, the detectability of errors, the reversibility of the outcome, model complexity, the quality and sensitivity of the data, and the availability of effective human review. The risk classification should drive the strength of governance, validation, monitoring, and QA independence.
It means the depth of evidence and oversight is deliberately concentrated where an AI failure could materially affect a patient, a product, a regulated record, or a regulatory decision.
AI quality depends on data quality and on the conditions under which the model is supplied and operated. Data governance should address provenance, ownership, integrity, completeness, representativeness, transformations, access controls, retention, privacy, and versioning.
For models developed or tuned internally, the separation and suitability of training, validation, and test data should be documented. For externally hosted models, the organization should understand which data were used to develop the model to the extent necessary to evaluate suitability and limitations.
Supplier qualification is especially important when an organization relies on a third-party foundation model, cloud service, or embedded AI feature. Due diligence should establish the vendor’s development and testing practices, security controls, incident notification process, subcontractor use, data use terms, retention practices, model update process, performance evidence, and support for investigation and audit.
The sponsor should also determine whether model versions can be pinned, whether updates can be delayed or tested before release, and how to exit the service without losing records or traceability. A vendor’s general assurance that a model is “validated” is not a substitute for evidence that the configured system is fit for the sponsor’s intended use.
Configuration, prompts, data sources, thresholds, integrations, and human workflow can materially change the risk profile even when the underlying model remains the same.
Validation should preserve core GxP principles, documented requirements, objective evidence, traceability, approved acceptance criteria, controlled release, and retained records, while adapting test design to AI-specific behavior. The goal is not to force every AI tool into a traditional deterministic software test model.
The goal is to demonstrate that the system and the human-AI workflow perform reliably enough for the defined use case. For lower-risk drafting or productivity tools, an appropriate package may include a documented GxP impact assessment, acceptable use restrictions, training, data-handling controls, representative user testing, and mandatory review of every output.
For moderate- and high-risk systems, the package may expand to include supplier qualification, detailed requirements, data qualification, traceability, challenge testing, subgroup or edge-case evaluation, role-based security, interface testing, release approval, periodic review, and independent QA assessment. Testing should reflect the model type and the failure modes that matter.
For generative AI, this may include accuracy, completeness, relevance, unsupported claims, fabricated references, response variability, prompt injection, leakage of sensitive information, prohibited content, and the effectiveness of guardrails. A controlled prompt library and a representative set of expected answers can improve repeatability, but testing should also challenge the system with ambiguous, incomplete, and adversarial inputs.
For predictive or classification models, evaluation may include sensitivity, specificity, precision, recall, calibration, false-positive and false-negative rates, subgroup performance, robustness to missing or shifted data, and stability over time.
Acceptance criteria should be linked to the clinical, quality, or regulatory consequence of error. A generic accuracy percentage is rarely sufficient without explaining which errors are acceptable, which are critical, and how the user is expected to respond.
For many GxP use cases, the most meaningful performance unit is not the model alone but the human-AI team. Validation should therefore test whether qualified users can recognize limitations, review output efficiently, detect errors, document decisions, and escalate exceptions under realistic workload conditions.3,7
Transparency does not require every user to understand the mathematics of a model. It requires the organization to provide sufficient information for the relevant audience to use, review, challenge, and govern the system appropriately.
The depth of explanation should be proportional to risk and tailored to system owners, end users, QA reviewers, data scientists, auditors, and regulators. A system factsheet or model card can serve as a controlled summary of the intended use, prohibited uses, model and software versions, data sources, known limitations, performance measures, relevant subpopulations, human-review requirements, change history, monitoring thresholds, and owner contacts.
Health Canada’s advisory discussions have specifically highlighted model cards as a means of maintaining accessible information on indications, limitations, known biases, failure conditions, and expected performance.7 Human oversight should be defined as a workflow control, not a general statement that “a human remains involved.”
The procedure should specify who reviews the output, what qualifications are required, what source evidence must be checked, what decisions cannot be delegated to the AI, when a second review is required, how disagreements are resolved, when use must be suspended, and how the final decision is recorded. Organizations should also address automation bias.
Reviewers may accept fluent or confident output without sufficient challenge, particularly when workloads are high. Training, interface design, sampling of approved outputs, and analysis of user overrides can help determine whether human review is functioning as an effective control rather than a nominal approval step.
AI governance continues after go-live. A model may change because the vendor releases a new version, the organization changes prompts or thresholds, source data shifts, integrations are modified, workflows evolve, or the model is retrained.
Each of these changes can affect validated performance even when the user interface appears unchanged. The change-control process should distinguish among locked, periodically updated, and adaptive models.
Uncontrolled continuous learning in a GxP production environment should not be assumed acceptable. If future changes are anticipated, the organization can define an approved change protocol describing the types of permitted modifications, supporting evidence, performance boundaries, review responsibilities, rollback criteria, and conditions requiring partial or full revalidation.
This approach is conceptually consistent with regulators’ emphasis on planned change control for ML systems.1,7,8 Monitoring should be risk-based and should combine technical and process measures.
Relevant indicators may include output accuracy, error categories, false-negative and false-positive rates, user overrides, unresolved exceptions, drift with inputs or outputs, changes in case or product mix, vendor releases, latency, failed interfaces, complaints, safety escapes, and incidents. Monitoring should include predefined thresholds that trigger investigation, corrective and preventive action (CAPA), retraining, revalidation, temporary restriction, rollback, or retirement.
Periodic review should confirm that the intended use remains current, the model and configuration are known, the supplier remains suitable, required training and access are current, monitoring is effective, deviations and incidents have been addressed, and the system continues to provide a favorable benefit-risk profile. Retirement should include record retention, data disposition, interface shutdown, access removal, and migration of any continuing business process.
Most organizations do not need a separate AI quality management system (QMS). They need to extend existing QMS processes to make AI-specific risks visible and consistently controlled.
The AI governance framework should connect to procedures for computerized systems, supplier quality, data integrity, information security, privacy, records management, change control, deviations, CAPA, training, periodic review, business continuity, and pharmacovigilance. A sustainable operating model typically includes executive sponsorship, a cross-functional AI risk or governance board, a business or model owner, a data owner, an IT system owner, and independent QA oversight for higher-risk applications.
Regulatory affairs, pharmacovigilance, legal, privacy, information security, and process subject-matter experts should participate according to the use case. Clear accountability is essential because AI systems often span organizational boundaries managed separately under conventional system ownership models.
An inspection-ready documentation set should allow an auditor or regulator to reconstruct the lifecycle and understand why the organization concluded that the system was suitable for use. At a minimum, higher-impact applications should have readily available:
The audit narrative should be straightforward: what the system does, why it is used, what could go wrong, what evidence supports its use, who is accountable, how changes are controlled, and how the organization knows the system is still performing acceptably. If the organization cannot tell that story with objective evidence, the governance framework is not yet mature.
Pharmacovigilance illustrates why AI governance must connect technology controls with end-to-end process accountability. AI can assist with intake, duplicate detection, language translation, coding suggestions, case summarization, narrative drafting, literature screening, case prioritization, and signal detection.
These uses can improve consistency and capacity, but they also create a risk that safety information may be delayed, misclassified, or lost inside an unmonitored channel. A chatbot, virtual agent, or medical-information tool that can receive a potential adverse event or product complaint should not become a “black box.”
The organization must either prevent the tool from accepting reportable information or ensure that information is captured, transferred, reconciled, reviewed, and reported within required timelines. The route from source information to the safety database should be testable and traceable.
Human accountability should remain explicit for decisions such as case validity, seriousness, expectedness, reportability, medical assessment, causality, signal validation, and regulatory submission. AI may recommend, prioritize, or draft, but the pharmacovigilance system must identify the qualified individual who reviewed the underlying evidence and made the final decision.
Monitoring should pay particular attention to false negatives, missed minimum criteria, duplicate errors, miscoding, delayed routing, and differences in performance across languages, products, populations, and data sources.4,10 Reconciliation is equally important.
Interfaces among chatbots, medical information, product-quality complaint systems, clinical systems, vendor platforms, and the safety database should be validated and periodically reconciled. Supplier agreements should define access to records, incident escalation, model changes, performance reporting, retention, and audit rights.
The effectiveness measure is not simply how many cases the AI processes; it is whether the overall safety system continues to produce complete, accurate, timely, and verifiable information.
Risk-based AI oversight does not mean applying less rigor, it means applying the right rigor to the decisions and failure modes that matter most. A low-impact drafting tool should not require the same evidence package as a model that influences manufacturing controls or safety reporting.
At the same time, a tool should not be classified as low risk merely because a human is nominally present at the end of the workflow. The most robust framework begins with intended use, assesses GxP and patient impact, qualifies data and suppliers, validates the configured human-AI system, makes limitations visible, controls changes, monitors performance, and retains evidence across the lifecycle.
QA should be embedded early enough to influence design, rather than asked to approve controls after deployment. AI-enabled systems can create meaningful value in regulated operations, but innovation is sustainable only when the organization can demonstrate reliability, explainability appropriate to the use, accountable human oversight, and continued control.
An AI system that cannot be governed should not be used for GxP purposes. An AI system governed by a proportionate, lifecycle-based QMS can be both innovative and inspection-ready.
References
About the Author
James Meckstroth is a senior quality assurance executive with more than 25 years of leadership experience in FDA-regulated GxP environments across commercial manufacturing and clinical research. He provides strategic oversight for global compliance service teams encompassing Global Auditing, AI/ML Governance, and Clinical Quality.
He is recognized for his expertise in global regulatory compliance and quality systems, with deep knowledge of FDA, EMA, ICH, GCP, and GMP standards. Mr. Meckstroth has successfully led organizations through FDA general and pre-approval inspections and has extensive experience managing regulatory interactions, enterprise audit programs, and independent Data and Safety Monitoring Boards.
His leadership focuses on building and optimizing Quality Management Systems (QMS), advancing risk-based quality strategies, and strengthening inspection readiness and remediation at scale. He brings a strong track record in governance, vendor oversight, CAPA, and complex deviation and OOS investigations, as well as operational excellence through process optimization and organizational capability building.
Mr. Meckstroth has broad product and manufacturing experience, including small molecule and potent/cytotoxic compounds in both sterile and non-sterile environments. He holds a B.S. in Microbiology from The University of Kansas and a Master of Business Administration from Avila University.