Artificial intelligence is already inside most GMP sites, whether or not the quality system knows about it. Vendors are adding machine learning to inspection and laboratory software, and staff use chat assistants to draft procedures. Until recently no GMP text addressed this directly. Draft Annex 22 changes that, and even before adoption it shows the questions a site should be able to answer about its AI.
This guide explains what Annex 22 is, where it stands today, what the draft expects, and how a site can prepare now: a risk-based way to categorise AI use, an inventory, supplier questions, and how much validation each category needs.
What Annex 22 is
Annex 22 is a proposed new annex to EU GMP (EudraLex Volume 4) on artificial intelligence. It is aimed at computerised systems used in the manufacture of medicinal products and active substances where an AI model is used in a critical application, meaning one with a direct impact on patient safety, product quality or data integrity. It does not replace Annex 11. An AI tool is still a computerised system, so Annex 11 applies in full, and Annex 22 adds what is specific to the model: its intended use, its testing and the control of its performance.
Current status: a draft, not yet in force
The European Commission and the EMA published draft Annex 22 for public consultation in July 2025, together with a revised Annex 11 and a revised Chapter 4 on documentation. The consultation closed in October 2025. As of autumn 2026, Annex 22 has not been adopted and is not part of EudraLex Volume 4. The target for a final text is the fourth quarter of 2026, but no adoption date or effective date has been announced.
The scope may still change. The 2025 draft excluded generative AI and large language models from critical GMP use, and that position is under review. On 30 June and 1 July 2026 the EMA held a multistakeholder workshop on a risk-based approach to generative AI, so the final text could treat these models differently.
What the draft expects
The draft is short and principle based, and most of it is good computerised system validation applied to a model. Its main expectations are:
- Static and deterministic models. For critical use, the model must not change while in use, and the same input must give the same output. Dynamic models that keep learning in operation, and generative models such as large language models, are not to be used in critical applications under the draft.
- A documented intended use. Before testing, describe what the model does, the process step it supports, and the characteristics of its input data, including expected variation, limitations and edge cases. Subject matter experts from the process should be involved.
- Pre-defined acceptance criteria. Choose performance metrics that fit the task, for example sensitivity, specificity, accuracy or a confusion matrix for a classifier, and fix the acceptance criteria before testing. The model should perform at least as well as the process it replaces.
- Independent, representative test data. Test data must be separate from the data used to train and tune the model, large enough to support the conclusion, correctly labelled, and cover the full range of real inputs, including rare and worst-case examples.
- Explainability and confidence. Where relevant, the system should show which features of the input drove a result, and record a confidence score, so that low-confidence outputs are flagged or treated as undecided rather than accepted.
- Human in the loop for non-critical use. Where AI is used outside critical applications, for example a generative tool that drafts text, a qualified person reviews the output and remains accountable for it.
- Control in operation. The model, its configuration and its input processing are under change control and configuration control. Performance is monitored during routine use, and input data is monitored for drift away from the data the model was tested on.
Step 1: categorise every AI use by risk
Not every use of AI is an Annex 22 matter. A four-level categorisation puts effort where it matters and answers the inspector's question 'how do you decide?'. Write the categories into a procedure and record the reasoning for each assignment.
- No GMP impact. The output never enters a GMP record or decision, such as drafting a newsletter. Controlled through an acceptable use policy that keeps GMP and confidential data out of unapproved tools.
- Assistive, low impact. The AI drafts something a qualified person checks and owns before it enters the quality system, such as a first draft of an SOP that then goes through normal document review. Controlled through documented human review against the source and a fitness-for-use assessment.
- Decision support. The AI proposes and a person decides: suggesting a deviation classification, flagging audit trail entries for review, prioritising environmental monitoring trends. The tool needs a defined intended use, testing on independent data, acceptance criteria and monitoring of how often its proposals are overridden.
- Autonomous critical. The AI output decides or directly determines a GMP outcome without item-by-item human review, for example an automated visual inspection system rejecting containers. Full validation to the expectations of the draft annex, with a static, deterministic model only.
Be honest about decision support. If reviewers accept the model's proposal almost every time without looking at the source, the decision has in practice been handed to the model, and the use should be treated as autonomous critical.
Step 2: build an AI register and find the shadow AI
You cannot control what you have not listed. Add AI to the computerised system inventory, or keep an AI register linked to it, recording the system, model and version, supplier, intended use, risk category, owner, validation status and the data the tool can access.
Then look for the AI nobody registered. Shadow AI includes public chat assistants on personal phones, AI features switched on by default in upgrades of office, LIMS or QMS software, and browser plug-ins. A staff survey and a review of release notes usually find more than expected. Each finding is registered and categorised, or stopped.
Step 3: ask suppliers the right questions
Most sites will buy AI rather than build it, so add these questions to the supplier questionnaire and quality agreement for any AI-enabled system:
- Model versioning: is the model version identifiable, can it be locked, and will it change only through a release you can see and test?
- Change notice: how much notice do you get before a model update or retraining, and what performance evidence comes with it?
- Training and test data: what was the model trained on, how representative is that of your products and processes, and how was the test data kept independent?
- Data location: where are your inputs processed and stored, and under which jurisdiction?
- Use of your data: are your inputs, records and outputs excluded from training the vendor's models, and is that written into the contract?
- Logging and export: are prompts, inputs, outputs, confidence scores and model versions logged with user and time, and can you export them in a readable form for review and inspection?
- AI Act role: how does the supplier classify the system under the EU AI Act, and is it the provider with you as the deployer? This is a separate regime from GMP, but the answer shows what documentation the supplier should hold.
Step 4: scale verification and validation to risk
Validation effort follows the category. For assistive tools, verification is mostly procedural: the tool is approved for a defined use, users are trained on its limits, and a qualified person checks every output against the source and signs for it. For decision support and autonomous critical uses, validation follows the draft annex: intended use, acceptance criteria, independent test data, documented results and a report against criteria set before testing.
Two points catch sites out. First, testing on the vendor's data is not enough: the test set must represent your inputs, including your rare events. Second, a retrained or updated model is a new model. Treat it as a change and retest it before use, as described in our guide 'Change control risk assessment: how to make it defensible to an inspector'.
Human review against the source record
Where a human is in the loop, the review must be real. The reviewer compares the output with the source record, not with what looks plausible: the summary against the batch record, the extracted value against the certificate. Generative tools produce fluent, confident text that can contain invented facts and numbers. The reviewer signs for the content, so the record should show what was checked.
Keep the input, output, model version and reviewer's decision together, so anyone can later reconstruct what the AI proposed and what the person did with it. The principles in our audit trail review checklist apply to these logs too.
Automation bias: test the humans too
Automation bias is the tendency to accept a system's output without enough scrutiny, and it grows as the system proves reliable. A human review that never disagrees with the model is not a control, so measure it. Track the override rate, and run seeded-error checks: periodically insert known wrong outputs, such as a misclassified deviation, and confirm reviewers catch them. A falling override rate combined with missed seeded errors shows the review has become a signature.
The classification itself still follows the logic in our guide 'Deviation management in pharma: the process step by step'. The model proposes; QA decides.
Annex 22 readiness checklist
- A procedure defines AI risk categories and how each use is assigned to one.
- An AI register is linked to the computerised system inventory, a shadow AI review has been done, and an acceptable use policy covers public AI tools.
- Each critical or decision support use has an approved intended use, including the characteristics of its input data.
- Acceptance criteria were set before testing, against a measured baseline of the current process.
- Test data is independent of training data, representative, and includes rare and worst cases.
- Models in critical use are static and deterministic, with the version locked and under change control.
- Confidence scores are recorded where available, and low-confidence outputs are handled manually.
- Human review is documented against the source record, and the reviewer is accountable.
- Override rates, seeded-error results and input drift are monitored, with defined triggers.
- Supplier agreements cover versioning, change notice, data location, use of your data and log export.
None of this needs the final text of Annex 22. It is computerised system validation, data integrity and change control applied to a new kind of system, and a site with an honest inventory and clear risk categories will be ready whatever the final annex says.
Frequently asked questions
What is EU GMP Annex 22?
A proposed new annex to EU GMP on artificial intelligence. It sets expectations for AI models used in critical GMP applications that affect patient safety, product quality or data integrity, covering intended use, testing, acceptance criteria, explainability and control in operation. It works alongside Annex 11, not instead of it.
Is Annex 22 in force?
Not as of autumn 2026. The draft was published for consultation in July 2025 together with a revised Annex 11 and Chapter 4, and the consultation closed in October 2025. A final text is targeted for the fourth quarter of 2026, but no adoption date has been announced, so check the current status before relying on it.
Can generative AI be used in GMP?
Under the 2025 draft, generative AI and large language models should not be used in critical GMP applications. They can be used in non-critical applications where a qualified person reviews the output and remains accountable. The EMA's 2026 workshop on generative AI means this position may be refined in the final text.
How do you validate an AI model for GMP use?
Define the intended use and the input data the model will see, set acceptance criteria before testing based on the performance of the current process, test on independent and representative data including rare cases, document the results, then keep the locked model under change control with ongoing performance and drift monitoring.
