← hristosbilis.ai
Note July 11, 2026

Security is now a GMP requirement

The 2025 Annex 11 rewrite puts pen-testing, MFA, and audit trails into GMP text — so for AI systems, the security controls and the validation controls are now the same controls.

annex-11 annex-22 gmp validation ai-security part-11

Ask the AI-platform lead and the CSV lead who owns the LLM gateway. How often do both point at the other team? Security assumes validation has it, validation assumes security has it, and the gateway sits unowned in the gap between them — which is exactly what the 2025 EU GMP rewrite just closed.

For years, “security” and “validation” lived in different buildings. QA owned IQ/OQ/PQ; security owned firewalls and pen-tests; each assumed the other had the AI system covered. The revised EU GMP Annex 11 (2025 draft) quietly ends that arrangement: it writes cybersecurity controls into GMP text. For an AI system, the security controls and the validation controls are now the same controls.

Why this matters to you

If you run AI in a GxP workflow, a line item you used to call “security best practice” is becoming an inspection finding waiting to happen. Pen-testing and MFA now are cheap; a cybersecurity-driven GMP deviation discovered later is not. The cheapest way to read the Annex 11 rewrite is as a pricing signal: security rework is validation rework — pay for it on the validation side of the ledger, before the deployment, not after a 483.

What actually changed

Three drafts moved together in the 2025 EudraLex Vol. 4 cycle (consultation closed 7 Oct 2025):

  • Revised Annex 11 — computerised systems.
  • New Annex 22 — AI in GMP. Its current line: static + deterministic models only; GenAI/LLMs excluded from critical GMP.
  • Revised Chapter 4 — documentation.

The plot twist is the pairing. While Annex 22 fences GenAI out of critical use, the revised Annex 11 pulls cybersecurity in.

My read: the EU is doing two things at once that only look contradictory. Annex 22 gates the models it can’t yet trust; Annex 11 hardens the plumbing everything else runs on. Together the message is “we’ll be conservative about what you deploy and uncompromising about how you secure it” — and “how you secure it” is now validation scope, not an IT afterthought.

The headline: Annex 11 made security a validation deliverable

The revised draft adds an entire §15 Security, plus §11 Identity & Access and §12 Audit Trails. The named clauses are the credibility proof — these are GMP text now, not OWASP advice:

  • §15.19 — penetration testing / “ethical hacking… at regular intervals” for critical internet-facing systems.
  • §11.6 — MFA for remote access to critical systems; §11.10 — least-privilege + segregation of duties; §11.7 — auto-lock.
  • §15.8 — network segmentation + firewalls; §15.13 — timely patching (immediate for critical vulnerabilities); §15.20 — encrypted remote access.
  • §12 — audit trails: who / what / when / why, locked and peer-reviewed.
  • §7 — supplier & service management for outsourcing and cloud: the regulated user stays responsible, runs a risk-based supplier audit, and keeps an exit strategy to retain data (§7.5.viii).

Here’s the thesis, the way I say it out loud:

For an AI system in a GxP workflow, your security controls and your validation controls are now the same controls. A red-team attack pack isn’t security theater — it’s OQ evidence against Annex 11 §15.19.

The crosswalk (the part to save and share)

Annex 22 reads foreign to a security person and familiar to a CSV validator. The translation: it’s mostly the GAMP 5 / Part 11 / Annex 11 lifecycle with AI-specific test-data and explainability bolted on. A trimmed map:

Annex 22 (draft)What it requiresGAMP 5 / Part 11 analogRevised Annex 11 (2025)
3 Intended Usetask, input sample space, subgroups, HITL responsibilityURS / intended use (V-model)§6 System Requirements
5 Test Datarepresentative, sized, labelled; synthetic discouragedtest-data integrity; §11.10(b)§10 Handling of Data
6 Test-Data Independencytrain/test separation, access control, audit trail, 4-eyessegregation of duties; §11.10(d)§11 Identity & Access; §12 Audit Trails
7 Test Executionapproved test plan, deviations investigated, evidence retainedOQ/PQ, RTM; §11.10(a)§9 Qualification & Validation
8 Explainabilityfeature attribution (SHAP/LIME), justification at approvalno clean CSV analognew / AI-specific
9 Confidencelog a confidence score, set a threshold, flag “undecided” outputsboundary/alarm design (loose fit)§8 Alarms; §10.1 input plausibility
10 Operationchange control, drift monitoring, human-review recordsperiodic review; §11.10(e)§14 Periodic Review; §12; §15 Security

The two AI-specific rows are the new territory: §8 Explainability has no clean GAMP precedent, and §9 Confidence maps only loosely (to alarm and threshold design). Everything else is your existing validation muscle, re-pointed at a model.

Why this hits the FDA-side reader too

This isn’t EU-parochial. The US FDA guidance never banned GenAI — it’s risk-based (Context of Use + influence × consequence). After industry pushback (ISPE, EFPIA), EMA reopened the GenAI question at a multistakeholder workshop on 30 Jun–1 Jul 2026, reframing it from whether to allow GenAI to how to control it. Joint FDA–EMA AI principles (Jan 2026) reinforce the same direction. So whichever framework you live under, you end up owing the same thing: a documented, risk-scaled evidence package — and security controls are now part of it. Nobody gets a free pass; that proof burden is the work.

A prediction, on the record: by the end of 2027, the final Annex 22 will not carry a categorical “GenAI banned from critical GMP” line — it’ll be replaced by a risk-based, guardrail-plus-evidence pathway much closer to the FDA credibility model. Confidence: medium. The operative draft still bans it, and “reassessing” is not “permitting,” so I may be wrong on timing — but between the industry comment letters, the reframed EMA workshop, and the joint FDA–EMA principles, the current only runs one way. Hold me to it.

What the EMA workshop actually signaled

The workshop closed days before this went out, and the direction it set matters more than any soundbite. EMA ran it as two days — an open, publicly-broadcast expert session on 30 June, then a closed Annex 22 drafting-group review on 1 July — explicitly to “shape a risk-based approach to the use of generative AI in medicines manufacturing.” Read the phrasing: the agency’s own language has moved from whether to allow GenAI to how to control it.

The substance is the evidence question EMA put to the room: what type and level of evidence — validation data, stress-testing results, failure analyses — would justify guardrails as reliable risk mitigation for GenAI in GMP functions? That shifts the standard from tool adoption to control substantiation — almost word for word, the “security controls are validation evidence” argument this post makes.

Two caveats keep this honest. The output report is still pending — EMA expects to publish a summary of expert contributions that feeds the next Annex 22 draft, and that report is the real checkpoint. And nothing is law: the operative draft still excludes GenAI from critical GMP. But the current runs toward risk-based credibility, and under either regime you owe the same package.

Show, don’t tell

Two working artifacts on this site make the point in code, not claims:

  • An ALCOA+ audit trail for LLM calls — a tamper-evident, hash-chained record for every model interaction. It’s a security control and, read the other way, it’s your Annex 11 §12 audit trail (who / what / when / why, locked) and Part 11 §11.10(e) evidence. Same artifact, two ledgers.
  • An LLM gateway from scratch — the provider key moves server-side and every call becomes an attributable, budgeted, policy-checked record. That’s Annex 11 §11 identity-and-access and least-privilege, enforced at the chokepoint instead of requested in a code review.

Neither started life as a “validation” project. Both emit validation evidence as a byproduct of being built securely — which is the whole thesis in miniature.

So what — for Monday

A short checklist you can act on this week:

  1. Treat your LLM gateway as a computerised system under Annex 11 — URS, validation, periodic review. Not “just infrastructure.”
  2. Inventory which §15 / §11 / §12 controls your AI stack already meets versus doesn’t.
  3. Reframe red-team and pen-test results as OQ evidence, traceable to the specific clause.
  4. For SaaS AI vendors, ask the §7 questions: supplier audit, audit-trail export, data exit strategy, and cloud auditability.

Caveats

All four documents are draft / non-final. Annex 22, the revised Annex 11, and Chapter 4 closed consultation on 7 Oct 2025; Annex 22’s GenAI line is being reopened. “Critical GMP application” is still undefined — the gate everything hinges on.


Deploying GenAI in a regulated workflow and stuck between an AI team that doesn’t speak validation and a CSV team that doesn’t speak AI? That gap is the work I do — get in touch.