Reconstructed from your brief, not from source code. No repo was available for this project, so the specifics below are a plausible, defensible architecture written from what you described (GDPR-compliant, Azure-native, RAG-based, for validation and audit support), built on the same real patterns as your EventQuoter system. Review it against what you actually built before using it in an interview — adjust anything that doesn't match, especially names, document types, and specific Azure resource choices.
Validation & Audit Assistant
A GDPR-compliant AI system built for a Cork-based pharmaceutical manufacturer, hosted entirely inside their own Azure tenant, that retrieves grounded context from the company's validation protocols, SOPs and audit history to help quality teams draft and review validation summaries, deviation reports, and audit findings — with every output traceable to source documents and every action logged for audit trail purposes.
What It Is
Pharmaceutical manufacturing runs on documentation: equipment and process validation protocols, standard operating procedures, batch records, deviation and CAPA (corrective and preventive action) reports, and the audit trails regulators expect to see behind all of it. Writing and reviewing that documentation is slow, specialist work, and quality teams were spending disproportionate time drafting validation summaries and cross-referencing prior deviations rather than on the judgment calls only a qualified person can make.
This system gives quality and validation staff an AI assistant that's grounded entirely in the company's own controlled documents — SOPs, prior validation protocols, equipment qualification records, and historical deviation/CAPA reports — retrieved via Azure AI Search and used as context for drafting, never as a source of invented facts. Every draft the AI produces cites the specific document and section it drew from, and nothing it produces is final: a qualified reviewer signs off before anything enters the company's actual quality management system.
The non-negotiable constraint going in was data residency and control: because the system touches regulated manufacturing data, it was built to run entirely inside the client's own Azure tenant and subscription, in an EU region, with the client's own security team retaining full control over access, logging and retention — not a third-party SaaS sitting in front of their documents.
Retrieve → Draft → Cite → Human Sign-Off
Retrieval — scoped to the right document class before anything else
Controlled documents are chunked (respecting section and table boundaries — a validation protocol's acceptance criteria table is exactly the kind of thing that must never be split mid-row) and indexed in Azure AI Search with metadata fields for document type, site, equipment ID, and approval status. A query for a deviation report never retrieves from an unrelated SOP category by accident — retrieval is filtered by document class and status (approved/current only, not superseded versions) before relevance ranking runs.
Drafting — Azure OpenAI, with retrieved chunks as the only permitted source of fact
The system prompt is explicit that the model may only state facts present in the retrieved context, must flag anything it cannot support from a cited source, and must never state a conclusion about compliance or acceptance criteria being met — that judgment belongs to the qualified person reviewing the draft, not the model.
Citation — every claim traces to a specific document and section
Each retrieved chunk carries its source document ID, title, version and section back through to the generated draft, so a reviewer can click through from any sentence in the AI's draft to the exact controlled-document passage it came from — the same discipline EventQuoter uses to trace a quoted price back to the catalogue line it matched, applied here to a regulated document instead of a price sheet.
Human sign-off — the only step that makes anything official
No AI output is written into the client's quality management system directly. A qualified reviewer edits, approves, or rejects each draft, and that approval — not the AI's output — is the record of record. The system's job is to save drafting time, not to replace the review step regulators actually require.
Tech Stack
Architecture & Design Decisions
Deployed inside the client's own Azure tenant, not a shared SaaS instance
Unlike EventQuoter's multi-tenant model, this system runs as a single-tenant deployment inside the pharma company's own Azure subscription. Their security and compliance team retains administrative control over the resource group, access policies, and retention settings — the system is built to their existing governance rather than asking them to trust a third party with regulated manufacturing data.
Document-class-aware retrieval, not one flat index
SOPs, validation protocols, equipment qualification records, and deviation/CAPA reports are indexed with explicit document-type and status metadata rather than as one undifferentiated pool of text. A query drafting a new deviation report retrieves prior deviations and relevant SOPs, filtered to currently-approved versions — never a superseded protocol revision, which in a regulated environment is exactly the kind of mistake that has to be structurally prevented, not just discouraged by a prompt.
The model is explicitly barred from making the compliance call
The system prompt and the UI both make it structurally clear that drafting is the AI's job and judgment is the reviewer's. The assistant can summarize what a protocol's acceptance criteria say, surface a relevant prior deviation, or draft the prose of a report — but it never states that a batch passed, that a deviation is closed, or that a validation is complete. That line exists because a model confidently asserting a compliance conclusion is a liability a regulated company can't accept, however good the retrieval grounding is.
Every retrieval and every draft is logged for audit purposes
Each query, the documents retrieved in response to it, the draft produced, and the reviewer's eventual approval/edit/rejection are all written to an append-only audit log — who asked, what was retrieved, what was generated, who signed off and when. This mirrors the audit-trail requirement pharma QA already has for every other step in their process; the AI assistant's involvement had to be exactly as traceable as a human's.
Private networking end to end — nothing publicly reachable
Azure AI Search, Azure OpenAI, the database and Blob Storage all sit behind private endpoints inside the client's own virtual network, with public network access disabled at the resource level. A leaked connection string or API key doesn't translate into exposure, because none of these services accept connections from outside the VNet in the first place.
Role-based access tied to the client's existing Entra directory
Rather than standing up a separate user database, authentication runs through Microsoft Entra ID against the client's own directory, so access to the assistant follows the same groups and role assignments the company already manages for its other systems — a validation engineer sees validation-document retrieval; a QA reviewer sees the approval queue; access changes when someone's role changes in the directory they already maintain, not in a second system that can drift out of sync.
EU data residency as an architectural default, not a policy statement
Every resource — the OpenAI deployment, the Search service, storage, compute — is pinned to an EU Azure region, and Azure OpenAI's terms mean prompts and completions aren't used for model training. For a GDPR-covered company handling employee and process data (and, depending on scope, data that touches clinical or batch records), keeping everything inside one EU region with private networking is what makes the compliance claim actually true, rather than a claim layered on top of an otherwise ordinary architecture — the same principle EventQuoter applied for its own GDPR story.
GDPR and GxP, Side by Side
| Requirement | How the architecture addresses it |
|---|---|
| Data residency (GDPR) | All Azure resources pinned to one EU region; no cross-border data transfer by default |
| No training on customer data | Azure OpenAI's enterprise terms keep prompts/completions out of model training |
| Access control & least privilege | Entra ID role-based access mapped to the client's existing directory groups |
| Audit trail (GxP / Annex 11 style expectations) | Append-only log of every query, retrieval, draft and human sign-off decision |
| Document version control | Retrieval filtered to current-approved-version metadata; superseded documents excluded by default |
| Human accountability for conclusions | AI never states a compliance conclusion; a qualified person signs off on every output |
| Secrets handling | Key Vault references + Managed Identity; no credentials in app config or source control |