Reconstructed from your brief, not from source code. No repo was available for this project, so the specifics below are a plausible, defensible architecture written from what you described (GDPR-compliant, Azure-native, RAG-based, for validation and audit support), built on the same real patterns as your EventQuoter system. Review it against what you actually built before using it in an interview — adjust anything that doesn't match, especially names, document types, and specific Azure resource choices.

// Azure OpenAI + Azure AI Search · pharmaceutical validation & audit support · EU data residency

Validation & Audit Assistant

A GDPR-compliant AI system built for a Cork-based pharmaceutical manufacturer, hosted entirely inside their own Azure tenant, that retrieves grounded context from the company's validation protocols, SOPs and audit history to help quality teams draft and review validation summaries, deviation reports, and audit findings — with every output traceable to source documents and every action logged for audit trail purposes.

Azure OpenAI (customer tenant) Azure AI Search (RAG) Microsoft Entra ID Azure Blob Storage Azure Key Vault + Managed Identity EU data residency · GDPR / GxP

What It Is

Pharmaceutical manufacturing runs on documentation: equipment and process validation protocols, standard operating procedures, batch records, deviation and CAPA (corrective and preventive action) reports, and the audit trails regulators expect to see behind all of it. Writing and reviewing that documentation is slow, specialist work, and quality teams were spending disproportionate time drafting validation summaries and cross-referencing prior deviations rather than on the judgment calls only a qualified person can make.

This system gives quality and validation staff an AI assistant that's grounded entirely in the company's own controlled documents — SOPs, prior validation protocols, equipment qualification records, and historical deviation/CAPA reports — retrieved via Azure AI Search and used as context for drafting, never as a source of invented facts. Every draft the AI produces cites the specific document and section it drew from, and nothing it produces is final: a qualified reviewer signs off before anything enters the company's actual quality management system.

The non-negotiable constraint going in was data residency and control: because the system touches regulated manufacturing data, it was built to run entirely inside the client's own Azure tenant and subscription, in an EU region, with the client's own security team retaining full control over access, logging and retention — not a third-party SaaS sitting in front of their documents.

Retrieve → Draft → Cite → Human Sign-Off

query (protocol / deviation / audit question)→ Azure AI Search retrieval, scoped by document type & site→ grounded draft with inline citations→ qualified reviewer approves / edits / rejects
1

Retrieval — scoped to the right document class before anything else

Controlled documents are chunked (respecting section and table boundaries — a validation protocol's acceptance criteria table is exactly the kind of thing that must never be split mid-row) and indexed in Azure AI Search with metadata fields for document type, site, equipment ID, and approval status. A query for a deviation report never retrieves from an unrelated SOP category by accident — retrieval is filtered by document class and status (approved/current only, not superseded versions) before relevance ranking runs.

2

Drafting — Azure OpenAI, with retrieved chunks as the only permitted source of fact

The system prompt is explicit that the model may only state facts present in the retrieved context, must flag anything it cannot support from a cited source, and must never state a conclusion about compliance or acceptance criteria being met — that judgment belongs to the qualified person reviewing the draft, not the model.

3

Citation — every claim traces to a specific document and section

Each retrieved chunk carries its source document ID, title, version and section back through to the generated draft, so a reviewer can click through from any sentence in the AI's draft to the exact controlled-document passage it came from — the same discipline EventQuoter uses to trace a quoted price back to the catalogue line it matched, applied here to a regulated document instead of a price sheet.

4

Human sign-off — the only step that makes anything official

No AI output is written into the client's quality management system directly. A qualified reviewer edits, approves, or rejects each draft, and that approval — not the AI's output — is the record of record. The system's job is to save drafting time, not to replace the review step regulators actually require.

Tech Stack

AI platform
Azure OpenAIDeployed in client's own tenant, EU region
Azure AI SearchHybrid (keyword + vector) retrieval, metadata-filtered
text-embedding modelChunk embeddings for vector search
Storage & data
Azure Blob StorageSource document landing zone, versioned
Azure Database for PostgreSQLReview/approval state, citation records
Identity & secrets
Microsoft Entra IDSSO into client's existing directory, role-based access
Azure Key VaultAll API keys and connection strings as references
Managed IdentityApp → Key Vault / Search / OpenAI, no stored credentials
Infrastructure
Azure App ServiceHosted inside client's own subscription
Virtual Network + Private EndpointsNo public access to Search, OpenAI, DB, Blob
Azure Monitor / Log AnalyticsAudit trail: who, what, when, which document

Architecture & Design Decisions

01

Deployed inside the client's own Azure tenant, not a shared SaaS instance

Unlike EventQuoter's multi-tenant model, this system runs as a single-tenant deployment inside the pharma company's own Azure subscription. Their security and compliance team retains administrative control over the resource group, access policies, and retention settings — the system is built to their existing governance rather than asking them to trust a third party with regulated manufacturing data.

02

Document-class-aware retrieval, not one flat index

SOPs, validation protocols, equipment qualification records, and deviation/CAPA reports are indexed with explicit document-type and status metadata rather than as one undifferentiated pool of text. A query drafting a new deviation report retrieves prior deviations and relevant SOPs, filtered to currently-approved versions — never a superseded protocol revision, which in a regulated environment is exactly the kind of mistake that has to be structurally prevented, not just discouraged by a prompt.

03

The model is explicitly barred from making the compliance call

The system prompt and the UI both make it structurally clear that drafting is the AI's job and judgment is the reviewer's. The assistant can summarize what a protocol's acceptance criteria say, surface a relevant prior deviation, or draft the prose of a report — but it never states that a batch passed, that a deviation is closed, or that a validation is complete. That line exists because a model confidently asserting a compliance conclusion is a liability a regulated company can't accept, however good the retrieval grounding is.

04

Every retrieval and every draft is logged for audit purposes

Each query, the documents retrieved in response to it, the draft produced, and the reviewer's eventual approval/edit/rejection are all written to an append-only audit log — who asked, what was retrieved, what was generated, who signed off and when. This mirrors the audit-trail requirement pharma QA already has for every other step in their process; the AI assistant's involvement had to be exactly as traceable as a human's.

05

Private networking end to end — nothing publicly reachable

Azure AI Search, Azure OpenAI, the database and Blob Storage all sit behind private endpoints inside the client's own virtual network, with public network access disabled at the resource level. A leaked connection string or API key doesn't translate into exposure, because none of these services accept connections from outside the VNet in the first place.

06

Role-based access tied to the client's existing Entra directory

Rather than standing up a separate user database, authentication runs through Microsoft Entra ID against the client's own directory, so access to the assistant follows the same groups and role assignments the company already manages for its other systems — a validation engineer sees validation-document retrieval; a QA reviewer sees the approval queue; access changes when someone's role changes in the directory they already maintain, not in a second system that can drift out of sync.

07

EU data residency as an architectural default, not a policy statement

Every resource — the OpenAI deployment, the Search service, storage, compute — is pinned to an EU Azure region, and Azure OpenAI's terms mean prompts and completions aren't used for model training. For a GDPR-covered company handling employee and process data (and, depending on scope, data that touches clinical or batch records), keeping everything inside one EU region with private networking is what makes the compliance claim actually true, rather than a claim layered on top of an otherwise ordinary architecture — the same principle EventQuoter applied for its own GDPR story.

GDPR and GxP, Side by Side

RequirementHow the architecture addresses it
Data residency (GDPR)All Azure resources pinned to one EU region; no cross-border data transfer by default
No training on customer dataAzure OpenAI's enterprise terms keep prompts/completions out of model training
Access control & least privilegeEntra ID role-based access mapped to the client's existing directory groups
Audit trail (GxP / Annex 11 style expectations)Append-only log of every query, retrieval, draft and human sign-off decision
Document version controlRetrieval filtered to current-approved-version metadata; superseded documents excluded by default
Human accountability for conclusionsAI never states a compliance conclusion; a qualified person signs off on every output
Secrets handlingKey Vault references + Managed Identity; no credentials in app config or source control

Q&A

Q
Why did this client need something built specifically for them rather than an off-the-shelf tool?
Their documents — validation protocols, SOPs, deviation and CAPA reports — can't leave their own Azure tenant, and a generic SaaS AI tool wasn't going to satisfy their quality and compliance team on data residency or audit-trail grounds. The requirement was an AI assistant that lived entirely inside infrastructure they already controlled, with every action as traceable as the paper-and-signature process it was supplementing.
Q
How did you prevent the AI from making an actual compliance judgment?
Structurally, not just by prompting nicely. The assistant's outputs are drafts with inline citations back to source documents; it's explicitly instructed never to assert that a batch passed or a validation is complete, and the application layer treats every output as unapproved until a qualified reviewer signs off. The AI's value is cutting drafting time, not making the regulatory call — that stays human, on purpose.
Q
How does the RAG retrieval avoid pulling the wrong document version?
Every indexed chunk carries document-type, site and approval-status metadata, and retrieval filters on current-approved-version before ranking runs — so a superseded SOP revision or a draft protocol never surfaces as if it were authoritative. That's a harder constraint here than in a typical RAG use case, because citing the wrong version of a controlled document isn't just a wrong answer, it's a potential finding in the client's own audits.
Q
How is this different from what you built for EventQuoter?
The RAG and Azure-native pattern is the same family of architecture, but the trust model is different. EventQuoter is multi-tenant SaaS with per-client isolation inside shared infrastructure; this is single-tenant, deployed entirely inside the client's own subscription, because the data sensitivity and regulatory stakes meant the client's own security team needed to retain direct control rather than trust a shared platform.
Q
What would a real audit of this system look for, and does it hold up?
An auditor would want to see exactly what was retrieved and generated for any given draft, who approved it, and proof the system can't silently use a superseded document. The append-only audit log and the document-status-filtered retrieval are built specifically to answer that question — the goal was a system that makes the client's own audits easier, not one that introduces a new thing to audit nervously.
Q
What would you improve or add next?
A structured feedback loop where reviewer edits to AI drafts get captured and analyzed over time — not to retrain a model on customer data, which their terms rule out, but to tune retrieval relevance and prompt instructions based on where human reviewers most often correct the draft. That's the highest-leverage next step once there's a real body of review history to learn from.