The failure that shaped the whole system
EventQuoter is a multi-tenant AI quotation platform for AV event-production companies: a sales rep describes a job in plain English and gets a priced, client-ready proposal back in seconds, grounded in that client's own equipment catalogue via RAG over Azure AI Search. The first version did everything — reading the brief, pricing it, writing the email — in a single Azure OpenAI prompt. It shipped, and it occasionally hallucinated: a plausible-sounding price, or a piece of equipment that didn't exist in the catalogue at all. For a product whose entire value is a quote a rep can send without double-checking, that's worse than no product.
Why a better prompt wasn't the fix
A large language model is very good at understanding messy, ambiguous language and writing fluent prose back. It has no structural guarantee of arithmetic correctness, or of faithfulness to one specific dataset it was never trained on — your price list. Asking it to also do the pricing math was asking it to do the one thing it's least suited to, and no amount of prompt tweaking changes that underlying fact.
The fix: three stages, two of them AI, one of them plain code
The pipeline now runs Extract → Decide → Price → Write, and only two of those steps are AI at all.
- Extract — Azure OpenAI reads the free-text brief and turns it into structured data: event type, attendance, equipment keywords. Nothing here is trusted as final; a strict schema check runs on the output before anything downstream sees it.
- Decide & Price — plain Python, no AI involved at all. It matches each requested item against the client's real inventory table, pulls the real price, and calculates the total. An item that doesn't match isn't silently dropped or guessed at — it's flagged, and the quote carries a confidence score so a rep can see exactly which lines need a manual check.
- Write — Azure OpenAI gets handed the already-priced line items, plus similar past quotes retrieved from that client's own RAG index for tone and context, and writes the final client-facing email. By the time the model is writing sentences, every number has already been decided by code.
The RAG layer that keeps it grounded
Each client's equipment catalogue and historical quotes are chunked, embedded and indexed per-tenant in Azure AI Search, with every query scoped to a hard tenant filter so one client's pricing can never surface in another client's quote — and hybrid search so an exact model number typed by a rep is never missed by a looser semantic match. This is the same RAG pattern behind the AI compliance work I do, applied to a different problem: grounding generated text in your real documents instead of the model's general training.
Plain-English takeaway
Use AI for the ambiguous parts. Use code for the parts that must be right.
The AI's job narrowed to exactly two things: turn messy English into structured data, and turn already-correct structured data back into good English. Everything in between — the actual pricing — is deterministic, testable, debuggable Python.
Where else this pattern applies
Any AI system that produces a number someone will act on — a quote, an estimate, a risk score — benefits from the same split: let the model handle language, let code handle arithmetic and lookups against real data, and make the model's output traceable back to a real source rather than a guess.