The shape every automation ends up taking

I've now built enough of these — an executive assistant that manages scheduling and correspondence, an email triage system, a Slack support agent grounded in internal docs, and ClauseGuard AI's contract risk engine — to notice they all converge on the same underlying architecture, even though they solve completely different problems.

That's not a coincidence. It's because the failure modes are the same every time: the model says something confidently wrong, a slow API call blocks something that shouldn't be blocked, or an action gets taken that a human really should have seen first. Once you've been burned by those a couple of times, you start designing around them by default.

Why the LLM never sits directly in the main request path

The most important decision, and the one I make first on every automation: the LLM call does not happen synchronously in the same request that triggered it.

Take ClauseGuard AI. Parsing and classifying a 20-page contract synchronously risks the request just timing out — nobody wants to stare at a loading spinner for a minute wondering if it crashed. So that entire pipeline is offloaded to an event-driven n8n workflow: a webhook kicks it off, the work happens asynchronously, and the result comes back through aggregation and notification once it's ready.

Trigger
Webhook or scheduled event fires
n8n orchestration
Routes the job, handles retries
Retrieval
Relevant context pulled before generation
LLM call
Structured, grounded generation
Human checkpoint
Review before anything ships

This same shape shows up in Email Triage (classify asynchronously, then route) and the AI Executive Assistant (intent gets parsed, a draft gets prepared, and a human reviews it before anything actually sends). The pattern isn't really about AI at all — it's standard practice for any slow, external, occasionally-unreliable dependency. LLMs just happen to be exactly that kind of dependency.

Human-in-the-loop as a design decision, not an afterthought

Every one of my automations has a point where a human looks at the output before it becomes real — before an email sends, before a calendar event gets created, before a risk score gets treated as final. This isn't caution for its own sake. It's a direct response to how these systems actually fail: not by crashing, but by confidently producing something plausible-sounding and wrong.

The AI Executive Assistant flow makes this explicit: request comes in, intent gets parsed, a draft gets prepared — and then it waits for review before the calendar, email, or task actually gets touched. That checkpoint is cheap to build and it's the single highest-leverage safety decision in the whole system.

Grounding the model instead of trusting it

The Slack RAG Support Agent and ClauseGuard AI both solve a version of the same problem: I need the model's answer to be based on something real, not on whatever it remembers from training. That's what retrieval-augmented generation is actually for.

In ClauseGuard's case, standard legal rubrics are embedded into a pgvector database ahead of time. Before Gemini scores a contract clause, the pipeline retrieves the relevant "good" and "bad" clause examples from that database and hands them to the model as context. The model isn't being asked "is this risky?" from a cold start — it's being asked to compare against grounded reference points. That's a meaningful difference, and it's the main lever for reducing hallucinated risk scores.

The Slack agent works the same way structurally: a question gets embedded, relevant internal docs get retrieved, and only then does generation happen — grounded in what was actually retrieved, not in whatever the model assumes.

Structured output over free text

Free-form LLM text is brittle the moment you need to actually do something with the output — store it, score it, route it, display it in a specific UI component. So wherever an automation's output needs to be used programmatically, I enforce strict structured (usually JSON) output from the model, rather than parsing prose after the fact.

ClauseGuard AI is the clearest example: risk scores need to map deterministically into PostgreSQL columns. Asking the model to "explain the risk in a paragraph" and then trying to regex a score out of that paragraph is fragile in a way that eventually breaks in production. Asking it to return a specific JSON shape, validated against a schema, isn't.

Security is part of the architecture, not a checkbox

Once an automation is handling anything sensitive — contracts, internal docs, personal scheduling — security decisions need to be made at the same time as the architecture, not bolted on afterward. For ClauseGuard AI specifically, that means designing around the CIA triad from the start: strict MIME-type validation on anything uploaded, prompt-injection sanitization before a document ever reaches the model, and secure storage with real data-retention lifecycle rules rather than keeping everything indefinitely by default.

The reason this has to be architectural rather than a later patch: prompt injection sanitization, for instance, has to happen before untrusted document content gets anywhere near the prompt. If that boundary isn't designed in from the start, it's very easy to accidentally build a pipeline where sanitization happens too late to matter.

Where this breaks down

I want to be honest about the limits of this pattern rather than presenting it as solved. Async orchestration adds real complexity — more moving parts, more places for something to silently fail, and a genuine engineering cost in observability (you need good logging and error handling around every hop, or debugging becomes guesswork). Retrieval only helps if what you're retrieving from is actually good; a RAG pipeline grounded in outdated or incomplete documentation will confidently produce wrong answers just as easily as an ungrounded model, just with more convincing citations. And human checkpoints only work if a human is actually paying attention to them — a review step that gets rubber-stamped every time provides the illusion of safety without the substance.

None of that makes the pattern wrong. It just means the architecture is a starting point that has to be paired with real operational discipline, not a guarantee on its own.