How I Structure Large AI Projects
AI-heavy projects don't usually fail because the model is bad. They fail because the system around the model was never designed to change, and AI features change constantly. My approach: architecture and layer boundaries before features, sprints as a forcing function rather than a formality, tests as a safety net for the parts that don't involve a model, and documentation as a running habit instead of a deliverable at the end.
The complexity problem is different with AI in the mix
Every growing codebase accumulates complexity. That's not new. What's different about AI-heavy projects is how fast the requirements shift underneath you — a prompt that worked well last month needs restructuring, a model gets swapped for a cheaper or better one, a feature that used to be a simple API call now needs retrieval, grounding, and a fallback path for when the model returns something unusable.
If the system isn't structured to absorb that kind of change, every one of those shifts turns into a small rewrite. The projects I've seen hold up well under this — ViaNova and ClauseGuard AI both, so far — share a few habits in common, and they're less about AI specifically and more about basic engineering discipline applied consistently.
Architecture before code, every time
On every project past a certain size, I spend real time on architecture before writing feature code — usually Clean Architecture with clear domain, application, and presentation layers, sometimes formalized as Domain-Driven Design when the problem domain is complex enough to warrant it (ViaNova's job-application domain, for instance).
The specific pattern matters less than the discipline behind it: domain logic should not know that a specific LLM provider or a specific UI framework exists. When that boundary is respected, swapping OpenAI for Gemini, or adding a new AI feature, is a change contained to one layer — not a ripple through the whole system.
Understand the actual problem before naming any technology
Layer boundaries defined before feature work starts
Sprint-scoped, reviewable, never a giant unreviewable diff
Catch regressions before they reach anyone relying on the system
This is essentially the same nine-step process I follow on every project, laid out in more detail on Mission Control — it isn't AI-specific, but it matters more on AI projects because the cost of skipping it compounds faster.
Sprints as a forcing function, not a ritual
I plan and build in sprints, but not because "agile" is the expected word to say. Sprints are useful here for a specific reason: they force me to decide, on a fixed cadence, what's actually load-bearing right now versus what can wait.
On ViaNova, that meant foundation, workspace UI, architecture, and persistence came first — deliberately before any AI feature — because a shaky data model would make every later AI feature harder to build correctly. On ClauseGuard AI, foundation and file processing came before the AI pipeline and vectorization, for the same reason: retrieval and scoring only work if the data underneath them is already trustworthy.
The pattern in both cases: sprint boundaries tend to fall along architectural boundaries, not just feature boundaries. That's a sign the sprints are doing their actual job.
Separating concerns so AI can be added without a rewrite
This is worth calling out on its own, because it's the single habit that's saved me the most rework. If AI logic — prompts, model calls, retrieval — lives mixed into the same code as UI rendering or database access, then every change to the AI side risks breaking something unrelated, and every change to the UI risks breaking the AI integration.
Keeping AI logic behind the same service-layer boundary as everything else means it can be added, swapped, or removed without destabilizing the rest of the system. I didn't fully appreciate how much this mattered until I was partway through ClauseGuard AI's risk engine and realized I could change the scoring formula, the embedding model, or the prompt structure independently of each other — because none of them knew the others existed.
Testing discipline, even on the parts that feel "unimportant"
I've come to treat testing as non-negotiable, not because every project demands full coverage, but because AI systems are already the least predictable part of the stack — the deterministic parts around them shouldn't add to that uncertainty. On a full-stack project I worked on at CITL — Bukidnon State University's Syllabus Automation system — I instituted Test-Driven Development using Cypress and Jest specifically because two developers touching the same codebase without a safety net is how regressions slip through unnoticed.
That experience shaped how I think about testing on AI projects generally: you can't easily write a deterministic test for "did the LLM give a good answer," but you absolutely can and should test everything around it — the validation layer, the data mapping, the retrieval logic, the API contracts. Tighten what you can control so the model is the only unpredictable variable left, not one of several.
Documentation as a habit, not a deliverable
I keep a running Knowledge Base and update project case studies as I go, rather than trying to write everything up after the fact from memory. Two reasons. First, decisions made under pressure in week three make a lot more sense to future-me if I wrote down *why* at the time, not just *what*. Second, on projects with more than one contributor, undocumented architecture decisions become tribal knowledge that only lives in one person's head — which is exactly the kind of single point of failure I'm trying to design out of the system in the first place.
Mistakes I've made
Being honest about this matters more than making it sound solved. Early on, I've been tempted to skip the architecture step "just this once" because a deadline was close — and every time, the AI feature I bolted on afterward took longer to integrate than it would have if the boundary had existed from the start. I've also under-tested a retrieval pipeline once, assumed the embeddings were good because the demo looked fine, and only found the gaps once real, messier data went through it. Both mistakes taught me the same lesson from different angles: the parts of an AI system that feel "boring" — data validation, layering, tests — are usually the parts that determine whether the AI parts actually work reliably.
The short version
Structure the system so the AI is a replaceable component, not the foundation everything else depends on. Use sprints to force honest sequencing decisions. Test everything deterministic so the model isn't the only source of uncertainty. Write things down as you go. None of this is unique to AI projects — it's just regular engineering discipline, applied to a class of project that punishes skipping it faster than most.