Building an AI product is mostly not an AI problem. The model is a component you rent. What decides whether the product works is everything around it: evaluation, retrieval quality, latency budgets, cost control, and the interface that sets user expectations honestly.
This is the guide we wish we had when we started shipping these systems for clients. It is opinionated, and it is written from the position of having had several of these opinions changed by production.
Decide whether you need a model at all
The first question is whether the problem genuinely needs a language model. A surprising amount of what gets scoped as AI work is better served by a database query, a rules engine, or a well-designed form. Models are a good fit when the input is unstructured, the output tolerates variation, and being right most of the time is genuinely useful.
If your problem needs to be right every time (invoice totals, compliance decisions, anything financial) the model can draft it, but something deterministic has to check it. Building that check is the work.
Architecture that does not paint you in
Treat the model as a replaceable dependency behind your own interface. Providers change pricing, deprecate versions and ship better models on their own schedule, and you want switching to be a config change rather than a rewrite.
- An abstraction layer over the provider, thin enough that it does not become its own problem
- Prompt versioning in source control, treated like migrations rather than like configuration
- Structured outputs validated against a schema, so downstream code never parses prose
- A cache on the boring repeated calls, which is usually where the bill actually comes from
- Graceful degradation when the provider is slow or down, because it will be
Retrieval is where quality comes from
For any product answering questions over your own data, retrieval quality dominates model quality. A strong model with poor retrieval produces confident nonsense. A modest model with excellent retrieval produces something people trust.
Chunking deserves more thought than it gets
Fixed-size chunks are the default and are almost always wrong. Split on semantic boundaries (sections, clauses, procedures) and keep enough surrounding context that a retrieved chunk is intelligible on its own. Overlap helps; so does storing the document title and section heading alongside each chunk.
Hybrid search beats pure vectors
Vector search alone fails on exact identifiers: part numbers, error codes, names. Combining it with keyword search and reranking the merged results is more work, and it is the difference between a demo and a product.
Cite everything
Every answer should point at the source it came from, and the system should say plainly when the corpus does not cover a question. Users forgive not knowing. They do not forgive being confidently misled, and one bad answer poisons trust in fifty good ones.
Evaluation before features
You cannot improve what you cannot measure, and vibes do not survive a model upgrade. Build an evaluation set early (a hundred real inputs with known-good outputs is enough to start) and run it before anything ships.
Pair automated scoring with a weekly human review of a sample. Automated metrics catch regressions; humans catch the failure modes you did not think to score for. Both are cheap compared with finding out from a customer.
Cost and latency are product decisions
Token costs look trivial in development and become the second-largest line item at scale. Instrument cost per request from day one, route simple requests to smaller models, and cache aggressively.
Latency is a design problem as much as an engineering one. Streaming makes a slow response feel acceptable. Showing intermediate steps makes a very slow one tolerable. A spinner for eleven seconds does not.
Designing for a system that is sometimes wrong
This is the part teams skip, and it determines adoption more than accuracy does. Set expectations in the interface, make it obvious how to verify an answer, and make correction cheap.
- Show confidence honestly, and route low-confidence cases to a human
- Make sources one click away, not buried behind a disclosure
- Let users correct output, and feed those corrections back into evaluation
- Never present a generated answer with the same visual authority as a verified fact
What we would do differently
On our first few builds we over-invested in prompt engineering and under-invested in evaluation. Prompts feel productive because feedback is immediate; evaluation feels like overhead until the day you need to upgrade a model and have no way to know whether it got worse.
We also underestimated data preparation every single time. Turning a messy archive into a clean, citable corpus is usually the majority of the work and almost never the majority of the estimate. Budget for it honestly, and start it before anything else.
If you are scoping an AI product and want a second opinion on the architecture, talk to us. We would rather tell you it is a database problem than build you a model that cannot fix it.



