Getting a model to work is the easy part. Keeping it working, in production, against data that shifts under you, with a team that needs to understand why it did what it did, that is the job, and it is mostly unglamorous engineering rather than modelling.
The gap between notebook and production
A model that scores well in a notebook has been evaluated on data that was collected, cleaned and split under ideal conditions. Production has none of those properties. Features arrive late or not at all, distributions drift, and the thing you were predicting changes meaning when the business changes a process.
The practical consequence is that most of the code around a production model is not modelling code. It is validation, monitoring, fallbacks and reproducibility.
Feature pipelines are the real system
Training-serving skew (where features are computed differently in training and at inference) is the most common and most costly bug in production machine learning. It is also silent: the model simply gets worse, and nobody knows why.
- Compute features once, in one place, used by both training and serving
- Version feature definitions the way you version schema migrations
- Validate inputs at serving time and reject or flag out-of-range values rather than quietly clamping them
- Store the exact feature values used for each prediction, so you can reproduce any decision later
Monitoring what actually matters
Uptime and latency monitoring will tell you the service is running. They will not tell you it has become useless. You need model-level monitoring alongside the infrastructure kind.
Input drift
Track the distribution of incoming features against the training distribution. Sustained shift is a warning even before accuracy moves, and it usually has a business explanation worth knowing about.
Prediction drift
Watch the distribution of outputs. A classifier that suddenly predicts one class 90% of the time is telling you something has broken upstream, and it will tell you days before the accuracy metrics catch up.
Delayed ground truth
Most real problems only reveal whether a prediction was right weeks later. Build the join between prediction and outcome deliberately, and accept that your accuracy dashboard is always looking at the past.
Retraining on purpose
Automatic retraining on a schedule sounds responsible and often is not. Retraining on drifted or contaminated data bakes the problem in, and a pipeline that deploys automatically will do it at three in the morning without anyone noticing.
We prefer triggered retraining with a human gate: drift or performance thresholds raise a flag, the new model is evaluated against a held-out set and the current production model, and a person approves the promotion. It is slower and it has never once shipped a worse model by accident.
Making decisions explainable
Somebody will eventually ask why the model rejected a particular case, and 'the model decided' is not an answer that survives a customer complaint or a regulator. Log the inputs, the model version, the prediction and the contributing factors for every decision.
This is worth doing even where you are not obliged to. The teams who trust a model are the ones who have looked at its reasoning on cases they know the answer to and found it sensible.
The boring practices that pay
- Pin everything. Model artefacts, data snapshots, library versions and preprocessing code, together, immutable
- Shadow deploy first. Run the new model alongside the old one on live traffic without acting on it
- Keep a fallback. A simple heuristic that runs when the model is unavailable is better than an outage
- Set a floor. Agree the performance level below which you roll back, before you need it
- Write the runbook. The person on call at 2am is not going to be the person who built it
Start simpler than feels right
A logistic regression in production, monitored and understood, beats a gradient-boosted ensemble that nobody can debug. It also gives you a real baseline, which is the only way to know whether the complicated version is actually earning its keep.
We have replaced more complex models with simpler ones than the other way round, and the business outcome improved every time, not because the simple model predicted better, but because the team started trusting it enough to act on it.
If you have models in production that nobody is quite watching, we can help put the monitoring and retraining discipline around them.



