Reliability & risk

What is model drift?

An AI feature is not done at launch. Without monitoring, performance decays silently, which is why running AI is an ongoing operating cost, not a project that ends.

In plain terms

Model drift is the gradual decline in an AI system's accuracy as the real world moves away from the data it learned from.

Model drift means the system gets worse because the world moved and the model did not. Customers phrase requests differently. Products change. A vendor updates the model behind your account. Last quarter's accuracy is not a promise about next quarter.

Watch the outcome, not the launch score. Pick a few numbers you already trust, such as how often staff undo the system's suggestion, and look at them on a schedule. When they slip, the fix might be newer documents, a new test, or a held model update. A system with no owner will drift until a customer notices.

Think of it this way: Model drift is what happens when you train a weather forecasting model in summer and try to use it in winter. The world changed; the model did not.

A demand forecasting model trained before a major market shift underperforms after it. The team that monitors drift catches it in weeks. The team that does not discovers it from a stockout.

A retailer trained a returns classifier on last year's reasons. A new product line created a new reason code. The classifier dumped those returns into other, and the warehouse lost a week before anyone saw the other bucket grow. A weekly count would have shown it on day two.

Monitoring for drift is part of every production AI deployment. Budget for it at design time, not as a reactive fix after quality has already declined in production. Do not assume a model that was accurate at launch will remain accurate without monitoring. All production AI systems require ongoing evaluation as the world they predict keeps changing.

RaftLabs treats this as part of the build: a source on the answer, a test set, and a record of what the system did. The launch is the start of that work, not the end. The related work on our side is MLOps.

This sits with the other reliability & risk terms on the glossary. Why a confident answer can still be wrong, and how you catch it. Worth reading next: Hallucination, Guardrails, and Evaluation (evals).

Common questions

On a rhythm that matches the cost of being wrong, and after every model or document change. A weekly look at a short list of cases is enough for many internal tools. Customer-facing or regulated decisions deserve a closer watch. The failure mode is checking once at launch and never again.
That is one cause. Your own data changing is another. Both show up as worse answers on cases that used to pass. Keep those cases, rerun them after vendor updates, and watch the business numbers in between. You need both checks.

Work with us

Tell us what's broken.

Tell us what's not working in your business. We'll find the real problem and tell you exactly what it would take to fix it.

  • Scope and cost agreed before work starts. No surprises. No obligation.
  • Working prototype within 3 weeks of kickoff.
  • Pay by milestone. You see progress before each invoice.
  • 60-day post-launch warranty. Bug fixes, UI tweaks, and deployment support. No retainer.
  • All conversations are NDA-protected.