Model drift means the system gets worse because the world moved and the model did not. Customers phrase requests differently. Products change. A vendor updates the model behind your account. Last quarter's accuracy is not a promise about next quarter.
Watch the outcome, not the launch score. Pick a few numbers you already trust, such as how often staff undo the system's suggestion, and look at them on a schedule. When they slip, the fix might be newer documents, a new test, or a held model update. A system with no owner will drift until a customer notices.
Think of it this way: Model drift is what happens when you train a weather forecasting model in summer and try to use it in winter. The world changed; the model did not.
A demand forecasting model trained before a major market shift underperforms after it. The team that monitors drift catches it in weeks. The team that does not discovers it from a stockout.
A retailer trained a returns classifier on last year's reasons. A new product line created a new reason code. The classifier dumped those returns into other, and the warehouse lost a week before anyone saw the other bucket grow. A weekly count would have shown it on day two.
Monitoring for drift is part of every production AI deployment. Budget for it at design time, not as a reactive fix after quality has already declined in production. Do not assume a model that was accurate at launch will remain accurate without monitoring. All production AI systems require ongoing evaluation as the world they predict keeps changing.
RaftLabs treats this as part of the build: a source on the answer, a test set, and a record of what the system did. The launch is the start of that work, not the end. The related work on our side is MLOps.
This sits with the other reliability & risk terms on the glossary. Why a confident answer can still be wrong, and how you catch it. Worth reading next: Hallucination, Guardrails, and Evaluation (evals).