Answer · updated

Our fraud model only learns it was wrong months later. How do we know it still works?

Use three layers. Monitor input and score distributions and threshold traffic immediately. Track leading indicators that resolve in days, such as override rate and early contact. And reserve a small random sample for full outcome review regardless of what the model decided, because rejected cases never produce outcomes on their own.

Credit default, fraud recovery, churn, claims leakage, recidivism, clinical outcome: in each, the label that tells you whether a decision was right arrives long after the decision. Waiting for it means finding out about a failure two quarters late. Three layers of monitoring close most of that gap.

Layer 1: watch what you can see immediately

You do not need labels to detect that something has changed. Track the distribution of every input feature against the training window, the distribution of the model’s own output scores, the rate at which cases fall near the decision threshold, and the volume and mix arriving per segment and channel. A shift in inputs is not proof of a problem, but it is the earliest available signal and it is free.

This matters because the pipeline around the model is where reliability is lost. Work presented at NeurIPS in 2015 made the point that only a small fraction of a real machine learning system is the model code, with the surrounding data and serving infrastructure far larger and carrying continuing maintenance cost. Most “model drift” turns out to be an upstream field that changed meaning.

Layer 2: find proxies that resolve sooner

Between the decision and the true outcome there are usually intermediate events that correlate with it: an early missed payment, a customer contact, a manual override by an investigator, a case reopened. These arrive in days rather than months. They are not the ground truth and should never be scored as if they were, but as a leading indicator they are usually enough to raise a question early.

Override rate is particularly informative: when the people downstream start disagreeing with the model more often than they used to, something has moved.

Layer 3: buy real labels on a small sample

Reserve a small, randomly selected slice of cases for full outcome review regardless of what the model said, including some it scored as low risk. This is the only layer that yields unbiased truth, because the cases the model rejected never generate outcomes otherwise — a system that declines an applicant never learns whether they would have repaid. Without a reserved sample, your feedback data is shaped by your own past decisions and the model will look better than it is.

Build it into the pipeline, not into a report

Checks that live in a document are read after an incident. Checks that live in the pipeline can hold a release. An interview study of machine learning engineers published in 2022 described production ML as a continual loop governed by velocity, validation and versioning — the monitoring belongs inside that loop, with a named owner who is alerted, and with the data and model versions recorded so a regression can be traced to what changed.

Related questions

Why not just wait for the real outcomes?

Because you would discover a failure a quarter or two after it started. Input and score distributions, threshold traffic and segment mix are observable immediately and give the earliest available signal.

Why reserve a random sample if we already get outcomes?

Cases the model rejected never generate outcomes, so the feedback you receive is shaped by your own past decisions. A reserved random sample, reviewed regardless of the model’s decision, is the only unbiased measurement.

Is drift usually the model?

Often it is the pipeline. NeurIPS 2015 work observed that only a small fraction of a real ML system is the model code, with the surrounding data and serving infrastructure far larger; an upstream field changing meaning is a common cause.

Related Qylis capability

AI Training Platforms

Platforms that make data preparation, model training and evaluation a repeatable pipeline.

Discuss your requirement