03 · Data and labels
Labels and delayed outcomes
3 min read
Logging connects a decision to later events, but those events still need to be interpreted before they become training labels. A label is the value assigned to an example to represent the outcome the model should learn to predict. Its reliability depends on who observed that outcome, how long observation takes, and which cases remain uncertain.
Define positive, negative, and unresolved
For payment fraud, a confirmed fraudulent chargeback can provide a positive label. A recent payment without a chargeback is not automatically a trustworthy negative, because disputes may arrive later. Some outcomes may remain ambiguous or never be observed. Preserve an unresolved state instead of converting uncertainty into zero.
A proxy is an observable signal used to stand in for an outcome that is harder to establish directly. User reports, automated rules, reviewer decisions, and confirmed disputes can therefore provide different evidence about the same payment. Recording where each signal came from and when it arrived helps determine whether it is suitable for training, evaluation, or a separate supporting task.
Allow enough time to observe the outcome
For an illustrative dataset with a 60-day dispute observation window, a payment made yesterday remains unresolved because its outcome may still arrive. An older payment with no observed dispute can become a negative under this policy, although later disputes and unreported fraud remain possible. Measuring those limitations helps establish how much confidence to place in the resulting labels.
A cohort is a group of examples from the same period, and it is considered mature once the chosen observation window has passed. Using mature cohorts gives outcomes more time to arrive, although the data is consequently older. Faster signals such as review decisions can provide a more recent view, but their reliability and selection process need to be assessed separately before combining them with confirmed outcomes.
Understand who gets labeled
A fraud system that blocks a payment may prevent the very outcome it wants to observe. A moderation queue sees the items an earlier model flagged. User reports overrepresent people who noticed a problem and chose to report it. These are selection mechanisms, not random samples of the population.
Audits or independently reviewed samples can help estimate errors outside the examples selected by the current system. For subjective labels, recording reviewer disagreements and how they were resolved helps reveal unclear policy boundaries. Clarifying those boundaries improves the target definition before asking a model to learn from inconsistent decisions.
For human-reviewed labels, a shared definition and a process for resolving disagreements help establish which examples are dependable enough to use. Once those labels are available, we also need to decide how many examples of each outcome should enter training. The next topic examines that choice when some outcomes occur much less frequently than others.