03 · Data and labels
Logging training examples
3 min read
The system’s requirements now tell us what to predict and how much traffic to expect. The next question is how its activity will become training data. Logging records what happened when the application made a decision, so that information can later be connected to an outcome and used to construct a training example.
Record what the user actually saw
An impression is a recorded exposure to an item, such as a recommendation appearing on screen. An item considered by the model may never be displayed, so the absence of a click does not tell you whether the user rejected it. Defining what counts as an exposure, including whether the item was actually visible, helps establish which interactions can support a meaningful training label.
An impression event can include request ID, pseudonymous user ID, item ID, position, event time, experiment assignment, model version, and the relevant serving features or a reference to their snapshot. A later interaction refers to the same impression. This lets you distinguish the recommendation that led to an action from other appearances of the same item.
| Record | Example fields | Purpose |
|---|---|---|
| Impression | request_id, item_id, position, served_at | Establish exposure and context. |
| Prediction | score, model_version, feature_version | Reconstruct the decision. |
| Outcome | impression_id, action, occurred_at | Create a label after observation. |
Expect late and duplicate events
Connecting an exposure to its outcome requires accounting for how events arrive. A mobile client may upload an interaction after reconnecting, while a retry may deliver an event more than once. Stable event identifiers let the pipeline remove duplicates, and separate occurrence and arrival timestamps let it recognize late records. The dataset builder can then wait for an agreed period and revise affected data when necessary.
For a toy watch target, you might label a displayed video positive when a qualifying watch is observed within a defined session window. An exposure with no qualifying watch becomes negative only after that window and an allowance for late events. The window is part of the target definition, so changing it creates a new dataset version.
The join also needs validation through counts of outcomes without matching impressions, impressions without predictions, and duplicate identifiers. If fewer outcomes can be joined after a client release, the resulting dataset may suggest a change in model quality even though user behavior has remained the same. Tracking that match rate helps distinguish a logging problem from a genuine change in outcomes.
Keep the log useful and bounded
Logging every raw payload indefinitely creates cost and privacy risk. Identify the information needed for training, evaluation, incident investigation, and deletion requests. Prefer stable references or minimized fields where they are sufficient, and set access and retention rules for sensitive content.
Define an event schema with an owner and version. Validate required fields and plausible ranges at ingestion, then monitor missingness downstream. A schema check catches a removed field; a distribution check can catch a field that still exists but has changed meaning, such as milliseconds being replaced with seconds.
Tracing one impression through these records shows whether the pipeline can reconstruct the circumstances of a prediction and its later outcome. The remaining question is how to interpret that outcome when it arrives late, is ambiguous, or is never observed. That is the role of the label definition, which we will examine next.