02 · Framing the problem
ML objectives and product metrics
3 min read
Requirements describe what the application should accomplish, but a model needs a more specific task that can be learned from data. This means defining the information it receives, the outcome it predicts, and how the application will use the prediction. The connection between these choices determines whether better model predictions are likely to improve the experience you intended.
Translate the product goal into a target
The target is the outcome you ask the model to predict, while a label records the observed outcome for a training example. For a video feed, a click may be easy to observe without indicating that the person enjoyed the video. Watching for a defined duration could be a more useful target, although a fixed duration may favor longer videos and still needs to be checked against the intended experience.
A training example would then contain the user and session information available when the video was displayed, the video itself, and a label recording whether the viewing condition was met afterward. The model learns to estimate the probability of that outcome. Its training loss measures agreement with those labels, while separate product measures are needed to establish whether the resulting feed is useful.
A prediction needs a decision policy
A payment model can estimate fraud probability, while a decision policy determines whether that estimate leads to approval, blocking, or additional verification. Each action has different costs and consequences, so the policy needs more than a score alone. Similarly, predicted click probability can inform an advertising decision alongside bids and eligibility rules, with the combined policy determining what is shown.
A simplified cost calculation shows why the action needs its own policy. Assume that missing fraud costs 100 units, blocking a legitimate payment costs 5, and correct decisions have no cost in this example. For a fraud probability p, allowing the payment has expected cost 100p and blocking it has expected cost 5(1−p), so blocking costs less when p exceeds 5/105, about 4.76%. This calculation relies on calibrated probabilities, meaning that predicted risks agree with observed outcome frequencies, and omits other costs that a real policy would need to consider.
Measure the product effect separately
Choose a primary online outcome and guardrails before launching an experiment. For a video feed, one might test a satisfaction-related engagement measure while watching negative feedback, latency, and creator coverage. State how each measure is defined and whether it can move for an undesirable reason.
When several outcomes matter, the policy needs to explain how they are balanced. Combining click, watch, and dislike scores with arbitrary weights can produce unintended behavior, so begin with a primary outcome and explicit constraints before testing the tradeoffs. Controlled experiments can then assess the user effect without allowing an improvement in the average to conceal a violation of a firm requirement.
A longer-term outcome such as retention may be closer to the product goal but slower to observe and harder to attribute to one prediction. A faster signal can support training while experiments test whether it improves that longer-term outcome. With the input, target, and decision defined, we can next estimate how much work the system must perform and how quickly it needs to finish.