03 · Data and labels
Sampling and class imbalance
3 min read
A reliable label definition does not guarantee that the dataset contains enough examples of every outcome. When one outcome is rare, training on all available records can spend most of its work on the common case. Sampling selects which examples enter training, while weighting controls their contribution, so both choices need to preserve a meaningful connection to the population the application will serve.
Start with the base rate
The base rate is the proportion of examples with the outcome of interest. In one million payments containing 1,000 fraud cases, it is 0.1%, so always predicting a legitimate payment gives 99.9% accuracy without detecting any fraud. If a model instead catches 800 fraud cases and flags 1,998 legitimate payments, it detects 80% of the fraud cases, called recall, while only about 28.6% of its 2,798 alerts are genuine fraud, called precision.
The false-positive rate in this example is just 0.2%. That small percentage still creates more false alarms than true detections because the negative population is so large. Always bring a rate back to counts at the expected workload, particularly when a human review queue must handle the output.
| Actual outcome | Flagged | Not flagged |
|---|---|---|
| 1,000 fraudulent payments | 800 true positives | 200 false negatives |
| 999,000 legitimate payments | 1,998 false positives | 997,002 true negatives |
Reduce training cost deliberately
Retaining all positives while sampling some negatives can reduce training cost, while weighting examples or changing batch composition can increase the contribution of rare cases during optimization. These methods affect learning differently, so the choice should follow a specific objective such as reducing computation or improving rare-pattern coverage. An equal number of examples from each class is not, by itself, evidence that the training setup matches the application.
Record the sampling probability or weight with each example so the selection can be accounted for later. Uniform sampling gives each negative example the same chance of inclusion. Hard-negative mining instead selects negatives that the current model finds difficult, which can help it learn distinctions but produces a deliberately unusual sample that may include more ambiguous or mislabeled cases.
Do not read sampled scores as population probabilities
Suppose you keep all positives but retain a fraction r of negatives uniformly. If q is a calibrated positive probability within the sampled distribution, the corresponding population probability is r q / (1 − q + r q), under that sampling assumption. With r = 0.1 and q = 0.5, the population probability is about 0.091, not 0.5.
This correction does not apply unchanged to arbitrary hard-negative selection, feature-dependent sampling, or a poorly calibrated model. In practice, validate probability meaning on representative held-out data and use an appropriate calibration or weighting strategy. A ranking use case and a cost-sensitive decision use case may require different levels of probability accuracy.
Keeping evaluation data representative of the intended workload lets you check both prediction quality and the number of alerts at the chosen threshold. The class balance is only one part of that check, because examples can also contain information that would be unavailable when a real prediction is made. We will examine that problem next through data leakage and the boundaries between training and evaluation.
Population probability = r × q / (1 − q + r × q)
This assumes all positives are retained and negatives are sampled uniformly at rate r.