02 · Framing the problem
Scale and latency estimates
3 min read
After defining what the model predicts, we need to establish how often predictions are requested and how much work each request involves. Scale estimates translate the requirements into approximate demand for computation, storage, and response time. They help you identify which parts of a proposed design need measurement before you can rely on them.
Count work rather than users alone
Suppose an illustrative service has two million daily active users, each making 20 recommendation requests a day. That is 40 million requests per day, or about 463 requests per second on average. A five-times peak assumption gives about 2,315 requests per second. Neither the peak multiplier nor the activity estimate is a fact about a real company; confirm the assumptions in an interview.
If each request scores 500 possible recommendations, called candidates, peak demand is roughly 1.16 million candidate scores per second. A request can send its candidates together as a batch, so the number of scores does not imply the same number of network calls. Measuring how quickly the model processes representative batches gives a more useful capacity estimate than multiplying an isolated prediction time by the full catalog size.
40,000,000 requests / 86,400 seconds ≈ 463 requests/s
463 × 5 peak factor ≈ 2,315 requests/s
2,315 × 500 candidates ≈ 1.16 million candidate scores/sBudget the critical path
For an illustrative 100 ms deadline, the request needs time to fetch context, retrieve candidates, rank them, and return the response, including communication and waiting. Independent lookups may run in parallel, while ranking must wait until candidates are available. Drawing those dependencies shows which operations must finish in sequence and therefore determine the critical path through the request.
A budget is an allocation, not a prediction of measured p95. Percentiles of individual stages do not generally add to the percentile of their sum. Validate the end-to-end distribution under realistic concurrency, cache hit rates, request sizes, and partial failures. Average latency can hide the requests users find most frustrating.
Estimate data and resource needs
If each of five million items is represented by 128 numbers stored as four-byte floating-point values, the representations alone occupy about 2.56 GB. These numeric representations are called embeddings, which we will examine in the features topics. A search index, item identifiers, and additional copies consume memory beyond this initial estimate, so the vector size only establishes a starting point for measuring storage needs.
At 40 million impressions per day and an illustrative 500 bytes per event, raw event data grows by about 20 GB per day before replication, indexes, or compression. Retention and aggregation requirements determine whether you need to keep the complete log or a narrower derived dataset.
Capacity also needs room for bursts and for the work that remaining instances inherit when one fails. The estimates tell you which measurements matter, including batch processing speed, response times under peak load, and memory use. With the workload understood, we can turn to the data behind it and follow how an application event becomes a usable training example.