Concept lesson · Foundations
Capacity estimation: throughput, latency, concurrency and storage
Start here
Definition
Capacity estimation translates an assumed workload into the compute, memory, storage and network resources needed to meet performance and failure targets. Throughput is work completed per unit time, latency is time per operation, and concurrency is work in progress.
Why it matters: Without a workload and units, “millions of users” cannot tell you how many servers or how much storage a design needs.
Estimate request rate and bandwidth from the photo workload. Apply Little’s law separately at the metadata-service boundary.
Read the diagram step by step
- One million daily users each view twenty photos: twenty million views per day, about 231.5/s on average.
- The assumed ten-times peak is 2,315 views/s. At 100 KB per thumbnail it requires about 231 MB/s before overhead.
- Separately, a metadata service at 2,000/s and mean time 0.05 s has about 100 requests in flight in steady state.
- A daily user count is not a simultaneous connection count. State decimal bytes and the measurement boundary.
Worked example
One million users making twenty requests a day create 20,000,000 / 86,400 = about 231.5 requests/s on average. A stated 10x peak is about 2,315 requests/s; it is an assumption to validate, not something implied by the user count.
Key takeaways
- Count requests, bytes and retained data separately.
- Use peak load and surviving capacity when sizing.
- Average concurrency = average arrival rate × average time spent in the same measured system, assuming stable operation.
You will learn to
- Calculate average and peak rates with units.
- Distinguish throughput, latency, concurrency, and storage.
- Use an estimate to justify a design change.
Practice in this chapter
8 interview questions with model answers and follow-ups.
Go to interview practiceUseful foundations: System design interview framework
Workload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01Capacity estimation and its units
Capacity estimation translates a workload into the resources needed to meet its targets. A workload specifies what users do, how often, how large their requests are and how concentrated the traffic becomes. The estimate should be accurate enough to choose a design; a load test must later measure the actual implementation. Begin with four separate ideas. Throughput is completed work per unit time, such as 500 uploads per second. Latency is how long one operation takes. Concurrency is the number of operations in progress at once. Storage is how much retained data exists at a point in time.
A restaurant can serve many meals an hour while one customer's meal takes a long time. A batch service can likewise have high throughput and high latency. More concurrent work helps use idle resources; once the limiting resource is fully busy, extra work mostly waits. A claim of “10,000 users” needs to say what they do and when.
QPS means queries per second; in an API discussion people often use it for requests per second, so state whether you are counting API requests or database queries. Network capacity, often called bandwidth, is the maximum data rate a link or path can carry under stated conditions, usually measured in bits/s. Network throughput is the rate actually achieved. Requests/s × bytes/request estimates the required transfer rate; provision capacity above that demand, including protocol overhead and headroom. Peak means the busiest declared interval, while an average spreads all work over the entire measured period. These are different quantities even when one calculation produces another.
| Quantity | Example and unit | Decision it informs | What it cannot establish alone |
|---|---|---|---|
| Throughput | 2,000 completed requests/s | Required processing rate | How long one user waits |
| Latency | 50 ms per request on average | Response-time objective | Total sustainable traffic |
| Concurrency | 100 requests in progress | Connections and memory | Whether queues are stable |
| Storage | 73 TB retained originals | Disk/object capacity | Read/write operations per second |
| Required transfer rate | 231 MB/s at peak | Network and delivery path | CPU cost of producing each byte |
02Worked example: average and peak QPS
Define a workload before estimating resources. For this photo-service example, assume one million daily active users, 20 photo views and 0.1 uploads per user per day, a 2 MB average upload, a 100 KB thumbnail, and 1 KB of metadata per photo. Use decimal units: 1 KB = 1,000 bytes, 1 MB = 1,000,000 bytes, and one day = 86,400 seconds.
| Quantity | Calculation | Approximate result |
|---|---|---|
| Uploads per day | 1,000,000 × 0.1 | 100,000 |
| Average uploads/s | 100,000 / 86,400 | 1.16 |
| Photo views per day | 1,000,000 × 20 | 20,000,000 |
| Average views/s | 20,000,000 / 86,400 | 231.5 |
| Assumed 10× peak | 231.5 × 10 | 2,315 views/s |
The calculation is a sequence: first count actions per day, then divide by seconds per day, then apply an explicitly assumed peak factor. For these inputs, 1,000,000 × 20 = 20,000,000 image views/day, 20,000,000 ÷ 86,400 ≈ 231.5 views/s average, and 231.5 × 10 ≈ 2,315 views/s peak. We size thumbnail delivery against the last rate, then verify it against measured bursts.
03Storage, retention and network bandwidth
The photo workload creates two different demands: storage for retained objects and network capacity for repeated delivery. Logical data counts one copy of each retained object; replicas and backups consume additional physical storage. Keep those counts separate from bytes sent to viewers:
| Category | Calculation | Result and scope |
|---|---|---|
| New originals | 100,000 × 2 MB |
200 GB/day |
| One year of originals | 200 GB/day × 365 |
73 TB before deletion, compression, indexes or redundancy |
| Three full copies | 73 TB × 3 |
219 TB for originals alone |
| Metadata growth | 100,000 × 1 KB |
100 MB/day; 36.5 GB/year before indexes |
| Thumbnail delivery | 20 million × 100 KB |
2 TB/day |
| Average transfer demand | 2 × 10^12 / 86,400 |
About 23.1 MB/s, or 185 megabits/s |
| Assumed 10× delivery peak | 23.1 MB/s × 10 |
About 231 MB/s before headers and retransmissions |
Keep these quantities separate:
- Count other stored data separately. Thumbnail variants and backups are additional categories; the replication multiplier does not include them.
- Bytes and metadata scale differently. Their large size difference is a reason to store them separately. The metadata database need not carry every byte transferred to viewers.
- Convert units explicitly. Network links are often rated in bits/s: multiply bytes by eight. State whether you mean MB or MiB.
04Little’s law and latency percentiles
Request rate alone does not tell us how many connections or request buffers are occupied. A request continues using some resources while it waits for storage or another service. To size those resources, relate the completion rate to the time each request remains in the service.
Each square represents one request. These are long-run averages for a stable service.
Remember: 2,000 requests/s x 0.050 seconds = 100 requests in flight.
Read the diagram
- Count five rows of twenty request squares inside the service.
- Arrivals and completions average 2,000 requests per second; mean time inside is 50 ms.
- Little’s law gives average in-flight work of 100, not a tail-latency prediction.
Try from memoryIf mean time doubles at the same stable throughput, what happens to average in-flight requests?
It doubles from 100 to 200: L = 2,000/s × 0.100 s. This assumes the service remains stable at that throughput.
If average time rises to 0.5 seconds while admitted traffic stays at 2,000/s, concurrency becomes about 1,000. The extra 900 requests need memory, sockets, and possibly database connections. An unbounded queue hides overload briefly while increasing latency. It does not create processing capacity.
The stable-system condition matters. If arrivals stay at 1,200/s while only 1,000/s complete, an unbounded backlog grows by 200 requests/s, or 12,000 requests in one minute. There is no steady finite average latency to insert into this calculation. Bound the queue and reduce admissions, or increase the bottleneck’s measured service capacity.
05Bottlenecks and failure headroom
A bottleneck is the resource that first limits the workload: for example, CPU, database writes or network transfer. Headroom is spare capacity reserved for bursts, uneven load and failures. Once a load test identifies the limiting resource, size enough instances to meet the target even with the chosen failures.
Assume a load test measures 800 requests/s per application instance while meeting the latency objective, and the target peak is 2,315 requests/s.
Each server block represents 800 requests/s at the measured latency target. The lower bar compares peak demand with surviving capacity.
Remember: Four servers can hide a problem that appears after one fails.
Read the diagram
- Remove one 800 requests/s block from the fleet and compare demand with what remains.
- Four instances supply 3,200 requests/s; three supply 2,400 requests/s.
- A peak of 2,315 uses 96.5% of surviving capacity, leaving 85 requests/s.
Try from memoryIs 2,400 requests/s enough for a peak of 2,315?
It covers the point estimate but leaves only 85 requests/s, about 3.5% of surviving capacity. That is little room for workload variance or measurement error.
| Fleet | Normal capacity | Capacity after one loss | Assessment |
|---|---|---|---|
| Three instances | 2,400 requests/s | 1,600 requests/s | Almost no normal spare capacity; insufficient after failure |
| Four instances | 3,200 requests/s | 2,400 requests/s | Little failure headroom for uneven load |
The notation ceil(x) means the smallest whole number at least as large as x; a partial server cannot satisfy the remaining load. For a chosen maximum of 70% of tested capacity after one failure:
- Budget each survivor:
800 × 0.7 = 560 requests/s. - Find the survivors needed:
ceil(2,315 / 560) = 5. - Add failure capacity: five survivors require six instances.
This is illustrative sizing, not a universal 70% rule. Real benchmarks, cost, autoscaling lag and failure domains determine the target.
Check downstream amplification
The database, network, and object store must support the same workload. Six application servers do not help if they all wait for one slow query. Estimate the read/write amplification: if each API call issues five database queries, 2,315 API calls/s can become 11,575 database queries/s.
Size CPU from CPU time
Compute demand has a different unit from elapsed latency. If a measured request uses 2 ms of CPU time:
- CPU demand:
2,315/s × 0.002 CPU-seconds = 4.63 CPU-seconds/s, about 4.63 fully busy cores. - Utilization headroom: at a chosen 70% limit,
ceil(4.63 / 0.7) = 7usable cores before additional failure capacity.
06Cache working set and cost model
A cache stores copies of reused data. Its size depends on distinct hot entries, not total requests. Suppose 500,000 frequently viewed photo records occupy 1.4 KB each including key and bookkeeping overhead. That is about 700 MB per full cache copy. Ten million reads of those same entries do not require ten million stored entries.
A cache hit finds the requested value in the cache; a miss must fetch it from the underlying database or storage service, called the origin. The request hit rate is the fraction of cacheable requests served as hits. This rate turns the delivery estimate into an estimate of work still reaching the origin.
A 95% request hit rate reduces 2,315 cacheable lookups/s to about 2,315 × 0.05 = 116 misses/s under the same workload. But when the cache is empty, the origin can suddenly see all 2,315/s. Protect that origin and warm popular entries gradually. Track byte hit rate separately: a few missed large images may dominate bandwidth despite a high request hit rate.
For cost, write a symbolic model before using current provider prices: storage GB-month + read/write operations + delivered GB + compute time + replication/backup. A cheaper storage tier can have retrieval fees and slower access. The interview value is identifying the dominant cost and a way to measure it, not memorizing a vendor price that may change.
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
What is capacity estimation? Estimate QPS for one million users making ten requests a day.
Reveal a model answer
“Capacity estimation converts a workload into rates and resource needs. Here one million users × ten requests is ten million requests/day. Dividing by 86,400 seconds gives about 116 requests/s average. I still need peak concentration, bytes per request, latency targets and failure headroom before choosing server capacity.”
Interviewer follow-up
Is 116 QPS the server capacity target?
What the answer must demonstrate: Show denominator and units.
Can a system have high throughput and high latency?
Reveal a model answer
“Yes. A batch worker may finish thousands of items a second while each item waits minutes in a queue. Throughput describes the completion rate; latency measures one item’s elapsed time. I would measure queue wait and processing time separately.”
Interviewer follow-up
Would more queued work always raise throughput?
Reveal the follow-up answer
Only until it helps saturate usable capacity. Beyond the bottleneck, more queued work generally raises wait time and resource pressure.
What the answer must demonstrate: Distinguish work rate from wait time.
How much storage do 200 GB/day of uploads need after a year?
Reveal a model answer
“Without deletion, 200 × 365 is 73,000 GB, or 73 TB in decimal units. That is logical originals. I would separately add derived images, indexes, copies, and backups, then apply the retention policy.”
Interviewer follow-up
Should three replicas be included twice for multi-region?
Reveal the follow-up answer
No. Enumerate the actual physical copies and locations. Multipliers represent specific copies, not labels to stack without a physical model.
What the answer must demonstrate: Separate logical data from physical overhead.
What happens at 2,000 requests/s if average latency grows from 50 to 500 ms?
Reveal a model answer
“Assuming both measurements cover the same system in stable operation, the average number of requests in progress grows from about 100 to 1,000. That can exhaust memory or connection pools even without a traffic increase. I would inspect downstream latency and bound admitted work.”
Interviewer follow-up
Can you substitute p99 into Little’s law?
Reveal the follow-up answer
Not to calculate average concurrency. The relation uses averages; percentiles need separate distribution analysis.
What the answer must demonstrate: Use seconds and matching averages.
Should a cache hold 20% of yesterday’s requests?
Reveal a model answer
“Requests are not stored objects. I estimate distinct hot keys and bytes per entry. If a million requests hit one record, that is one cache entry. I use observed reuse and eviction behavior to choose the working set, then account for replication and overhead.”
Interviewer follow-up
How can a 99% hit ratio be misleading?
Reveal the follow-up answer
The remaining 1% may be huge objects or expensive queries. Measure byte hits, expensive misses, and cold-cache behavior.
What the answer must demonstrate: Count distinct retained entries.
Three servers can just meet peak. Is that a resilient design?
Reveal a model answer
“Not if the requirement includes surviving a server failure at that peak. I calculate the remaining capacity after the failure and keep headroom for imbalance. If two survivors cannot meet the objective, I add capacity, reduce admitted work, or agree on degraded behavior.”
Interviewer follow-up
Can autoscaling replace all spare capacity?
Reveal the follow-up answer
Autoscaling has detection and startup delay, and dependencies may scale more slowly. A sudden failure needs capacity or load shedding during that interval.
What the answer must demonstrate: Calculate surviving capacity.
An API runs five database queries. Which QPS matters?
Reveal a model answer
“Count both. At 2,315 API requests/s and five queries per request, the database receives about 11,575 operations/s before retries or cache effects. I would check whether each query is needed, indexed and independent of the others.”
Interviewer follow-up
What if all five run in parallel?
What the answer must demonstrate: Explain amplification rather than hiding it.
A photo service serves 2,315 peak views/s at 100 KB each. Which measurements would change the storage or delivery design?
Reveal a model answer
“The stated peak is 2,315 × 100 KB = 231.5 MB/s, about 1.85 Gb/s before overhead. I would measure repeated-key reuse and permission constraints to evaluate a CDN, and measure metadata and CPU costs separately to decide where scaling helps. Peak QPS alone cannot determine daily delivered bytes or metadata growth; those require daily volume and stored bytes per upload.”
Interviewer follow-up
What observation could invalidate your CDN assumption?
Reveal the follow-up answer
If most images are viewed only once or authorization prevents useful sharing, cache reuse may be low. I would measure the access distribution and policy constraints.
What the answer must demonstrate: Use a number to justify a decision.
Blank-page exercise · 15 minutes
Build the answer yourself
Estimate a file-sharing service with 2 million daily users, five 200 KB downloads each, and a 6× peak. Defend one architecture decision.
- Compute average and peak requests/s.
- Compute delivered bytes/day and peak bytes/s.
- Explain one failure-headroom calculation.
- Distinguish assumptions from measurements.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
Capacity estimation: throughput, latency, concurrency and storageRate conversionRecall first, then reveal
Daily operations ÷ 86,400 gives average operations per second.
Actions → daily count → seconds.
Return to lessonCapacity estimation: throughput, latency, concurrency and storageAt 2,000 requests/s and 0.05 seconds per request, how many are in progress?Recall first, then reveal
About 100 on average: 2,000 × 0.05. Little’s law requires a stable workload and matching measurement boundaries.
Rate × time = work in progress.
Return to lessonCapacity estimation: throughput, latency, concurrency and storageCache sizeRecall first, then reveal
Distinct hot entries × bytes per entry, then copies and headroom.
Keys, not requests.
Return to lessonFinal revision
Summary and interview notes
Capacity estimates translate a declared workload into rates, retained bytes, concurrent work and resource demand. Size each component for its peak load and for the capacity it must retain after the failures you plan to tolerate, then validate the assumptions against a load test.
Remember these points
- Average requests/s = daily requests / 86,400; a peak multiplier is a separate assumption.
- Logical storage, replicas, derived objects, indexes and backups are separate physical categories.
- Little’s law uses average arrival rate and average time for the same system in stable operation; substituting a latency percentile does not give average concurrency.
- CPU-seconds per request differ from elapsed request time; both affect sizing in different ways.
- A warm-cache miss rate is not the capacity requirement after cache loss.
Interview tips
- Write units at every conversion, especially bits versus bytes and MB versus MiB.
- Show the surviving capacity after the required failure, rather than counting only healthy servers.
- End an estimate by naming the architectural decision it changes.
Important qualifications
- The 10× peak and 70% utilization figures are example assumptions, not universal defaults.
- If accepted requests keep arriving faster than they finish, the queue keeps growing. A stable-workload concurrency estimate no longer describes that overload.
Technical references
- Google SRE service-level objectivesDefinitions of throughput, latency, measurements, and the limitations of averages.
- Google SRE handling overloadCapacity limits and overload behavior; numerical examples here are original practice assumptions.
- RFC 5136: Defining Network CapacityDistinguishes capacity from actual usage and explains why the informal term bandwidth is ambiguous.
Practice marks stay in this browser.