Concept lesson · Foundations
Load balancing: definition, algorithms and failover
Start here
Definition
Load balancing distributes incoming network connections or application requests across eligible backend servers. A load balancer selects a destination using a routing policy and available health or load information.
Why it matters: When one server cannot handle the workload or fails, callers need a way to reach other servers without choosing them manually.
Round robin ignores work already in flight. Least connections is useful only when those connections are comparable.
Read the diagram step by step
- For equal-cost short requests, round robin sends R1...R6 to A,B,C,A,B,C.
- The chapter snapshot has A=10, B=2, C=5 active connections. Least connections selects B if all are eligible and work is comparable.
- Counts can mislead when one connection carries many expensive streams.
- Remove unhealthy targets from new routing and use bounded, retry-safe recovery for failed requests.
Worked example
With healthy equal-capacity servers A, B and C, round robin sends requests R1–R6 to A, B, C, A, B, C. If B fails and is removed, later requests use A and C; those survivors still need enough capacity.
Key takeaways
- Choose what to balance: connections, requests, bytes or work.
- Health checks detect failure after a delay.
- Routing elsewhere does not preserve state stored only on the failed server.
You will learn to
- Explain what a load balancer does and where it sits.
- Replay round robin, weighted routing, and least-connections choices.
- Handle health detection, draining, failover, and overload.
Practice in this chapter
8 interview questions with model answers and follow-ups.
Go to interview practiceUseful foundations: Distributed systems: scalability, reliability, availability and efficiency · HTTP APIs and request lifecycle
Workload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01Load balancing: definition and purpose
Load balancing distributes connections or requests across eligible servers. A load balancer selects a backend for traffic addressed to one logical service. The client uses that service address; routing determines whether application instance A, B, or C receives the work.
Horizontal application scaling introduces multiple instances behind the same service address. A load balancer distributes work among them using a routing policy. Essential session state must be available to every eligible instance: rerouting a request cannot recover a cart stored only in a failed process.
02Layer 4 versus Layer 7 load balancing
“Layer” refers to the kind of network information the intermediary understands. Layer 4 (L4) is the transport layer: TCP/UDP connections or flows identified by network addresses and port numbers. A port identifies a service endpoint on a machine. L4 balancing commonly chooses a backend using this connection information. It can forward a connection without interpreting each application request. Layer 7 (L7) is the application layer. L7 balancing understands an application protocol such as HTTP, the request/response protocol used by web applications. It can route /images to an image service and /checkout to a checkout service, or use a hostname or header.
TLS termination means the encrypted client connection ends at the balancer, which can inspect the decrypted HTTP message. The balancer may then create a separately encrypted connection to the backend. If TLS passes through untouched, a transport balancer does not get the same HTTP routing information. Say where encryption ends and which network segments remain protected.
| Choice | Information available for routing | Choose it when | Limit to explain |
|---|---|---|---|
| L4 connection/flow balancing | Addresses, ports and transport state | Distribute transport traffic without needing HTTP paths | A long connection can carry uneven application work |
| L7 HTTP balancing | Hostname, path and permitted headers after HTTP is visible | /images and /checkout need different service pools |
Parsing and TLS termination add processing work; the balancer must be trusted with decrypted request data |
For GET /images/P7.jpg, an L7 rule can first choose the image pool; a balancing algorithm then chooses A or B inside that pool. Choosing the right service and distributing work among its instances are two separate decisions.
03Worked example: round robin, weights and least connections
After choosing the service pool, the balancer still needs a rule for selecting an instance. The main choices use a fixed schedule, a configured share of capacity, a measurement of current work, or a stable caller identity. Compare them by asking which signal best represents the work in this pool.
Assume A, B, and C are healthy and serve equal-cost short requests. Round robin visits them in order. Requests R1 through R6 go to A, B, C, A, B, C. Step one: R1 arrives and A is next. Step two: R2 advances the cursor to B. Step three: R3 advances to C. Step four: R4 wraps to A; R5 and R6 repeat B and C. Each server receives two requests, but equal counts imply equal work only under our equal-cost assumption. It is easy to operate, but it ignores work already running.
Follow each request arrow to A, B or C. All three backends are eligible and equally weighted in this example.
Remember: A, B, C, then repeat; equal request counts need not mean equal work.
Read the diagram
- Map requests R1 through R6 onto three backends.
- A receives R1 and R4; B receives R2 and R5; C receives R3 and R6.
Try from memoryWhich backend receives R7?
A, provided the same three backends remain eligible and the rotation continues.
Now A has twice the tested capacity of B and C. A weighted schedule such as A, B, A, C repeats, giving A about half the requests and B/C a quarter each. Weights express capacity assumptions; they do not detect a new slow dependency.
For long-lived connections, suppose A has 10 active connections, B has 2, and C has 5. Least connections sends the next comparable connection to B. But if B's two connections each carry many expensive streams, the count may misrepresent actual load.
| Policy | What guides the choice | What can go wrong |
|---|---|---|
| Round robin | Next eligible server | Unequal request cost is ignored |
| Weighted round robin | Configured capacity share | Old weights miss changing capacity |
| Least connections | Active connections, sometimes weighted | Connections can carry unequal work |
| Power of two choices | Sample two eligible servers, choose the less loaded | The load signal can be stale or unrepresentative |
| Least response time | Observed latency, often plus active work | Stale measurements or oscillation |
| Least bandwidth | Current transferred bytes | CPU-heavy small requests look cheap |
| IP hash | A stable function of client IP | Many users behind one NAT share a target |
Implementations differ, so explain the signal instead of promising an exact universal algorithm.
Power of two choices reduces the need to compare every backend: randomly sample two eligible servers and choose the one with less measured work. With A=10, B=2 and C=5 comparable active connections, sampling A/C chooses C; sampling A/B chooses B. It need not find the global minimum to reduce imbalance. A least-request implementation counts active requests instead of transport connections, which can better match multiplexed HTTP work; neither count captures arbitrary CPU cost.
- 1 → 2one service addressHTTP request → Redundant HTTP balancers
- 2 → 3two shares of new trafficRedundant HTTP balancers → A: ready, weight 2
- 2 → 5one share of new trafficRedundant HTTP balancers → C: ready, weight 1
- 3 → 6load authoritative cartA: ready, weight 2 → Shared cart store
- 5 → 6same cart after rerouteC: ready, weight 1 → Shared cart store
- 4 → 2health failure stops assignmentHealth check: B failed → Redundant HTTP balancers
04Session affinity and persistent state
Affinity, or a sticky session, tries to send a caller back to the same backend. It can improve reuse of a local cache. IP hashing is one way; a routing cookie is another. If many students share one school's public IP through network address translation (NAT), IP hashing can concentrate them on one server.
Hashing a stable key can also place cached objects consistently. That is useful when the same key should reach the same owner, but the design must explain how keys move when servers join or leave and how it handles a key that receives unusually heavy traffic. A balancer cannot divide one expensive request simply by hashing it.
05Health checks and backend failover
At 12:00:00, B's process stops. A health check is a small probe used to decide whether B should receive new work. A readiness check asks whether it can serve new requests; a liveness check asks whether restarting the process might be necessary.
- Requests already sent to B may fail before a health probe detects the crash. A health system does not make detection instantaneous.
- After the configured failure threshold, the balancer removes B from new selection. A and C inherit its traffic.
- Safe retries use a deadline and an operation identity where needed. Retrying a purchase blindly may duplicate it if B committed just before losing the response.
- When B restarts, readiness remains false until required initialization completes. Reintroduce it gradually so a cold cache does not create a surge of database work.
A shallow probe can say “healthy” while every database query fails. An overly broad probe can remove all servers when one shared optional dependency fails. Design probes around the work each pool must actually serve.
Distinguish the source of health information. Active checks send dedicated probes even when no user traffic arrives. Passive checks infer trouble from real request failures, so an idle backend can remain untested. Thresholds reduce transient ejections but increase detection time. Neither proves future success, and an application error caused by the caller is not automatically evidence that the server is unhealthy.
06Connection draining, overload and balancer redundancy
Draining stops assigning new work while allowing existing requests to finish within a deadline. For a deployment, mark C unready, let short requests finish, then stop it. Long-lived sockets need a reconnect protocol or explicit migration; draining does not preserve in-memory conversation state by itself.
The balancer also needs a replacement if it fails. Active/passive keeps a standby and a way to redirect traffic; active/active runs several balancers. Cached DNS answers can delay redirection. Even a managed balancer needs enough surviving capacity for the failures you plan to tolerate. Existing TCP or TLS connections may break when their balancer fails: another balancer does not automatically inherit them. Clients therefore need reconnect and retry limits.
Interview answer: “For similar short HTTP calls I begin with weighted round robin over ready instances. I move essential session state out of individual servers. I calculate surviving capacity, drain during changes, and make retries safe. For long-lived or uneven work, I change the routing signal after measuring which resource is saturated.”
These policies become routing configuration in a proxy. In NGINX, an upstream group names the available backends, while proxy_pass forwards matching requests to that group. Weights and selection rules then control how the group distributes work.
A concrete starting implementation is an NGINX HTTP proxy with an upstream group, explicit backend weights and proxy_pass; choose least_conn when comparable active connections are a better signal. Its upstream module documents passive failure handling through max_fails and fail_timeout. Dedicated active HTTP checks require the documented health-check module/product support, so verify the installed edition rather than assuming all capabilities follow from the NGINX name. This configuration routes requests. The application and storage design must separately preserve carts, determine which database node may accept writes, and handle retries without losing or duplicating an operation. See the upstream reference and health-check guide.
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
What does a load balancer do?
Reveal a model answer
“It chooses a healthy backend for incoming service traffic. The client uses one service address; the balancer can route its request to A or B. It distributes work and helps route around detected failures, but shared data and correct write ownership still need their own design.”
Interviewer follow-up
Does it automatically make the application stateless?
Reveal the follow-up answer
No. If the cart exists only in A’s memory, switching to B can lose it. The application must place essential state where another instance can recover or access it.
What the answer must demonstrate: Explain routing separately from state.
Where do six equal requests go across A, B, and C?
Reveal a model answer
“Under simple round robin: A, B, C, A, B, C. If A has twice the capacity, I can use weights giving A roughly half the traffic. These choices assume requests are similar enough that request count represents work.”
Interviewer follow-up
What breaks that assumption?
Reveal the follow-up answer
A large export can use far more CPU or time than a small read. Equal request counts may create uneven load. I would separate pools or consider a work-sensitive signal.
What the answer must demonstrate: Demonstrate a schedule before discussing limitations.
When should a service use L7 routing instead of L4 balancing?
Reveal a model answer
“I need to route by HTTP path, such as /images versus /checkout. An L7 balancer understands those fields, usually after TLS termination. I would also protect the backend connection. An L4 connection balancer is sufficient when I only need transport-level distribution.”
Interviewer follow-up
Can one HTTP/2 connection represent many requests?
Reveal the follow-up answer
Yes. Multiplexing means connection counts are not request counts, so a connection-level policy can still produce uneven application work.
What the answer must demonstrate: Describe what information the layer can inspect.
A has 10 connections, B 2, C 5. Who gets the next one?
Reveal a model answer
“Least connections chooses B if the servers and connection costs are comparable. I would not assume B is least busy if those two connections each contain many expensive streams. I validate the signal against CPU, queueing, and latency.”
Interviewer follow-up
Why not always use least response time?
Reveal the follow-up answer
It relies on measurements that may lag or react poorly to small samples. A recently idle slow node may look deceptively good, and routing can oscillate.
What the answer must demonstrate: Qualify the unit of work.
Would sticky sessions solve cart persistence?
Reveal a model answer
“They reduce movement while the chosen server works, but they do not preserve a cart when it fails. I store the authoritative cart durably and use stickiness only if locality improves performance. Then another backend can continue the session.”
Interviewer follow-up
Why is IP hashing risky behind a school network?
Reveal the follow-up answer
Many users can share one public NAT address and hash to the same backend. A client IP is not a unique user identifier.
What the answer must demonstrate: Separate locality and durability.
A backend B crashes before its next health probe. What happens until the balancer removes it?
Reveal a model answer
“Some requests may still be sent to B until failure detection crosses its threshold. I use bounded timeouts and safe retries. After removal, A and C must have capacity for the redirected work; otherwise detection can turn one crash into a broader overload.”
Interviewer follow-up
Should readiness check every downstream system?
Reveal the follow-up answer
Only dependencies necessary for that pool’s promised work. Active probes exercise a selected path; passive checks observe actual failures. Checking an optional shared service can eject the entire pool unnecessarily, while a shallow process probe can miss failed critical operations.
What the answer must demonstrate: Acknowledge detection delay and correlated failures.
How do you update an instance serving WebSockets?
Reveal a model answer
“I stop new assignments, signal clients to reconnect where the protocol allows, and enforce a drain deadline. Message state lives in durable storage so reconnecting to another instance can resume from a cursor. I cannot assume a balancer transfers the old socket’s process memory.”
Interviewer follow-up
What if a client never disconnects?
Reveal the follow-up answer
The deadline eventually closes it. The application protocol must make reconnection and replay a supported path.
What the answer must demonstrate: Explain the long-lived session explicitly.
Does adding a load balancer eliminate all single points of failure?
Reveal a model answer
“No. The balancer, discovery and shared database are separate dependencies. I use redundant balancers with supported traffic failover and enough surviving capacity. If the failed balancer terminated a client connection, that connection may still break: the client reconnects and retries safely. Routing a new request is different from preserving an old connection.”
Interviewer follow-up
Why is DNS failover not immediate for everyone?
Reveal the follow-up answer
Resolvers and clients may keep cached answers until their expiration behavior permits a refresh. I would avoid promising universal instant traffic movement.
What the answer must demonstrate: Trace the full failure path.
Blank-page exercise · 15 minutes
Build the answer yourself
Route six requests across three servers, then remove one during peak. Explain state, retries, and remaining capacity.
- Compare L4 and L7 with actual request fields.
- Show round-robin and weighted assignments.
- Trace detection, removal, and reintroduction.
- Give a failure-safe session design.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
Load balancing: definition, algorithms and failoverChoose the balancing unitRecall first, then reveal
Connection count, request count, bytes, and CPU work can differ.
Measure the work being balanced.
Return to lessonLoad balancing: definition, algorithms and failoverIf one server fails, what must the remaining servers have?Recall first, then reveal
Spare capacity for redirected traffic, or admission limits that reject excess work. Removing a failed server from routing alone cannot prevent overload.
Detect failure → redirect → check spare capacity.
Return to lessonLoad balancing: definition, algorithms and failoverAffinityRecall first, then reveal
Sending a caller to the same server can reuse its local cache; essential session data must also survive that server failing.
Sticky is not durable.
Return to lessonFinal revision
Summary and interview notes
Load balancing selects an eligible destination for a connection or request; the chosen unit and load signal determine how well it distributes work. Safe operation also requires health detection, enough surviving capacity, recoverable session state and bounded retries.
Remember these points
- L4 routing uses transport information; L7 routing can use visible application fields such as an HTTP path.
- Round robin spreads request counts; weights reflect server capacity. Least-connections uses active connections as a load estimate, which can mislead when connections carry different amounts of work.
- Power of two choices compares two sampled eligible backends instead of finding a global minimum.
- Affinity improves locality but cannot recover state lost with a process.
- A failed server or balancer may interrupt requests. Reusing a saved operation ID can make retries safe for operations designed to support it.
Interview tips
- Replay a short routing schedule, then explain how requests with different processing costs could make the load uneven.
- Calculate the load on survivors after removing one node.
- Separate active probes, passive error observations, readiness, restart policy and overload controls.
Important qualifications
Technical references
- NGINX HTTP load balancingPrimary implementation reference for request routing and balancing signals.
- NGINX upstream moduleDetails of weights, least connections, health behavior, and upstream configuration.
- NGINX HTTP health checksActive and passive health-check mechanisms and product requirements; checked 2026-09-23.
Practice marks stay in this browser.