System designby Learnastra

Concept lesson · Foundations

Load balancing: definition, algorithms and failover

By Anup Rai

Start here

Definition

Load balancing distributes incoming network connections or application requests across eligible backend servers. A load balancer selects a destination using a routing policy and available health or load information.

Why it matters: When one server cannot handle the workload or fails, callers need a way to reach other servers without choosing them manually.

The visual modelLoad-balancing algorithm: least connections

Round robin ignores work already in flight. Least connections is useful only when those connections are comparable.

Load-balancing algorithm: least connectionsRound robin ignores work already in flight. Least connections is useful only when those connections are comparable. For equal-cost short requests, round robin sends R1...R6 to A,B,C,A,B,C. The chapter snapshot has A=10, B=2, C=5 active connections. Least connections selects B if all are eligible and work is comparable. Counts can mislead when one connection carries many expensive streams. Remove unhealthy targets from new routing and use bounded, retry-safe recovery for failed requests.Least connections: A = 10, B = 2, C = 5One bar segment = one active connectionA10 activeB2 activeC5 activeeligiblechoose BeligibleHealth decides eligibility first. Two busy streams may cost more than ten idle connections.
Read the diagram step by step
  1. For equal-cost short requests, round robin sends R1...R6 to A,B,C,A,B,C.
  2. The chapter snapshot has A=10, B=2, C=5 active connections. Least connections selects B if all are eligible and work is comparable.
  3. Counts can mislead when one connection carries many expensive streams.
  4. Remove unhealthy targets from new routing and use bounded, retry-safe recovery for failed requests.

Worked example

With healthy equal-capacity servers A, B and C, round robin sends requests R1–R6 to A, B, C, A, B, C. If B fails and is removed, later requests use A and C; those survivors still need enough capacity.

Key takeaways

  • Choose what to balance: connections, requests, bytes or work.
  • Health checks detect failure after a delay.
  • Routing elsewhere does not preserve state stored only on the failed server.

You will learn to

  • Explain what a load balancer does and where it sits.
  • Replay round robin, weighted routing, and least-connections choices.
  • Handle health detection, draining, failover, and overload.

Practice in this chapter

8 interview questions with model answers and follow-ups.

Go to interview practice

Useful foundations: Distributed systems: scalability, reliability, availability and efficiency · HTTP APIs and request lifecycle

Workload and timing examples are interview assumptions.

01Load balancing: definition and purpose

Load balancing distributes connections or requests across eligible servers. A load balancer selects a backend for traffic addressed to one logical service. The client uses that service address; routing determines whether application instance A, B, or C receives the work.

Horizontal application scaling introduces multiple instances behind the same service address. A load balancer distributes work among them using a routing policy. Essential session state must be available to every eligible instance: rerouting a request cannot recover a cart stored only in a failed process.

02Layer 4 versus Layer 7 load balancing

“Layer” refers to the kind of network information the intermediary understands. Layer 4 (L4) is the transport layer: TCP/UDP connections or flows identified by network addresses and port numbers. A port identifies a service endpoint on a machine. L4 balancing commonly chooses a backend using this connection information. It can forward a connection without interpreting each application request. Layer 7 (L7) is the application layer. L7 balancing understands an application protocol such as HTTP, the request/response protocol used by web applications. It can route /images to an image service and /checkout to a checkout service, or use a hostname or header.

TLS termination means the encrypted client connection ends at the balancer, which can inspect the decrypted HTTP message. The balancer may then create a separately encrypted connection to the backend. If TLS passes through untouched, a transport balancer does not get the same HTTP routing information. Say where encryption ends and which network segments remain protected.

Choice Information available for routing Choose it when Limit to explain
L4 connection/flow balancing Addresses, ports and transport state Distribute transport traffic without needing HTTP paths A long connection can carry uneven application work
L7 HTTP balancing Hostname, path and permitted headers after HTTP is visible /images and /checkout need different service pools Parsing and TLS termination add processing work; the balancer must be trusted with decrypted request data

For GET /images/P7.jpg, an L7 rule can first choose the image pool; a balancing algorithm then chooses A or B inside that pool. Choosing the right service and distributing work among its instances are two separate decisions.

03Worked example: round robin, weights and least connections

After choosing the service pool, the balancer still needs a rule for selecting an instance. The main choices use a fixed schedule, a configured share of capacity, a measurement of current work, or a stable caller identity. Compare them by asking which signal best represents the work in this pool.

Assume A, B, and C are healthy and serve equal-cost short requests. Round robin visits them in order. Requests R1 through R6 go to A, B, C, A, B, C. Step one: R1 arrives and A is next. Step two: R2 advances the cursor to B. Step three: R3 advances to C. Step four: R4 wraps to A; R5 and R6 repeat B and C. Each server receives two requests, but equal counts imply equal work only under our equal-cost assumption. It is easy to operate, but it ignores work already running.

Concept in focusTrace six requests through round robin

Follow each request arrow to A, B or C. All three backends are eligible and equally weighted in this example.

Trace six requests through round robinFollow each request arrow to A, B or C. All three backends are eligible and equally weighted in this example. Map requests R1 through R6 onto three backends. A receives R1 and R4; B receives R2 and R5; C receives R3 and R6.Round robin with three eligible backendsR1R2R3R4R5R6AR1 + R4BR2 + R5CR3 + R6The pattern repeats A, B, C; it does not measure request cost.

Remember: A, B, C, then repeat; equal request counts need not mean equal work.

Read the diagram
  1. Map requests R1 through R6 onto three backends.
  2. A receives R1 and R4; B receives R2 and R5; C receives R3 and R6.
Try from memoryWhich backend receives R7?

A, provided the same three backends remain eligible and the rotation continues.

Now A has twice the tested capacity of B and C. A weighted schedule such as A, B, A, C repeats, giving A about half the requests and B/C a quarter each. Weights express capacity assumptions; they do not detect a new slow dependency.

For long-lived connections, suppose A has 10 active connections, B has 2, and C has 5. Least connections sends the next comparable connection to B. But if B's two connections each carry many expensive streams, the count may misrepresent actual load.

Policy What guides the choice What can go wrong
Round robin Next eligible server Unequal request cost is ignored
Weighted round robin Configured capacity share Old weights miss changing capacity
Least connections Active connections, sometimes weighted Connections can carry unequal work
Power of two choices Sample two eligible servers, choose the less loaded The load signal can be stale or unrepresentative
Least response time Observed latency, often plus active work Stale measurements or oscillation
Least bandwidth Current transferred bytes CPU-heavy small requests look cheap
IP hash A stable function of client IP Many users behind one NAT share a target

Implementations differ, so explain the signal instead of promising an exact universal algorithm.

Power of two choices reduces the need to compare every backend: randomly sample two eligible servers and choose the one with less measured work. With A=10, B=2 and C=5 comparable active connections, sampling A/C chooses C; sampling A/B chooses B. It need not find the global minimum to reduce imbalance. A least-request implementation counts active requests instead of transport connections, which can better match multiplexed HTTP work; neither count captures arbitrary CPU cost.

Worked example diagramAfter B is removed, new requests reach A and C. Shared cart storage lets either recover the session; surviving capacity still must be checked.
Load balancing: definition, algorithms and failover: architecture diagram1. HTTP request to 2. Redundant HTTP balancers: one service address; 2. Redundant HTTP balancers to 3. A: ready, weight 2: two shares of new traffic; 2. Redundant HTTP balancers to 5. C: ready, weight 1: one share of new traffic; 3. A: ready, weight 2 to 6. Shared cart store: load authoritative cart; 5. C: ready, weight 1 to 6. Shared cart store: same cart after reroute; 4. Health check: B failed to 2. Redundant HTTP balancers: health failure stops assignment1 → 2: one service address2 → 3: two shares of new traffic2 → 5: one share of new traffic3 → 6: load authoritative cart5 → 6: same cart after reroute4 → 2: health failure stops assignment01HTTP request02Redundant HTTPbalancers03A: ready, weight 204Health check: Bfailed05C: ready, weight 106Shared cart store
  1. 1 → 2one service addressHTTP request → Redundant HTTP balancers
  2. 2 → 3two shares of new trafficRedundant HTTP balancers → A: ready, weight 2
  3. 2 → 5one share of new trafficRedundant HTTP balancers → C: ready, weight 1
  4. 3 → 6load authoritative cartA: ready, weight 2 → Shared cart store
  5. 5 → 6same cart after rerouteC: ready, weight 1 → Shared cart store
  6. 4 → 2health failure stops assignmentHealth check: B failed → Redundant HTTP balancers

04Session affinity and persistent state

Affinity, or a sticky session, tries to send a caller back to the same backend. It can improve reuse of a local cache. IP hashing is one way; a routing cookie is another. If many students share one school's public IP through network address translation (NAT), IP hashing can concentrate them on one server.

Hashing a stable key can also place cached objects consistently. That is useful when the same key should reach the same owner, but the design must explain how keys move when servers join or leave and how it handles a key that receives unusually heavy traffic. A balancer cannot divide one expensive request simply by hashing it.

05Health checks and backend failover

At 12:00:00, B's process stops. A health check is a small probe used to decide whether B should receive new work. A readiness check asks whether it can serve new requests; a liveness check asks whether restarting the process might be necessary.

  1. Requests already sent to B may fail before a health probe detects the crash. A health system does not make detection instantaneous.
  2. After the configured failure threshold, the balancer removes B from new selection. A and C inherit its traffic.
  3. Safe retries use a deadline and an operation identity where needed. Retrying a purchase blindly may duplicate it if B committed just before losing the response.
  4. When B restarts, readiness remains false until required initialization completes. Reintroduce it gradually so a cold cache does not create a surge of database work.

A shallow probe can say “healthy” while every database query fails. An overly broad probe can remove all servers when one shared optional dependency fails. Design probes around the work each pool must actually serve.

Distinguish the source of health information. Active checks send dedicated probes even when no user traffic arrives. Passive checks infer trouble from real request failures, so an idle backend can remain untested. Thresholds reduce transient ejections but increase detection time. Neither proves future success, and an application error caused by the caller is not automatically evidence that the server is unhealthy.

06Connection draining, overload and balancer redundancy

Draining stops assigning new work while allowing existing requests to finish within a deadline. For a deployment, mark C unready, let short requests finish, then stop it. Long-lived sockets need a reconnect protocol or explicit migration; draining does not preserve in-memory conversation state by itself.

The balancer also needs a replacement if it fails. Active/passive keeps a standby and a way to redirect traffic; active/active runs several balancers. Cached DNS answers can delay redirection. Even a managed balancer needs enough surviving capacity for the failures you plan to tolerate. Existing TCP or TLS connections may break when their balancer fails: another balancer does not automatically inherit them. Clients therefore need reconnect and retry limits.

Interview answer: “For similar short HTTP calls I begin with weighted round robin over ready instances. I move essential session state out of individual servers. I calculate surviving capacity, drain during changes, and make retries safe. For long-lived or uneven work, I change the routing signal after measuring which resource is saturated.”

These policies become routing configuration in a proxy. In NGINX, an upstream group names the available backends, while proxy_pass forwards matching requests to that group. Weights and selection rules then control how the group distributes work.

A concrete starting implementation is an NGINX HTTP proxy with an upstream group, explicit backend weights and proxy_pass; choose least_conn when comparable active connections are a better signal. Its upstream module documents passive failure handling through max_fails and fail_timeout. Dedicated active HTTP checks require the documented health-check module/product support, so verify the installed edition rather than assuming all capabilities follow from the NGINX name. This configuration routes requests. The application and storage design must separately preserve carts, determine which database node may accept writes, and handle retries without losing or duplicating an operation. See the upstream reference and health-check guide.

Practise the interview questions

Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.

Foundation · Question 1

What does a load balancer do?

Reveal a model answer

“It chooses a healthy backend for incoming service traffic. The client uses one service address; the balancer can route its request to A or B. It distributes work and helps route around detected failures, but shared data and correct write ownership still need their own design.”

What the answer must demonstrate: Explain routing separately from state.

Foundation · Question 2

Where do six equal requests go across A, B, and C?

Reveal a model answer

“Under simple round robin: A, B, C, A, B, C. If A has twice the capacity, I can use weights giving A roughly half the traffic. These choices assume requests are similar enough that request count represents work.”

What the answer must demonstrate: Demonstrate a schedule before discussing limitations.

Applied · Question 3

When should a service use L7 routing instead of L4 balancing?

Reveal a model answer

“I need to route by HTTP path, such as /images versus /checkout. An L7 balancer understands those fields, usually after TLS termination. I would also protect the backend connection. An L4 connection balancer is sufficient when I only need transport-level distribution.”

What the answer must demonstrate: Describe what information the layer can inspect.

Applied · Question 4

A has 10 connections, B 2, C 5. Who gets the next one?

Reveal a model answer

“Least connections chooses B if the servers and connection costs are comparable. I would not assume B is least busy if those two connections each contain many expensive streams. I validate the signal against CPU, queueing, and latency.”

What the answer must demonstrate: Qualify the unit of work.

Applied · Question 5

Would sticky sessions solve cart persistence?

Reveal a model answer

“They reduce movement while the chosen server works, but they do not preserve a cart when it fails. I store the authoritative cart durably and use stickiness only if locality improves performance. Then another backend can continue the session.”

What the answer must demonstrate: Separate locality and durability.

Applied · Question 6

A backend B crashes before its next health probe. What happens until the balancer removes it?

Reveal a model answer

“Some requests may still be sent to B until failure detection crosses its threshold. I use bounded timeouts and safe retries. After removal, A and C must have capacity for the redirected work; otherwise detection can turn one crash into a broader overload.”

What the answer must demonstrate: Acknowledge detection delay and correlated failures.

Applied · Question 7

How do you update an instance serving WebSockets?

Reveal a model answer

“I stop new assignments, signal clients to reconnect where the protocol allows, and enforce a drain deadline. Message state lives in durable storage so reconnecting to another instance can resume from a cursor. I cannot assume a balancer transfers the old socket’s process memory.”

What the answer must demonstrate: Explain the long-lived session explicitly.

Applied · Question 8

Does adding a load balancer eliminate all single points of failure?

Reveal a model answer

“No. The balancer, discovery and shared database are separate dependencies. I use redundant balancers with supported traffic failover and enough surviving capacity. If the failed balancer terminated a client connection, that connection may still break: the client reconnects and retries safely. Routing a new request is different from preserving an old connection.”

What the answer must demonstrate: Trace the full failure path.

Blank-page exercise · 15 minutes

Build the answer yourself

Route six requests across three servers, then remove one during peak. Explain state, retries, and remaining capacity.

  • Compare L4 and L7 with actual request fields.
  • Show round-robin and weighted assignments.
  • Trace detection, removal, and reintroduction.
  • Give a failure-safe session design.

Check that each component and design decision follows from your requirements and workload.

Recall the key ideas

Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.

Load balancing: definition, algorithms and failoverChoose the balancing unitRecall first, then reveal

Connection count, request count, bytes, and CPU work can differ.

Measure the work being balanced.

Return to lesson
Load balancing: definition, algorithms and failoverIf one server fails, what must the remaining servers have?Recall first, then reveal

Spare capacity for redirected traffic, or admission limits that reject excess work. Removing a failed server from routing alone cannot prevent overload.

Detect failure → redirect → check spare capacity.

Return to lesson
Load balancing: definition, algorithms and failoverAffinityRecall first, then reveal

Sending a caller to the same server can reuse its local cache; essential session data must also survive that server failing.

Sticky is not durable.

Return to lesson

Final revision

Summary and interview notes

Load balancing selects an eligible destination for a connection or request; the chosen unit and load signal determine how well it distributes work. Safe operation also requires health detection, enough surviving capacity, recoverable session state and bounded retries.

Remember these points

  • L4 routing uses transport information; L7 routing can use visible application fields such as an HTTP path.
  • Round robin spreads request counts; weights reflect server capacity. Least-connections uses active connections as a load estimate, which can mislead when connections carry different amounts of work.
  • Power of two choices compares two sampled eligible backends instead of finding a global minimum.
  • Affinity improves locality but cannot recover state lost with a process.
  • A failed server or balancer may interrupt requests. Reusing a saved operation ID can make retries safe for operations designed to support it.

Interview tips

  • Replay a short routing schedule, then explain how requests with different processing costs could make the load uneven.
  • Calculate the load on survivors after removing one node.
  • Separate active probes, passive error observations, readiness, restart policy and overload controls.

Important qualifications

  • Algorithm names and health features vary by implementation and edition; verify the actual configuration.
  • HTTP/2 multiplexing means one connection may carry many requests.
  • The balancing layer does not determine which database replica owns a write.

Technical references

Practice marks stay in this browser.