The Interview Edge Blog
← Back to all guides
Systems · System design

Design a Load Balancer: the Traffic Cop Every Backend Interview Expects

One server melts when you go viral. The fix isn’t a bigger server — it’s a fleet of them, with a load balancer in front deciding which request goes where. Here’s how the choosing works, how it learns a server died, and why its real job is absorbing failure.

Explain it like I’m five

Picture a theme park on the hottest Saturday of the summer. Ten thousand people want ice cream, so the park opens six windows. Left alone, everyone stampedes to the first window they see — one exhausted scooper does all the work while five stand idle. So the park posts a friendly guide at the entrance. She watches every line, waves each newcomer toward the shortest one, and when the machine at window four breaks down, she quietly stops sending people there. Nobody in line ever notices the broken machine.

A load balancer is that guide, but for your servers. Requests pour in; the balancer hands each one to a server with room to spare, keeps checking that every server is still breathing, and reroutes around the dead ones in seconds. Everything in this guide is the guide’s playbook, written out: how she chooses, how she notices a breakdown, and the arithmetic of the moment a server dies.

Intuition: the front door of every serious backend

“Design a load balancer” shows up in system-design interviews because it is secretly two questions: how do you spread work across many machines — and what happens when one of those machines dies at 2am? The candidates who pass answer the second question first.

One server has a hard ceiling — CPU, memory, network — and it is a single point of failure wearing a trench coat. So production systems scale horizontally: many identical servers behind one address. But that creates two new problems. Somebody has to decide which server gets each request, and somebody has to notice when a server stops answering. That somebody is the load balancer: the single front door that makes a fleet of machines look like one reliable server.

It has three jobs. Distribute — spread incoming requests so no server drowns while others idle. Detect — probe every server continuously and cut the dead ones from rotation. Route smart — send each request where it will be handled best: the emptiest server, the closest region, or the machine already holding the right cache. And here is the line worth carrying into the interview: a load balancer exists to absorb failure, not just to split traffic. Splitting traffic is the easy part. Absorbing a dead server at 2am without paging anyone — that is the job.

One sentence worth memorizing

Put one address in front of N servers; the balancer spreads the load, health-checks dead servers out of rotation in seconds, and routes each request where it lands best — so the fleet survives individual machines dying.

How it works: layers, algorithms, and the death notice

Layer 4 vs Layer 7: what the balancer is allowed to see

Layer 4 works at the transport layer — IP addresses and TCP/UDP ports. It sees connections, not content: fast, cheap, and blind. It forwards packets without understanding them, which makes it the right tool for raw speed — millions of long-lived connections, static IPs, cases where the bytes must stay encrypted end to end. AWS’s Network Load Balancer lives here.

Layer 7 understands HTTP. It reads the host, the path, the headers, the cookies — so /api goes to the API fleet, /static to the file servers, /checkout to the payment-tier machines. It can terminate TLS once at the edge instead of on every server, inject headers, and pin a user to a server with a cookie. AWS’s Application Load Balancer, NGINX, and HAProxy live here. Most web apps sit behind Layer 7 — the moment you route by URL, you have left Layer 4 behind.

L4 vs L7 — say it in one line

Layer 4 balances connections it cannot read — fast and cheap. Layer 7 balances requests it understands — a little slower per decision, but it routes by URL, terminates TLS, and can hold a user’s session.

The choosing: five algorithms and their temperaments

Round robin cycles 1-2-3-1 — dead simple, perfectly fair, and blind: it cannot see that server 3 is twice the size of the others, or that server 2 is melting. Weighted round robin fixes the blindness — give the big machine weight 4 and the small ones weight 1, and traffic splits 4:1:1. Least connections sends each request to the server with the fewest active ones — the right call when work is uneven, like slow database queries mixed with fast cache hits. Least response time watches latency instead of counts and favors whichever server answers fastest. And IP hash / consistent hashing pins each client to the same server, so caches stay warm — at the cost of lopsided load when one client is a whale.

The interview move is to match the algorithm to the workload: even, stateless requests → round robin; mixed server sizes → weighted; uneven request costs → least connections; warm caches matter → consistent hashing. And to name the failure mode each choice invites — which is exactly what the next section is about.

Health checks: the death notice

A balancer that cannot detect death is just a traffic splitter with ambitions. So it probes. Active health checks send synthetic requests — GET /health every few seconds — and a server that misses several in a row is cut from rotation. Passive checks watch real traffic instead: a streak of timeouts or 500s marks the server sick with no extra probes. Most production setups run both — active checks catch the server that accepts connections but never answers, passive checks react to real user pain.

Coming back is gated too: a recovered server must pass several consecutive checks before it rejoins, so a flapping machine doesn’t yo-yo in and out of rotation. The shape to remember is interval × threshold — check every 5 seconds, require 3 failures, and a dead server is cut in at most 15. The body count inside that window is worked by hand below.

Sticky sessions vs stateless: where the shopping cart lives

Some apps keep per-user state in server memory — the classic example is a shopping cart. If request one lands on server A and request two lands on server B, the cart vanishes. The quick fix is sticky sessions: the balancer drops a cookie pinning each user to one server. It works — until that server dies and every pinned user’s cart dies with it, or until one server collects all the whales and the “balanced” load is anything but.

The grown-up answer is stateless servers plus a shared session store: keep the cart in Redis, let any server serve any user, and let the balancer route purely on load. A dead server becomes a non-event — its replacement reads the same Redis and the user never notices. Sticky sessions trade a hard problem (shared state) for a fragile one (pinned fate).

Clients send traffic to one load balancer, which spreads it across three servers and cuts the one that stops answeringclients1,600 rpsload balancerhealth checks every 5sserver A533 rps · healthyserver B533 rps · healthyserver C ✕3 missed checks → cut

The balancer’s real job, drawn: server C stops answering, misses three checks, and is cut — its share redistributes to A and B. Whether they survive that moment is the capacity math worked below.

Front doors of the real internet

Every load balancer you will ever meet is the same three jobs — distribute, detect, route smart — wearing different uniforms. Here are the uniforms that matter in interviews.

01 · Cloud

AWS Elastic Load Balancing is the default front door for apps on AWS. The Application Load Balancer works at Layer 7: it routes by host and path, terminates TLS, and spreads traffic across target groups — named sets of servers with their own health checks, so the API fleet and the web fleet fail independently. The Network Load Balancer works at Layer 4 for the extreme cases: millions of TCP connections, static IPs, and TLS passthrough when the backend must see the encrypted bytes itself.

02 · Edge

Cloudflare runs load balancing as a global anycast network: one IP address is announced from hundreds of cities, and each request lands at the nearest healthy datacenter. Health checks run from many vantage points at once, so when a region goes dark, traffic shifts continents before users notice. Same three jobs — at planetary scale, with geography as a routing input.

03 · Cluster

Inside Kubernetes, a Service load-balances across pods — the cluster’s own small Layer-4 balancer — while Ingress controllers (NGINX, HAProxy) do Layer-7 routing by host and path to the right Service. Your app gets the same front-door pattern twice: once at the cloud edge, once inside the cluster.

Source: AWS Elastic Load Balancing docs — ALB vs NLB, target groups, health checks ↗
Source: NGINX HTTP load balancing docs — upstream methods: round robin, least connections, IP hash ↗

The uniforms change — cloud, edge, cluster — but the playbook doesn’t: one address in front, health-checked servers behind it, and any single machine’s failure absorbed before users notice.

The failure math, by hand

Four servers, each comfortably handling 500 requests per second. Fleet capacity: 2,000 rps. Tonight’s peak: 1,600 rps — 80% of capacity, 400 per server. Then server D dies.

Redistribute1,600 rps across 3 survivors = 533 rps each — past the 500 comfort line on every remaining machine. The failure didn’t just remove 25% of capacity; it pushed the survivors over their limit. This is the cascade: the next server to tip takes its 533 with it, and suddenly 800 rps land on 2 servers.
HeadroomThe fix is N+1 capacity planning: size the fleet so one death is absorbable. Three survivors × 500 = 1,500 rps of survivable peak — 75% of the four-server capacity. At a 1,400 rps peak, one death means 467 per server: uncomfortable, alive. Run hotter than 75% and a single failure is a cascade wearing a timetable.
WeightedServers weighted 4:2:1:1. Each 8-request cycle deals 4 to the big machine, 2 to the medium, 1 and 1 to the small pair — smooth weighted round robin spaces them evenly instead of dumping four in a row. Weights are how heterogeneous fleets stay fair.
RehashOne million cached sessions pinned by client across 4 servers; you add a fifth. Naive hash % N remaps nearly all million — every cache goes cold at once. Consistent hashing remaps only ~1/5: about 200,000 sessions move, 800,000 stay warm. That fraction — 1/(N+1) — is the whole reason the ring exists.
DetectHealth checks every 5 seconds, 3 consecutive failures to cut: a dead server is out in at most 15 seconds. During those 15 seconds, at 400 rps, roughly 6,000 requests arrive at a corpse — they fail or retry. Faster checks shrink that window; too fast and a merely slow server gets executed by mistake. interval × threshold is the dial between blindness and friendly fire.

Three numbers run every load-balancer interview: the surviving share after one death, the remap fraction when the fleet changes, and the detection window your health checks buy you.

Three algorithms in Python

Smooth weighted round robin (the nginx algorithm), a least-connections picker, and a consistent hash ring with virtual nodes — the three patterns behind almost every “how does the balancer choose” follow-up.

python · smooth weighted round robin, least connections, consistent hash ring
import hashlib
from bisect import bisect


class SmoothWeightedRoundRobin:
    # nginx's algorithm: deals 4:2:1:1 evenly, not in clumps
    def __init__(self, weights):
        self.servers = [
            {"name": n, "weight": w, "current": 0}
            for n, w in weights.items()
        ]
        self.total = sum(weights.values())

    def pick(self):
        best = None
        for s in self.servers:
            s["current"] += s["weight"]
            if best is None or s["current"] > best["current"]:
                best = s
        best["current"] -= self.total
        return best["name"]


class LeastConnections:
    def __init__(self, servers):
        self.load = {s: 0 for s in servers}

    def pick(self):
        # fewest in-flight requests wins
        return min(self.load, key=self.load.get)

    def acquire(self, server):
        self.load[server] += 1

    def release(self, server):
        self.load[server] -= 1


class HashRing:
    # consistent hashing: adding a node moves ~1/(N+1) of keys
    def __init__(self, nodes, replicas=100):
        self.ring = {}
        for node in nodes:
            for i in range(replicas):  # virtual nodes even out the slices
                h = self._hash(f"{node}:{i}")
                self.ring[h] = node
        self.keys = sorted(self.ring)

    @staticmethod
    def _hash(key):
        return int(hashlib.md5(key.encode()).hexdigest(), 16)

    def node_for(self, key):
        h = self._hash(key)
        i = bisect(self.keys, h) % len(self.keys)
        return self.ring[self.keys[i]]


lb = SmoothWeightedRoundRobin({"big": 4, "med": 2, "sm-a": 1, "sm-b": 1})
print([lb.pick() for _ in range(8)])
# ['big', 'med', 'big', 'sm-a', 'sm-b', 'big', 'med', 'big'] — 4:2:1:1, evenly spaced

Three shapes, three answers: smooth weighted round robin for heterogeneous fleets, least connections when request costs vary, and the hash ring when warm caches matter more than perfect balance.

Six questions that test the real understanding

Meta
“Design a load balancer for a million requests a second.” Walk me through it.

What to say (≈90 sec): “A million rps means the balancer tier itself has to scale — one box can’t do it, so the front door gets its own fleet. DNS and anycast spread clients across regions; in each region a fleet of Layer-7 balancers terminates TLS and routes by path into target groups — the API fleet, the web fleet, the static fleet — each with its own health checks, so they fail independently. Choosing per group: least connections where request costs vary, consistent hashing where cache warmth matters. Health checks every few seconds, three strikes to cut, gated re-entry so flapping servers don’t yo-yo. And the capacity plan: N+1 headroom, fleet sized so one dead server redistributes without cascading — at a million rps that means running at 75% max. The whole design is the failure story: a balancer dies and anycast reroutes, a server dies and health checks cut it, a region dies and traffic shifts.”

Likely follow-up: “Where does TLS terminate?” → At the balancer fleet — one place to manage certificates, backends skip the crypto cost. Passthrough only when compliance demands end-to-end encryption.

The answer that sinks you: “One big load balancer in front of N servers.” Why it fails: the balancer is now the single point of failure and the bottleneck — at a million rps, the front door needs its own fleet.

Google
Layer 4 or Layer 7 — and where do you terminate TLS?

What to say (≈90 sec): “Layer 7 for anything HTTP: it reads host, path, and headers, so /api goes to the API fleet and /static to the file servers — and it terminates TLS once at the edge, so certificates live in one place and backends skip the crypto cost. Layer 4 when I need raw speed or can’t see inside the bytes: millions of long-lived TCP connections, or TLS passthrough where compliance requires end-to-end encryption. The tradeoff is visibility versus cost — L7 pays per-request parsing to make smarter decisions, L4 forwards packets nearly for free but routes blind. Most web apps are L7; the moment you route by URL, you’ve left L4 behind.”

Likely follow-up: “What breaks when you terminate at the balancer?” → The balancer-to-backend hop is plaintext inside your network — fine in your own VPC, a compliance problem across trust boundaries. Then you re-encrypt to the backend or passthrough.

The answer that sinks you: “Always terminate at Layer 7 — it’s strictly better.” Why it fails: it ignores the passthrough cases and the per-request cost. “Strictly better” is never the right answer in systems design.

Amazon
The shopping cart lives in server memory. Sticky sessions, or stateless servers plus Redis?

What to say (≈90 sec): “Stateless plus Redis — and I’d say why out loud. Sticky sessions pin each user to one server with a cookie, so the cart survives — until that server dies and every pinned cart dies with it, or one server collects all the whales and my ‘balanced’ load isn’t. Deployments hurt too: draining a server means migrating pinned users. Stateless servers read the cart from Redis on every request, so any server can serve any user; a dead server is a non-event and the balancer routes purely on load. The cost is one Redis lookup per request — a millisecond against a cart that must survive server death. Sticky sessions trade a hard problem, shared state, for a fragile one: pinned fate.”

Likely follow-up: “What if Redis becomes the bottleneck?” → Shard it by user ID — with the same consistent hashing from this guide. The session store gets its own little load-balancing story.

The answer that sinks you: “Sticky sessions — simpler, no extra hop.” Why it fails: “simpler” ignores fate-sharing. The interviewer was waiting to hear what happens to pinned users when their server dies.

OpenAI
How fast can you cut a dead server — and what does “faster” cost you?

What to say (≈90 sec): “Detection time is interval × threshold: probe every 5 seconds, require 3 consecutive failures, and a dead server is cut in at most 15 seconds. During those 15 seconds at 400 rps, roughly 6,000 requests arrive at a corpse — they fail or retry, so the window has a real body count. Faster checks shrink the window but raise the false-positive rate: a slow server in a GC pause looks dead and gets executed, flapping in and out of rotation. So I’d run both: active synthetic probes for the baseline, passive checks watching real traffic — a streak of timeouts or 500s reacts to actual user pain faster than any probe interval. And recovered servers rejoin only after consecutive passes, so a flapping machine doesn’t yo-yo.”

Likely follow-up: “Active, passive, or both — defend the cost.” → Both. Active catches the server that accepts connections but never answers — passive never fires there, because no real traffic completes. The cost is complexity, and it’s worth it.

The answer that sinks you: “Check every second — faster is always better.” Why it fails: it names no cost — probe load, false positives, flapping — and “always better” dodges the actual engineering tradeoff.

Anthropic
You add a fifth server to the fleet. Why does consistent hashing move only a fifth of the keys?

What to say (≈90 sec): “Naive hashing does hash(key) % N: change N from 4 to 5 and nearly every key’s slot changes — a million cached sessions go cold at once. Consistent hashing puts servers and keys on the same ring, and each key walks clockwise to the first server. Adding a fifth server only steals the key ranges that now land on its ring positions — about a fifth of the ring, so ~200,000 of the million sessions move and 800,000 stay warm. Virtual nodes — each physical server claims, say, 100 ring positions — smooth out the uneven slices, so no server inherits a whale’s worth of keys. That fraction, 1/(N+1), is the whole reason the ring exists.”

Likely follow-up: “What breaks if you skip virtual nodes?” → Uneven ring slices: random server positions mean one server can own 40% of the ring. Consistent hashing without virtual nodes balances fate, not load.

The answer that sinks you: “Just rehash everything — caches warm up again.” Why it fails: it treats a million cold caches as free. At scale, the thundering refill is the outage.

Apple
Four servers at 85% load. One dies. Walk me through the next sixty seconds.

What to say (≈90 sec): “Four servers at 425 rps each against a 500 comfort line — 1,700 total. One dies: 1,700 over 3 survivors is 567 each, past the limit on every machine. Health checks cut the corpse within ~15 seconds, but the math doesn’t care: every survivor is now overloaded, latency climbs, timeouts start, clients retry — and retries are new load, which tips the next server. That’s the cascade, and the sixty-second story is the balancer failing to absorb it because there was no headroom. The fixes, in order: run at N+1 — 75% max, so one death is survivable; shed load at the balancer — fail fast with 503s instead of queueing into the overload; and break the retry storm with jittered backoff so clients don’t hammer the survivors in lockstep. The lesson I’d lead with: the balancer absorbed the failure only on paper. Absorption is a capacity plan, not a feature.”

Likely follow-up: “How do you size N+1 in practice?” → Peak traffic divided by (N−1) must stay under the per-server comfort line. Capacity planning is division, not vibes.

The answer that sinks you: “The balancer redistributes the load — problem solved.” Why it fails: redistribution without headroom just spreads the overload thinner. The cascade is the whole point of the question.

Key takeaways

  1. A load balancer has three jobs — distribute, detect, route smart — and the interview is really about the second one: what happens when a server dies.
  2. Layer 4 balances connections it can’t read; Layer 7 understands HTTP — route by URL, terminate TLS, hold sessions. Most web apps live at Layer 7.
  3. Match the algorithm to the workload: round robin for even traffic, weighted for mixed fleets, least connections for uneven costs, consistent hashing when caches must stay warm.
  4. Health checks are interval × threshold — 5-second probes with 3 strikes cut a dead server in ≤15s. Run active and passive together, and gate re-entry on consecutive passes.
  5. Sticky sessions pin a user’s fate to one server; stateless servers plus a shared session store survive server death — the grown-up answer to “where does the cart live.”
  6. Absorption is a capacity plan: keep peak under (N−1)/N of fleet capacity, or one dead server becomes a cascade — at 85% on four servers, the math already fails.

Sources & further reading

Every claim in this guide traces to one of these — the wording is ours, the ideas are credited.

Back toAll guides →