Skip to main content
Ahmed Salama

Architecture Lab

Load balancing

Spread requests across identical instances, and stop sending them to the sick one.

The shape of the decision

System diagram: Load balancing: Spread requests across identical instances, and stop sending them to the sick one.Address one endpoint and know nothing about what is behind it.CLIENTClientsDistributes requests and health-checks every instance.EDGELoad balancerIdentical to B and C. Holds no state that another instance would need.SERVICEInstance ACan be replaced mid-deployment without a client noticing.SERVICEInstance BRemoved from rotation the moment its health check fails.SERVICEInstance C
Load balancing: Spread requests across identical instances, and stop sending them to the sick one.
Read this diagram as text
ClientsClient
Address one endpoint and know nothing about what is behind it.
Load balancerEdge
Distributes requests and health-checks every instance.
Instance AService
Identical to B and C. Holds no state that another instance would need.
Instance BService
Can be replaced mid-deployment without a client noticing.
Instance CService
Removed from rotation the moment its health check fails.

Connections

  • Clients → Load balancer
  • Load balancer → Instance A
  • Load balancer → Instance B
  • Load balancer → Instance C
The problem it solves
One instance is a capacity ceiling and a single point of failure at the same time. You cannot deploy without downtime, and one bad process takes the whole service with it.
How it works
A balancer in front of several identical instances distributes requests and health-checks each one. Failing instances are pulled out of rotation; new ones join it. Deployments become a rolling replacement rather than an outage.
What it costs
Your application must become stateless, which is a real constraint: no in-process sessions, no local file writes, no assuming two requests from one user land in the same place. It also moves the single point of failure to the balancer, so that has to be redundant too.
When not to reach for it
A single instance is comfortably within capacity and the business can tolerate the restart window. Two instances behind a balancer is a meaningful step up in operational complexity. Take that step when you need availability, not because it looks more professional on a diagram.

Layer 4 or layer 7

A layer 4 balancer works at the transport level. It sees an IP address, a port and a stream of packets, and it routes purely on that: no idea whether the payload is HTTP, what path is being requested, or what the request even means. That ignorance is what makes it fast, because there is no parsing and no buffering a full request before deciding where it goes.

A layer 7 balancer terminates the connection and reads the application protocol itself, the host, the path, the method, the cookies. That understanding is what lets it route by content rather than by connection, and it is not free: every request is parsed, and in most designs it is proxied rather than forwarded, meaning the balancer opens its own connection to the chosen backend instead of simply redirecting packets. That costs CPU and adds a hop most people never notice until it shows up in a latency breakdown.

TLS termination placement follows the same split. Terminate at the layer 7 balancer and it can see the decrypted request well enough to route on path or host, while backends run plain HTTP internally and carry no certificate management of their own. Terminate at layer 4, or pass the connection through untouched, and every backend needs its own certificate and its own rotation, in exchange for the traffic never appearing as plaintext anywhere the balancer runs. Re-encrypting to the backend after an L7 termination sits between the two, and it earns its cost when the internal network cannot be trusted outright.

Path-based routing is what actually earns layer 7's cost. Splitting one hostname across several backends by path, sending a canary release to a percentage of requests by cookie, or routing by content type are all decisions a layer 4 balancer cannot make, because it never sees that information. I would not reach for a full HTTP proxy for a single backend with a single purpose. Paying for one anyway is added latency and an added certificate surface bought for a feature nothing is using.

Health checks that tell the truth

A shallow health check confirms that a process is running and responding, usually a static endpoint that returns success without touching anything else. That tells you the process has not crashed. It tells you nothing about whether the instance can do its actual job: an instance whose database connection pool is exhausted will still answer a shallow check in milliseconds while every real request behind it times out.

A deep health check calls into the dependencies that matter: a database, a queue, a cache, and reports healthy only if those respond too. That is a far more honest signal of whether the instance can serve real traffic. It also has a real cost. The check itself consumes a connection against every dependency it touches, on every instance, on whatever interval the balancer polls at, and each dependency it checks becomes something the instance's own health now depends on.

That dependency is also the cascade risk. If a downstream service is merely slow rather than down, a deep check against it can time out on every instance at once, since they are all checking the same downstream on roughly the same schedule. The balancer then sees the entire fleet fail its health check simultaneously and pulls all of it from rotation, turning one degraded dependency into a complete outage the balancer itself caused, which is a worse outcome than the slow downstream would have produced alone.

Separating liveness from readiness stops two different questions from being collapsed into one check. Liveness asks whether the process should be restarted, and failing it instructs the orchestrator to kill and replace it. Readiness asks whether the instance should currently receive traffic, and failing it only pulls the instance from the balancer's rotation without touching the process at all. Wiring a downstream dependency into liveness rather than readiness is how a slow dependency turns into a restart storm across the whole fleet, destroying warm caches and in-flight work that a simple rotation pull would have left alone.

The price of sticky sessions

Stickiness buys the simplest possible migration from a single instance. Pin each client to one backend by cookie or by IP hash, and session state can stay exactly where it already lives, in that process's memory, without standing up a shared store first. Nothing about the application has to change, which is why it is usually the first thing teams reach for.

The cost shows up at deploy. A rolling release eventually has to take every instance out of rotation, including the ones holding sessions nothing else knows about, and there is no gentle way to hand that state to the instance replacing it. Either the deploy waits for stuck sessions to expire on their own, which slows every release down, or it accepts that some sessions get dropped exactly when a deploy runs, which is the worst possible moment for that to happen.

Failover costs the same thing without any warning first. When the instance holding a session fails a health check, everything pinned to it loses that state at once, and the number of people affected scales with how much traffic that instance happened to be carrying. The instance serving the most active users is also the one whose failure does the most damage, the opposite of what you want from a component meant to add resilience.

Moving session state into a shared store, a small cache or database keyed by session id, removes both costs at once. Any instance can then serve any request, so retiring one for a deploy or losing one to failure costs nothing beyond whatever was in flight at that exact moment. The balancer is free to route however is most efficient rather than being constrained to keep every client on the instance it started on.

Deploys that drain

Draining is the mechanism that makes a rolling replacement actually safe. Before an instance is terminated, it is marked as failing readiness so the balancer stops sending it new connections, while requests already in flight are left to finish normally. Killing a process outright instead cuts every one of those requests off mid-response, turning a routine deploy into a burst of errors for whoever was unlucky enough to be mid-request.

That requires an explicit budget: a drain timeout stating how long to wait for in-flight work before terminating the instance regardless. Too short and slow requests get cut off anyway, which defeats the point of draining at all. Too long and a single long-lived connection, a websocket or a long poll that may never finish on its own, holds up the whole deploy. I set that number from my own measured request duration at a high percentile, not from whatever default shipped with the platform.

Rolling replacement, a few instances at a time behind the same balancer, is what turns a deploy from an event into a background process. Fleet capacity never drops below what traffic needs, because the instances being replaced are drained and removed one small batch at a time while the rest keep serving, rather than every instance going down together and taking capacity with it.

This is where the concept's own caution about taking on that operational complexity only when you need the availability actually gets tested. Standing up a balancer and then deploying by killing every instance at once throws away exactly the availability it was built to provide, so draining and the rolling shape are not optional extras. They are the part of the design that makes the earlier decision to add a balancer worth having made, and I would rather test that draining happens the way the platform claims than assume it because a setting exists somewhere.

Let’s talk

Have a product, platform or delivery challenge? Let’s talk about turning it into a structured, scalable solution.

Open to technical leadership, product delivery and senior engineering roles, and available for architecture consulting, technical reviews and mentorship. Engagements run as project-based work, contracts, consulting, freelance engagements, remote collaboration and long-term partnerships.

Based in Cairo, Egypt, working remotely with clients across the MENA region and internationally.

Also on LinkedIn (opens in a new tab)