A load balancer is a piece of infrastructure that distributes incoming network traffic across a group of servers. Instead of sending every request to one machine until it collapses, a load balancer spreads the work. The result is a faster, more reliable application, even when thousands of users hit it at once.
You've relied on load balancers without knowing it. Every time you log into your bank, stream a video, or complete a purchase on a large e-commerce site, a load balancer made that feel smooth. Understanding how they work helps anyone running a web application, not just network engineers.
Why a single server isn't enough
A single web server has a ceiling. It can handle a certain number of requests per second before response times slow and connections start dropping. For low-traffic sites that ceiling rarely matters. For anything serving thousands of concurrent users, it matters a great deal.
The obvious fix is to add more servers. But adding servers creates a new problem: how do you decide which server gets which request? Without coordination, users would need to know which server to contact. That's where the load balancer steps in. It's the single address that users connect to, while the real work happens across a pool of servers sitting behind it.
This is conceptually similar to how a reverse proxy works. In fact, many reverse proxy tools like Nginx and HAProxy can act as load balancers too. The key distinction is intent: a reverse proxy is primarily about forwarding and protecting requests, while a load balancer is primarily about distributing them for capacity and resilience.
How a load balancer distributes traffic
Load balancers use algorithms to decide where each request goes. The most common ones are:
- Round robin. Requests are sent to each server in turn, one after another. Simple, and works well when servers are identical in capacity.
- Least connections. The request goes to whichever server currently has the fewest active connections. Better for workloads where some requests take longer than others.
- IP hash. The user's IP address determines which server they're sent to, every time. Useful when you need the same user to always hit the same server (called "sticky sessions").
- Weighted round robin. Servers are assigned a weight based on their capacity. A server with twice the RAM might receive twice the traffic.
The right algorithm depends on your workload. A simple blog with uniform request sizes suits round robin. A checkout system where sessions need to stay consistent suits IP hash or sticky sessions.
Layer 4 vs layer 7 load balancing
Load balancers operate at different layers of the network stack, and the layer determines how much they understand about your traffic.
Layer 4 load balancers work at the transport layer. They see the source and destination IP address and port, and route traffic accordingly. They don't read the actual content of the request. This makes them extremely fast, but also limited. They can't make routing decisions based on what's in the HTTP request itself.
Layer 7 load balancers work at the application layer. They read the full HTTP request, including the URL, headers, and cookies. This means they can route traffic based on content. A request for /api/images can go to one server pool, while a request for /api/checkout goes to another, purpose-built cluster. Layer 7 load balancing is more powerful and more common in modern cloud architectures, though it requires more processing.
Health checks: how load balancers know when a server is down
A load balancer constantly monitors the servers in its pool. It sends periodic health checks, which are small test requests to each server. If a server fails to respond within a threshold, the load balancer marks it as unhealthy and stops sending traffic to it. When the server recovers, the load balancer brings it back into rotation automatically.
This is the mechanism that gives load balancers their high availability guarantee. A single server failing doesn't take down your application. Users are silently rerouted to healthy machines without ever seeing an error. In practice, this is how large services achieve uptime figures above 99.9%.
Hardware vs software vs cloud load balancers
Load balancers come in three forms. Hardware load balancers, like those sold by F5 Networks, are physical appliances designed specifically for high-throughput traffic distribution. They're powerful and expensive, and still used in large enterprise data centres.
Software load balancers run on standard servers. HAProxy is the most widely used open-source option, capable of handling millions of requests per second on commodity hardware. Nginx also acts as a software load balancer and is deeply embedded in the Australian startup and developer ecosystem.
Cloud load balancers are managed services offered by AWS, Google Cloud, and Azure. You configure rules through a web console; the provider handles capacity, maintenance, and redundancy. For most Australian businesses building on cloud infrastructure today, a managed cloud load balancer is the default choice. It's cheaper than hardware, simpler than self-managing software, and scales automatically.
SSL termination and why it saves resources
When users connect over HTTPS, their traffic is encrypted. Decrypting that traffic (called SSL termination) is computationally expensive. Load balancers can handle SSL termination at the edge, decrypting each request once before forwarding it to backend servers over plain HTTP within a private network. Backend servers skip the decryption step entirely, which reduces CPU load across the entire pool.
This is one of the less-discussed reasons load balancers appear in nearly every serious web architecture. The capacity benefits extend beyond traffic distribution.
Where load balancers sit in a broader architecture
In a typical modern setup, a user's request travels from their browser through a CDN (which serves cached static files), then hits a load balancer, which forwards the request to one of several application servers. Those servers may query a database or call internal APIs, and the response travels back the same way.
If you've read about how a CDN works, you'll notice that CDNs and load balancers solve adjacent problems. A CDN reduces the number of requests that ever reach your infrastructure by caching content at edge locations. A load balancer handles the requests that do reach your infrastructure, distributing them across servers. Most production systems use both.
For anyone building a web application that expects more than light traffic, a load balancer isn't optional. It's the mechanism that turns a fragile single-server setup into something that can grow, recover from failures, and serve users reliably.

