Skip to content
[menu][close]

FOUNDRYNETNo. 001SEPTEMBER 2026LOAD BALANCING

Original page: foundrynet.com/services/documentation/sixl/slb.html (ServerIron SLB configuration guide, around 2000)

Server Load Balancing: VIPs, Real Servers and Health Checks

A server load balancer presents one virtual IP for a pool of real servers, spreads connections across the pool and drops any member that fails a health check.

On this page

01 What server load balancing does

Server load balancing advertises one virtual IP address (VIP) and distributes the connections arriving at it across two or more real servers. Clients only see the VIP. The balancer owns the address and keeps a table that maps every active flow to the real server handling it. The rest of the load balancing hub builds on this idea.

Three objects appear in almost every implementation. A real server is an address and port on a machine running the application. A virtual server is the VIP and port clients connect to. A bind ties a virtual port to one or more real ports. Foundry sold this as an appliance from around 1998: the ServerIron was marketed as a Layer 4 through 7 "web switch"; the original Foundry SLB configuration guide was among the most linked documents on this domain.

Basic SLB topologyBASIC SLB TOPOLOGYClient AClient BClient CVIP 192.0.2.10:80web1 10.1.1.11web2 10.1.1.12web3 10.1.1.13CLIENTSREAL SERVERSBasic SLB topologyBASIC SLB TOPOLOGYCLIENTSClient AClient BClient CVIP 192.0.2.10:80REAL SERVERSweb1 10.1.1.11web2 10.1.1.12web3 10.1.1.13
Clients connect to the VIP; the balancer maps each new connection to a real server and tracks the flow.

02 Layer 4 balancing methods

The selection method (Foundry called it the predictor) decides which real server receives the next new connection.

Most appliances default to round robin.
MethodHow it choosesFits when
Round robinNext server in a fixed rotationIdentical servers, short uniform requests
Weighted round robinRotation biased by a per-server weightMixed hardware generations in one pool
Least connectionsServer with the fewest active connectionsLong-lived or uneven requests
Response timeFastest recent health check or handshakeLatency-sensitive services on shared hosts

None of these methods remember a client between connections; per-user server state needs a persistence rule.

03 Health checks at Layer 3, 4 and 7

Checks are layered:

  • Layer 3: ICMP echo to the server. Proves the host is up.
  • Layer 4: a TCP handshake to the bound port. Proves a process is listening.
  • Layer 7: an HTTP request such as GET /health, passing only on an expected status code or body string. Proves the application answers.

Each check has an interval, a failure count before the server is marked down and a recovery count before it returns. A server that passes Layer 4 but fails Layer 7 is the classic hung process: the socket accepts, the application never answers.

RULE OF THUMB

Check what the user needs, not what is cheap to check. A 200 from a health URL that touches the database beats a completed TCP handshake.

04 Worked example: server load balancing one VIP across three real servers

Illustrative values, with the VIP from the RFC 5737 documentation range and the servers from a private range. The VIP is 192.0.2.10 on TCP 443; real servers 10.1.1.11 to 10.1.1.13 listen on 8443. The health check is an HTTPS GET /health every 5 seconds, down after 3 consecutive failures, up after 2 consecutive passes. The method is least connections, with web3 weighted at 2 because it has twice the cores.

Follow one failure. web2 hangs at 10:00:00. Its socket still accepts, so a Layer 4 check would keep passing; the Layer 7 check times out at 10:00:05, 10:00:10 and 10:00:15, and web2 is marked down at 10:00:15. New connections go to web1 and web3 from that moment; connections already open to web2 are forwarded until the client gives up, or reset, depending on the implementation.

Illustrative balancer state during the failure
lb1# show server real web2
Real server web2 10.1.1.12  state: FAILED  last check: 10:00:15 GET /health timeout
  port 8443  active conns 37  weight 1  fails 3/3
lb1# show server virtual www
Virtual www 192.0.2.10:443  method least-conn  binds: web1 8443 UP, web2 8443 FAILED, web3 8443 UP
$ curl -sk -o /dev/null -w '%{http_code} %{remote_ip}\n' https://192.0.2.10/
200 192.0.2.10

Detection takes interval times failure count, 15 seconds here; halving the interval halves it but doubles check traffic on every real server. Recovery takes interval times recovery count plus application warm-up.

05 NAT mode versus direct server return

In NAT mode the balancer rewrites the destination of each inbound packet from the VIP to the chosen real server. Return traffic must pass back through it so the source can be rewritten to the VIP, so the balancer is either the default gateway of the servers or it also source-NATs the client. NAT mode sees both directions and is required for any Layer 7 feature.

In direct server return (DSR, or transparent mode) the balancer changes only the destination MAC address. The real server carries the VIP on a loopback interface and replies straight to the client. Responses never touch the balancer, so a modest appliance can front heavy outbound traffic. The cost: no Layer 7 inspection and no server-side NAT. Foundry marketed its DSR mode under the name SwitchBack, around the early 2000s.

06 NAT, DSR and full proxy compared

Three ways to forward a balanced connection. Full proxy trades client address visibility for placement freedom.
PropertyNAT modeDirect server returnFull proxy (source NAT)
Client address seen by serverOriginal clientOriginal clientBalancer address, plus a header if Layer 7
Return pathThrough the balancerDirect to clientThrough the balancer
Server requirementBalancer is default gatewayVIP on loopback, ARP suppressedNone
Layer 7 featuresPossibleNot possibleFull

Full proxy mode makes the balancer the only client the real servers ever see, which is what makes client certificates between balancer and origin practical: one trusted peer, and any connection arriving without its certificate is not from the balancer.

07 Config shape in ServerIron-style CLI

The general shape of an SLB configuration on IronWare-era ServerIron software is below. It is illustrative: define the real servers, define the virtual server, bind the ports.

Illustrative SLB configuration shape
server real web1 10.1.1.11
 port http
 port http url "HEAD /"
server real web2 10.1.1.12
 port http
server virtual www 192.0.2.10
 port http
 bind http web1 http web2 http
!
show server real
show server virtual
show server bind

The url line attaches a Layer 7 check to the real port. The bind line maps the virtual HTTP port to each real HTTP port. Second-hand units still boot into this syntax (CLI reference).

08 High availability of the balancer itself

The balancer is a single point of failure unless paired. The usual pattern is active/standby: two units share the VIP through a VRRP-style election (see VRRP and VSRP) and the standby takes the address within seconds of the active unit disappearing. Session synchronisation copies the flow table to the standby so established connections survive failover.

09 Capacity terms that matter

Three numbers describe a balancer. Connections per second (CPS): how many new flows it can set up, the limit for busy web front ends. Concurrent sessions: the size of the flow table, exhausted first by long-lived connections. Throughput only matters in NAT mode, because DSR keeps the return direction off the box. Layer 7 processing lowers all three, the trade discussed under Layer 4 versus Layer 7. Automated clients that open many short connections spend CPS and session table entries far faster than people browsing, which makes throttling crawlers at the balancer a capacity question as much as a policy one.

10 Failure modes and pitfalls

  • ARP for the VIP in DSR. A real server that answers ARP for its loopback VIP steals traffic from the balancer for every host on the segment. Suppress ARP on the loopback (RFC 826 defines the request the server must ignore) and check the ARP table after every server build.
  • Asymmetric return in NAT mode. A server with a route around the balancer replies with its real address; the client resets a packet from an address it never spoke to. The symptom is connections that complete the handshake and then die.
  • Connection table exhaustion. Idle flows expire by timer, not by FIN, when a client vanishes. Read the table high-water mark by polling the balancer's counters over SNMP before choosing an idle timeout.

11 Expert tips

  • Give every real server a response header naming itself, stripped at the balancer for outsiders; every persistence and failover test depends on it.
  • Set the failure count so a single lost check never removes a server and the recovery count so a single pass never restores one; two and three are the usual minimums.
  • Drain before maintenance: weight zero or graceful removal, wait for the old connections to finish, then take the server down.

Server load balancing rewards conservative timers and honest checks. When the pool behind one VIP is not enough and a second site is needed, DNS-based site selection is the next tier.

Model-serving pools stretch these methods in new ways, described in balancing large language model inference servers.

12 Questions

What is the difference between a VIP and a real server?

The VIP is the address clients connect to and the balancer owns. A real server is an actual host and port behind it. One VIP maps to several real servers through a bind; clients never learn the real addresses.

Which load balancing method should be the default?

Round robin for pools of identical servers handling short requests. Switch to least connections when request duration varies or when some connections are long-lived, because round robin piles new work onto a busy server.

Why do health checks need Layer 7 at all?

A TCP handshake succeeds as long as something is listening. A hung application still listens. Only a Layer 7 check that expects a specific status code or body string proves the application is producing correct answers.

When is direct server return worth the complexity?

When responses dwarf requests, as on media or download servers, and no Layer 7 features are needed. Return traffic bypasses the balancer, so its throughput stops being the ceiling. Servers must carry the VIP on a loopback and not answer ARP for it.

Do sessions survive a balancer failover?

Only if the pair synchronises its flow table. Without session sync the standby takes the VIP but has no record of established connections, so clients see resets and reconnect. Appliance-class balancers, including the ServerIron generation, offered synchronisation as an option.

How long does a load balancer take to notice a failed server?

Check interval multiplied by the failure count, plus the check timeout on the last attempt. A 5 second interval with 3 failures gives roughly 15 to 20 seconds. Shorter intervals detect sooner but add check load on every real server and make a transient blip more likely to remove a healthy one.