Load-balancer health checks: a healthy port is not a healthy app
By the Just Create App team A server can accept a TCP connection and still fail every useful request. Our load-balancer explanation covers the routing layer; this guide focuses on the checks that decide which backends s
By the Just Create App team
A server can accept a TCP connection and still fail every useful request. Our load-balancer explanation covers the routing layer; this guide focuses on the checks that decide which backends should receive work. Its process may be running while its configuration is wrong, its cache is cold or the route people need is broken. That is why a load balancer's health check deserves its own design review.
This guide is a documentation-based explanation, not a benchmark or a report from a production incident. The example is deliberately small: an HTTP application with three interchangeable backends behind one traffic entry point.
Start with what the check actually proves
A TCP check answers a narrow question: can the balancer establish a connection to the configured port? That is useful for detecting a closed socket or an unreachable target. It does not tell you whether the application returns a valid response to a real user.
An HTTP check can inspect an endpoint and its response. HAProxy documents active checks that connect to a server or send it an HTTP request at regular intervals. Consecutive failure and success thresholds determine when a server leaves and rejoins rotation. Those thresholds matter as much as the endpoint itself.
Think of each check as a contract. What exact condition makes this instance eligible to receive new work? Write that condition down before choosing a path or copying a configuration snippet.
A three-server example
Suppose requests can go to A, B or C. All three open their HTTP port. A and C have completed startup; B is still loading required configuration. A port-only check might accept all three even though B cannot serve the application yet.
A readiness endpoint should report the eligibility condition you designed for that application. In this teaching example, B is not eligible until startup has completed. The balancer can then direct new traffic to A and C while continuing to check B. Once B meets the configured success threshold, it can return to rotation.
This is a conceptual sequence, not a promise about a particular product's timing. Real behavior depends on the probe interval, thresholds, timeouts and the handling of connections already in flight. Verify the exact product documentation and test the sequence in staging.
Keep readiness separate from restart decisions
Readiness asks whether an instance should receive traffic now. Liveness asks whether a process should be restarted. Kubernetes documents them as different probes with different effects: a failed readiness probe removes a pod from service traffic, while a liveness failure can lead to a restart.
Conflating the two can turn a temporary inability to serve traffic into repeated restarts. A process still warming up may need time rather than another restart. Conversely, an instance can be alive indefinitely and still be unready. Do not assume that one endpoint covers every lifecycle state.
Choose dependency checks deliberately
It is tempting to make the health endpoint call every downstream service. That can create a shared failure signal: a dependency problem may make every backend report failure together. It is equally tempting to return 200 unconditionally, which proves little about useful work.
The better question is specific: if this dependency is unavailable, can this backend still serve the traffic that will be routed to it? A read-only route and a payment route may have different requirements. Your application's fallback behavior determines the answer; a generic checklist cannot supply it.
Keep the endpoint cheap enough to run repeatedly. Avoid an expensive report query or a destructive operation. Decide what status code or response represents eligibility, and make sure the check is not accidentally testing a cached response instead of the backend state you intended.
Plan for failures between checks
Health checks sample state. A backend may fail just after a successful probe and receive traffic until detection completes. Existing connections may have their own behavior. Removing a target from new-request selection does not mean every in-flight operation finishes successfully.
That gap belongs in your failure test. Send representative requests, take one backend out of useful service, and observe which requests fail, when the target leaves rotation and what happens during recovery. Record the result for the configuration you actually deploy. Do not describe it as instant failover because a dashboard eventually turns red.
Application retries also need care. Retrying a read and retrying a payment are different actions. A routing layer does not know whether replaying an operation creates a duplicate business effect unless the application and protocol were designed for it.
Checks do not make backends interchangeable
Routing only helps if the remaining backends can serve the same workload. If sessions or uploaded files live only on B, requests moved to A may still fail. Likewise, a broken shared database is not repaired by selecting a different web server.
Before calling the design highly available, review session storage, file storage, configuration, job execution and the availability of the traffic entry point itself. A health check is one part of that design, not its replacement.
A review checklist you can use
- State what "ready" means for this workload.
- Identify whether a TCP connection is enough or an application response is required.
- Separate traffic eligibility from process-restart logic.
- Name the dependencies that actually prevent useful work, and the behavior during a shared outage.
- Set and review intervals, failure thresholds, success thresholds and timeouts together.
- Test failure, recovery and planned draining with representative traffic.
- Check session and file access on every eligible backend.
Start with that contract, then choose the configuration. It is much easier to diagnose a check whose purpose is explicit than one inherited from an unrelated sample.
Sources and disclosure
Written by the Just Create App team. No production throughput or availability claim is made.
- HAProxy, Health checks: https://www.haproxy.com/documentation/haproxy-configuration-tutorials/reliability/health-checks/
- Kubernetes, Liveness, Readiness, and Startup Probes: https://kubernetes.io/docs/concepts/configuration/liveness-readiness-startup-probes/
- AWS, What is Elastic Load Balancing?: https://docs.aws.amazon.com/elasticloadbalancing/latest/userguide/what-is-load-balancing.html
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.