Skill

HAProxy

haproxy · current version v1

Download v1

Troubleshoot and support HAProxy — validate config, read the stats/runtime API, and resolve problems on Docker, bare metal, systemd, or Kubernetes. Covers haproxy -c/-vv, the runtime API socket, the stats page, backend health, and the common 503/504/SSL/connection failures of a load balancer. Use when someone reports HAProxy will not start or reload, returns 503/504, marks backends down, drops connections, or fails a TLS handshake.

11 downloads · published 2026-09-02

What this grants

Skill Card

Security Audits

Version history

VersionPublishedStatus
v1 2026-09-02 published

Files

SKILL.md

raw | preview

---
name: haproxy
description: Troubleshoot and support HAProxy — validate config, read the stats/runtime API, and resolve problems on Docker, bare metal, systemd, or Kubernetes. Covers haproxy -c/-vv, the runtime API socket, the stats page, backend health, and the common 503/504/SSL/connection failures of a load balancer. Use when someone reports HAProxy will not start or reload, returns 503/504, marks backends down, drops connections, or fails a TLS handshake.
---

# HAProxy — Troubleshooting & Support

A support runbook, not a tutorial. Establish the host with the **detect-platform** skill first,
then work top to bottom: **validate before you reload**, and remember HAProxy is almost never the
fault — it reports the backend's health, and a 503/504 is usually the backend. Commands and output
below were run against HAProxy 3.4.4.

HAProxy versions by year (2.8 LTS, 3.0 LTS, 3.1/3.2…, 3.4 current). There is no "7/8/9". The LTS
lines (even-ish, marked LTS on haproxy.org) are what most deployments run; diagnose against
`haproxy -v`.

## 0. Identify and validate

```bash
haproxy -v                       # HAProxy version 3.4.4-… 2026/08/27
haproxy -vv                      # build options + which features (OpenSSL, Lua, threads) are compiled in
haproxy -c -f /etc/haproxy/haproxy.cfg   # CONFIG CHECK — "Configuration file is valid" (exit 0). Always before reload.
```
`haproxy -vv` matters: a config using `ssl` or `lua` fails if the binary was built without them —
`-vv` shows what is actually present.

Deployment (detect-platform gave you bare metal vs container vs pod). Config lives at
`/etc/haproxy/haproxy.cfg` in the package and the official image.

## 1. Connect / observe, per deployment

**Docker**
```bash
docker exec -it <c> haproxy -c -f /usr/local/etc/haproxy/haproxy.cfg   # image config path
docker logs --tail 200 -f <c>          # HAProxy logs to stdout in the image
```
**systemd (bare metal or VM)**
```bash
systemctl status haproxy
journalctl -u haproxy -n 200 --no-pager
# On a package install HAProxy usually logs via rsyslog to /var/log/haproxy.log — check there too.
```
**Bare metal** — the LB-specific host limits (detect-platform found the substrate):
```bash
# A load balancer holds ~2 sockets per client (front+back); fd limits and ephemeral ports bite first.
cat /proc/$(pgrep -o haproxy)/limits | grep 'open files'   # vs `maxconn` in the global section
sysctl net.ipv4.ip_local_port_range     # exhaustion here = intermittent connection failures to backends
sysctl net.core.somaxconn                # listen backlog under a connection burst
```
**Kubernetes** — `kubectl exec -it <pod> -- haproxy -c -f <cfg>`; if it is an ingress controller the
config is generated — edit the source, not the file in the pod.

## 2. The runtime API and stats (how you see live state)

HAProxy exposes a **runtime API** on a unix (or TCP) socket declared with `stats socket` in the
`global` section, and optionally an HTML **stats page** via a `stats` listener.
```bash
# Runtime API (adjust the socket path to your global `stats socket`):
echo "show info"      | socat stdio /run/haproxy/admin.sock     # version, uptime, current conns, maxconn
echo "show stat"      | socat stdio /run/haproxy/admin.sock     # per frontend/backend/server CSV: status, sessions, errors
echo "show servers state" | socat stdio /run/haproxy/admin.sock # per-server operational state
echo "show sess"      | socat stdio /run/haproxy/admin.sock     # live sessions
# If socat is absent: `nc -U /run/haproxy/admin.sock` then type the command.
```
The HTML stats page (if a `stats` listener is configured, e.g. on `:8404/stats`) shows the same
`show stat` data in a browser — the fastest way to see which backend server is red.

## 3. Diagnostics — read-only first

```bash
haproxy -c -f <cfg>                       # does the config even load?
echo "show stat" | socat stdio <sock> | awk -F, '$18!="UP"&&$2!=""{print $1"/"$2" = "$18}'  # non-UP servers
echo "show info" | socat stdio <sock> | grep -E 'CurrConns|CumConns|Maxconn|Idle_pct'
ss -ltnp | grep haproxy                   # listening on the bind you expect?
```
The `show stat` columns to read: `status` (UP/DOWN/MAINT), `check_status` (why a health check
failed — L4CON, L7STS, L7TOUT), `econ`/`eresp` (connection/response errors), `qcur` (queued).

## 4. Common problems → resolution

**Won't start / reload refused**
- `haproxy -c -f <cfg>` names the file:line. Reload with validation: a reload of a bad config is
  refused and the running instance keeps serving; a blind restart of a bad config exits.
- `cannot bind socket … :443` — another process holds the port, or (bare metal, non-root) HAProxy
  lacks `CAP_NET_BIND_SERVICE`; the packaged systemd unit grants it — `systemctl cat haproxy`.
- A directive rejected as unknown: the binary lacks that feature — `haproxy -vv` (e.g. an `ssl`
  keyword on a no-SSL build).

**503 Service Unavailable**
This is HAProxy's own answer meaning "no available server to send this to":
- All servers in the backend are DOWN — `show stat` and the `check_status` column say why the
  health check fails (L4CON = TCP refused/backend down; L7STS = HTTP check got a bad status; L7TOUT
  = check timed out). Fix the backend or the check definition.
- No backend matched the request (a `use_backend`/ACL gap) and no `default_backend` — the request
  has nowhere to go.
- `maxconn` reached and the queue is full (`qcur` high) — the backends cannot keep up.

**504 Gateway Timeout**
The backend accepted but did not answer within `timeout server`. Confirm the backend is genuinely
slow (not HAProxy) before raising `timeout server` / `timeout connect`.

**Backends flap UP/DOWN**
The health check is too aggressive or checks the wrong thing — a `check` hitting `/` that the app
sometimes 500s will flap. Point the check at a real health endpoint and tune `inter`/`fall`/`rise`.

**Connections dropped under load**
`maxconn` (global and per-listener) caps concurrency; raise it, but raise the OS `ulimit -n` with it
(HAProxy needs fds for both sides) or it silently caps lower. Ephemeral-port exhaustion to a single
backend shows as intermittent `econ` — spread across more backend addresses or enable reuse.

**TLS handshake fails**
- `haproxy -vv | grep -i ssl` — is SSL compiled in? Then the `bind … ssl crt <pem>` must point at a
  PEM that is **cert+key+chain concatenated** (HAProxy wants them in one file), not separate files.
- `openssl s_client -connect host:443 -servername <name>` shows what is served for that SNI; a
  wrong default cert means the SNI did not match a `crt`.

## 5. Before you run anything that changes serving

Reloads, runtime-API `set`/`disable`/`enable` commands, and config edits change what users get —
on a nanoinfra deployment a `mutate.remote` capability (approval interactively, a standing grant
unattended).

- **`haproxy -c -f <cfg>` before every reload.** Modern reloads are hitless (the new process takes
  over the listeners), but only a valid config reloads.
- Runtime-API writes are live and un-persisted: `disable server b/s1`, `set server b/s1 state maint`,
  `set weight` change traffic immediately and are lost on the next reload — mirror the intent in the
  config if it should survive.
- Prefer `set server … state maint` to drain a backend gracefully over yanking it from the config.