Skill

caddy

caddy · current version v2

Download v2

Troubleshoot and support the Caddy web server / reverse proxy — validate config, use the admin API, and resolve problems on Docker, bare metal, systemd, or Kubernetes. Covers caddy validate/reload/adapt/fmt, the admin API on :2019, the Caddyfile vs JSON, automatic HTTPS/ACME, and 502/TLS/permission failures. Use when someone reports Caddy will not start or reload, a certificate will not issue, a reverse proxy returns 502, or a site serves the wrong thing.

12 downloads · published 2026-09-02

What this grants

Skill Card

Security Audits

Version history

VersionPublishedStatus
v2 2026-09-02 published
v1 2026-09-02 published

Files

SKILL.md

raw | preview

---
name: caddy
description: Troubleshoot and support the Caddy web server / reverse proxy — validate config, use the admin API, and resolve problems on Docker, bare metal, systemd, or Kubernetes. Covers caddy validate/reload/adapt/fmt, the admin API on :2019, the Caddyfile vs JSON, automatic HTTPS/ACME, and 502/TLS/permission failures. Use when someone reports Caddy will not start or reload, a certificate will not issue, a reverse proxy returns 502, or a site serves the wrong thing.
---

# Caddy — Troubleshooting & Support

A support runbook, not a tutorial. Work top to bottom: know the config form you are on,
**validate before you reload**, and remember Caddy's defining feature — it gets certificates by
itself, which is both the magic and the thing that fails. Commands and output below were run
against Caddy v2.11.4.

Caddy is on **v2**; there is no Caddy 7/8/9, and **v1 is EOL and incompatible** — a v1 config
does not run on v2. If someone has a `Caddyfile` that "stopped working after an upgrade," check
they are not carrying v1 syntax into v2.

Caddy configures two ways, and this shapes every other answer:
- a **Caddyfile** (the human format), which Caddy *adapts* into
- **JSON** (the native config the server and the admin API actually run).

`caddy adapt` shows you the JSON a Caddyfile becomes — reach for it when a Caddyfile behaves in a
way the directives do not obviously explain.

## 0. Identify version, config form, and deployment

```bash
caddy version                                  # v2.11.4 …
caddy validate --config /etc/caddy/Caddyfile   # or --config caddy.json --adapter '' for native JSON
```
Deployment: establish it with the **detect-platform** skill (bare metal vs VM vs container vs
pod, OS/package family, init system); this skill assumes that profile and keys §1 off it.

Default paths in the official image / package: config `/etc/caddy/Caddyfile`, **certificates and
ACME state `/data`** (this is the one to persist — losing it means re-issuing every cert), and
autosaved running config `/config/caddy/autosave.json`.

## 1. Where things are, per deployment

**Docker**
```bash
docker exec -it <container> caddy validate --config /etc/caddy/Caddyfile
docker logs --tail 200 -f <container>          # Caddy logs structured JSON to stderr
```
Persist `/data` with a volume, or every container recreate re-runs ACME and can hit Let's
Encrypt rate limits.

**systemd (bare metal or VM)**
```bash
systemctl status caddy
journalctl -u caddy -n 200 --no-pager          # Caddy's JSON logs go to the journal here
caddy validate --config /etc/caddy/Caddyfile
```

**Bare metal — what Caddy specifically needs**
```bash
# Binding :80/:443 as a non-root service: the packaged unit grants the capability, verify it.
systemctl cat caddy | grep -E 'AmbientCapabilities|User'   # expect CAP_NET_BIND_SERVICE
# Outbound 443 must be open for ACME (the CA is reached over HTTPS) AND inbound 80/443 reachable
# from the internet for the HTTP-01/TLS-ALPN challenge — a firewall here is the top cert failure.
ss -ltnp | grep caddy
# /data must be writable and persistent — it holds the account key and issued certs.
ls -ld /data /data/caddy 2>/dev/null
```

**Kubernetes**
```bash
kubectl exec -it <pod> -- caddy validate --config /etc/caddy/Caddyfile
kubectl logs <pod> --tail 200 -f
```
If config comes from a ConfigMap or the admin API, edit the source — a change written into the
pod is lost on restart, and Caddy may re-pull from its configured source anyway.

## 2. The commands and arguments you will reach for

```bash
caddy validate --config <file>          # load + check without serving. ALWAYS before reload. -> "Valid configuration"
caddy reload   --config <file>          # apply new config with ZERO downtime, via the admin API
caddy adapt    --config <Caddyfile>     # print the JSON a Caddyfile becomes (does not apply it)
caddy fmt --overwrite <Caddyfile>       # canonical-format a Caddyfile (fixes brace/indent surprises)
caddy run  --config <file>              # run in the foreground (containers/systemd)
caddy start / caddy stop                # background start / stop (dev; prefer run under a supervisor)
--adapter caddyfile|''                  # force the config adapter; '' means the file is native JSON
--resume                                # start from the last autosaved config (/config/caddy/autosave.json)
```
Verified: `caddy validate` → `Valid configuration` (exit 0); `caddy reload` logs
`using config from file` / `adapted config to JSON` (exit 0); `caddy adapt` emits the JSON.

**The admin API** (this is how a running Caddy is inspected and changed) listens on
`localhost:2019` by default — confirmed in the startup log
(`admin endpoint started … localhost:2019`):
```bash
curl -s localhost:2019/config/        | jq .    # the full running config (JSON)
curl -s localhost:2019/reverse_proxy/upstreams | jq .   # upstream health as Caddy sees it
```
`caddy reload` drives this same endpoint; the API is why a reload never drops a connection.

## 3. Diagnostics — read-only first

```bash
caddy validate --config /etc/caddy/Caddyfile     # does it even load?
caddy adapt --config /etc/caddy/Caddyfile | jq . # what does the Caddyfile actually become?
curl -s localhost:2019/config/ | jq '.apps.http.servers'   # what is running right now
journalctl -u caddy -n 100 --no-pager | jq -R 'fromjson? | {level,msg,error}' 2>/dev/null
ss -ltnp | grep caddy                             # listening on 80/443/2019?
```
Caddy's logs are JSON — pipe them through `jq` and read `level`, `msg`, and `error`. A cert
problem logs under the `tls`/`tls.obtain` logger with the CA's own error text.

## 4. Common problems → resolution

**A certificate will not issue (the signature Caddy problem)**
Automatic HTTPS needs three things, and the log names which failed under the `tls` logger:
- **Reachability**: the ACME challenge needs the public internet to reach this host on **80**
  (HTTP-01) or **443** (TLS-ALPN-01). A cloud firewall/security group closing 80 is the most
  common cause. For a split-horizon or internal host, use the **DNS-01** challenge instead.
- **The name must resolve to this host.** If DNS points elsewhere, the challenge validates
  elsewhere and fails. `dig +short <name>` from outside.
- **Rate limits**: repeated failed issuance (often from a non-persistent `/data` recreating the
  container) hits the CA's weekly cap. Persist `/data`; while debugging, point at the CA's
  **staging** endpoint so failures do not burn the real quota.
- To confirm it is issuance and not serving, `caddy validate` passes but the site serves Caddy's
  internal cert / a browser warning — the log will show the pending or failed `tls.obtain`.

**502 Bad Gateway (reverse proxy)**
The upstream failed. `curl -s localhost:2019/reverse_proxy/upstreams | jq .` shows what Caddy
thinks of each backend (healthy, fails). Then `curl` the upstream directly from the Caddy host;
`dial tcp … connection refused` in the log = backend down or wrong address in `reverse_proxy`.

**Won't start / reload refused**
- `caddy validate` names the error. A frequent one is a Caddyfile that is actually v1 syntax on
  v2, or a global-options block placed after site blocks (order matters). `caddy fmt` surfaces
  brace/indentation mistakes that change how blocks nest.
- `loading initial config … listen tcp :443: bind: permission denied` — the process lacks
  `CAP_NET_BIND_SERVICE` (see §1 bare metal) or is not root.
- `address already in use` — something else holds 80/443; `ss -ltnp`.

**Config edit "did nothing"**
Caddy may be running from a different source than the file you edited (the admin API, a JSON
config, or `--resume` from autosave). `curl -s localhost:2019/config/` is ground truth for what
is actually running; reconcile the file with it, then `caddy reload`.

**Serving the wrong site / default**
`caddy adapt` shows the matched routes and the site addresses; a request that matches no site
address gets the default handler. Confirm the host label in the Caddyfile matches the `Host`
header the client sends.

## 5. Before you run anything that changes serving

`reload`, admin-API `POST`/`PATCH`, and config edits change what users get, and on a nanoinfra
deployment they resolve to a `mutate.remote` capability — approval interactively, a standing
grant unattended.

- **`caddy validate` before every reload.** A reload of an invalid config is refused and the old
  one keeps serving; validating first makes that explicit rather than discovered.
- Treat `/data` as precious: deleting it (or a non-persistent container losing it) forces
  re-issuance of every certificate and can hit CA rate limits — an outage that looks like Caddy
  "suddenly serving a bad cert".
- The admin API on `:2019` can rewrite the entire running config; if it is exposed beyond
  localhost, that is a control-plane surface — a `POST /load` replaces everything.
- While debugging issuance, use the CA staging endpoint so a mistake does not exhaust the real
  quota; switch back only once it issues clean.