Skill
caddy
caddy · current version v2
Troubleshoot and support the Caddy web server / reverse proxy — validate config, use the admin API, and resolve problems on Docker, bare metal, systemd, or Kubernetes. Covers caddy validate/reload/adapt/fmt, the admin API on :2019, the Caddyfile vs JSON, automatic HTTPS/ACME, and 502/TLS/permission failures. Use when someone reports Caddy will not start or reload, a certificate will not issue, a reverse proxy returns 502, or a site serves the wrong thing.
13 downloads · published 2026-09-02
What this grants
- skill caddy
Skill Card
- License or terms: check the skill's own repository for license details (caddy on GitHub).
Security Audits
- NanoInfra Scanner PASS no issues found
- VirusTotal PASS no engines flagged this file (full report)
Version history
| Version | Published | Status |
|---|---|---|
| v2 | 2026-09-02 | published |
| v1 | 2026-09-02 | published |
Files
SKILL.md(9098 bytes)
SKILL.md
raw | preview
---
name: caddy
description: Troubleshoot and support the Caddy web server / reverse proxy — validate config, use the admin API, and resolve problems on Docker, bare metal, systemd, or Kubernetes. Covers caddy validate/reload/adapt/fmt, the admin API on :2019, the Caddyfile vs JSON, automatic HTTPS/ACME, and 502/TLS/permission failures. Use when someone reports Caddy will not start or reload, a certificate will not issue, a reverse proxy returns 502, or a site serves the wrong thing.
---
# Caddy — Troubleshooting & Support
A support runbook, not a tutorial. Work top to bottom: know the config form you are on,
**validate before you reload**, and remember Caddy's defining feature — it gets certificates by
itself, which is both the magic and the thing that fails. Commands and output below were run
against Caddy v2.11.4.
Caddy is on **v2**; there is no Caddy 7/8/9, and **v1 is EOL and incompatible** — a v1 config
does not run on v2. If someone has a `Caddyfile` that "stopped working after an upgrade," check
they are not carrying v1 syntax into v2.
Caddy configures two ways, and this shapes every other answer:
- a **Caddyfile** (the human format), which Caddy *adapts* into
- **JSON** (the native config the server and the admin API actually run).
`caddy adapt` shows you the JSON a Caddyfile becomes — reach for it when a Caddyfile behaves in a
way the directives do not obviously explain.
## 0. Identify version, config form, and deployment
```bash
caddy version # v2.11.4 …
caddy validate --config /etc/caddy/Caddyfile # or --config caddy.json --adapter '' for native JSON
```
Deployment: establish it with the **detect-platform** skill (bare metal vs VM vs container vs
pod, OS/package family, init system); this skill assumes that profile and keys §1 off it.
Default paths in the official image / package: config `/etc/caddy/Caddyfile`, **certificates and
ACME state `/data`** (this is the one to persist — losing it means re-issuing every cert), and
autosaved running config `/config/caddy/autosave.json`.
## 1. Where things are, per deployment
**Docker**
```bash
docker exec -it <container> caddy validate --config /etc/caddy/Caddyfile
docker logs --tail 200 -f <container> # Caddy logs structured JSON to stderr
```
Persist `/data` with a volume, or every container recreate re-runs ACME and can hit Let's
Encrypt rate limits.
**systemd (bare metal or VM)**
```bash
systemctl status caddy
journalctl -u caddy -n 200 --no-pager # Caddy's JSON logs go to the journal here
caddy validate --config /etc/caddy/Caddyfile
```
**Bare metal — what Caddy specifically needs**
```bash
# Binding :80/:443 as a non-root service: the packaged unit grants the capability, verify it.
systemctl cat caddy | grep -E 'AmbientCapabilities|User' # expect CAP_NET_BIND_SERVICE
# Outbound 443 must be open for ACME (the CA is reached over HTTPS) AND inbound 80/443 reachable
# from the internet for the HTTP-01/TLS-ALPN challenge — a firewall here is the top cert failure.
ss -ltnp | grep caddy
# /data must be writable and persistent — it holds the account key and issued certs.
ls -ld /data /data/caddy 2>/dev/null
```
**Kubernetes**
```bash
kubectl exec -it <pod> -- caddy validate --config /etc/caddy/Caddyfile
kubectl logs <pod> --tail 200 -f
```
If config comes from a ConfigMap or the admin API, edit the source — a change written into the
pod is lost on restart, and Caddy may re-pull from its configured source anyway.
## 2. The commands and arguments you will reach for
```bash
caddy validate --config <file> # load + check without serving. ALWAYS before reload. -> "Valid configuration"
caddy reload --config <file> # apply new config with ZERO downtime, via the admin API
caddy adapt --config <Caddyfile> # print the JSON a Caddyfile becomes (does not apply it)
caddy fmt --overwrite <Caddyfile> # canonical-format a Caddyfile (fixes brace/indent surprises)
caddy run --config <file> # run in the foreground (containers/systemd)
caddy start / caddy stop # background start / stop (dev; prefer run under a supervisor)
--adapter caddyfile|'' # force the config adapter; '' means the file is native JSON
--resume # start from the last autosaved config (/config/caddy/autosave.json)
```
Verified: `caddy validate` → `Valid configuration` (exit 0); `caddy reload` logs
`using config from file` / `adapted config to JSON` (exit 0); `caddy adapt` emits the JSON.
**The admin API** (this is how a running Caddy is inspected and changed) listens on
`localhost:2019` by default — confirmed in the startup log
(`admin endpoint started … localhost:2019`):
```bash
curl -s localhost:2019/config/ | jq . # the full running config (JSON)
curl -s localhost:2019/reverse_proxy/upstreams | jq . # upstream health as Caddy sees it
```
`caddy reload` drives this same endpoint; the API is why a reload never drops a connection.
## 3. Diagnostics — read-only first
```bash
caddy validate --config /etc/caddy/Caddyfile # does it even load?
caddy adapt --config /etc/caddy/Caddyfile | jq . # what does the Caddyfile actually become?
curl -s localhost:2019/config/ | jq '.apps.http.servers' # what is running right now
journalctl -u caddy -n 100 --no-pager | jq -R 'fromjson? | {level,msg,error}' 2>/dev/null
ss -ltnp | grep caddy # listening on 80/443/2019?
```
Caddy's logs are JSON — pipe them through `jq` and read `level`, `msg`, and `error`. A cert
problem logs under the `tls`/`tls.obtain` logger with the CA's own error text.
## 4. Common problems → resolution
**A certificate will not issue (the signature Caddy problem)**
Automatic HTTPS needs three things, and the log names which failed under the `tls` logger:
- **Reachability**: the ACME challenge needs the public internet to reach this host on **80**
(HTTP-01) or **443** (TLS-ALPN-01). A cloud firewall/security group closing 80 is the most
common cause. For a split-horizon or internal host, use the **DNS-01** challenge instead.
- **The name must resolve to this host.** If DNS points elsewhere, the challenge validates
elsewhere and fails. `dig +short <name>` from outside.
- **Rate limits**: repeated failed issuance (often from a non-persistent `/data` recreating the
container) hits the CA's weekly cap. Persist `/data`; while debugging, point at the CA's
**staging** endpoint so failures do not burn the real quota.
- To confirm it is issuance and not serving, `caddy validate` passes but the site serves Caddy's
internal cert / a browser warning — the log will show the pending or failed `tls.obtain`.
**502 Bad Gateway (reverse proxy)**
The upstream failed. `curl -s localhost:2019/reverse_proxy/upstreams | jq .` shows what Caddy
thinks of each backend (healthy, fails). Then `curl` the upstream directly from the Caddy host;
`dial tcp … connection refused` in the log = backend down or wrong address in `reverse_proxy`.
**Won't start / reload refused**
- `caddy validate` names the error. A frequent one is a Caddyfile that is actually v1 syntax on
v2, or a global-options block placed after site blocks (order matters). `caddy fmt` surfaces
brace/indentation mistakes that change how blocks nest.
- `loading initial config … listen tcp :443: bind: permission denied` — the process lacks
`CAP_NET_BIND_SERVICE` (see §1 bare metal) or is not root.
- `address already in use` — something else holds 80/443; `ss -ltnp`.
**Config edit "did nothing"**
Caddy may be running from a different source than the file you edited (the admin API, a JSON
config, or `--resume` from autosave). `curl -s localhost:2019/config/` is ground truth for what
is actually running; reconcile the file with it, then `caddy reload`.
**Serving the wrong site / default**
`caddy adapt` shows the matched routes and the site addresses; a request that matches no site
address gets the default handler. Confirm the host label in the Caddyfile matches the `Host`
header the client sends.
## 5. Before you run anything that changes serving
`reload`, admin-API `POST`/`PATCH`, and config edits change what users get, and on a nanoinfra
deployment they resolve to a `mutate.remote` capability — approval interactively, a standing
grant unattended.
- **`caddy validate` before every reload.** A reload of an invalid config is refused and the old
one keeps serving; validating first makes that explicit rather than discovered.
- Treat `/data` as precious: deleting it (or a non-persistent container losing it) forces
re-issuance of every certificate and can hit CA rate limits — an outage that looks like Caddy
"suddenly serving a bad cert".
- The admin API on `:2019` can rewrite the entire running config; if it is exposed beyond
localhost, that is a control-plane surface — a `POST /load` replaces everything.
- While debugging issuance, use the CA staging endpoint so a mistake does not exhaust the real
quota; switch back only once it issues clean.