Skill

Docker

docker · current version v2

Download v2

Troubleshoot and support Docker Engine and containers — inspect the daemon, containers, images, volumes, and networks, and resolve problems on a bare-metal or VM host. Covers docker ps/logs/inspect/stats/events, exit codes, restart loops, OOM kills, disk exhaustion from images/logs, port and DNS failures, and daemon-down. Use when someone reports a container keeps restarting, exits immediately, is OOM-killed, cannot reach the network, the host is out of disk, or the Docker daemon will not start.

15 downloads · published 2026-09-02

What this grants

Skill Card

Security Audits

Version history

VersionPublishedStatus
v2 2026-09-02 published
v1 2026-09-02 published

Files

SKILL.md

raw | preview

---
name: docker
description: Troubleshoot and support Docker Engine and containers — inspect the daemon, containers, images, volumes, and networks, and resolve problems on a bare-metal or VM host. Covers docker ps/logs/inspect/stats/events, exit codes, restart loops, OOM kills, disk exhaustion from images/logs, port and DNS failures, and daemon-down. Use when someone reports a container keeps restarting, exits immediately, is OOM-killed, cannot reach the network, the host is out of disk, or the Docker daemon will not start.
---

# Docker — Troubleshooting & Support

A support runbook, not a tutorial. Establish the host with the **detect-platform** skill first —
whether you are on the bare-metal/VM host, and how much RAM/disk it has, decides most of what
follows. Then work top to bottom: observe before you change, and never docker-prune on a shared
host without knowing what you delete.

This is Docker Engine on a Linux host. A container is not a VM — it is a process in namespaces and
cgroups on this kernel, so the host's memory, disk, and pids are the container's limits.

## 0. Is the daemon even up?

```bash
docker version           # Client AND Server sections. Only a Client section => daemon is down/unreachable.
docker info              # daemon state: storage driver, cgroup version, root dir, live/total containers
systemctl status docker  # the daemon is a systemd service on a normal host
journalctl -u docker -n 200 --no-pager    # why the daemon failed to start (bad /etc/docker/daemon.json is common)
```
If `docker version` prints only the client, or `Cannot connect to the Docker daemon at
unix:///var/run/docker.sock`: the daemon is down (start it), or your user is not in the `docker`
group (permission denied on the socket), or you are pointing at a remote `DOCKER_HOST`.

## 1. See what is running and what just happened

```bash
docker ps -a             # -a includes stopped ones — with their STATUS and exit code
docker logs --tail 200 -f <c>          # stdout/stderr of the main process; --since 10m to bound it
docker events --since 15m              # daemon timeline: kills, OOMs, restarts, health transitions
docker stats --no-stream               # live CPU/mem/net/block per container; mem vs its limit
docker inspect <c>                     # the whole truth: State, RestartCount, Mounts, NetworkSettings, Config
```
Targeted `inspect` reads (Go templates) answer most questions without scrolling:
```bash
docker inspect -f '{{.State.Status}} exit={{.State.ExitCode}} oom={{.State.OOMKilled}} restarts={{.RestartCount}}' <c>
docker inspect -f '{{.State.Health.Status}}' <c>            # if a HEALTHCHECK is defined
docker inspect -f '{{json .Config.Env}}' <c>                # env the container actually has
```

## 2. Diagnostics — read-only first

```bash
docker ps -a --format 'table {{.Names}}\t{{.Status}}\t{{.Image}}'   # who is up/restarting/exited
docker inspect -f '{{.State.ExitCode}} {{.State.OOMKilled}}' <c>    # exit code + was it OOM-killed
docker system df                          # space by images / containers / volumes / build cache
docker top <c>                            # processes inside the container (PID as seen on the host)
docker port <c>                           # published port mappings
```
Exit codes worth knowing: **0** clean; **1**/app-specific app error; **125** the daemon/`docker run`
itself failed (bad flag); **126** command not executable; **127** command not found in the image;
**137** = 128+9, SIGKILL — usually OOM (`OOMKilled:true`) or `docker kill`; **139** = SIGSEGV;
**143** = 128+15, SIGTERM (a normal stop).

## 3. Common problems → resolution

**Container keeps restarting / exits immediately**
- `docker ps -a` shows `Restarting` or `Exited (N)`. Read `docker logs <c>` first — the app usually
  says why (missing env var, bad config, cannot reach a dependency).
- Exit **127**/**126**: the `command`/entrypoint path is wrong for this image, or not executable.
- It runs then exits **0**: the main process is not a foreground long-running process — a container
  lives only as long as PID 1. A backgrounded daemon means PID 1 exits and the container stops.
- `restart: always` turns any of the above into a fast restart loop — `RestartCount` climbs; fix the
  underlying exit, do not just remove the restart policy.

**OOM-killed (exit 137, `OOMKilled:true`)**
- The container hit its cgroup memory limit (`docker run -m`, or the host ran out). `docker stats`
  shows usage vs limit; `docker inspect -f '{{.HostConfig.Memory}}' <c>` shows the cap (0 = none, so
  it can take the whole host down). Raise the limit only if the host has the RAM (detect-platform),
  else the app is leaking or genuinely needs more.

**Host out of disk**
- `docker system df` shows where it went. The usual culprits:
  - **Container logs** with the default `json-file` driver grow unbounded — a chatty container fills
    `/var/lib/docker/containers/*/*-json.log`. Fix with `max-size`/`max-file` in the container's
    log-opts or the daemon default; the space is only freed when the container is recreated.
  - **Dangling images / build cache** — `docker system df` counts them; `docker image prune` /
    `docker builder prune` reclaim (§4 — this deletes).
  - **Unused volumes** hold data that survives the container — never blind-prune volumes.
- `/var/lib/docker` may be its own filesystem; `df -h /var/lib/docker`, not just `df -h /`.

**Networking: cannot reach / cannot be reached**
- Published port not answering: `docker port <c>` and `ss -ltnp | grep <hostport>` — is it published
  to `0.0.0.0` or only `127.0.0.1`? A `-p 127.0.0.1:8080:80` is unreachable from other hosts.
- Container→container by name fails on the **default bridge** (no built-in DNS there) but works on a
  user-defined network — put them on the same `docker network create` network.
- Container cannot resolve external DNS: `docker exec <c> cat /etc/resolv.conf`; a restrictive host
  firewall or a broken `daemon.json` `dns` entry is the usual cause.
- Port publish fails at start: another process holds the host port (`ss -ltnp`), exit 125.

**"No space left" but disk looks free** — likely **inodes** (`df -i`) or a full `/var/lib/docker`
filesystem specifically, not `/`.

**Cannot exec / image debugging** — a distroless or scratch image has no shell, so `docker exec
<c> sh` fails with "no such file". Inspect from the host instead (`docker inspect`, `docker logs`,
`docker top`), or use `docker debug`/an ephemeral sidecar if available.

## 4. Before you run anything destructive

`prune`, `rm`, `rmi`, `volume rm`, `docker kill`/`stop`, and editing `daemon.json` + restarting the
daemon change or destroy state — on a nanoinfra deployment a `mutate.remote` capability (approval
interactively, a standing grant unattended). Name the container/volume and the host.

- **`docker system prune`** removes stopped containers, unused networks, dangling images and build
  cache; add `-a` and it removes *all* unused images; add `--volumes` and it deletes unused
  **volumes** — that is data loss. On a shared host, list first (`docker ps -a`, `docker volume ls`,
  `docker system df`) and prune narrowly (`docker image prune`, not `system prune --volumes`).
- `docker rm -v` / `docker volume rm` delete the volume's data irreversibly — confirm nothing needs it.
- Restarting the daemon (`systemctl restart docker`) restarts every container without
  `restart:always`? No — modern Docker with `live-restore` can keep containers running, but a plain
  restart interrupts them. Validate `daemon.json` first: `dockerd --validate` (or check
  `journalctl -u docker` after) — a malformed `daemon.json` leaves the daemon down.
- `docker kill` sends SIGKILL (no graceful shutdown); prefer `docker stop` (SIGTERM then timeout).