Skill

Docker

docker · current version v2

Download v2

Troubleshoot and support Docker Engine and containers — inspect the daemon, containers, images, volumes, and networks, and resolve problems on a bare-metal or VM host. Covers docker ps/logs/inspect/stats/events, exit codes, restart loops, OOM kills, disk exhaustion from images/logs, port and DNS failures, and daemon-down. Use when someone reports a container keeps restarting, exits immediately, is OOM-killed, cannot reach the network, the host is out of disk, or the Docker daemon will not start.

14 downloads · published 2026-09-02

What this grants

Skill Card

Security Audits

Version history

VersionPublishedStatus
v2 2026-09-02 published
v1 2026-09-02 published

Files

SKILL.md

raw | preview

Docker — Troubleshooting & Support

A support runbook, not a tutorial. Establish the host with the detect-platform skill first — whether you are on the bare-metal/VM host, and how much RAM/disk it has, decides most of what follows. Then work top to bottom: observe before you change, and never docker-prune on a shared host without knowing what you delete.

This is Docker Engine on a Linux host. A container is not a VM — it is a process in namespaces and cgroups on this kernel, so the host's memory, disk, and pids are the container's limits.

0. Is the daemon even up?

docker version           # Client AND Server sections. Only a Client section => daemon is down/unreachable.
docker info              # daemon state: storage driver, cgroup version, root dir, live/total containers
systemctl status docker  # the daemon is a systemd service on a normal host
journalctl -u docker -n 200 --no-pager    # why the daemon failed to start (bad /etc/docker/daemon.json is common)

If docker version prints only the client, or Cannot connect to the Docker daemon at unix:///var/run/docker.sock: the daemon is down (start it), or your user is not in the docker group (permission denied on the socket), or you are pointing at a remote DOCKER_HOST.

1. See what is running and what just happened

docker ps -a             # -a includes stopped ones — with their STATUS and exit code
docker logs --tail 200 -f <c>          # stdout/stderr of the main process; --since 10m to bound it
docker events --since 15m              # daemon timeline: kills, OOMs, restarts, health transitions
docker stats --no-stream               # live CPU/mem/net/block per container; mem vs its limit
docker inspect <c>                     # the whole truth: State, RestartCount, Mounts, NetworkSettings, Config

Targeted inspect reads (Go templates) answer most questions without scrolling:

docker inspect -f '{{.State.Status}} exit={{.State.ExitCode}} oom={{.State.OOMKilled}} restarts={{.RestartCount}}' <c>
docker inspect -f '{{.State.Health.Status}}' <c>            # if a HEALTHCHECK is defined
docker inspect -f '{{json .Config.Env}}' <c>                # env the container actually has

2. Diagnostics — read-only first

docker ps -a --format 'table {{.Names}}\t{{.Status}}\t{{.Image}}'   # who is up/restarting/exited
docker inspect -f '{{.State.ExitCode}} {{.State.OOMKilled}}' <c>    # exit code + was it OOM-killed
docker system df                          # space by images / containers / volumes / build cache
docker top <c>                            # processes inside the container (PID as seen on the host)
docker port <c>                           # published port mappings

Exit codes worth knowing: 0 clean; 1/app-specific app error; 125 the daemon/docker run itself failed (bad flag); 126 command not executable; 127 command not found in the image; 137 = 128+9, SIGKILL — usually OOM (OOMKilled:true) or docker kill; 139 = SIGSEGV; 143 = 128+15, SIGTERM (a normal stop).

3. Common problems → resolution

Container keeps restarting / exits immediately

OOM-killed (exit 137, OOMKilled:true)

Host out of disk

Networking: cannot reach / cannot be reached

"No space left" but disk looks free — likely inodes (df -i) or a full /var/lib/docker filesystem specifically, not /.

Cannot exec / image debugging — a distroless or scratch image has no shell, so docker exec <c> sh fails with "no such file". Inspect from the host instead (docker inspect, docker logs, docker top), or use docker debug/an ephemeral sidecar if available.

4. Before you run anything destructive

prune, rm, rmi, volume rm, docker kill/stop, and editing daemon.json + restarting the daemon change or destroy state — on a nanoinfra deployment a mutate.remote capability (approval interactively, a standing grant unattended). Name the container/volume and the host.