Skill

Linux Host

linux-host · current version v1

Download v1

Troubleshoot and support a Linux host — triage CPU, memory, disk, I/O, network, processes, systemd services, logs, and the OOM killer on a bare-metal server, VM, or container host. The general host-triage skill the service skills build on. Use when someone reports a server is slow, unresponsive, out of disk or memory, load is high, a service will not start, the network is down, or "something is wrong with the box" and you do not yet know what.

16 downloads · published 2026-09-02

What this grants

Skill Card

Security Audits

Version history

VersionPublishedStatus
v1 2026-09-02 published

Files

SKILL.md

raw | preview

Linux Host — Troubleshooting & Support

A support runbook for the box itself, not a tutorial. Run the detect-platform skill first — it tells you the distro/init, whether this is bare metal, a VM, or a container, and the CPU/RAM/disk budget. Everything below is read-only triage until §5; the goal is to name the bottleneck (CPU, memory, disk space, disk I/O, or network) before touching anything.

Assumes systemd and a mainstream distro. On a minimal container many of these tools are absent — that is itself the finding (inspect from the host instead).

0. The 60-second triage

uptime                 # load averages: 1/5/15 min. Compare to core count (nproc) — load≈cores is busy, ≫cores is saturated.
nproc                  # cores, so the load numbers mean something
top -b -n1 | head -20  # or `htop`: top CPU/mem consumers right now
free -h                # memory: look at "available", not "free" (Linux uses free RAM as cache)
df -h                  # disk space per filesystem — a full / or /var breaks almost everything
df -i                  # inodes — "no space" with free bytes is exhausted inodes (many tiny files)

Load average counts processes both running (CPU) and in uninterruptible I/O wait (D state), so high load with low CPU usage means you are I/O- or network-blocked, not CPU-bound.

1. Narrow the bottleneck

CPU

top -b -n1 | head -20               # %Cpu line: us(user) sy(system) wa(io-wait) st(stolen)
mpstat -P ALL 1 3 2>/dev/null       # per-core; high %iowait = disk, high %steal = noisy neighbour (VM)
ps -eo pid,ppid,%cpu,%mem,comm --sort=-%cpu | head

High st (steal) on a VM means the hypervisor is giving CPU to others — not your fault, not fixable from inside. High wa (io-wait) sends you to the disk section.

Memory & the OOM killer

free -h
ps -eo pid,rss,comm --sort=-rss | head            # top RSS consumers
cat /proc/meminfo | grep -Ei 'MemAvailable|Swap|Dirty|Committed'
dmesg -T | grep -iE 'oom|killed process' | tail    # did the OOM killer strike, and who did it kill?
journalctl -k | grep -i 'out of memory' | tail

The OOM killer fires when the kernel cannot reclaim enough memory; it logs Out of memory: Killed process <pid> (<name>). A container gets OOM-killed at its cgroup limit even when the host has free RAM — check /sys/fs/cgroup/…/memory.max (v2) / memory.limit_in_bytes (v1).

Disk space (the most common "server broken")

df -h                                     # which filesystem is full
du -x -h -d1 / 2>/dev/null | sort -rh | head    # -x stays on one fs; drill into the full one
lsof +L1 2>/dev/null | head               # DELETED files still held open by a process — space not freed until it restarts

A file deleted while a process holds it open is not reclaimed until that process closes it or restarts — df stays full while du shows less. Restart the holder (a logging daemon is typical).

Disk I/O

iostat -xz 1 3 2>/dev/null    # %util near 100 = the device is the bottleneck; await = latency per IO
iotop -b -n2 2>/dev/null | head    # which process is doing the I/O (needs root)

Network

ss -s                         # socket summary: totals, TIME-WAIT, orphaned
ss -ltnp                      # what is listening on which port (and the pid)
ss -tnp state established | head
ip -br a; ip -br link         # interfaces up? addresses assigned?
ping -c3 <gateway>; ping -c3 1.1.1.1; getent hosts example.com   # L3 vs DNS: isolate where it breaks

Split the failure: gateway reachable but 1.1.1.1 not = routing/firewall; IP works but getent fails = DNS (/etc/resolv.conf).

2. Processes and services

systemctl --failed                      # every unit that failed — start here for "a service is down"
systemctl status <unit>                 # state, last exit, recent log lines
journalctl -u <unit> -n 200 --no-pager  # that unit's log; -b this boot; -f follow
journalctl -p err -b --no-pager         # all error-priority messages this boot
ps auxf                                  # process tree — parentage of a runaway

systemctl status "Active: failed (Result: …)" + the last journal lines usually name the cause (config, port in use, missing dependency, permission). A service that flaps has a rising nRestarts under a Restart= policy — fix the exit, not the policy.

3. Logs

journalctl -b -p warning --no-pager     # this boot, warning and worse
journalctl --since '30 min ago' --no-pager
journalctl --disk-usage                 # the journal itself can be what filled /var/log
dmesg -T | tail -50                      # kernel ring buffer: OOM, disk errors, link up/down, segfaults
# non-journald logs still matter: /var/log/syslog|messages, and app logs under /var/log/*

dmesg for hardware/kernel-level events (I/O errors blk_update_request, filesystem remounted read-only, link flaps); the journal for userspace/services.

4. Common problems → resolution

5. Before you change anything

Killing processes, restarting services, editing sysctl/limits/firewall, clearing logs, and rebooting change the host's behaviour or availability — on a nanoinfra deployment a mutate.remote capability (approval interactively, a standing grant unattended). Name the host and the change.