Skill
Linux Host
linux-host · current version v1
Troubleshoot and support a Linux host — triage CPU, memory, disk, I/O, network, processes, systemd services, logs, and the OOM killer on a bare-metal server, VM, or container host. The general host-triage skill the service skills build on. Use when someone reports a server is slow, unresponsive, out of disk or memory, load is high, a service will not start, the network is down, or "something is wrong with the box" and you do not yet know what.
16 downloads · published 2026-09-02
What this grants
- skill linux-host
Skill Card
- License or terms: check the skill's own repository for license details (linux-host on GitHub).
Security Audits
- NanoInfra Scanner PASS no issues found
- VirusTotal PASS no engines flagged this file (full report)
Version history
| Version | Published | Status |
|---|---|---|
| v1 | 2026-09-02 | published |
Files
SKILL.md(7850 bytes)
SKILL.md
raw | preview
--- name: linux-host description: Troubleshoot and support a Linux host — triage CPU, memory, disk, I/O, network, processes, systemd services, logs, and the OOM killer on a bare-metal server, VM, or container host. The general host-triage skill the service skills build on. Use when someone reports a server is slow, unresponsive, out of disk or memory, load is high, a service will not start, the network is down, or "something is wrong with the box" and you do not yet know what. --- # Linux Host — Troubleshooting & Support A support runbook for the box itself, not a tutorial. Run the **detect-platform** skill first — it tells you the distro/init, whether this is bare metal, a VM, or a container, and the CPU/RAM/disk budget. Everything below is read-only triage until §5; the goal is to name the bottleneck (CPU, memory, disk space, disk I/O, or network) before touching anything. Assumes systemd and a mainstream distro. On a minimal container many of these tools are absent — that is itself the finding (inspect from the host instead). ## 0. The 60-second triage ```bash uptime # load averages: 1/5/15 min. Compare to core count (nproc) — load≈cores is busy, ≫cores is saturated. nproc # cores, so the load numbers mean something top -b -n1 | head -20 # or `htop`: top CPU/mem consumers right now free -h # memory: look at "available", not "free" (Linux uses free RAM as cache) df -h # disk space per filesystem — a full / or /var breaks almost everything df -i # inodes — "no space" with free bytes is exhausted inodes (many tiny files) ``` Load average counts processes both running (CPU) **and** in uninterruptible I/O wait (D state), so high load with low CPU usage means you are I/O- or network-blocked, not CPU-bound. ## 1. Narrow the bottleneck **CPU** ```bash top -b -n1 | head -20 # %Cpu line: us(user) sy(system) wa(io-wait) st(stolen) mpstat -P ALL 1 3 2>/dev/null # per-core; high %iowait = disk, high %steal = noisy neighbour (VM) ps -eo pid,ppid,%cpu,%mem,comm --sort=-%cpu | head ``` High `st` (steal) on a VM means the hypervisor is giving CPU to others — not your fault, not fixable from inside. High `wa` (io-wait) sends you to the disk section. **Memory & the OOM killer** ```bash free -h ps -eo pid,rss,comm --sort=-rss | head # top RSS consumers cat /proc/meminfo | grep -Ei 'MemAvailable|Swap|Dirty|Committed' dmesg -T | grep -iE 'oom|killed process' | tail # did the OOM killer strike, and who did it kill? journalctl -k | grep -i 'out of memory' | tail ``` The OOM killer fires when the kernel cannot reclaim enough memory; it logs `Out of memory: Killed process <pid> (<name>)`. A container gets OOM-killed at its **cgroup** limit even when the host has free RAM — check `/sys/fs/cgroup/…/memory.max` (v2) / `memory.limit_in_bytes` (v1). **Disk space (the most common "server broken")** ```bash df -h # which filesystem is full du -x -h -d1 / 2>/dev/null | sort -rh | head # -x stays on one fs; drill into the full one lsof +L1 2>/dev/null | head # DELETED files still held open by a process — space not freed until it restarts ``` A file deleted while a process holds it open is not reclaimed until that process closes it or restarts — `df` stays full while `du` shows less. Restart the holder (a logging daemon is typical). **Disk I/O** ```bash iostat -xz 1 3 2>/dev/null # %util near 100 = the device is the bottleneck; await = latency per IO iotop -b -n2 2>/dev/null | head # which process is doing the I/O (needs root) ``` **Network** ```bash ss -s # socket summary: totals, TIME-WAIT, orphaned ss -ltnp # what is listening on which port (and the pid) ss -tnp state established | head ip -br a; ip -br link # interfaces up? addresses assigned? ping -c3 <gateway>; ping -c3 1.1.1.1; getent hosts example.com # L3 vs DNS: isolate where it breaks ``` Split the failure: gateway reachable but `1.1.1.1` not = routing/firewall; IP works but `getent` fails = DNS (`/etc/resolv.conf`). ## 2. Processes and services ```bash systemctl --failed # every unit that failed — start here for "a service is down" systemctl status <unit> # state, last exit, recent log lines journalctl -u <unit> -n 200 --no-pager # that unit's log; -b this boot; -f follow journalctl -p err -b --no-pager # all error-priority messages this boot ps auxf # process tree — parentage of a runaway ``` `systemctl status` "Active: failed (Result: …)" + the last journal lines usually name the cause (config, port in use, missing dependency, permission). A service that flaps has a rising `nRestarts` under a `Restart=` policy — fix the exit, not the policy. ## 3. Logs ```bash journalctl -b -p warning --no-pager # this boot, warning and worse journalctl --since '30 min ago' --no-pager journalctl --disk-usage # the journal itself can be what filled /var/log dmesg -T | tail -50 # kernel ring buffer: OOM, disk errors, link up/down, segfaults # non-journald logs still matter: /var/log/syslog|messages, and app logs under /var/log/* ``` `dmesg` for hardware/kernel-level events (I/O errors `blk_update_request`, filesystem remounted read-only, link flaps); the journal for userspace/services. ## 4. Common problems → resolution - **"Server is slow"** — triage §0/§1 to a bottleneck first. High load + high `wa` = disk I/O; high `us` = an app burning CPU (find it in `ps --sort=-%cpu`); high `st` = the hypervisor (VM); low CPU but high load = D-state processes stuck on I/O or NFS. - **"Out of memory"** — is it the host or a cgroup? `free -h` vs the cgroup limit. Check `dmesg` for the OOM kill and its victim. Swap thrashing (high `si/so` in `vmstat 1`) feels like a hang. - **"Disk full"** — `df -h` for which fs, `du -x` to drill in; check `df -i` for inodes; check `lsof +L1` for deleted-but-open files and `journalctl --disk-usage` for a runaway journal. - **"Service won't start"** — `systemctl status` + `journalctl -u`; common causes are a port already bound (`ss -ltnp`), a permission/SELinux/AppArmor denial (detect-platform flagged which), or a failed dependency (`systemctl list-dependencies <unit>`). - **"Can't reach the network"** — walk L2→L3→DNS with §1's ladder; a host firewall (`nft list ruleset` / `iptables -S`) or a wrong default route (`ip route`) is often it. - **"Load is high but CPU looks idle"** — uninterruptible I/O wait; `iostat -xz 1` and look for a saturated device or an NFS/stuck mount (`mount`, and processes in D state in `ps`). ## 5. Before you change anything Killing processes, restarting services, editing sysctl/limits/firewall, clearing logs, and rebooting change the host's behaviour or availability — on a nanoinfra deployment a `mutate.remote` capability (approval interactively, a standing grant unattended). Name the host and the change. - `kill`/`kill -9`: prefer a graceful `SIGTERM`; `-9` (SIGKILL) gives the process no chance to flush or clean up. Confirm the PID is what you think (`ps -p <pid> -o comm=`). - `systemctl restart` interrupts service; drain/announce first for anything user-facing. - `sysctl -w` and `ulimit` changes are runtime-only — persist in `/etc/sysctl.d/` or `/etc/security/limits.d/` or a reboot loses them. - Deleting under `/var/log` frees space now, but truncate a file a process holds open (`: > file`) rather than `rm` it, or the space is not returned until the holder restarts. - Firewall edits can lock you out of an SSH session — have out-of-band console access before you change rules on a remote host.