Skill
Linux Host
linux-host · current version v1
Troubleshoot and support a Linux host — triage CPU, memory, disk, I/O, network, processes, systemd services, logs, and the OOM killer on a bare-metal server, VM, or container host. The general host-triage skill the service skills build on. Use when someone reports a server is slow, unresponsive, out of disk or memory, load is high, a service will not start, the network is down, or "something is wrong with the box" and you do not yet know what.
16 downloads · published 2026-09-02
What this grants
- skill linux-host
Skill Card
- License or terms: check the skill's own repository for license details (linux-host on GitHub).
Security Audits
- NanoInfra Scanner PASS no issues found
- VirusTotal PASS no engines flagged this file (full report)
Version history
| Version | Published | Status |
|---|---|---|
| v1 | 2026-09-02 | published |
Files
SKILL.md(7850 bytes)
SKILL.md
raw | preview
Linux Host — Troubleshooting & Support
A support runbook for the box itself, not a tutorial. Run the detect-platform skill first — it tells you the distro/init, whether this is bare metal, a VM, or a container, and the CPU/RAM/disk budget. Everything below is read-only triage until §5; the goal is to name the bottleneck (CPU, memory, disk space, disk I/O, or network) before touching anything.
Assumes systemd and a mainstream distro. On a minimal container many of these tools are absent — that is itself the finding (inspect from the host instead).
0. The 60-second triage
uptime # load averages: 1/5/15 min. Compare to core count (nproc) — load≈cores is busy, ≫cores is saturated.
nproc # cores, so the load numbers mean something
top -b -n1 | head -20 # or `htop`: top CPU/mem consumers right now
free -h # memory: look at "available", not "free" (Linux uses free RAM as cache)
df -h # disk space per filesystem — a full / or /var breaks almost everything
df -i # inodes — "no space" with free bytes is exhausted inodes (many tiny files)
Load average counts processes both running (CPU) and in uninterruptible I/O wait (D state), so high load with low CPU usage means you are I/O- or network-blocked, not CPU-bound.
1. Narrow the bottleneck
CPU
top -b -n1 | head -20 # %Cpu line: us(user) sy(system) wa(io-wait) st(stolen)
mpstat -P ALL 1 3 2>/dev/null # per-core; high %iowait = disk, high %steal = noisy neighbour (VM)
ps -eo pid,ppid,%cpu,%mem,comm --sort=-%cpu | head
High st (steal) on a VM means the hypervisor is giving CPU to others — not your fault, not fixable
from inside. High wa (io-wait) sends you to the disk section.
Memory & the OOM killer
free -h
ps -eo pid,rss,comm --sort=-rss | head # top RSS consumers
cat /proc/meminfo | grep -Ei 'MemAvailable|Swap|Dirty|Committed'
dmesg -T | grep -iE 'oom|killed process' | tail # did the OOM killer strike, and who did it kill?
journalctl -k | grep -i 'out of memory' | tail
The OOM killer fires when the kernel cannot reclaim enough memory; it logs Out of memory: Killed process <pid> (<name>). A container gets OOM-killed at its cgroup limit even when the host has
free RAM — check /sys/fs/cgroup/…/memory.max (v2) / memory.limit_in_bytes (v1).
Disk space (the most common "server broken")
df -h # which filesystem is full
du -x -h -d1 / 2>/dev/null | sort -rh | head # -x stays on one fs; drill into the full one
lsof +L1 2>/dev/null | head # DELETED files still held open by a process — space not freed until it restarts
A file deleted while a process holds it open is not reclaimed until that process closes it or
restarts — df stays full while du shows less. Restart the holder (a logging daemon is typical).
Disk I/O
iostat -xz 1 3 2>/dev/null # %util near 100 = the device is the bottleneck; await = latency per IO
iotop -b -n2 2>/dev/null | head # which process is doing the I/O (needs root)
Network
ss -s # socket summary: totals, TIME-WAIT, orphaned
ss -ltnp # what is listening on which port (and the pid)
ss -tnp state established | head
ip -br a; ip -br link # interfaces up? addresses assigned?
ping -c3 <gateway>; ping -c3 1.1.1.1; getent hosts example.com # L3 vs DNS: isolate where it breaks
Split the failure: gateway reachable but 1.1.1.1 not = routing/firewall; IP works but getent
fails = DNS (/etc/resolv.conf).
2. Processes and services
systemctl --failed # every unit that failed — start here for "a service is down"
systemctl status <unit> # state, last exit, recent log lines
journalctl -u <unit> -n 200 --no-pager # that unit's log; -b this boot; -f follow
journalctl -p err -b --no-pager # all error-priority messages this boot
ps auxf # process tree — parentage of a runaway
systemctl status "Active: failed (Result: …)" + the last journal lines usually name the cause
(config, port in use, missing dependency, permission). A service that flaps has a rising
nRestarts under a Restart= policy — fix the exit, not the policy.
3. Logs
journalctl -b -p warning --no-pager # this boot, warning and worse
journalctl --since '30 min ago' --no-pager
journalctl --disk-usage # the journal itself can be what filled /var/log
dmesg -T | tail -50 # kernel ring buffer: OOM, disk errors, link up/down, segfaults
# non-journald logs still matter: /var/log/syslog|messages, and app logs under /var/log/*
dmesg for hardware/kernel-level events (I/O errors blk_update_request, filesystem remounted
read-only, link flaps); the journal for userspace/services.
4. Common problems → resolution
- "Server is slow" — triage §0/§1 to a bottleneck first. High load + high
wa= disk I/O; highus= an app burning CPU (find it inps --sort=-%cpu); highst= the hypervisor (VM); low CPU but high load = D-state processes stuck on I/O or NFS. - "Out of memory" — is it the host or a cgroup?
free -hvs the cgroup limit. Checkdmesgfor the OOM kill and its victim. Swap thrashing (highsi/soinvmstat 1) feels like a hang. - "Disk full" —
df -hfor which fs,du -xto drill in; checkdf -ifor inodes; checklsof +L1for deleted-but-open files andjournalctl --disk-usagefor a runaway journal. - "Service won't start" —
systemctl status+journalctl -u; common causes are a port already bound (ss -ltnp), a permission/SELinux/AppArmor denial (detect-platform flagged which), or a failed dependency (systemctl list-dependencies <unit>). - "Can't reach the network" — walk L2→L3→DNS with §1's ladder; a host firewall
(
nft list ruleset/iptables -S) or a wrong default route (ip route) is often it. - "Load is high but CPU looks idle" — uninterruptible I/O wait;
iostat -xz 1and look for a saturated device or an NFS/stuck mount (mount, and processes in D state inps).
5. Before you change anything
Killing processes, restarting services, editing sysctl/limits/firewall, clearing logs, and
rebooting change the host's behaviour or availability — on a nanoinfra deployment a mutate.remote
capability (approval interactively, a standing grant unattended). Name the host and the change.
kill/kill -9: prefer a gracefulSIGTERM;-9(SIGKILL) gives the process no chance to flush or clean up. Confirm the PID is what you think (ps -p <pid> -o comm=).systemctl restartinterrupts service; drain/announce first for anything user-facing.sysctl -wandulimitchanges are runtime-only — persist in/etc/sysctl.d/or/etc/security/limits.d/or a reboot loses them.- Deleting under
/var/logfrees space now, but truncate a file a process holds open (: > file) rather thanrmit, or the space is not returned until the holder restarts. - Firewall edits can lock you out of an SSH session — have out-of-band console access before you change rules on a remote host.