Skill

Linux Host

linux-host · current version v1

Download v1

Troubleshoot and support a Linux host — triage CPU, memory, disk, I/O, network, processes, systemd services, logs, and the OOM killer on a bare-metal server, VM, or container host. The general host-triage skill the service skills build on. Use when someone reports a server is slow, unresponsive, out of disk or memory, load is high, a service will not start, the network is down, or "something is wrong with the box" and you do not yet know what.

15 downloads · published 2026-09-02

What this grants

Skill Card

Security Audits

Version history

VersionPublishedStatus
v1 2026-09-02 published

Files

SKILL.md

raw | preview

---
name: linux-host
description: Troubleshoot and support a Linux host — triage CPU, memory, disk, I/O, network, processes, systemd services, logs, and the OOM killer on a bare-metal server, VM, or container host. The general host-triage skill the service skills build on. Use when someone reports a server is slow, unresponsive, out of disk or memory, load is high, a service will not start, the network is down, or "something is wrong with the box" and you do not yet know what.
---

# Linux Host — Troubleshooting & Support

A support runbook for the box itself, not a tutorial. Run the **detect-platform** skill first — it
tells you the distro/init, whether this is bare metal, a VM, or a container, and the CPU/RAM/disk
budget. Everything below is read-only triage until §5; the goal is to name the bottleneck (CPU,
memory, disk space, disk I/O, or network) before touching anything.

Assumes systemd and a mainstream distro. On a minimal container many of these tools are absent —
that is itself the finding (inspect from the host instead).

## 0. The 60-second triage

```bash
uptime                 # load averages: 1/5/15 min. Compare to core count (nproc) — load≈cores is busy, ≫cores is saturated.
nproc                  # cores, so the load numbers mean something
top -b -n1 | head -20  # or `htop`: top CPU/mem consumers right now
free -h                # memory: look at "available", not "free" (Linux uses free RAM as cache)
df -h                  # disk space per filesystem — a full / or /var breaks almost everything
df -i                  # inodes — "no space" with free bytes is exhausted inodes (many tiny files)
```
Load average counts processes both running (CPU) **and** in uninterruptible I/O wait (D state), so
high load with low CPU usage means you are I/O- or network-blocked, not CPU-bound.

## 1. Narrow the bottleneck

**CPU**
```bash
top -b -n1 | head -20               # %Cpu line: us(user) sy(system) wa(io-wait) st(stolen)
mpstat -P ALL 1 3 2>/dev/null       # per-core; high %iowait = disk, high %steal = noisy neighbour (VM)
ps -eo pid,ppid,%cpu,%mem,comm --sort=-%cpu | head
```
High `st` (steal) on a VM means the hypervisor is giving CPU to others — not your fault, not fixable
from inside. High `wa` (io-wait) sends you to the disk section.

**Memory & the OOM killer**
```bash
free -h
ps -eo pid,rss,comm --sort=-rss | head            # top RSS consumers
cat /proc/meminfo | grep -Ei 'MemAvailable|Swap|Dirty|Committed'
dmesg -T | grep -iE 'oom|killed process' | tail    # did the OOM killer strike, and who did it kill?
journalctl -k | grep -i 'out of memory' | tail
```
The OOM killer fires when the kernel cannot reclaim enough memory; it logs `Out of memory: Killed
process <pid> (<name>)`. A container gets OOM-killed at its **cgroup** limit even when the host has
free RAM — check `/sys/fs/cgroup/…/memory.max` (v2) / `memory.limit_in_bytes` (v1).

**Disk space (the most common "server broken")**
```bash
df -h                                     # which filesystem is full
du -x -h -d1 / 2>/dev/null | sort -rh | head    # -x stays on one fs; drill into the full one
lsof +L1 2>/dev/null | head               # DELETED files still held open by a process — space not freed until it restarts
```
A file deleted while a process holds it open is not reclaimed until that process closes it or
restarts — `df` stays full while `du` shows less. Restart the holder (a logging daemon is typical).

**Disk I/O**
```bash
iostat -xz 1 3 2>/dev/null    # %util near 100 = the device is the bottleneck; await = latency per IO
iotop -b -n2 2>/dev/null | head    # which process is doing the I/O (needs root)
```

**Network**
```bash
ss -s                         # socket summary: totals, TIME-WAIT, orphaned
ss -ltnp                      # what is listening on which port (and the pid)
ss -tnp state established | head
ip -br a; ip -br link         # interfaces up? addresses assigned?
ping -c3 <gateway>; ping -c3 1.1.1.1; getent hosts example.com   # L3 vs DNS: isolate where it breaks
```
Split the failure: gateway reachable but `1.1.1.1` not = routing/firewall; IP works but `getent`
fails = DNS (`/etc/resolv.conf`).

## 2. Processes and services

```bash
systemctl --failed                      # every unit that failed — start here for "a service is down"
systemctl status <unit>                 # state, last exit, recent log lines
journalctl -u <unit> -n 200 --no-pager  # that unit's log; -b this boot; -f follow
journalctl -p err -b --no-pager         # all error-priority messages this boot
ps auxf                                  # process tree — parentage of a runaway
```
`systemctl status` "Active: failed (Result: …)" + the last journal lines usually name the cause
(config, port in use, missing dependency, permission). A service that flaps has a rising
`nRestarts` under a `Restart=` policy — fix the exit, not the policy.

## 3. Logs

```bash
journalctl -b -p warning --no-pager     # this boot, warning and worse
journalctl --since '30 min ago' --no-pager
journalctl --disk-usage                 # the journal itself can be what filled /var/log
dmesg -T | tail -50                      # kernel ring buffer: OOM, disk errors, link up/down, segfaults
# non-journald logs still matter: /var/log/syslog|messages, and app logs under /var/log/*
```
`dmesg` for hardware/kernel-level events (I/O errors `blk_update_request`, filesystem remounted
read-only, link flaps); the journal for userspace/services.

## 4. Common problems → resolution

- **"Server is slow"** — triage §0/§1 to a bottleneck first. High load + high `wa` = disk I/O; high
  `us` = an app burning CPU (find it in `ps --sort=-%cpu`); high `st` = the hypervisor (VM); low CPU
  but high load = D-state processes stuck on I/O or NFS.
- **"Out of memory"** — is it the host or a cgroup? `free -h` vs the cgroup limit. Check `dmesg` for
  the OOM kill and its victim. Swap thrashing (high `si/so` in `vmstat 1`) feels like a hang.
- **"Disk full"** — `df -h` for which fs, `du -x` to drill in; check `df -i` for inodes; check
  `lsof +L1` for deleted-but-open files and `journalctl --disk-usage` for a runaway journal.
- **"Service won't start"** — `systemctl status` + `journalctl -u`; common causes are a port already
  bound (`ss -ltnp`), a permission/SELinux/AppArmor denial (detect-platform flagged which), or a
  failed dependency (`systemctl list-dependencies <unit>`).
- **"Can't reach the network"** — walk L2→L3→DNS with §1's ladder; a host firewall
  (`nft list ruleset` / `iptables -S`) or a wrong default route (`ip route`) is often it.
- **"Load is high but CPU looks idle"** — uninterruptible I/O wait; `iostat -xz 1` and look for a
  saturated device or an NFS/stuck mount (`mount`, and processes in D state in `ps`).

## 5. Before you change anything

Killing processes, restarting services, editing sysctl/limits/firewall, clearing logs, and
rebooting change the host's behaviour or availability — on a nanoinfra deployment a `mutate.remote`
capability (approval interactively, a standing grant unattended). Name the host and the change.

- `kill`/`kill -9`: prefer a graceful `SIGTERM`; `-9` (SIGKILL) gives the process no chance to flush
  or clean up. Confirm the PID is what you think (`ps -p <pid> -o comm=`).
- `systemctl restart` interrupts service; drain/announce first for anything user-facing.
- `sysctl -w` and `ulimit` changes are runtime-only — persist in `/etc/sysctl.d/` or
  `/etc/security/limits.d/` or a reboot loses them.
- Deleting under `/var/log` frees space now, but truncate a file a process holds open
  (`: > file`) rather than `rm` it, or the space is not returned until the holder restarts.
- Firewall edits can lock you out of an SSH session — have out-of-band console access before you
  change rules on a remote host.