Skill
redis
redis · current version v2
Troubleshoot and support a Redis server — connect, diagnose, and resolve problems across 6.2/7.0/7.2/7.4/8.0, on Docker, bare metal, systemd, a managed service, or a Kubernetes operator. Covers redis-cli arguments, INFO/CONFIG, memory and eviction, persistence (RDB/AOF), replication, latency, and slowlog. Use when someone reports Redis is down, slow, unreachable, out of memory, evicting keys, failing to save, rejecting connections, or lagging a replica.
14 downloads · published 2026-09-02
What this grants
- skill redis
Skill Card
- License or terms: check the skill's own repository for license details (redis on GitHub).
Security Audits
- NanoInfra Scanner PASS no issues found
- VirusTotal PASS no engines flagged this file (full report)
Version history
| Version | Published | Status |
|---|---|---|
| v2 | 2026-09-02 | published |
| v1 | 2026-09-02 | published |
Files
SKILL.md(15404 bytes)
SKILL.md
raw | preview
---
name: redis
description: Troubleshoot and support a Redis server — connect, diagnose, and resolve problems across 6.2/7.0/7.2/7.4/8.0, on Docker, bare metal, systemd, a managed service, or a Kubernetes operator. Covers redis-cli arguments, INFO/CONFIG, memory and eviction, persistence (RDB/AOF), replication, latency, and slowlog. Use when someone reports Redis is down, slow, unreachable, out of memory, evicting keys, failing to save, rejecting connections, or lagging a replica.
---
# Redis — Troubleshooting & Support
A support runbook, not a tutorial. Work top to bottom: identify what you are on, read before
you write, and treat every mutating step as one that needs a reason. Every command and its
output below was run against Redis 8.10.1; where a field or default differs by version it is
called out. Confirm lifecycle claims against <https://redis.io/docs/latest/> — knowledge here
is current to early 2026, where the newest major is **8.x** and there is no Redis 9.
## 0. Identify what you are actually on
```bash
redis-cli INFO server | grep -E 'redis_version|redis_mode|os|run_id|config_file|tcp_port'
# redis_version:8.10.1
# redis_mode:standalone <- standalone | sentinel | cluster
redis-cli --version # the client's version, which need not match the server
```
Establish the host first with the **detect-platform** skill — OS/package family, init system,
bare metal vs VM vs container vs pod, the cgroup memory limit, and the security module. This
skill assumes you have that profile: it decides the unit names below, whether the bare-metal
host checks apply, and the memory ceiling to size against. (An endpoint you connect to but
cannot get a host shell on is a managed service — §2.)
Standalone, replicated, Sentinel, or Cluster changes the whole picture:
```bash
redis-cli INFO replication | grep -E 'role|connected_slaves|master_link_status'
redis-cli INFO cluster | grep cluster_enabled # cluster_enabled:1 on a cluster
```
## 1. Versions at a glance (why the version matters for support)
| Major | GA | Support-relevant fact |
|---|---|---|
| 6.2 | 2021 | Base of many long-lived installs; ACLs present since 6.0 |
| 7.0 | 2022 | Functions, sharded pub/sub, ACL selectors |
| 7.2 | 2023 | Last on the original BSD license |
| 7.4 | 2024 | License changed to dual RSALv2 / SSPL; hash-field TTLs |
| 8.0 | 2025 | AGPLv3 option; the query, JSON, time-series and probabilistic (bloom) modules ship in core |
Two facts that shape a ticket:
- **Persistence and eviction are config, not defaults you can assume.** Whether this instance
saves to disk, and what it does when it fills up, are `CONFIG` values — read them (§3) before
you reason about data loss or evictions.
- The client and server versions are independent; a new `redis-cli` talks to an old server and
vice versa. Diagnose against `INFO server`'s `redis_version`, not `redis-cli --version`.
## 2. Getting a shell / connecting, per deployment
**Self-managed — Docker**
```bash
docker ps --filter ancestor=redis --format '{{.Names}}\t{{.Image}}\t{{.Status}}'
docker exec -it <container> redis-cli
docker logs --tail 200 -f <container> # startup warnings + "Ready to accept connections"
```
**Self-managed — systemd (bare metal or VM)**
```bash
systemctl status redis-server # unit is redis-server on Debian/Ubuntu, redis on RHEL
journalctl -u redis-server -n 200 --no-pager
grep -E '^(bind|port|requirepass|maxmemory|maxmemory-policy|appendonly|save|dir|tcp-backlog)' /etc/redis/redis.conf
redis-cli # local socket/loopback
```
**Bare metal — the host checks nothing else needs**
Redis prints its host problems as warnings in the first log lines. Read those before anything
inside the server — they were the two that fired on a stock start here:
```
# WARNING Memory overcommit must be enabled! Without it, a background save or replication may
# fail under low memory condition. … add 'vm.overcommit_memory = 1' to /etc/sysctl.conf
# WARNING: Redis does not require authentication and is not protected by network restrictions.
```
The host settings that cause the classic bare-metal failures:
```bash
# Memory overcommit — BGSAVE/AOF-rewrite and replication fork; without this the fork can fail.
cat /proc/sys/vm/overcommit_memory # want 1
# Transparent Huge Pages — enabled THP causes latency spikes and fork stalls; disable it.
cat /sys/kernel/mm/transparent_hugepage/enabled # want [never] (or madvise), not [always]
# Listen backlog — Redis asks for tcp-backlog (default 511); the kernel silently truncates it
# to somaxconn, so a low somaxconn drops connections under a burst.
sysctl net.core.somaxconn # raise to >= tcp-backlog
redis-cli CONFIG GET tcp-backlog # 1) tcp-backlog 2) 511
# Open files vs maxclients — Redis lowers maxclients if the fd limit is low.
redis-cli INFO clients | grep maxclients # maxclients:10000 here
ulimit -n
# Swappiness — swapping out the dataset destroys latency.
cat /proc/sys/vm/swappiness # 1 is the usual DB-host recommendation
```
Fixes are host config (`sysctl`, a THP `tuned` profile or boot flag, `LimitNOFILE` in the unit),
not `redis.conf`. All deliberate host changes — see §6.
**Managed — a cloud-hosted Redis service**
No host shell, no `redis.conf`, no `docker logs`; you connect to an endpoint and diagnose from
the provider's console and metrics.
```bash
redis-cli -h <endpoint> -p <port> --tls -a <password> INFO
```
`CONFIG SET` may be restricted or disabled; persistence, `maxmemory`, and eviction are set in
the provider's parameter group, and restarts/failovers/upgrades are console actions. Use
`INFO`, `SLOWLOG`, and `--latency` (below) — those still work — plus the provider's dashboards.
**Kubernetes — a Redis operator**
```bash
kubectl get pods -l app=redis # match the operator's own label
kubectl exec -it <pod> -- redis-cli
kubectl logs <pod> --tail 200 -f
kubectl get secret <name> -o jsonpath='{.data.redis-password}' | base64 -d
```
Config is the operator's custom resource, not a file on the pod — edit the CR and let it
reconcile; changes made inside the pod are reverted.
## 3. The tools and the arguments you will reach for
**`redis-cli`** — connection and one-shot
```
-h <host> -p <port> -n <db> -a <password> --user <name> (ACL, 6.0+)
--tls --cacert <f> --cert <f> --key <f> (TLS/mTLS)
--eval <script.lua> <numkeys> key… arg… (run a Lua script)
```
Passing `-a` on the command line warns it is visible in the process list; prefer `REDISCLI_AUTH`.
**`redis-cli` — diagnostic modes (safe, read-only)**, all confirmed against 8.10.1:
```bash
redis-cli --scan --pattern 'user:*' # iterate keys without blocking, unlike KEYS
redis-cli --bigkeys # sample the biggest key per type; e.g.:
# Biggest string found "k1" has 2 bytes
# 1 lists with 5 items (50.00% of keys, avg size 5.00)
redis-cli --memkeys # like --bigkeys but by memory, not element count
redis-cli --latency # live min/avg/max ms + samples: "0.024 0.084 0.043 101"
redis-cli --latency-history # the same, one line per interval
redis-cli --intrinsic-latency 5 # the host's own scheduling latency, Redis excluded
redis-cli --stat # one line/second of keys, mem, clients, ops
```
**`redis-server`** — the daemon (you mostly read its config, rarely launch it by hand)
```
redis-server /etc/redis/redis.conf
redis-server --port 6380 --maxmemory 512mb --maxmemory-policy allkeys-lru
```
**Offline data-file checks**
```bash
redis-check-rdb /data/dump.rdb # validate an RDB snapshot
redis-check-aof --fix /data/appendonly.aof # --fix truncates a corrupt tail: destructive, see §6
```
## 4. Diagnostics — read-only, run these first
`INFO` is the spine. Real fields from 8.10.1:
```bash
redis-cli INFO memory | grep -E 'used_memory_human|used_memory_rss_human|maxmemory_human|maxmemory_policy|mem_fragmentation_ratio|evicted_keys'
# used_memory_human:1.35M used_memory_rss_human:24.00M
# maxmemory:0 (no limit) maxmemory_policy:noeviction
# mem_fragmentation_ratio: rss/used; a high value under low load is often just a fresh process
redis-cli INFO stats | grep -E 'instantaneous_ops_per_sec|keyspace_hits|keyspace_misses|rejected_connections|expired_keys|evicted_keys'
redis-cli INFO clients | grep -E 'connected_clients|blocked_clients|maxclients'
redis-cli INFO persistence | grep -E 'rdb_last_bgsave_status|rdb_changes_since_last_save|aof_enabled|aof_last_write_status|aof_last_bgrewrite_status'
redis-cli INFO replication # role, connected_slaves, master_link_status, master_repl_offset
redis-cli INFO keyspace # db0:keys=2,expires=0,avg_ttl=0
redis-cli DBSIZE # key count in the selected db
```
Slow commands and stalls:
```bash
redis-cli CONFIG GET slowlog-log-slower-than # microseconds; default 10000 (=10ms)
redis-cli SLOWLOG GET 10 # the last 10 slow entries (id, time, μs, argv)
redis-cli SLOWLOG RESET # clear it after you have read it
# LATENCY DOCTOR needs monitoring enabled first, or it refuses:
redis-cli CONFIG SET latency-monitor-threshold 100
redis-cli LATENCY DOCTOR # plain-English latency summary once events exist
redis-cli MEMORY DOCTOR # memory advice; says nothing on an near-empty instance
```
Who is connected, and who is running what:
```bash
redis-cli CLIENT LIST # one line per client: addr, age, idle, flags, db, cmd=, user=, tot-mem…
redis-cli ACL WHOAMI # the user this connection authenticated as (default if none)
redis-cli ACL LIST # e.g. "user default on nopass sanitize-payload ~* &* +@all"
```
A `MONITOR` shows every command live but **doubles server load** — use it briefly, never leave
it running on a busy instance.
## 5. Common problems → resolution
**Cannot connect / auth**
- `NOAUTH Authentication required.` — the server has `requirepass`/ACL; pass `-a` (or
`REDISCLI_AUTH`) and, for a named ACL user, `--user`.
- `ERR AUTH <password> called without any password configured for the default user.` — the
opposite: you sent a password to an instance that has none. Drop `-a`.
- `Connection refused` — check `bind` (`CONFIG GET bind` returned `* -::*` here, i.e. all
interfaces; a value of `127.0.0.1` is unreachable from another host) and the port. Note the
startup warning above: an instance bound to the world with no `requirepass` is exposed — the
fix is auth + network restriction, not widening the bind further.
- TLS: the server must be built/started with TLS and a port for it; `--tls` on the client alone
is not enough.
**Out of memory / evicting keys / OOMKilled**
- `maxmemory_policy:noeviction` (the default) means writes start failing with `OOM command not
allowed` once `maxmemory` is reached, rather than evicting. If you want a cache, set a policy:
```bash
redis-cli CONFIG SET maxmemory 512mb
redis-cli CONFIG SET maxmemory-policy allkeys-lru # or allkeys-lfu / volatile-* / noeviction
```
Make it durable in `redis.conf` too, or a restart loses it.
- Rising `evicted_keys` (INFO stats) = the instance is at `maxmemory` and shedding data; either
the working set outgrew the limit or the policy is wrong for the access pattern.
- In a **container**, `maxmemory:0` (no limit) plus a container memory cap = the kernel OOM-kills
Redis instead of Redis evicting. Set `maxmemory` below the container limit, leaving headroom for
RSS and a save-time fork.
- `mem_fragmentation_ratio` well above ~1.5 under real load points at fragmentation; activedefrag
(`CONFIG SET activedefrag yes`) can help, but confirm it is fragmentation and not a fresh, nearly
empty process (it read 17.96 on an idle instance here, which is meaningless at 1.35M used).
**Persistence: saves failing / AOF problems**
```bash
redis-cli INFO persistence | grep -E 'rdb_last_bgsave_status|aof_last_write_status|aof_last_bgrewrite_status'
```
- Any status other than `ok` — the usual cause is a failed background fork, which comes back to
`vm.overcommit_memory` (§2 bare metal) or no disk space in `dir` (`CONFIG GET dir` → `/data`
here). Free space or enable overcommit; a server that cannot fork cannot snapshot or replicate.
- `appendonly no` with a `save` schedule (`3600 1 300 100 60 10000` is the shipped default) means
point-in-time RDB only — the window since the last save is at risk on a crash. If durability
matters, enable AOF: `CONFIG SET appendonly yes`.
- `rdb_changes_since_last_save` large and growing while status is `ok` just means it has not hit a
save trigger yet — force one with `BGSAVE` if you need a snapshot now.
**Replication: replica not syncing / lag**
```bash
redis-cli INFO replication # on the replica: master_link_status:up? on the master: connected_slaves
```
- `master_link_status:down` — the replica cannot reach or authenticate to the master; check
network, `masterauth`, and the master's `requirepass`.
- A replica stuck in full resync loops usually means the master cannot fork to produce the RDB
(overcommit/disk again) or the replication buffer is too small under write load
(`client-output-buffer-limit replica`).
**Latency spikes**
```bash
redis-cli --intrinsic-latency 5 # first rule out the host: if this is high, it is not Redis
redis-cli --latency # then measure round-trip to this instance
redis-cli SLOWLOG GET 20 # a single O(N) command (KEYS, big SORT, large SMEMBERS) stalls all
```
- `KEYS` on a large keyspace blocks the single thread — replace with `--scan`/`SCAN`.
- THP and a save-time fork are the classic host causes; see §2 bare metal.
**Connections rejected**
- `rejected_connections` climbing (INFO stats) with `connected_clients` near `maxclients` — raise
`maxclients`, but check `ulimit -n` first, because Redis caps `maxclients` to the fd limit minus
its own reserve. On a burst, also raise `net.core.somaxconn` to match `tcp-backlog`.
## 6. Before you run anything that writes
These change data, durability, availability, or config, and on a nanoinfra deployment they
resolve to a `mutate.remote` capability — they ask for approval interactively and need a standing
grant to run unattended. Name the instance and the change before you make it:
- `CONFIG SET …` — live config change; not persisted unless `CONFIG REWRITE` (or the file) follows,
and some are disabled on a managed service.
- `FLUSHALL` / `FLUSHDB` — delete everything / the current db. Rarely the right fix; almost never.
- `DEBUG …`, `MONITOR` — `MONITOR` doubles load; `DEBUG SLEEP`/`DEBUG SEGFAULT` are foot-guns.
- `redis-check-aof --fix`, `BGREWRITEAOF` — rewrite the append log; `--fix` truncates a corrupt tail.
- `REPLICAOF` / `SLAVEOF`, `FAILOVER`, `CLUSTER FAILOVER` — change who serves writes.
- `SHUTDOWN [NOSAVE]` — stops the server; `NOSAVE` drops unsaved data on purpose.
Capture state first — `INFO`, `CONFIG GET *`, `SLOWLOG GET`, a copy of the RDB/AOF — when the
next step is a write. A read that tells you what is wrong is cheaper than a write that was the
wrong fix.