Skill

redis

redis · current version v2

Download v2

Troubleshoot and support a Redis server — connect, diagnose, and resolve problems across 6.2/7.0/7.2/7.4/8.0, on Docker, bare metal, systemd, a managed service, or a Kubernetes operator. Covers redis-cli arguments, INFO/CONFIG, memory and eviction, persistence (RDB/AOF), replication, latency, and slowlog. Use when someone reports Redis is down, slow, unreachable, out of memory, evicting keys, failing to save, rejecting connections, or lagging a replica.

14 downloads · published 2026-09-02

What this grants

Skill Card

Security Audits

Version history

VersionPublishedStatus
v2 2026-09-02 published
v1 2026-09-02 published

Files

SKILL.md

raw | preview

---
name: redis
description: Troubleshoot and support a Redis server — connect, diagnose, and resolve problems across 6.2/7.0/7.2/7.4/8.0, on Docker, bare metal, systemd, a managed service, or a Kubernetes operator. Covers redis-cli arguments, INFO/CONFIG, memory and eviction, persistence (RDB/AOF), replication, latency, and slowlog. Use when someone reports Redis is down, slow, unreachable, out of memory, evicting keys, failing to save, rejecting connections, or lagging a replica.
---

# Redis — Troubleshooting & Support

A support runbook, not a tutorial. Work top to bottom: identify what you are on, read before
you write, and treat every mutating step as one that needs a reason. Every command and its
output below was run against Redis 8.10.1; where a field or default differs by version it is
called out. Confirm lifecycle claims against <https://redis.io/docs/latest/> — knowledge here
is current to early 2026, where the newest major is **8.x** and there is no Redis 9.

## 0. Identify what you are actually on

```bash
redis-cli INFO server | grep -E 'redis_version|redis_mode|os|run_id|config_file|tcp_port'
# redis_version:8.10.1
# redis_mode:standalone           <- standalone | sentinel | cluster
redis-cli --version               # the client's version, which need not match the server
```

Establish the host first with the **detect-platform** skill — OS/package family, init system,
bare metal vs VM vs container vs pod, the cgroup memory limit, and the security module. This
skill assumes you have that profile: it decides the unit names below, whether the bare-metal
host checks apply, and the memory ceiling to size against. (An endpoint you connect to but
cannot get a host shell on is a managed service — §2.)

Standalone, replicated, Sentinel, or Cluster changes the whole picture:

```bash
redis-cli INFO replication | grep -E 'role|connected_slaves|master_link_status'
redis-cli INFO cluster | grep cluster_enabled      # cluster_enabled:1 on a cluster
```

## 1. Versions at a glance (why the version matters for support)

| Major | GA | Support-relevant fact |
|---|---|---|
| 6.2 | 2021 | Base of many long-lived installs; ACLs present since 6.0 |
| 7.0 | 2022 | Functions, sharded pub/sub, ACL selectors |
| 7.2 | 2023 | Last on the original BSD license |
| 7.4 | 2024 | License changed to dual RSALv2 / SSPL; hash-field TTLs |
| 8.0 | 2025 | AGPLv3 option; the query, JSON, time-series and probabilistic (bloom) modules ship in core |

Two facts that shape a ticket:

- **Persistence and eviction are config, not defaults you can assume.** Whether this instance
  saves to disk, and what it does when it fills up, are `CONFIG` values — read them (§3) before
  you reason about data loss or evictions.
- The client and server versions are independent; a new `redis-cli` talks to an old server and
  vice versa. Diagnose against `INFO server`'s `redis_version`, not `redis-cli --version`.

## 2. Getting a shell / connecting, per deployment

**Self-managed — Docker**
```bash
docker ps --filter ancestor=redis --format '{{.Names}}\t{{.Image}}\t{{.Status}}'
docker exec -it <container> redis-cli
docker logs --tail 200 -f <container>        # startup warnings + "Ready to accept connections"
```

**Self-managed — systemd (bare metal or VM)**
```bash
systemctl status redis-server        # unit is redis-server on Debian/Ubuntu, redis on RHEL
journalctl -u redis-server -n 200 --no-pager
grep -E '^(bind|port|requirepass|maxmemory|maxmemory-policy|appendonly|save|dir|tcp-backlog)' /etc/redis/redis.conf
redis-cli                            # local socket/loopback
```

**Bare metal — the host checks nothing else needs**

Redis prints its host problems as warnings in the first log lines. Read those before anything
inside the server — they were the two that fired on a stock start here:

```
# WARNING Memory overcommit must be enabled! Without it, a background save or replication may
#         fail under low memory condition. … add 'vm.overcommit_memory = 1' to /etc/sysctl.conf
# WARNING: Redis does not require authentication and is not protected by network restrictions.
```

The host settings that cause the classic bare-metal failures:

```bash
# Memory overcommit — BGSAVE/AOF-rewrite and replication fork; without this the fork can fail.
cat /proc/sys/vm/overcommit_memory          # want 1
# Transparent Huge Pages — enabled THP causes latency spikes and fork stalls; disable it.
cat /sys/kernel/mm/transparent_hugepage/enabled   # want [never] (or madvise), not [always]
# Listen backlog — Redis asks for tcp-backlog (default 511); the kernel silently truncates it
# to somaxconn, so a low somaxconn drops connections under a burst.
sysctl net.core.somaxconn                    # raise to >= tcp-backlog
redis-cli CONFIG GET tcp-backlog             # 1) tcp-backlog 2) 511
# Open files vs maxclients — Redis lowers maxclients if the fd limit is low.
redis-cli INFO clients | grep maxclients     # maxclients:10000 here
ulimit -n
# Swappiness — swapping out the dataset destroys latency.
cat /proc/sys/vm/swappiness                  # 1 is the usual DB-host recommendation
```

Fixes are host config (`sysctl`, a THP `tuned` profile or boot flag, `LimitNOFILE` in the unit),
not `redis.conf`. All deliberate host changes — see §6.

**Managed — a cloud-hosted Redis service**
No host shell, no `redis.conf`, no `docker logs`; you connect to an endpoint and diagnose from
the provider's console and metrics.
```bash
redis-cli -h <endpoint> -p <port> --tls -a <password> INFO
```
`CONFIG SET` may be restricted or disabled; persistence, `maxmemory`, and eviction are set in
the provider's parameter group, and restarts/failovers/upgrades are console actions. Use
`INFO`, `SLOWLOG`, and `--latency` (below) — those still work — plus the provider's dashboards.

**Kubernetes — a Redis operator**
```bash
kubectl get pods -l app=redis                          # match the operator's own label
kubectl exec -it <pod> -- redis-cli
kubectl logs <pod> --tail 200 -f
kubectl get secret <name> -o jsonpath='{.data.redis-password}' | base64 -d
```
Config is the operator's custom resource, not a file on the pod — edit the CR and let it
reconcile; changes made inside the pod are reverted.

## 3. The tools and the arguments you will reach for

**`redis-cli`** — connection and one-shot
```
-h <host>  -p <port>  -n <db>  -a <password>  --user <name>   (ACL, 6.0+)
--tls --cacert <f> --cert <f> --key <f>                        (TLS/mTLS)
--eval <script.lua> <numkeys> key… arg…                        (run a Lua script)
```
Passing `-a` on the command line warns it is visible in the process list; prefer `REDISCLI_AUTH`.

**`redis-cli` — diagnostic modes (safe, read-only)**, all confirmed against 8.10.1:
```bash
redis-cli --scan --pattern 'user:*'    # iterate keys without blocking, unlike KEYS
redis-cli --bigkeys                     # sample the biggest key per type; e.g.:
#   Biggest string found "k1" has 2 bytes
#   1 lists with 5 items (50.00% of keys, avg size 5.00)
redis-cli --memkeys                     # like --bigkeys but by memory, not element count
redis-cli --latency                     # live min/avg/max ms + samples: "0.024 0.084 0.043 101"
redis-cli --latency-history             # the same, one line per interval
redis-cli --intrinsic-latency 5         # the host's own scheduling latency, Redis excluded
redis-cli --stat                        # one line/second of keys, mem, clients, ops
```

**`redis-server`** — the daemon (you mostly read its config, rarely launch it by hand)
```
redis-server /etc/redis/redis.conf
redis-server --port 6380 --maxmemory 512mb --maxmemory-policy allkeys-lru
```

**Offline data-file checks**
```bash
redis-check-rdb /data/dump.rdb          # validate an RDB snapshot
redis-check-aof --fix /data/appendonly.aof   # --fix truncates a corrupt tail: destructive, see §6
```

## 4. Diagnostics — read-only, run these first

`INFO` is the spine. Real fields from 8.10.1:

```bash
redis-cli INFO memory | grep -E 'used_memory_human|used_memory_rss_human|maxmemory_human|maxmemory_policy|mem_fragmentation_ratio|evicted_keys'
#   used_memory_human:1.35M   used_memory_rss_human:24.00M
#   maxmemory:0 (no limit)    maxmemory_policy:noeviction
#   mem_fragmentation_ratio:  rss/used; a high value under low load is often just a fresh process
redis-cli INFO stats | grep -E 'instantaneous_ops_per_sec|keyspace_hits|keyspace_misses|rejected_connections|expired_keys|evicted_keys'
redis-cli INFO clients | grep -E 'connected_clients|blocked_clients|maxclients'
redis-cli INFO persistence | grep -E 'rdb_last_bgsave_status|rdb_changes_since_last_save|aof_enabled|aof_last_write_status|aof_last_bgrewrite_status'
redis-cli INFO replication      # role, connected_slaves, master_link_status, master_repl_offset
redis-cli INFO keyspace         # db0:keys=2,expires=0,avg_ttl=0
redis-cli DBSIZE                # key count in the selected db
```

Slow commands and stalls:
```bash
redis-cli CONFIG GET slowlog-log-slower-than    # microseconds; default 10000 (=10ms)
redis-cli SLOWLOG GET 10                         # the last 10 slow entries (id, time, μs, argv)
redis-cli SLOWLOG RESET                          # clear it after you have read it
# LATENCY DOCTOR needs monitoring enabled first, or it refuses:
redis-cli CONFIG SET latency-monitor-threshold 100
redis-cli LATENCY DOCTOR                         # plain-English latency summary once events exist
redis-cli MEMORY DOCTOR                          # memory advice; says nothing on an near-empty instance
```

Who is connected, and who is running what:
```bash
redis-cli CLIENT LIST     # one line per client: addr, age, idle, flags, db, cmd=, user=, tot-mem…
redis-cli ACL WHOAMI      # the user this connection authenticated as (default if none)
redis-cli ACL LIST        # e.g. "user default on nopass sanitize-payload ~* &* +@all"
```
A `MONITOR` shows every command live but **doubles server load** — use it briefly, never leave
it running on a busy instance.

## 5. Common problems → resolution

**Cannot connect / auth**
- `NOAUTH Authentication required.` — the server has `requirepass`/ACL; pass `-a` (or
  `REDISCLI_AUTH`) and, for a named ACL user, `--user`.
- `ERR AUTH <password> called without any password configured for the default user.` — the
  opposite: you sent a password to an instance that has none. Drop `-a`.
- `Connection refused` — check `bind` (`CONFIG GET bind` returned `* -::*` here, i.e. all
  interfaces; a value of `127.0.0.1` is unreachable from another host) and the port. Note the
  startup warning above: an instance bound to the world with no `requirepass` is exposed — the
  fix is auth + network restriction, not widening the bind further.
- TLS: the server must be built/started with TLS and a port for it; `--tls` on the client alone
  is not enough.

**Out of memory / evicting keys / OOMKilled**
- `maxmemory_policy:noeviction` (the default) means writes start failing with `OOM command not
  allowed` once `maxmemory` is reached, rather than evicting. If you want a cache, set a policy:
  ```bash
  redis-cli CONFIG SET maxmemory 512mb
  redis-cli CONFIG SET maxmemory-policy allkeys-lru   # or allkeys-lfu / volatile-* / noeviction
  ```
  Make it durable in `redis.conf` too, or a restart loses it.
- Rising `evicted_keys` (INFO stats) = the instance is at `maxmemory` and shedding data; either
  the working set outgrew the limit or the policy is wrong for the access pattern.
- In a **container**, `maxmemory:0` (no limit) plus a container memory cap = the kernel OOM-kills
  Redis instead of Redis evicting. Set `maxmemory` below the container limit, leaving headroom for
  RSS and a save-time fork.
- `mem_fragmentation_ratio` well above ~1.5 under real load points at fragmentation; activedefrag
  (`CONFIG SET activedefrag yes`) can help, but confirm it is fragmentation and not a fresh, nearly
  empty process (it read 17.96 on an idle instance here, which is meaningless at 1.35M used).

**Persistence: saves failing / AOF problems**
```bash
redis-cli INFO persistence | grep -E 'rdb_last_bgsave_status|aof_last_write_status|aof_last_bgrewrite_status'
```
- Any status other than `ok` — the usual cause is a failed background fork, which comes back to
  `vm.overcommit_memory` (§2 bare metal) or no disk space in `dir` (`CONFIG GET dir` → `/data`
  here). Free space or enable overcommit; a server that cannot fork cannot snapshot or replicate.
- `appendonly no` with a `save` schedule (`3600 1 300 100 60 10000` is the shipped default) means
  point-in-time RDB only — the window since the last save is at risk on a crash. If durability
  matters, enable AOF: `CONFIG SET appendonly yes`.
- `rdb_changes_since_last_save` large and growing while status is `ok` just means it has not hit a
  save trigger yet — force one with `BGSAVE` if you need a snapshot now.

**Replication: replica not syncing / lag**
```bash
redis-cli INFO replication    # on the replica: master_link_status:up? on the master: connected_slaves
```
- `master_link_status:down` — the replica cannot reach or authenticate to the master; check
  network, `masterauth`, and the master's `requirepass`.
- A replica stuck in full resync loops usually means the master cannot fork to produce the RDB
  (overcommit/disk again) or the replication buffer is too small under write load
  (`client-output-buffer-limit replica`).

**Latency spikes**
```bash
redis-cli --intrinsic-latency 5     # first rule out the host: if this is high, it is not Redis
redis-cli --latency                 # then measure round-trip to this instance
redis-cli SLOWLOG GET 20            # a single O(N) command (KEYS, big SORT, large SMEMBERS) stalls all
```
- `KEYS` on a large keyspace blocks the single thread — replace with `--scan`/`SCAN`.
- THP and a save-time fork are the classic host causes; see §2 bare metal.

**Connections rejected**
- `rejected_connections` climbing (INFO stats) with `connected_clients` near `maxclients` — raise
  `maxclients`, but check `ulimit -n` first, because Redis caps `maxclients` to the fd limit minus
  its own reserve. On a burst, also raise `net.core.somaxconn` to match `tcp-backlog`.

## 6. Before you run anything that writes

These change data, durability, availability, or config, and on a nanoinfra deployment they
resolve to a `mutate.remote` capability — they ask for approval interactively and need a standing
grant to run unattended. Name the instance and the change before you make it:

- `CONFIG SET …` — live config change; not persisted unless `CONFIG REWRITE` (or the file) follows,
  and some are disabled on a managed service.
- `FLUSHALL` / `FLUSHDB` — delete everything / the current db. Rarely the right fix; almost never.
- `DEBUG …`, `MONITOR` — `MONITOR` doubles load; `DEBUG SLEEP`/`DEBUG SEGFAULT` are foot-guns.
- `redis-check-aof --fix`, `BGREWRITEAOF` — rewrite the append log; `--fix` truncates a corrupt tail.
- `REPLICAOF` / `SLAVEOF`, `FAILOVER`, `CLUSTER FAILOVER` — change who serves writes.
- `SHUTDOWN [NOSAVE]` — stops the server; `NOSAVE` drops unsaved data on purpose.

Capture state first — `INFO`, `CONFIG GET *`, `SLOWLOG GET`, a copy of the RDB/AOF — when the
next step is a write. A read that tells you what is wrong is cheaper than a write that was the
wrong fix.