Skill

redis

redis · current version v2

Download v2

Troubleshoot and support a Redis server — connect, diagnose, and resolve problems across 6.2/7.0/7.2/7.4/8.0, on Docker, bare metal, systemd, a managed service, or a Kubernetes operator. Covers redis-cli arguments, INFO/CONFIG, memory and eviction, persistence (RDB/AOF), replication, latency, and slowlog. Use when someone reports Redis is down, slow, unreachable, out of memory, evicting keys, failing to save, rejecting connections, or lagging a replica.

15 downloads · published 2026-09-02

What this grants

Skill Card

Security Audits

Version history

VersionPublishedStatus
v2 2026-09-02 published
v1 2026-09-02 published

Files

SKILL.md

raw | preview

Redis — Troubleshooting & Support

A support runbook, not a tutorial. Work top to bottom: identify what you are on, read before you write, and treat every mutating step as one that needs a reason. Every command and its output below was run against Redis 8.10.1; where a field or default differs by version it is called out. Confirm lifecycle claims against https://redis.io/docs/latest/ — knowledge here is current to early 2026, where the newest major is 8.x and there is no Redis 9.

0. Identify what you are actually on

redis-cli INFO server | grep -E 'redis_version|redis_mode|os|run_id|config_file|tcp_port'
# redis_version:8.10.1
# redis_mode:standalone           <- standalone | sentinel | cluster
redis-cli --version               # the client's version, which need not match the server

Establish the host first with the detect-platform skill — OS/package family, init system, bare metal vs VM vs container vs pod, the cgroup memory limit, and the security module. This skill assumes you have that profile: it decides the unit names below, whether the bare-metal host checks apply, and the memory ceiling to size against. (An endpoint you connect to but cannot get a host shell on is a managed service — §2.)

Standalone, replicated, Sentinel, or Cluster changes the whole picture:

redis-cli INFO replication | grep -E 'role|connected_slaves|master_link_status'
redis-cli INFO cluster | grep cluster_enabled      # cluster_enabled:1 on a cluster

1. Versions at a glance (why the version matters for support)

| Major | GA | Support-relevant fact | |---|---|---| | 6.2 | 2021 | Base of many long-lived installs; ACLs present since 6.0 | | 7.0 | 2022 | Functions, sharded pub/sub, ACL selectors | | 7.2 | 2023 | Last on the original BSD license | | 7.4 | 2024 | License changed to dual RSALv2 / SSPL; hash-field TTLs | | 8.0 | 2025 | AGPLv3 option; the query, JSON, time-series and probabilistic (bloom) modules ship in core |

Two facts that shape a ticket:

2. Getting a shell / connecting, per deployment

Self-managed — Docker

docker ps --filter ancestor=redis --format '{{.Names}}\t{{.Image}}\t{{.Status}}'
docker exec -it <container> redis-cli
docker logs --tail 200 -f <container>        # startup warnings + "Ready to accept connections"

Self-managed — systemd (bare metal or VM)

systemctl status redis-server        # unit is redis-server on Debian/Ubuntu, redis on RHEL
journalctl -u redis-server -n 200 --no-pager
grep -E '^(bind|port|requirepass|maxmemory|maxmemory-policy|appendonly|save|dir|tcp-backlog)' /etc/redis/redis.conf
redis-cli                            # local socket/loopback

Bare metal — the host checks nothing else needs

Redis prints its host problems as warnings in the first log lines. Read those before anything inside the server — they were the two that fired on a stock start here:

# WARNING Memory overcommit must be enabled! Without it, a background save or replication may
#         fail under low memory condition. … add 'vm.overcommit_memory = 1' to /etc/sysctl.conf
# WARNING: Redis does not require authentication and is not protected by network restrictions.

The host settings that cause the classic bare-metal failures:

# Memory overcommit — BGSAVE/AOF-rewrite and replication fork; without this the fork can fail.
cat /proc/sys/vm/overcommit_memory          # want 1
# Transparent Huge Pages — enabled THP causes latency spikes and fork stalls; disable it.
cat /sys/kernel/mm/transparent_hugepage/enabled   # want [never] (or madvise), not [always]
# Listen backlog — Redis asks for tcp-backlog (default 511); the kernel silently truncates it
# to somaxconn, so a low somaxconn drops connections under a burst.
sysctl net.core.somaxconn                    # raise to >= tcp-backlog
redis-cli CONFIG GET tcp-backlog             # 1) tcp-backlog 2) 511
# Open files vs maxclients — Redis lowers maxclients if the fd limit is low.
redis-cli INFO clients | grep maxclients     # maxclients:10000 here
ulimit -n
# Swappiness — swapping out the dataset destroys latency.
cat /proc/sys/vm/swappiness                  # 1 is the usual DB-host recommendation

Fixes are host config (sysctl, a THP tuned profile or boot flag, LimitNOFILE in the unit), not redis.conf. All deliberate host changes — see §6.

Managed — a cloud-hosted Redis service No host shell, no redis.conf, no docker logs; you connect to an endpoint and diagnose from the provider's console and metrics.

redis-cli -h <endpoint> -p <port> --tls -a <password> INFO

CONFIG SET may be restricted or disabled; persistence, maxmemory, and eviction are set in the provider's parameter group, and restarts/failovers/upgrades are console actions. Use INFO, SLOWLOG, and --latency (below) — those still work — plus the provider's dashboards.

Kubernetes — a Redis operator

kubectl get pods -l app=redis                          # match the operator's own label
kubectl exec -it <pod> -- redis-cli
kubectl logs <pod> --tail 200 -f
kubectl get secret <name> -o jsonpath='{.data.redis-password}' | base64 -d

Config is the operator's custom resource, not a file on the pod — edit the CR and let it reconcile; changes made inside the pod are reverted.

3. The tools and the arguments you will reach for

redis-cli — connection and one-shot

-h <host>  -p <port>  -n <db>  -a <password>  --user <name>   (ACL, 6.0+)
--tls --cacert <f> --cert <f> --key <f>                        (TLS/mTLS)
--eval <script.lua> <numkeys> key… arg…                        (run a Lua script)

Passing -a on the command line warns it is visible in the process list; prefer REDISCLI_AUTH.

redis-cli — diagnostic modes (safe, read-only), all confirmed against 8.10.1:

redis-cli --scan --pattern 'user:*'    # iterate keys without blocking, unlike KEYS
redis-cli --bigkeys                     # sample the biggest key per type; e.g.:
#   Biggest string found "k1" has 2 bytes
#   1 lists with 5 items (50.00% of keys, avg size 5.00)
redis-cli --memkeys                     # like --bigkeys but by memory, not element count
redis-cli --latency                     # live min/avg/max ms + samples: "0.024 0.084 0.043 101"
redis-cli --latency-history             # the same, one line per interval
redis-cli --intrinsic-latency 5         # the host's own scheduling latency, Redis excluded
redis-cli --stat                        # one line/second of keys, mem, clients, ops

redis-server — the daemon (you mostly read its config, rarely launch it by hand)

redis-server /etc/redis/redis.conf
redis-server --port 6380 --maxmemory 512mb --maxmemory-policy allkeys-lru

Offline data-file checks

redis-check-rdb /data/dump.rdb          # validate an RDB snapshot
redis-check-aof --fix /data/appendonly.aof   # --fix truncates a corrupt tail: destructive, see §6

4. Diagnostics — read-only, run these first

INFO is the spine. Real fields from 8.10.1:

redis-cli INFO memory | grep -E 'used_memory_human|used_memory_rss_human|maxmemory_human|maxmemory_policy|mem_fragmentation_ratio|evicted_keys'
#   used_memory_human:1.35M   used_memory_rss_human:24.00M
#   maxmemory:0 (no limit)    maxmemory_policy:noeviction
#   mem_fragmentation_ratio:  rss/used; a high value under low load is often just a fresh process
redis-cli INFO stats | grep -E 'instantaneous_ops_per_sec|keyspace_hits|keyspace_misses|rejected_connections|expired_keys|evicted_keys'
redis-cli INFO clients | grep -E 'connected_clients|blocked_clients|maxclients'
redis-cli INFO persistence | grep -E 'rdb_last_bgsave_status|rdb_changes_since_last_save|aof_enabled|aof_last_write_status|aof_last_bgrewrite_status'
redis-cli INFO replication      # role, connected_slaves, master_link_status, master_repl_offset
redis-cli INFO keyspace         # db0:keys=2,expires=0,avg_ttl=0
redis-cli DBSIZE                # key count in the selected db

Slow commands and stalls:

redis-cli CONFIG GET slowlog-log-slower-than    # microseconds; default 10000 (=10ms)
redis-cli SLOWLOG GET 10                         # the last 10 slow entries (id, time, μs, argv)
redis-cli SLOWLOG RESET                          # clear it after you have read it
# LATENCY DOCTOR needs monitoring enabled first, or it refuses:
redis-cli CONFIG SET latency-monitor-threshold 100
redis-cli LATENCY DOCTOR                         # plain-English latency summary once events exist
redis-cli MEMORY DOCTOR                          # memory advice; says nothing on an near-empty instance

Who is connected, and who is running what:

redis-cli CLIENT LIST     # one line per client: addr, age, idle, flags, db, cmd=, user=, tot-mem…
redis-cli ACL WHOAMI      # the user this connection authenticated as (default if none)
redis-cli ACL LIST        # e.g. "user default on nopass sanitize-payload ~* &* +@all"

A MONITOR shows every command live but doubles server load — use it briefly, never leave it running on a busy instance.

5. Common problems → resolution

Cannot connect / auth

Out of memory / evicting keys / OOMKilled

Persistence: saves failing / AOF problems

redis-cli INFO persistence | grep -E 'rdb_last_bgsave_status|aof_last_write_status|aof_last_bgrewrite_status'

Replication: replica not syncing / lag

redis-cli INFO replication    # on the replica: master_link_status:up? on the master: connected_slaves

Latency spikes

redis-cli --intrinsic-latency 5     # first rule out the host: if this is high, it is not Redis
redis-cli --latency                 # then measure round-trip to this instance
redis-cli SLOWLOG GET 20            # a single O(N) command (KEYS, big SORT, large SMEMBERS) stalls all

Connections rejected

6. Before you run anything that writes

These change data, durability, availability, or config, and on a nanoinfra deployment they resolve to a mutate.remote capability — they ask for approval interactively and need a standing grant to run unattended. Name the instance and the change before you make it:

Capture state first — INFO, CONFIG GET *, SLOWLOG GET, a copy of the RDB/AOF — when the next step is a write. A read that tells you what is wrong is cheaper than a write that was the wrong fix.