Skill
redis
redis · current version v2
Troubleshoot and support a Redis server — connect, diagnose, and resolve problems across 6.2/7.0/7.2/7.4/8.0, on Docker, bare metal, systemd, a managed service, or a Kubernetes operator. Covers redis-cli arguments, INFO/CONFIG, memory and eviction, persistence (RDB/AOF), replication, latency, and slowlog. Use when someone reports Redis is down, slow, unreachable, out of memory, evicting keys, failing to save, rejecting connections, or lagging a replica.
15 downloads · published 2026-09-02
What this grants
- skill redis
Skill Card
- License or terms: check the skill's own repository for license details (redis on GitHub).
Security Audits
- NanoInfra Scanner PASS no issues found
- VirusTotal PASS no engines flagged this file (full report)
Version history
| Version | Published | Status |
|---|---|---|
| v2 | 2026-09-02 | published |
| v1 | 2026-09-02 | published |
Files
SKILL.md(15404 bytes)
SKILL.md
raw | preview
Redis — Troubleshooting & Support
A support runbook, not a tutorial. Work top to bottom: identify what you are on, read before you write, and treat every mutating step as one that needs a reason. Every command and its output below was run against Redis 8.10.1; where a field or default differs by version it is called out. Confirm lifecycle claims against https://redis.io/docs/latest/ — knowledge here is current to early 2026, where the newest major is 8.x and there is no Redis 9.
0. Identify what you are actually on
redis-cli INFO server | grep -E 'redis_version|redis_mode|os|run_id|config_file|tcp_port'
# redis_version:8.10.1
# redis_mode:standalone <- standalone | sentinel | cluster
redis-cli --version # the client's version, which need not match the server
Establish the host first with the detect-platform skill — OS/package family, init system, bare metal vs VM vs container vs pod, the cgroup memory limit, and the security module. This skill assumes you have that profile: it decides the unit names below, whether the bare-metal host checks apply, and the memory ceiling to size against. (An endpoint you connect to but cannot get a host shell on is a managed service — §2.)
Standalone, replicated, Sentinel, or Cluster changes the whole picture:
redis-cli INFO replication | grep -E 'role|connected_slaves|master_link_status'
redis-cli INFO cluster | grep cluster_enabled # cluster_enabled:1 on a cluster
1. Versions at a glance (why the version matters for support)
| Major | GA | Support-relevant fact | |---|---|---| | 6.2 | 2021 | Base of many long-lived installs; ACLs present since 6.0 | | 7.0 | 2022 | Functions, sharded pub/sub, ACL selectors | | 7.2 | 2023 | Last on the original BSD license | | 7.4 | 2024 | License changed to dual RSALv2 / SSPL; hash-field TTLs | | 8.0 | 2025 | AGPLv3 option; the query, JSON, time-series and probabilistic (bloom) modules ship in core |
Two facts that shape a ticket:
- Persistence and eviction are config, not defaults you can assume. Whether this instance
saves to disk, and what it does when it fills up, are
CONFIGvalues — read them (§3) before you reason about data loss or evictions. - The client and server versions are independent; a new
redis-clitalks to an old server and vice versa. Diagnose againstINFO server'sredis_version, notredis-cli --version.
2. Getting a shell / connecting, per deployment
Self-managed — Docker
docker ps --filter ancestor=redis --format '{{.Names}}\t{{.Image}}\t{{.Status}}'
docker exec -it <container> redis-cli
docker logs --tail 200 -f <container> # startup warnings + "Ready to accept connections"
Self-managed — systemd (bare metal or VM)
systemctl status redis-server # unit is redis-server on Debian/Ubuntu, redis on RHEL
journalctl -u redis-server -n 200 --no-pager
grep -E '^(bind|port|requirepass|maxmemory|maxmemory-policy|appendonly|save|dir|tcp-backlog)' /etc/redis/redis.conf
redis-cli # local socket/loopback
Bare metal — the host checks nothing else needs
Redis prints its host problems as warnings in the first log lines. Read those before anything inside the server — they were the two that fired on a stock start here:
# WARNING Memory overcommit must be enabled! Without it, a background save or replication may
# fail under low memory condition. … add 'vm.overcommit_memory = 1' to /etc/sysctl.conf
# WARNING: Redis does not require authentication and is not protected by network restrictions.
The host settings that cause the classic bare-metal failures:
# Memory overcommit — BGSAVE/AOF-rewrite and replication fork; without this the fork can fail.
cat /proc/sys/vm/overcommit_memory # want 1
# Transparent Huge Pages — enabled THP causes latency spikes and fork stalls; disable it.
cat /sys/kernel/mm/transparent_hugepage/enabled # want [never] (or madvise), not [always]
# Listen backlog — Redis asks for tcp-backlog (default 511); the kernel silently truncates it
# to somaxconn, so a low somaxconn drops connections under a burst.
sysctl net.core.somaxconn # raise to >= tcp-backlog
redis-cli CONFIG GET tcp-backlog # 1) tcp-backlog 2) 511
# Open files vs maxclients — Redis lowers maxclients if the fd limit is low.
redis-cli INFO clients | grep maxclients # maxclients:10000 here
ulimit -n
# Swappiness — swapping out the dataset destroys latency.
cat /proc/sys/vm/swappiness # 1 is the usual DB-host recommendation
Fixes are host config (sysctl, a THP tuned profile or boot flag, LimitNOFILE in the unit),
not redis.conf. All deliberate host changes — see §6.
Managed — a cloud-hosted Redis service
No host shell, no redis.conf, no docker logs; you connect to an endpoint and diagnose from
the provider's console and metrics.
redis-cli -h <endpoint> -p <port> --tls -a <password> INFO
CONFIG SET may be restricted or disabled; persistence, maxmemory, and eviction are set in
the provider's parameter group, and restarts/failovers/upgrades are console actions. Use
INFO, SLOWLOG, and --latency (below) — those still work — plus the provider's dashboards.
Kubernetes — a Redis operator
kubectl get pods -l app=redis # match the operator's own label
kubectl exec -it <pod> -- redis-cli
kubectl logs <pod> --tail 200 -f
kubectl get secret <name> -o jsonpath='{.data.redis-password}' | base64 -d
Config is the operator's custom resource, not a file on the pod — edit the CR and let it reconcile; changes made inside the pod are reverted.
3. The tools and the arguments you will reach for
redis-cli — connection and one-shot
-h <host> -p <port> -n <db> -a <password> --user <name> (ACL, 6.0+)
--tls --cacert <f> --cert <f> --key <f> (TLS/mTLS)
--eval <script.lua> <numkeys> key… arg… (run a Lua script)
Passing -a on the command line warns it is visible in the process list; prefer REDISCLI_AUTH.
redis-cli — diagnostic modes (safe, read-only), all confirmed against 8.10.1:
redis-cli --scan --pattern 'user:*' # iterate keys without blocking, unlike KEYS
redis-cli --bigkeys # sample the biggest key per type; e.g.:
# Biggest string found "k1" has 2 bytes
# 1 lists with 5 items (50.00% of keys, avg size 5.00)
redis-cli --memkeys # like --bigkeys but by memory, not element count
redis-cli --latency # live min/avg/max ms + samples: "0.024 0.084 0.043 101"
redis-cli --latency-history # the same, one line per interval
redis-cli --intrinsic-latency 5 # the host's own scheduling latency, Redis excluded
redis-cli --stat # one line/second of keys, mem, clients, ops
redis-server — the daemon (you mostly read its config, rarely launch it by hand)
redis-server /etc/redis/redis.conf
redis-server --port 6380 --maxmemory 512mb --maxmemory-policy allkeys-lru
Offline data-file checks
redis-check-rdb /data/dump.rdb # validate an RDB snapshot
redis-check-aof --fix /data/appendonly.aof # --fix truncates a corrupt tail: destructive, see §6
4. Diagnostics — read-only, run these first
INFO is the spine. Real fields from 8.10.1:
redis-cli INFO memory | grep -E 'used_memory_human|used_memory_rss_human|maxmemory_human|maxmemory_policy|mem_fragmentation_ratio|evicted_keys'
# used_memory_human:1.35M used_memory_rss_human:24.00M
# maxmemory:0 (no limit) maxmemory_policy:noeviction
# mem_fragmentation_ratio: rss/used; a high value under low load is often just a fresh process
redis-cli INFO stats | grep -E 'instantaneous_ops_per_sec|keyspace_hits|keyspace_misses|rejected_connections|expired_keys|evicted_keys'
redis-cli INFO clients | grep -E 'connected_clients|blocked_clients|maxclients'
redis-cli INFO persistence | grep -E 'rdb_last_bgsave_status|rdb_changes_since_last_save|aof_enabled|aof_last_write_status|aof_last_bgrewrite_status'
redis-cli INFO replication # role, connected_slaves, master_link_status, master_repl_offset
redis-cli INFO keyspace # db0:keys=2,expires=0,avg_ttl=0
redis-cli DBSIZE # key count in the selected db
Slow commands and stalls:
redis-cli CONFIG GET slowlog-log-slower-than # microseconds; default 10000 (=10ms)
redis-cli SLOWLOG GET 10 # the last 10 slow entries (id, time, μs, argv)
redis-cli SLOWLOG RESET # clear it after you have read it
# LATENCY DOCTOR needs monitoring enabled first, or it refuses:
redis-cli CONFIG SET latency-monitor-threshold 100
redis-cli LATENCY DOCTOR # plain-English latency summary once events exist
redis-cli MEMORY DOCTOR # memory advice; says nothing on an near-empty instance
Who is connected, and who is running what:
redis-cli CLIENT LIST # one line per client: addr, age, idle, flags, db, cmd=, user=, tot-mem…
redis-cli ACL WHOAMI # the user this connection authenticated as (default if none)
redis-cli ACL LIST # e.g. "user default on nopass sanitize-payload ~* &* +@all"
A MONITOR shows every command live but doubles server load — use it briefly, never leave
it running on a busy instance.
5. Common problems → resolution
Cannot connect / auth
NOAUTH Authentication required.— the server hasrequirepass/ACL; pass-a(orREDISCLI_AUTH) and, for a named ACL user,--user.ERR AUTH <password> called without any password configured for the default user.— the opposite: you sent a password to an instance that has none. Drop-a.Connection refused— checkbind(CONFIG GET bindreturned* -::*here, i.e. all interfaces; a value of127.0.0.1is unreachable from another host) and the port. Note the startup warning above: an instance bound to the world with norequirepassis exposed — the fix is auth + network restriction, not widening the bind further.- TLS: the server must be built/started with TLS and a port for it;
--tlson the client alone is not enough.
Out of memory / evicting keys / OOMKilled
maxmemory_policy:noeviction(the default) means writes start failing withOOM command not allowedoncemaxmemoryis reached, rather than evicting. If you want a cache, set a policy:
Make it durable inredis-cli CONFIG SET maxmemory 512mb redis-cli CONFIG SET maxmemory-policy allkeys-lru # or allkeys-lfu / volatile-* / noevictionredis.conftoo, or a restart loses it.- Rising
evicted_keys(INFO stats) = the instance is atmaxmemoryand shedding data; either the working set outgrew the limit or the policy is wrong for the access pattern. - In a container,
maxmemory:0(no limit) plus a container memory cap = the kernel OOM-kills Redis instead of Redis evicting. Setmaxmemorybelow the container limit, leaving headroom for RSS and a save-time fork. mem_fragmentation_ratiowell above ~1.5 under real load points at fragmentation; activedefrag (CONFIG SET activedefrag yes) can help, but confirm it is fragmentation and not a fresh, nearly empty process (it read 17.96 on an idle instance here, which is meaningless at 1.35M used).
Persistence: saves failing / AOF problems
redis-cli INFO persistence | grep -E 'rdb_last_bgsave_status|aof_last_write_status|aof_last_bgrewrite_status'
- Any status other than
ok— the usual cause is a failed background fork, which comes back tovm.overcommit_memory(§2 bare metal) or no disk space indir(CONFIG GET dir→/datahere). Free space or enable overcommit; a server that cannot fork cannot snapshot or replicate. appendonly nowith asaveschedule (3600 1 300 100 60 10000is the shipped default) means point-in-time RDB only — the window since the last save is at risk on a crash. If durability matters, enable AOF:CONFIG SET appendonly yes.rdb_changes_since_last_savelarge and growing while status isokjust means it has not hit a save trigger yet — force one withBGSAVEif you need a snapshot now.
Replication: replica not syncing / lag
redis-cli INFO replication # on the replica: master_link_status:up? on the master: connected_slaves
master_link_status:down— the replica cannot reach or authenticate to the master; check network,masterauth, and the master'srequirepass.- A replica stuck in full resync loops usually means the master cannot fork to produce the RDB
(overcommit/disk again) or the replication buffer is too small under write load
(
client-output-buffer-limit replica).
Latency spikes
redis-cli --intrinsic-latency 5 # first rule out the host: if this is high, it is not Redis
redis-cli --latency # then measure round-trip to this instance
redis-cli SLOWLOG GET 20 # a single O(N) command (KEYS, big SORT, large SMEMBERS) stalls all
KEYSon a large keyspace blocks the single thread — replace with--scan/SCAN.- THP and a save-time fork are the classic host causes; see §2 bare metal.
Connections rejected
rejected_connectionsclimbing (INFO stats) withconnected_clientsnearmaxclients— raisemaxclients, but checkulimit -nfirst, because Redis capsmaxclientsto the fd limit minus its own reserve. On a burst, also raisenet.core.somaxconnto matchtcp-backlog.
6. Before you run anything that writes
These change data, durability, availability, or config, and on a nanoinfra deployment they
resolve to a mutate.remote capability — they ask for approval interactively and need a standing
grant to run unattended. Name the instance and the change before you make it:
CONFIG SET …— live config change; not persisted unlessCONFIG REWRITE(or the file) follows, and some are disabled on a managed service.FLUSHALL/FLUSHDB— delete everything / the current db. Rarely the right fix; almost never.DEBUG …,MONITOR—MONITORdoubles load;DEBUG SLEEP/DEBUG SEGFAULTare foot-guns.redis-check-aof --fix,BGREWRITEAOF— rewrite the append log;--fixtruncates a corrupt tail.REPLICAOF/SLAVEOF,FAILOVER,CLUSTER FAILOVER— change who serves writes.SHUTDOWN [NOSAVE]— stops the server;NOSAVEdrops unsaved data on purpose.
Capture state first — INFO, CONFIG GET *, SLOWLOG GET, a copy of the RDB/AOF — when the
next step is a write. A read that tells you what is wrong is cheaper than a write that was the
wrong fix.