Software Engineering WikiSE Wiki

Redis

Inspect and fix a slow or full Redis: data structure costs, eviction, persistence, keyspace scans, locks, replication and cluster checks.

Reviewed MarkdownEdit

On this page

Cheatsheet#

TaskCommand
Connect with TLS and ACL userREDISCLI_AUTH="$REDIS_PASSWORD" redis-cli -h host.example.com --tls --cacert ca.pem --user app
Server summaryredis-cli info server (also memory, stats, replication, clients)
Slow commandsredis-cli slowlog get 10
Non-blocking key scanredis-cli --scan --pattern 'session:*' --count 1000
Biggest keys by element countredis-cli --bigkeys
Biggest keys by memoryredis-cli --memkeys
Memory of one keyredis-cli memory usage mykey
Keys per databaseredis-cli info keyspace
Round-trip latencyredis-cli --latency
Ops per second, liveredis-cli --stat
Cluster healthredis-cli --cluster check host.example.com:6379
Replication stateredis-cli info replication
Snapshot now (background)redis-cli bgsave
Empty one database, non-blockingredis-cli -n 3 flushdb async

keys *, flushall and monitor on production

KEYS walks the whole keyspace in one command and blocks every other client until it finishes. A synchronous FLUSHALL or FLUSHDB blocks the same way and deletes data irreversibly. MONITOR streams every command to your client; the Redis docs note a single monitor can cut throughput by more than half. Use --scan, flushdb async and the slowlog instead.

Pass passwords with REDISCLI_AUTH or --askpass, not -a, which leaves the password in shell history and the process list.

First checks for a slow instance#

redis-cli slowlog get 10                  # commands over slowlog-log-slower-than (default 10000 µs)
redis-cli --latency-history -i 5          # PING round trip, a new sample window every 5 s
redis-cli info stats | grep -E 'instantaneous_ops|keyspace_(hits|misses)|evicted_keys|expired_keys'
redis-cli info persistence | grep -E 'rdb_bgsave_in_progress|aof_rewrite_in_progress|latest_fork_usec'
redis-cli info memory | grep -E 'used_memory_human|maxmemory_human|mem_fragmentation_ratio'
redis-cli latency doctor                  # needs latency-monitor-threshold > 0

Latency in Redis almost always has one of four causes: a command whose cost grows with the size of a key, a fork for persistence, the host swapping, or the network. redis-cli --intrinsic-latency 30, run on the Redis host itself, measures the latency floor the kernel or hypervisor imposes. See Linux performance for host checks.

The model#

Redis executes commands one at a time on its main thread. Since Redis 6, I/O threads can handle socket reads and writes, but command execution remains single-threaded. That is why each command is atomic and latency is predictable, and also why one O(n) command against a key with a million elements stalls every other client.

The complexity of every command is documented. Check it before using a command on a large key: SMEMBERS, HGETALL, LRANGE 0 -1, ZRANGE 0 -1, SORT, SUNION and DEL on a large key are all proportional to size. UNLINK deletes in a background thread.

redis-cli info commandstats | sed 's/cmdstat_//' | sort -t= -k2 -rn | head   # most-called commands

Data structures and their costs#

TypeUseWatch out for
StringCounters, cached blobs, flagsValues up to 512 MB; large values block while transferred
HashObjects with individually addressed fieldsCompact (listpack) only while small; HGETALL is O(n)
ListQueues, recent-N listsPush and pop are O(1); LRANGE 0 -1 and LINDEX in the middle are O(n)
SetMembership, tagsSMEMBERS on a big set is O(n); use SSCAN
Sorted setLeaderboards, time indexes, rate limitsRanged reads with a LIMIT are cheap, unbounded ranges are not
StreamAppend-only event log with consumer groupsGrows until trimmed: XADD ... MAXLEN ~ 100000
Bitmap / HyperLogLogFlags per ID / approximate distinct countsHyperLogLog has a standard error of 0.81%
redis-cli set session:123 '{"u":7}' ex 3600 nx         # set with a 1 h TTL, only if absent
redis-cli hset user:7 name alice email alice@example.com
redis-cli lpush jobs '{"id":1}'
redis-cli brpop jobs 5                                  # block up to 5 s for an item
redis-cli zadd scores 42 alice
redis-cli zrange scores 0 9 rev withscores              # top 10; REV needs Redis 6.2+
redis-cli xadd events maxlen '~' 100000 '*' type deploy service my-app
redis-cli object encoding user:7    # listpack, hashtable, intset, skiplist, quicklist...

A hash, set or sorted set that crosses its *-max-listpack-entries or *-max-listpack-value threshold converts to a hashtable or skiplist and uses several times more memory. The thresholds are in redis.conf; Redis 7 renamed the ziplist settings to listpack.

TTL and eviction#

Expired keys are removed in two ways: lazily, when a command touches the key, and by an active cycle that samples keys with a TTL ten times a second. An expired key that is never touched can occupy memory for a while. Many keys expiring in the same second can make the active cycle run longer and add latency.

redis-cli ttl session:123          # -1 = no expiry, -2 = key does not exist
redis-cli expire session:123 600
redis-cli persist session:123      # remove the TTL
redis-cli config get maxmemory maxmemory-policy
redis-cli config set maxmemory-policy allkeys-lru   # runtime only; also update redis.conf or CONFIG REWRITE

When used memory exceeds maxmemory, writes trigger eviction according to maxmemory-policy. maxmemory 0 (the 64-bit default) means no limit, and the host OOM killer becomes the limit instead.

PolicyBehaviour
noevictionDefault. Commands that add data fail with OOM command not allowed; reads still work. Right for a queue or a store of record
allkeys-lruEvict least recently used keys. The usual choice for a cache
allkeys-lfuEvict least frequently used keys; better when a few keys are much hotter than the rest
allkeys-randomEvict at random; fits uniform access
volatile-lru / volatile-lfu / volatile-randomSame, but only among keys with a TTL
volatile-ttlEvict keys with the shortest remaining TTL
allkeys-lrm / volatile-lrmEvict least recently modified; reads do not refresh the key. Redis 8.6+

All volatile-* policies behave like noeviction when no keys have a TTL. The instance fills, nothing is evictable, and writes fail even though a policy is set. Check evicted_keys and expired_keys in info stats. See the eviction docs.

Persistence#

ModeWhat survives a crashCost
RDB snapshotData up to the last snapshotfork plus a full write, periodically
AOF, appendfsync everysecAll but about the last second of writesContinuous writes; periodic rewrite also forks
AOF, appendfsync alwaysEvery acknowledged writeMuch slower; needs fast fsync
RDB and AOFAOF is used on restartBoth costs
NeitherNothingPure cache
redis-cli config get save appendonly appendfsync
redis-cli bgsave
redis-cli info persistence | grep -E 'rdb_last_bgsave_status|aof_last_write_status|latest_fork_usec'
redis-cli bgrewriteaof
redis-cli --rdb /backup/dump.rdb      # pull an RDB snapshot to the local machine

BGSAVE and AOF rewrites fork the process. The fork copies page tables (about 48 MB for a 24 GB instance), which pauses the main thread, and pages written during the save are copied on write, so memory use can grow sharply on a write-heavy instance. Leave headroom below physical RAM for this; info memory reports buffer overhead as mem_not_counted_for_evict.

On Linux, set vm.overcommit_memory = 1 so the fork does not fail under memory pressure, and disable transparent huge pages, which make copy-on-write copy 2 MB pages and cause latency spikes after every fork:

cat /sys/kernel/mm/transparent_hugepage/enabled    # Redis docs recommend [never]
sysctl vm.overcommit_memory                        # expect 1

Keyspace inspection#

redis-cli --scan --pattern 'session:*' --count 1000 | head
redis-cli --bigkeys          # largest key per type by element count, uses SCAN
redis-cli --memkeys          # largest key per type by memory
redis-cli --keystats         # both, plus size distribution (recent redis-cli)
redis-cli memory usage session:123
redis-cli info keyspace
redis-cli dbsize

--scan, --bigkeys and --memkeys use SCAN, so they are safe on a live server. Add -i 0.01 to sleep between batches on a busy one. SCAN can return a key more than once if the keyspace changes during the scan, but it never misses a key present for the whole scan.

Locks and atomicity#

SET key token NX EX 30 acquires a lock only if nobody holds it, with an expiry so a crashed holder cannot keep it forever. Release must check ownership, or a holder whose lock already expired deletes the next holder’s lock. A Lua script runs atomically, so the check and delete happen as one step.

TOKEN=$(uuidgen)
redis-cli set lock:deploy "$TOKEN" nx ex 30      # OK = acquired, (nil) = held by someone else
-- release.lua: delete only if we still own the lock
if redis.call("get", KEYS[1]) == ARGV[1] then
  return redis.call("del", KEYS[1])
end
return 0
redis-cli --eval release.lua lock:deploy , "$TOKEN"   # keys before the comma, arguments after

A single-node lock disappears with the node, and replication is asynchronous, so a failover can hand the lock to a second client. Where two holders would cause real damage, use a coordination service with consensus (etcd, ZooKeeper) or a fencing token checked by the protected resource.

Replication and cluster#

redis-cli info replication                            # role, master_link_status, offsets, replicas
redis-cli --cluster check host.example.com:6379
redis-cli --cluster info host.example.com:6379
redis-cli -h host.example.com cluster info            # cluster_state:ok, slots assigned and failing
redis-cli -c -h host.example.com get user:7           # -c follows MOVED and ASK redirects

Replication is asynchronous: the primary acknowledges a write before replicas have it, and a failover can lose it. WAIT 1 100 blocks until at least one replica acknowledges or 100 ms pass, which narrows but does not close the window.

In cluster mode, keys map to 16384 hash slots. Multi-key commands, transactions and Lua scripts must touch keys in one slot. Hash tags force that: only the part inside {} is hashed, so {user:7}:profile and {user:7}:sessions share a slot. Overusing one tag puts all that data on one shard.

Streams and consumer groups#

A stream is an append-only log addressed by auto-generated IDs (<ms>-<seq>). Consumer groups let several workers share the load while each message is delivered to exactly one consumer until acknowledged, which makes streams a durable work queue rather than fire-and-forget pub/sub.

redis-cli xadd events '*' type deploy service my-app        # append; * = server-assigned ID
redis-cli xlen events
redis-cli xrange events - + count 5                         # oldest 5 (- and + are min/max IDs)
redis-cli xrevrange events + - count 5                      # newest 5

redis-cli xgroup create events workers '$' mkstream         # group reads only new messages; mkstream if absent
redis-cli xreadgroup group workers worker-1 count 10 block 5000 streams events '>'   # '>' = undelivered
redis-cli xack events workers 1700000000000-0              # acknowledge one processed message

Unacknowledged messages sit in the group’s Pending Entries List (PEL). A crashed worker leaves its messages there; another worker claims them once they are idle long enough.

redis-cli xpending events workers                          # summary: count, min/max ID, per-consumer
redis-cli xautoclaim events workers worker-2 60000 0       # claim messages idle > 60 s (Redis 6.2+)
redis-cli xinfo groups events                              # last-delivered ID, lag, pending per group
redis-cli xadd events maxlen '~' 100000 '*' k v            # cap length approximately on write; ~ is far cheaper

XADD ... MAXLEN ~ N trims lazily to roughly N entries; without a cap a stream grows until it fills memory. Track lag in XINFO GROUPS to see how far a group is behind the tail.

Transactions, pipelines and Lua#

Three ways to run several commands together, with different guarantees.

# MULTI/EXEC: queued commands run atomically, but there is no rollback and no reading a value mid-transaction
printf 'MULTI\nINCR a\nINCR b\nEXEC\n' | redis-cli

Pipelining sends many commands without waiting for each reply, cutting round trips; it is not atomic and other clients interleave. Use --pipe to load bulk data:

{ for i in $(seq 1 100000); do printf 'SET k:%d %d\n' "$i" "$i"; done; } | redis-cli --pipe

A Lua script runs atomically on the server: no other command interleaves, so read-modify-write is safe without WATCH. Keys the script touches must be passed in KEYS so Cluster can route it.

-- incr_with_cap.lua: increment KEYS[1], but never above ARGV[1]
local cur = tonumber(redis.call('get', KEYS[1]) or '0')
if cur >= tonumber(ARGV[1]) then return cur end
return redis.call('incr', KEYS[1])
redis-cli --eval incr_with_cap.lua counter , 100          # keys before comma, args after
redis-cli script load "$(cat incr_with_cap.lua)"          # returns a SHA; run later with EVALSHA

Keep scripts short: a long script blocks the whole server exactly like any other command. FUNCTION (Redis 7+) registers reusable server-side functions with the same atomicity.

Sentinel#

Sentinel provides automatic failover for a primary-replica setup without Cluster’s sharding. A quorum of Sentinel processes monitors the primary; when enough agree it is down, they elect a replica and promote it, and clients discover the new primary by asking Sentinel.

redis-cli -p 26379 sentinel masters                       # monitored primaries and their state
redis-cli -p 26379 sentinel master mymaster               # detail: flags, quorum, num-slaves
redis-cli -p 26379 sentinel replicas mymaster
redis-cli -p 26379 sentinel get-master-addr-by-name mymaster   # what clients should connect to now
redis-cli -p 26379 sentinel ckquorum mymaster             # can this Sentinel reach quorum to fail over?
redis-cli -p 26379 sentinel failover mymaster             # force a manual failover (for testing)

Run an odd number of Sentinels (three or five) across failure domains so a network partition cannot leave two halves each believing they have quorum. The client library must be Sentinel-aware; a client pointed straight at a host misses failovers. Sentinel and Cluster are alternatives, not layers: Cluster has its own failover built in.

Monitoring#

INFO is the primary source; it returns dozens of fields grouped into sections. Watch these over time rather than as absolutes.

redis-cli info stats     | grep -E 'total_commands|instantaneous_ops|keyspace_(hits|misses)|expired|evicted|rejected'
redis-cli info memory    | grep -E 'used_memory_human|used_memory_peak_human|maxmemory_human|mem_fragmentation_ratio|mem_clients'
redis-cli info clients   | grep -E 'connected_clients|blocked_clients|maxclients'
redis-cli info persistence | grep -E 'rdb_last_save_time|rdb_last_bgsave_status|aof_enabled|aof_last_bgrewrite_status'
redis-cli info replication | grep -E 'role|connected_slaves|master_repl_offset'
SignalFieldWhat it means
Cache effectivenesskeyspace_hits / keyspace_missesFalling hit ratio: TTLs too short or eviction pressure
Memory pressureevicted_keys risingAt maxmemory, evicting under the policy
Fragmentationmem_fragmentation_ratioWell above 1.0 wastes RAM; below 1.0 means swapping
Overloadrejected_connections, blocked_clientsConnection limit reached, or clients blocked on BLPOP etc.
Fork costlatest_fork_usecLong forks pause the main thread during saves

SLOWLOG records commands over slowlog-log-slower-than microseconds (measured excluding network I/O). MEMORY DOCTOR and LATENCY DOCTOR give plain-language diagnoses.

redis-cli slowlog get 10          # id, timestamp, microseconds, command, client
redis-cli slowlog reset
redis-cli memory doctor
redis-cli memory stats | head -30
redis-cli latency reset && redis-cli latency latest

Troubleshooting#

SymptomLikely causeCheck
Periodic latency spikesFork for BGSAVE / AOF rewrite, transparent huge pages, swappinglatest_fork_usec; THP setting; si/so in vmstat 1
One client stalls everyoneAn O(n) command on a big keyslowlog get; --bigkeys
OOM command not allowed when used memory > 'maxmemory'noeviction, or volatile-* with no TTL keysmaxmemory-policy; info keyspace expires count
Hit rate fallingEviction pressure or TTLs too shortevicted_keys, expired_keys
Memory far above dataset sizeFragmentationmem_fragmentation_ratio; activedefrag
MOVED / ASK errorsClient not in cluster moderedis-cli -c; client library cluster setting
CROSSSLOT Keys in request don't hash to the same slotMulti-key command across slotsUse hash tags
Replica far behind or resyncing repeatedlyNetwork, write burst, or replication backlog too smallmaster_link_status; repl-backlog-size; offsets in info replication
max number of clients reachedmaxclients or file descriptor limitinfo clients; LimitNOFILE in the unit
NOAUTH / WRONGPASSMissing or wrong credentials, wrong ACL userACL WHOAMI; ACL LOG
MISCONF ... stop-writes-on-bgsave-errorLast RDB save failed (disk full, permissions)rdb_last_bgsave_status; Redis log
Consumer group not drainingWorkers not acknowledging; messages stuck in the PELxpending; xautoclaim idle messages
Stream memory grows without boundNo MAXLEN cap on XADDxlen; add maxlen '~' N on writes
NOSCRIPT No matching scriptScript cache cleared (restart, SCRIPT FLUSH)Reload with SCRIPT LOAD or use EVAL
Failover did not happenSentinel quorum not reached, or client not Sentinel-awaresentinel ckquorum; client library config
BUSY Redis is busy running a scriptA long Lua script is blockingSCRIPT KILL (if no writes yet); shorten the script
LOADING Redis is loading the datasetInstance still reading RDB/AOF after restartWait; watch loading:1 in info persistence
redis-cli info clients | grep -E 'connected_clients|blocked_clients|maxclients'
redis-cli client list | head                       # one line per client: addr, age, idle, cmd
redis-cli client kill id 42                        # disconnect one client
redis-cli acl log 10                               # recent auth and permission failures

Oneliners#

# Key count by prefix, without blocking
redis-cli --scan --count 1000 | awk -F: '{print $1}' | sort | uniq -c | sort -rn | head

# Keys under a prefix with no TTL (possible leak); one TTL call per key, keep it bounded
redis-cli --scan --pattern 'cache:*' | head -10000 | while read -r k; do [ "$(redis-cli ttl "$k")" = -1 ] && echo "$k"; done | head

# Cache hit ratio
redis-cli info stats | tr -d '\r' | awk -F: '/^keyspace_(hits|misses)/ {a[$1]=$2} END {t=a["keyspace_hits"]+a["keyspace_misses"]; if (t) printf "%.2f%%\n", 100*a["keyspace_hits"]/t}'

# Delete keys matching a pattern in batches, reclaiming memory in the background
redis-cli --scan --pattern 'tmp:*' | xargs -r -L 500 redis-cli unlink

# Copy one key to another instance (db 0, 5000 ms timeout), keeping the source
redis-cli migrate target.example.com 6379 mykey 0 5000 copy replace

# Confirm a replica is healthy before promoting it
redis-cli -h replica.example.com info replication | grep -E 'role|master_link_status|master_last_io_seconds_ago'

# Benchmark GET/SET with pipelining (not against production)
redis-benchmark -h host.example.com -t get,set -n 100000 -P 16 -q

# Total keys across every database
redis-cli info keyspace | awk -F'[:,=]' '/^db/ {s+=$3} END {print s+0}'

# Memory used by every key under a prefix (bounded; one MEMORY USAGE per key)
redis-cli --scan --pattern 'cache:*' | head -5000 | while read -r k; do redis-cli memory usage "$k"; done | awk '{s+=$1} END {printf "%.1f MB\n", s/1048576}'

# Set a TTL on every key under a prefix that currently has none
redis-cli --scan --pattern 'session:*' | while read -r k; do [ "$(redis-cli ttl "$k")" = -1 ] && redis-cli expire "$k" 3600; done

# Top 10 commands by call count
redis-cli info commandstats | sed 's/cmdstat_//' | awk -F'[:,=]' '{print $3, $1}' | sort -rn | head

# Ops per second sampled once
redis-cli info stats | awk -F: '/instantaneous_ops_per_sec/ {print $2+0}'

# Watch memory fragmentation live
watch -n2 "redis-cli info memory | grep -E 'used_memory_human|mem_fragmentation_ratio'"

# Every client blocked on a blocking command
redis-cli client list | awk -F'[ =]' '{for(i=1;i<=NF;i++) if($i=="cmd" && $(i+1) ~ /b(lpop|rpop|zpopmin)/) print}'

# Persist the current config to redis.conf after a runtime CONFIG SET
redis-cli config rewrite

# Rename a dangerous command check (see if FLUSHALL is disabled)
redis-cli config get 'rename-command' 2>/dev/null; redis-cli command info flushall | head -2

# Slot a key maps to in Cluster mode
redis-cli -h host.example.com cluster keyslot user:7

# Which node owns a slot
redis-cli -h host.example.com cluster shards | head -20

# Approximate distinct count with HyperLogLog
redis-cli pfadd visitors user:1 user:2 user:3 && redis-cli pfcount visitors

Scripts#

A memory report that lists the largest keys by memory and the encoding of each, using SCAN so it never blocks the server. Cap the sample with the first argument on a huge keyspace.

#!/usr/bin/env bash
set -euo pipefail
limit=${1:-10000}
redis-cli --scan --count 1000 | head -n "$limit" |
while IFS= read -r key; do
  bytes=$(redis-cli memory usage "$key" 2>/dev/null || echo 0)
  enc=$(redis-cli object encoding "$key" 2>/dev/null || echo '-')
  printf '%d\t%s\t%s\n' "${bytes:-0}" "$enc" "$key"
done | sort -rn | head -20 |
awk -F'\t' '{printf "%10.1f KB  %-10s %s\n", $1/1024, $2, $3}'

A safe pattern delete that removes keys in bounded batches with UNLINK (background free) and sleeps between batches so it does not starve other clients. Destructive: it deletes matching keys.

#!/usr/bin/env bash
set -euo pipefail
pattern=${1:?usage: redis-purge.sh <pattern> [batch]}
batch=${2:-500}
total=0
redis-cli --scan --pattern "$pattern" --count "$batch" |
while mapfile -t -n "$batch" keys && ((${#keys[@]})); do
  redis-cli unlink "${keys[@]}" >/dev/null
  total=$((total + ${#keys[@]}))
  printf 'deleted %d\n' "$total"
  sleep 0.05
done

A pre-failover replica check: confirm a replica is connected, its link is up, and its lag is within a threshold before promoting it. Exits non-zero if any check fails.

#!/usr/bin/env bash
set -euo pipefail
host=${1:?usage: redis-replica-check.sh <replica-host> [max_lag_s]}
max_lag=${2:-5}
info=$(redis-cli -h "$host" info replication | tr -d '\r')

role=$(awk -F: '/^role:/ {print $2}' <<<"$info")
link=$(awk -F: '/^master_link_status:/ {print $2}' <<<"$info")
lag=$(awk -F: '/^master_last_io_seconds_ago:/ {print $2+0}' <<<"$info")

[[ $role == slave ]]      || { echo "not a replica (role=$role)"; exit 1; }
[[ $link == up ]]         || { echo "master link down"; exit 1; }
(( lag <= max_lag ))      || { echo "lag ${lag}s exceeds ${max_lag}s"; exit 1; }
echo "ok: replica healthy, last IO ${lag}s ago"

Further reading#