Production Deployment Guide¶
This guide covers deploying Moon in production, from a single Docker container to tuned multi-shard configurations with TLS, ACL, persistence, and monitoring.
Quick Start¶
Docker run (minimal)¶
This starts Moon on port 6379 with a single shard (the default), persistence directory at /data, and protected mode disabled (safe inside Docker networks). Pass --shards 0 to auto-detect from the CPU count.
Docker run (production)¶
docker run -d \
--name moon \
--restart unless-stopped \
-p 6379:6379 \
-p 9100:9100 \
-v moon-data:/data \
-v /etc/moon/certs:/certs:ro \
--ulimit nofile=65536:65536 \
--memory 8g \
moondb/moon:latest \
moon --bind 0.0.0.0 \
--port 6379 \
--admin-port 9100 \
--shards 0 \
--requirepass "$MOON_PASSWORD" \
--appendonly yes \
--appendfsync everysec \
--dir /data \
--maxmemory 6442450944 \
--maxmemory-policy allkeys-lfu \
--tls-port 6443 \
--tls-cert-file /certs/server.crt \
--tls-key-file /certs/server.key
Docker Compose¶
# Clone the repo (or just copy docker-compose.yml)
git clone https://github.com/pilotspace/moon.git
cd moon
# Start with defaults (AOF enabled, 4 CPU limit, 2 GB memory)
docker compose up -d
# Override settings via environment
MOON_SHARDS=8 MOON_MAXMEMORY=8589934592 docker compose up -d
# View logs
docker compose logs -f
# Stop
docker compose down
See docker-compose.yml for all configurable environment variables.
Build from source¶
docker build -t moon .
# Or for a specific runtime:
docker build --build-arg FEATURES=runtime-tokio,jemalloc -t moon:tokio .
# Multi-platform:
docker buildx build --platform linux/amd64,linux/arm64 -t moondb/moon:latest .
Configuration Reference¶
All options are command-line flags. Run moon --help for the full list.
Server¶
| Flag | Default | Description |
|---|---|---|
--bind |
127.0.0.1 |
Bind address. Use 0.0.0.0 in containers. |
--port / -p |
6379 |
Redis protocol port |
--admin-port |
0 (disabled) |
Admin/console HTTP port. Serves /metrics, /healthz, /readyz, and web UI at /ui/ |
--shards |
0 (auto) |
Number of shards. 0 = CPU core count. |
--databases |
16 |
Number of logical databases |
--maxclients |
10000 |
Maximum simultaneous client connections |
--timeout |
0 (disabled) |
Close idle connections after N seconds |
--tcp-keepalive |
300 |
TCP keepalive interval in seconds |
--protected-mode |
yes |
Reject non-loopback connections when no password set |
Persistence¶
| Flag | Default | Description |
|---|---|---|
--appendonly |
no |
Enable AOF persistence (yes/no) |
--appendfsync |
everysec |
AOF fsync policy: always, everysec, or no |
--appendfilename |
appendonly.aof |
AOF filename |
--save |
(none) | RDB auto-save rules (e.g., "3600 1 300 100") |
--dir |
. |
Directory for persistence files |
--dbfilename |
dump.rdb |
RDB snapshot filename |
Memory and Eviction¶
| Flag | Default | Description |
|---|---|---|
--maxmemory |
0 (unlimited) |
Maximum memory in bytes |
--maxmemory-policy |
noeviction |
Eviction policy when maxmemory is reached |
--maxmemory-samples |
5 |
Keys to sample per eviction cycle |
Eviction policies: noeviction, allkeys-lru, allkeys-lfu, allkeys-random, volatile-lru, volatile-lfu, volatile-random, volatile-ttl.
TLS¶
| Flag | Default | Description |
|---|---|---|
--tls-port |
0 (disabled) |
TLS listener port |
--tls-cert-file |
(none) | PEM certificate file |
--tls-key-file |
(none) | PEM private key file |
--tls-ca-cert-file |
(none) | CA cert for mTLS client verification |
--tls-ciphersuites |
(default) | TLS 1.3 cipher suites (comma-separated) |
ACL¶
| Flag | Default | Description |
|---|---|---|
--requirepass |
(none) | Require password for all clients |
--aclfile |
(none) | Path to ACL file (Redis-compatible format) |
--acllog-max-len |
128 |
Maximum ACL log entries |
Console/Admin Hardening¶
| Flag | Default | Description |
|---|---|---|
--console-auth-required |
false |
Require Bearer auth on admin API endpoints |
--console-auth-secret |
(auto-generated) | HMAC-SHA256 secret for token verification |
--console-cors-origin |
localhost:5173 |
CORS origin allowlist (repeatable) |
--console-rate-limit |
1000 |
Per-IP request rate limit (req/s) |
--console-rate-burst |
2000 |
Token-bucket burst capacity |
Performance¶
| Flag | Default | Description |
|---|---|---|
--uring-sqpoll |
(disabled) | io_uring SQPOLL idle timeout (ms). Requires CAP_SYS_NICE. |
--disk-offload |
enable |
Tiered storage: RAM to NVMe |
--slowlog-log-slower-than |
10000 |
Slowlog threshold in microseconds |
--slowlog-max-len |
128 |
Maximum slowlog entries |
Cluster¶
| Flag | Default | Description |
|---|---|---|
--cluster-enabled |
false |
Enable cluster mode |
--cluster-node-timeout |
15000 |
Node timeout in ms |
TLS Setup¶
Moon uses rustls with aws-lc-rs for TLS 1.3. No OpenSSL dependency.
Generate certificates (self-signed, for testing)¶
# Generate CA
openssl req -x509 -newkey ec -pkeyopt ec_paramgen_curve:prime256v1 \
-keyout ca.key -out ca.crt -days 365 -nodes -subj "/CN=Moon CA"
# Generate server cert
openssl req -newkey ec -pkeyopt ec_paramgen_curve:prime256v1 \
-keyout server.key -out server.csr -nodes -subj "/CN=moon"
openssl x509 -req -in server.csr -CA ca.crt -CAkey ca.key \
-CAcreateserial -out server.crt -days 365
# Clean up
rm server.csr ca.srl
Start with TLS¶
docker run -d \
-p 6443:6443 \
-v moon-data:/data \
-v $(pwd)/certs:/certs:ro \
moondb/moon:latest \
moon --bind 0.0.0.0 \
--tls-port 6443 \
--tls-cert-file /certs/server.crt \
--tls-key-file /certs/server.key \
--dir /data
Connect with TLS¶
mTLS (mutual TLS)¶
Add --tls-ca-cert-file /certs/ca.crt to require client certificates:
moon --bind 0.0.0.0 \
--tls-port 6443 \
--tls-cert-file /certs/server.crt \
--tls-key-file /certs/server.key \
--tls-ca-cert-file /certs/ca.crt
Disable plaintext port¶
Set --port 0 to accept only TLS connections:
moon --bind 0.0.0.0 --port 0 --tls-port 6443 \
--tls-cert-file /certs/server.crt \
--tls-key-file /certs/server.key
ACL Configuration¶
Moon supports Redis-compatible ACL files for fine-grained access control.
Simple password authentication¶
ACL file¶
Create users.acl:
# Default user (backward-compatible with --requirepass)
user default on >password ~* &* +@all
# Read-only user
user reader on >reader-pass ~* &* +@read -@write -@admin
# Application user with key restrictions
user app on >app-secret ~app:* ~cache:* &* +@all -@admin -@dangerous
# Admin user
user admin on >admin-secret ~* &* +@all
Start with the ACL file:
In Docker:
docker run -d \
-p 6379:6379 \
-v moon-data:/data \
-v $(pwd)/users.acl:/etc/moon/users.acl:ro \
moondb/moon:latest \
moon --bind 0.0.0.0 --dir /data --aclfile /etc/moon/users.acl
ACL categories¶
Moon supports Redis ACL categories: @all, @read, @write, @admin, @dangerous, @fast, @slow, @string, @hash, @list, @set, @sortedset, @stream, @pubsub, @scripting, @connection, @server, @generic, @keyspace, @hyperloglog, @bitmap, @geo.
Runtime ACL management¶
AUTH username password
ACL LIST
ACL SETUSER myuser on >pass ~key:* +get +set
ACL DELUSER myuser
ACL SAVE
Persistence Tuning¶
Moon provides per-shard WAL (Write-Ahead Log) for AOF persistence and forkless RDB snapshots.
AOF (recommended for production)¶
appendfsync |
Durability | Performance Impact |
|---|---|---|
always |
RPO = 0 (zero data loss) | fsync-per-batch under pipelining; P1 is fsync-device-bound (parity with any engine), P16 ~0.91x Redis |
everysec |
RPO <= 1 second | Recommended default; P16 SET ~1.32x Redis, P1 at parity |
no |
OS flush window (minutes) | Cache-mode only; do not use for primary storage |
Moon's per-shard WAL avoids the global serialization bottleneck that Redis's single AOF file creates. The AOF advantage over Redis grows with pipeline depth; the Linux measurement is the everysec P16 1.32x in the table above (BENCHMARK.md §7.3). A "2.75x at p=64" figure also circulates — it is an Apple M4 Pro development reference (§7.1), never reproduced on Linux.
Write-path internals (how the AOF writer keeps up). These are automatic — no tuning knobs — but understanding them explains the durability/throughput tradeoff:
- Group commit. Under
appendfsync always, pipelined writes are enqueued fire-and-forget and the whole batch is made durable by ONEfsync_barrier(not one fsync per command). A 16-deep pipeline therefore pays ~1 fsync, not 16. The fsync still runs strictly before the client's+OK(RPO = 0 is preserved); a failed batch returns an error to every write in it, never a silent success. - Coalesced batch write. Each group-commit batch is written to the AOF with a
single contiguous
write_all, not onewrite(2)per record — so the writer thread is not syscall-bound at high pipeline depth (this is what makeseverysecP16 beat Redis rather than trail it). - Writer poll: warm, then parked (
everysec/no). While writes flow, the AOF writer thread polls its channel every 500 µs instead of parking in a blocking receive: a parked receiver forces every shard thread to issue a futex wake on each write — at non-pipelinedeverysecload that was ~150k wakes/sec of pure overhead on the hot path. After 5 ms with nothing queued it parks, so the first write after an idle period wakes it at once (one futex wake) instead of waiting out a poll step (moon#1266: that step was up to 50 ms). Underalwaysthe writer always parks (the client is already blocked on the fsync ack, so receive latency there is client-visible RTT anyway). - The
everysecfsync runs on an agent thread (moon#1266). Each AOF writer has anaof-fsync-<n>thread; once a second the writer hands it the fsync and goes straight back to writing, so a slow disk no longer stops the writer from draining acknowledged writes into the file. At most one fsync is in flight per writer; a deadline that finds the previous one still running is postponed until it returns.alwayskeeps its fsync on the writer, before the acks. CONFIG SET appendfsyncapplies at once, as in redis. Every producer and AOF writer uses the new policy from its next write. Leavingeverysecfirst waits for an fsync still running on the agent (redis drains its background fsync the same way); a write already queued to be acknowledged after its fsync is fsynced before its ack whatever the switch. (Before the R1 review fix the command answeredOKandCONFIG GETshowed the new value while the writers kept their startup policy.)- A slow or hung
everysecfsync is loud, as in redis. While a writer's fsync has been running for 2 s or more, moon logs redis'sAsynchronous AOF fsync is taking too long (disk is busy?)(at most once every 2 s), and a WARN when the fsync finally returns.INFO persistenceshows it while it lasts: aof_pending_bio_fsync— writers with an fsync in flight. redis's field counts pendingBIO_AOF_FSYNCjobs and is 0 or 1 in practice; moon's is 0 or 1 with one writer (--shards 1) but counts WRITERS, so it reaches N at--shards N(per-shard writers, one agent each). Alert on> 0, never on== 1;aof_fsync_in_flight_ms— how long the oldest of them has been running (0 when none; moon-only);aof_delayed_fsync— counted as redis counts it: once for every 2 s an fsync stays in flight while written data waits for the next one, so an alert built for redis (rate(aof_delayed_fsync) > 0) fires the same way. Writes keep being acknowledged and reach the kernel meanwhile (a process crash loses nothing extra), but nothing written since the last completed fsync survives an OS crash or power loss until the fsync returns.
None of these weaken durability: always remains RPO = 0 (an ack is sent only
after its fsync — tests/aof_everysec_kill9_1266.rs holds the fsync open and
sees no reply until it returns), everysec remains RPO ≤ 1 s against an OS crash
or power loss. See BENCHMARK.md §7.3 for the measured before/after matrix.
What a process crash (kill -9, OOM kill, panic) can lose under everysec.
A SIGKILL does not touch the kernel page cache, so a record survives it once the
AOF writer has write(2)-n it; only an OS crash or power loss needs the fsync.
moon acknowledges a write when its record is queued to the shard's writer, so
the exposure to a process crash is the time from the ack to that write(2):
one poll step (500 µs) while writes flow, one thread wake-up after an idle
period, plus any time the writer thread is not scheduled or its write(2)
blocks. tests/aof_everysec_kill9_1266.rs measures it: 10,000 acked SETs
(unpipelined, or pipelined 100 deep) or one SET after an idle second, SIGKILL
1 ms after the last ack, restart, count what is missing. On a 4-vCPU Linux
container shared with other builds (2026-09-30, 20 reps per cell, --shards
1 and 4): before moon#1266 Option 3, 226 of 240 reps lost acked writes (median
rep 1–1,100 keys, worst 10,000); after it, 9 of 240 reps did in the run of
the final binaries (monoio 5 of 120: 3, 18, 400, 546 and 1,100 keys; tokio 4 of
120: 1, 1, 1 and 800). A second 20-rep tokio run lost in 7 of its 120 reps (up
to 40 keys), so across both tokio runs the total is 16 of 360 reps. Every lossy
rep was a writer stalled or descheduled for longer than the 1 ms kill delay. A kill inside that sub-millisecond window,
or while the writer thread is starved of CPU or its write(2) stalls, can
still lose the last acknowledged writes. redis
has no such window: it write(2)s its AOF buffer before it sends the replies
of an event-loop iteration (moon#1266 option 1A is the measured follow-up).
redis's own everysec is not absolutely kill-9-safe either. When the
previous background fsync is still running, redis postpones the write of its
AOF buffer (the write(2), not only the fsync) for up to 2 s, because on Linux
a write(2) to a file whose fsync is in progress would block behind it
anyway. Replies are still sent during those 2 s; past them redis writes anyway
and counts aof_delayed_fsync. So on a slow disk redis can lose up to ~2 s of
acknowledged writes to a process crash (0 on a healthy disk). moon never
postpones the write — only the fsync; its aof_delayed_fsync counts the same
2 s periods —
so a slow fsync by itself opens no window; a write(2) that the kernel makes
wait behind that fsync still delays the record, and that wait is part of the
window above.
RDB snapshots¶
# Snapshot every 3600 seconds if at least 1 key changed,
# or every 300 seconds if at least 100 keys changed
moon --save "3600 1 300 100" --dir /data --dbfilename dump.rdb
Moon uses forkless RDB snapshots -- no fork(), no COW memory spike.
Combined AOF + RDB¶
Recovery order: RDB snapshot, then WAL segments, then AOF tail.
Graceful shutdown and stop timeouts¶
A graceful stop — SIGTERM (systemctl stop, docker stop, a Kubernetes pod
deletion), SIGINT or SHUTDOWN — does two things before the process exits:
- The final RDB save, when save points are configured (
--save, orCONFIG SET save): aSIGTERM/SIGINTwaits for a save already running, then saves once more. There is no deadline on it, because a large dataset must not turn a stop into "keep running". Every 20 s without progress it logsErrors trying to shut down the serverand that the stop stays armed. A failed save keeps the server running (redis parity); a secondSIGINTexits at once without the save (status 1), andSHUTDOWN ABORTcancels. - The AOF drain (
--appendonly yes): after the shards stop, every AOF writer writes what is still queued and fsyncs. The wait is bounded at 60 s (plus a 2 s grace); after that the process exits non-zero and logs the writers it abandoned. A secondSIGINTcuts the wait to about 2 s.
The supervisor's stop timeout must cover both, or it SIGKILLs the save or the
drain half-way: the unsaved writes are then lost, as on a crash. Size it to the
final save of your dataset (roughly the time a BGSAVE takes: time one, from
BGSAVE until INFO persistence shows rdb_bgsave_in_progress:0) plus
60 s:
| Supervisor | Default | Setting |
|---|---|---|
| systemd | TimeoutStopSec=90s |
TimeoutStopSec= in the [Service] section (packaging/moon.service) |
| Docker | 10 s | docker stop -t <s>, docker run --stop-timeout <s>, Compose stop_grace_period |
| Kubernetes | 30 s | terminationGracePeriodSeconds in the pod spec |
For example, a dataset whose BGSAVE takes 2 minutes needs
TimeoutStopSec=180 (120 s + 60 s). With --appendonly yes and no save
points, the 60 s drain bound alone is the minimum.
Switching --appendonly yes → no¶
Take a snapshot first (BGSAVE, or SHUTDOWN SAVE). Under --appendonly no
boot loads only the RDB snapshot and ignores an AOF on disk (redis parity), so a
dataset that lived only in the AOF boots empty. moon logs a WARN naming the AOF
it did not load. See configuration.
AOF rewrite¶
Trigger manual compaction:
Key expiry during AOF replay¶
A restart after a key's TTL passed must not revive it, yet the writes logged while it was alive must replay onto it. moon judges expiry during replay by the time the log was last written: the newest modification time (mtime) of the AOF files being replayed (and of the WAL segments), capped at the wall clock (moon#1277). Keep those mtimes truthful:
- An mtime LATER than the last write (a file copied without preserving it) only brings back the pre-moon#1277 behaviour for keys whose TTL passed during the downtime.
- An mtime EARLIER than the last write — the system clock stepped back after
the write, a network or virtio filesystem whose server clock lags, or
touch -d/ a restore tool that resets it — makes keys that expired while the server ran look alive to the replay. A key that was read as expired, then rewritten before its expiryDELreached the log, then replays onto its OLD value (measured with the mtime set an hour back: 27–36 of 40 such keys, against 0 on a build that judged by the wall clock).
Since moon#1283 the AOF carries that time itself: the writer stamps the log
with MOON.TS <ms> records (the shard clock each record was judged under),
and a replay judges every stamped record by its own stamp, whatever the
file's mtime. The mtime still judges a log written before moon#1283 (and the
stamp-less start of one an older binary began), so keep restoring AOF files
with their original mtimes (cp -p, rsync -t, tar).
Downgrading and upgrading again. An older binary appends to the same AOF
without stamps. moon recognises those records by position — each orderly
stop of the newer binary ends the file with a clean-close marker
(MOON.TS <ms> CLOSE), and its first write after a restart is stamped — and
judges them as the older binary did, on every later boot
(STORAGE-FORMAT-V1.md §3.3). That works only if the
newer binary was stopped cleanly:
- Stop the newer binary with
SHUTDOWNor SIGTERM and check its log forAOF writers drained and syncedbefore starting the older one. A stop that logsWITHOUT their clean-close marker(a writer with a latched write error, a failed append or a panic) or an abandoned-writer error instead did not protect the downgrade: fix the cause, then start and stop the newer binary cleanly again. If it crashed (kill -9, OOM kill, power loss), start it once and stop it cleanly first. - If that was not done, run
BGREWRITEAOFon the older binary as its last action, and stop it once the rewrite completed, before upgrading again.
Otherwise keys the older binary saw expire and restarted (an INCR on an
expired counter) are judged by the newer binary's last stamp, which is stale,
and replay onto their old value and deadline: they are lost.
Persistence volume in Docker¶
Always mount a named volume or host directory for /data:
For host-path mounts (better for backup tooling):
Memory Management¶
Setting maxmemory¶
Reserve 20-25% of available RAM for OS, jemalloc overhead, and fragmentation:
# On a 32 GB host, allocate 24 GB to Moon
moon --maxmemory 25769803776 --maxmemory-policy allkeys-lfu
Human-readable suffixes are not supported; use bytes. Common values:
| Memory | Bytes |
|---|---|
| 1 GB | 1073741824 |
| 2 GB | 2147483648 |
| 4 GB | 4294967296 |
| 8 GB | 8589934592 |
| 16 GB | 17179869184 |
| 32 GB | 34359738368 |
Eviction policies¶
| Policy | Best for |
|---|---|
noeviction |
Primary storage (returns error on OOM) |
allkeys-lfu |
Cache workloads (evicts least frequently used) |
allkeys-lru |
Cache workloads (evicts least recently used) |
volatile-lfu |
Mixed: only evicts keys with TTL set |
volatile-ttl |
Evicts keys closest to expiration |
allkeys-random |
Simple random eviction |
Disk offload (tiered storage)¶
Moon can spill evicted keys to NVMe instead of deleting them:
Reads from the cold tier use async read-through with full crash recovery.
SWAPDB is refused while either database has keys in the cold tier or spill
files not yet reclaimed (a spill file records the database it was spilled
from, and a swap cannot re-tag it crash-consistently). To swap such a
database, delete its cold keys (or read them back into RAM), then run
BGREWRITEAOF (--appendonly yes) or BGSAVE (--appendonly no): once
that fold or snapshot covers the emptied files, the next orphan sweep
reclaims them and the swap is allowed. Waiting for the automatic rewrite can
take a long time. If another shard holds a database for the whole bounded
check (about 20,000 yields, a few ms), the reply is
ERR SWAPDB could not check the disk-offload cold tier of every shard, try again
and nothing is swapped. A replica whose own cold tier holds either database
does not apply its master's SWAPDB: it drops the link and resyncs in full
(moon#1278), which costs a full transfer per such swap.
Without an AOF (--appendonly no), a restart returns to the last snapshot
plus the cold tier on disk, and the cold tier records no removals:
- A spill file is removed only after a successful snapshot that started after
its last key left (moon#1260), because until then it may be the only copy
of a key read back into RAM. With no save rules (the default under
--appendonly no) nothing takes that snapshot: aDELorFLUSHALLof cold keys is undone by ANY restart — a cleanSHUTDOWNincluded — until a manualBGSAVE(orSHUTDOWN SAVE). Measured: 96 of 200 deleted keys back after a plainSHUTDOWN. That is consistent with the snapshot, but visible; runBGSAVEafter deleting cold data you want gone. - A cold key deleted, or overwritten with a TTL that then passes, can come
back after a LATER snapshot and a crash if its spill file still holds other
live keys: the rebuild re-indexes the dead slot (moon#1281). Use
--appendonly yeswhere deletes of cold data must be durable.
jemalloc tuning¶
Moon ships with jemalloc by default. The allocator is pre-tuned for the thread-per-core architecture. Key environment variables for advanced tuning:
# Per-CPU arenas (reduces cross-thread contention)
MALLOC_CONF="percpu_arena:percpu,background_thread:true,metadata_thp:auto"
In Docker Compose:
Monitoring memory¶
The admin port (--admin-port 9100) exposes /metrics with Prometheus-format memory gauges.
Shard Count Selection¶
Moon uses a thread-per-core, shared-nothing architecture. Each shard owns its event loop, data partition, WAL writer, and Pub/Sub registry.
| Scenario | Recommended --shards |
Rationale |
|---|---|---|
| Development/testing | 1 |
Simplest debugging, deterministic behavior |
| Small workload (<100K keys) | 1 |
Single shard avoids cross-shard dispatch overhead |
| General production | 1 (default) |
Best non-pipelined throughput; cross-shard dispatch dominates otherwise |
| Memory benchmarking | 1 |
Fair per-key comparison against Redis |
| High-throughput pipeline | CPU cores (or 0 = auto) |
Per-shard WAL eliminates global serialization bottleneck |
When using --shards 0 (auto-detect), set CPU limits in Docker to control the detected count:
Hash tags for key co-location¶
Use {tag} in key names to route related keys to the same shard:
SET user:{1234}:name "Alice"
SET user:{1234}:email "alice@example.com"
MGET user:{1234}:name user:{1234}:email # Same shard, no cross-shard dispatch
Health Checks and Monitoring¶
Admin port¶
Enable the admin port for HTTP-based monitoring:
| Endpoint | Purpose |
|---|---|
/healthz |
Liveness probe (returns 200 when server is running) |
/readyz |
Readiness probe (returns 200 when accepting connections) |
/metrics |
Prometheus metrics (QPS, latency, memory, clients, keyspace) |
/ui/ |
Web console (Dashboard, Browser, Console, Vectors, Graph, Memory) |
Kubernetes probes¶
livenessProbe:
httpGet:
path: /healthz
port: 9100
initialDelaySeconds: 5
periodSeconds: 10
readinessProbe:
httpGet:
path: /readyz
port: 9100
initialDelaySeconds: 5
periodSeconds: 5
Docker healthcheck¶
The default Docker healthcheck uses moon --check-config. For a more thorough probe when the admin port is available, override with curl:
healthcheck:
test: ["CMD", "curl", "-sf", "http://localhost:9100/healthz"]
interval: 10s
timeout: 3s
retries: 3
Note: The default distroless image does not include curl. Use a debian-slim based image or a sidecar for HTTP health checks.
Redis PING¶
From any Redis client:
INFO command¶
INFO server # Version, uptime, mode
INFO memory # RSS, peak, fragmentation ratio
INFO clients # Connected clients, blocked clients
INFO stats # Total commands, ops/sec, keyspace hits/misses
INFO replication # Master/replica status
INFO keyspace # Per-database key counts
INFO all # Everything
SLOWLOG¶
SLOWLOG GET 10 # Last 10 slow commands
SLOWLOG LEN # Number of entries
SLOWLOG RESET # Clear the log
CONFIG SET slowlog-log-slower-than 5000 # 5ms threshold
Scaling Guidelines¶
Single node¶
Moon's thread-per-core architecture scales vertically to the number of CPU cores. A single instance on an 8-core machine with monoio + io_uring can achieve:
- 4.8M+ GET/s at pipeline depth 64
- 3.6M+ SET/s at pipeline depth 64 (with AOF everysec)
Horizontal scaling¶
For datasets larger than a single node's memory or for high availability:
- Read replicas -- use
REPLICAOF <master-ip> <master-port>for read scaling - Client-side sharding -- partition keyspace across multiple Moon instances
- Cluster mode (experimental) --
--cluster-enabledwith gossip-based slot migration
Resource planning¶
| Metric | Guideline |
|---|---|
| CPU | 1 core per shard; add 1 core for admin/monitoring overhead |
| Memory | Dataset size + 25% for jemalloc overhead and fragmentation |
| Disk | 2x dataset size for AOF rewrite headroom |
| Network | 10 Gbps for >1M ops/sec workloads |
| File descriptors | ulimit -n 65536 for >1K concurrent clients |
Backup and Restore¶
RDB snapshot backup¶
# Trigger a background save
redis-cli -p 6379 BGSAVE
# Wait for completion
redis-cli -p 6379 LASTSAVE
# Copy the dump file
docker cp moon:/data/dump.rdb ./backup/dump.rdb
AOF backup¶
# Trigger AOF rewrite to compact the file
redis-cli -p 6379 BGREWRITEAOF
# Copy persistence directory
docker cp moon:/data/ ./backup/
Restore from backup¶
# Stop the server
docker compose down
# Replace persistence files
cp backup/dump.rdb /var/lib/moon/dump.rdb
# OR for AOF:
cp backup/appendonly.aof /var/lib/moon/appendonly.aof
# Start the server (will replay from persistence files)
docker compose up -d
Automated backup with cron¶
# /etc/cron.d/moon-backup
0 */6 * * * root docker exec moon redis-cli BGSAVE && sleep 5 && docker cp moon:/data/dump.rdb /backup/moon/dump-$(date +\%Y\%m\%d-\%H\%M).rdb
Security Checklist¶
Before deploying to production:
- Set
--requirepassor use an ACL file -- never run without authentication on a network-accessible port - Enable TLS (
--tls-port) for encrypted connections; consider--port 0to disable plaintext - Use mTLS (
--tls-ca-cert-file) for zero-trust environments - Set
--protected-mode yes(default) when running outside Docker - Restrict admin port access with
--console-auth-requiredand--console-auth-secret - Set
--console-cors-originto your specific domain (not wildcard) - Run as non-root (the Docker image uses UID 65534 by default)
- Set resource limits (CPU, memory) to prevent noisy-neighbor issues
- Mount secrets (passwords, TLS keys) as Docker secrets or read-only volumes, not environment variables
- Review ACL file permissions:
chmod 600 users.acl
Troubleshooting¶
"Connection refused" from outside Docker¶
Use --bind 0.0.0.0 inside containers. The default 127.0.0.1 is only reachable from within the container.
High latency with many shards¶
For small datasets (<100K keys), use --shards 1. Cross-shard dispatch overhead dominates local DashTable lookup for non-pipelined workloads.
"Too many open files"¶
Set ulimit -n 65536 on the host or use ulimits in Docker Compose. Testing with >1K concurrent clients requires this.
io_uring not available¶
Moon falls back gracefully. Set MOON_NO_URING=1 to explicitly disable io_uring, or use the tokio runtime:
WAL sync kills write throughput¶
appendfsync=always reduces write throughput by approximately 11x. Use appendfsync=everysec for the best durability/performance trade-off.
Container OOM killed¶
Set --maxmemory to 75-80% of the container memory limit. Moon's eviction policies will keep memory in bounds. Without maxmemory, the dataset grows until the container is killed.
Multi-Shard Scaling: When It Wins and When It Hurts¶
Moon shards the keyspace across N independent per-shard event loops, each with its own
SO_REUSEPORT listener, DashTable, and WAL writer. This section documents the empirically
validated operating envelope for multi-shard configurations.
When multi-shard wins¶
| Scenario | Recommended shards | Reason |
|---|---|---|
Pipeline depth ≥ 16, c ≥ 25 × shards |
4–8 | Parallel WAL, parallel DashTable, no contention |
| AOF enabled, write-heavy | 4–8 | Per-shard WAL eliminates the global lock Redis's single AOF file imposes |
Hash-tagged key workloads ({tag}:...) |
N (any) | Tagged keys co-locate on one shard; zero cross-shard dispatch |
c ≥ 200 clients, mixed read/write |
4–8 | RwLock contention becomes invisible at high client counts |
When multi-shard hurts¶
| Scenario | Use instead | Reason |
|---|---|---|
| Non-pipelined (p=1), small client counts | --shards 1 |
SPSC dispatch overhead dominates local lookup |
c < 25 × shards |
Reduce shard count or increase clients | Benchmark artifact: each shard sees fewer than 25 clients, causing measurable sub-linear scaling |
| Random keyspace, p=1, c=50 | --shards 1 |
Observed 14× SET throughput collapse at c=50 (benchmark artifact, vanishes at c=200) |
| Per-key memory comparison with Redis | --shards 1 |
Fair apples-to-apples comparison requires a single shard |
The c ≥ 25 × shards rule¶
Empirically validated 2026-04-22 through 2026-04-26 (GCloud c4a Axion):
- At
c=50with 8 shards (6 clients/shard), SET p=1 collapsed 14× vs single-shard. - At
c=200with 8 shards (25 clients/shard), the collapse vanished — throughput was flat across 1/4/8 shards. - Root cause: at
c < 25 × shards, each shard is statistically under-subscribed and the SPSC dispatch round-trip dominates.
Rule: ensure clients ≥ 25 × shards in production. For 8 shards, that means at least
200 concurrent connections. Use redis-benchmark -c 200 or higher for fair multi-shard
benchmarks.
Cross-shard fast path¶
For a read whose key hashes to a foreign shard (a GET on shard 3 arriving on a
connection pinned to shard 0), Moon can either hop the command through that shard's SPSC
channel or serve it in place under a shared read guard on the owner's database.
--cross-shard-fast-path |
Behaviour |
|---|---|
auto (default) |
Serve the read in place when the path can actually fire — monoio handler, --shards > 1. Declines elsewhere rather than lighting a switch the leg ignores. |
on |
Force it on regardless of shard count or runtime. |
off |
Route the read through the SPSC channel, exactly like a write. One extra channel round-trip and one park per read. The rollback. |
Measured at --shards 8, uniform keyspace, GET p=1, counter ratios from
INFO stats (total_dispatch_cross_read_fast vs total_dispatch_cross_spsc
vs total_remote_awaits_parked), same binary, one flag apart:
| mode | served in place | parks/cmd |
|---|---|---|
off |
0.0% | 0.87336 |
auto |
100.0% | 0.00023 |
docs/internal/cross-shard-cost-model.md prices a park at ~24.9 core-us and at
85% of p=1 cost, so this is the largest single lever on the cross-shard read
path. Reads only — a cross-shard write still parks.
The default shipped off while the only evidence was moon#768's -8.61% CPU/op
and a doubled s8 p16 variance. Both readings came from a half-populated
keyspace: the path declines a key that is not resident (dispatch_read cannot
consult the cold tier — the moon#610 class), and a declined read falls back to
the hop. So the "in-place rate" was tracking the benchmark's key HIT rate, and a
wandering hit rate is exactly the variance that held the default down. Against
DBSIZE: 63,114 keys resident gives 62.9% in place, 98,169 gives 98.2%, 100,000
gives 100.0%.
monoio only. The dispatch site is in the monoio connection handler — the runtime that ships on Linux and macOS. A build using
--features runtime-tokioaccepts the flag but ignores it: cross-shard reads still take the SPSC hop andtotal_dispatch_cross_read_faststays at 0. The server logs a warning at startup when the flag is enabled on that runtime.
The default is auto. off is the rollback: it restores the SPSC hop for every
cross-shard read without a rebuild.
The measurement that first held the default at off¶
This is the moon#768 run that kept the flag off until moon#777. It is kept because the
p=16 variance it shows is real; its low in-place rates are the half-populated-keyspace
artefact explained above, not a property of the path.
Two-box GCE ARM (t2a-standard-8 server, dedicated load generator), the same binary with
the flag off vs on, ABBA-ordered, n=10 reps per cell, server-side CPU/op from
/proc/<pid>/stat:
| cell | CPU/op | 95% CI | throughput | reps cheaper | p | reads served in place |
|---|---|---|---|---|---|---|
--shards 1, p=1 |
−0.05% | −0.82 … +0.72 | +0.07% | 4/10 | 0.75 | 0.0% |
--shards 8, p=1 |
−8.61% | −11.55 … −5.67 | +12.03% | 10/10 | 0.002 | 50.5% |
--shards 8, p=16 |
+6.46% | −10.05 … +22.98 | +6.71% | 4/10 | 0.75 | 30.3% |
The --shards 1 row is the negative control, not a result: every read there is already
local, so the path fires 0.0% of the time and must show nothing. It does. That is what
makes the --shards 8 row credible.
The p=16 row is what held the default at off, not any doubt about p=1. At depth 16 the
effect is not measurable at n=10, and the enabled leg's run-to-run variance roughly doubles
(sd 1.24 vs 0.55 µs/op) instead of shifting, which looks like contention rather than a
uniform regression. The populated-keyspace measurement above explains the variance as a
wandering hit rate. On a deep-pipeline workload, still compare against off on your own
data.
What the fast path declines¶
It is not a blanket bypass. Each of these falls back to SPSC, silently and correctly:
- The connection has in-flight remote work on that shard. Serving the read locally would let it overtake this connection's own queued write — read-your-own-writes (moon#507 / moon#512). Ordering wins; the read takes the hop.
- The command spans shards. Placement is read from
multikey_placement, the same function the router uses, so the two cannot drift apart. - The command is multi-key. Conservative for now: an all-on-one-shard
MGETcan still have some keys in the cold tier, and the residency probe only covers the primary key. - The key is not resident. The read-only dispatch path does not consult the cold tier, so a cold key takes the SPSC path where the full dispatch promotes it.
- The owner is mid-write. The guard is attempted, never waited on.
Consequently the fast-path counter tracks well under the total cross-shard read count, and a workload that is write-heavy on the same keys may see almost no fast-path traffic. That is the design working, not a fault.
Observability¶
moon_dispatch_path_total{path="cross_read_fast"}— reads served in place.moon_dispatch_path_total{path="cross_spsc"}— cross-shard commands that took the hop (reads and writes both).INFO statsexposes the same two astotal_dispatch_cross_read_fastandtotal_dispatch_cross_spsc, so they are readable withadmin_port=0.
total_dispatch_cross_read_fast is 0 whenever the flag is off. If it is 0 with the flag
on, a gate above is declining every read — check whether the workload pipelines writes
and reads of the same keys together.
Future: shared-nothing architecture (Phase 4)¶
The --cross-shard-fast-path flag is a Phase 0 observability and safety valve for the
5-phase shared-nothing migration documented in
.planning/shared-nothing-migration/PLAN.md.
That plan (last revised May 2026) predates moon#777, which made the fast path the default,
and it has not been revised since. Its Phase 3 step, deleting the fast path and routing all
cross-shard reads through SPSC, is a proposal, not a scheduled change. In Phase 4, RwLock<Database> is replaced with per-thread owned state (ShardSlice
+ thread_local!), eliminating lock contention entirely. The --cross-shard-fast-path=off
flag previews that future behavior and can be used to validate SPSC-only routing in
staging today.