The Moon JourneyΒΆ
The goal, from day one: a Redis-compatible server whose thread-per-core architecture measurably out-scales Redis on multi-core hardware β without sacrificing protocol compatibility or durability semantics.
Not "faster in a microbenchmark." Faster where it counts, proven on real hardware, with the durability and correctness a production datastore is required to keep. This page is the story of how Moon got there β and, just as importantly, the ideas it measured and threw away along the road, because "real efficiency" is what survives a benchmark, not what sounds fast.
Every number below is a measured result. Absolute figures come from the Linux
production reference (GCloud c3-standard-8 x86_64 and t2a-standard-8 ARM64,
vs Redis 8.6.1); the Benchmarks page and
BENCHMARK.md
carry the full methodology and raw data.
fsync=always durable
write throughput both landed with the v0.5.x busy-poll + group-commit work
and held through v0.8.1, where the O3 governor made the p=1 spin safe to
deploy on any host. Same-instance A/B vs Redis 8.6.1.
The thesis: efficiency is a shared-nothing problemΒΆ
Redis is single-threaded by design. On a modern many-core box that leaves throughput on the table, and its one global AOF file serializes every durable write. Moon's bet was that a thread-per-core, shared-nothing design β each shard owning its slice of the keyspace, its own event loop, and its own write-ahead log β could turn those cores into linear throughput and durable writes into a parallel, not a serial, operation.
The bet paid off, but only after the hard parts were solved: cross-shard dispatch overhead, the single-connection latency floor, allocator pressure, and making durability free instead of an 11Γ tax. The releases below are that work.
Milestone arcΒΆ
timeline
title Moon's road to a production-efficient database
v0.4.x : Shared-nothing core
: Lock-free cross-shard dispatch
: Zero-copy protocol + FTS/graph foundations
v0.5.x : KV correctness + WAL v3
: Won p=1 GET/SET vs Redis
: Durable write-path 2x campaign
v0.6.0 : Multi-db isolation
: Tiered engine offload (-26% RSS)
: --profile standalone preset
v0.7.x : Replication GA (6 planes)
: WAIT / ACK durability
: 24h kill-9 soak, zero loss
v0.8.0 : One Storage Kernel
: 46-cell kill-9 crash matrix
: 10x RAM via disk offload
v0.8.1 : Deploy-safe busy-poll (O3)
: p=1 win auto-gates on shared cores
: Single-shard tuning everywhere
| Release | Headline | The efficiency it bought |
|---|---|---|
| v0.4.x | Shared-nothing core + performance + FTS/graph foundations | Per-shard keyspace, lock-free cross-shard flume dispatch, zero-copy protocol parsing |
| v0.5.x | KV correctness + the p=1 win + WAL v3 | Beat Redis at single-connection GET/SET (its historical best case); durability write-path 2Γ campaign |
| v0.6.0 | Multi-db isolation + tiered engine offload | db-scoped indexes + per-db quotas; vector segments demote HOTβWARMβCOLD for β26% RSS; --profile standalone |
| v0.7.x | Replication GA for multi-shard masters | WAIT/ACK across all six data planes; 24 h kill-9 soak, zero acked-write loss |
| v0.8.0 | One Storage Kernel β kill-9-lossless + 10Γ RAM | Every plane crash-durable (46-cell matrix); datasets 10Γ RAM via disk offload with truthful accounting |
| v0.8.1 | Deploy-safe busy-poll + single-shard preset | The p=1 latency win auto-gates on shared cores β safe on any host, not just pinned ones |
Full detail per release: RELEASES.md Β· CHANGELOG.md.
Capabilities accreted without a module loader β every engine compiles into the one binary and shares the keyspace, durability, and crash matrix:
From a shared-nothing KV core with FTS/vector/graph foundations to a crash-lossless storage kernel with replication β one binary, growing.
What was achieved β by dimensionΒΆ
Throughput: turning cores into ops/secΒΆ
| Metric | Moon | vs Redis | Conditions |
|---|---|---|---|
| Peak GET | 5.11M ops/sec | 1.72Γ | GCloud x86_64, p=64 |
| Peak SET | 3.50M ops/sec | 1.92Γ | GCloud x86_64, p=64 |
| Peak GET (ARM64) | 3.47M ops/sec | 2.20Γ | Neoverse-N1, p=64 |
| 8-shard SET p=16 | 2.52M ops/sec | 1.99Γ | c=50 |
Peak throughput vs Redis 8.6.1 on the GCloud production reference (pipeline=64).
The shared-nothing payoff shows up at pipeline depth, not at raw shard count: for a uniform single-key workload, 1β8 shards is flat-to-slightly negative (most keys route cross-shard, so SPSC dispatch cost dominates the local lookup β single-shard is best). The multi-shard win comes from the per-shard WAL and independent event loops parallelizing pipelined and durable work, plus hash-tag co-location:
At pipeline depth, the per-shard WAL parallelizes work that Redis's single event loop serializes β 1.48β1.99Γ across multi-shard configs.
CPU efficiency: more work, far less siliconΒΆ
The number that best captures the architecture:
At pipeline=64, Moon delivers 1.71Γ the throughput of Redis while using 23Γ less CPU (1.9% vs 43.9% of a core for the same offered load).
Same offered load, pipeline=64: Redis burns 43.9% of a core, Moon 1.9% β and still serves 1.71Γ the throughput.
That efficiency is engineered, not incidental: lock-free oneshot channels
removed ~12% CPU of pthread_mutex contention, the cached shard clock removed
~4% of clock_gettime syscalls, and zero-copy parsing keeps the hot path
allocation-free. Idle cost was hunted just as hard β the adaptive idle-park and
an O(1) page-cache resident counter trimmed idle-shard CPU from 3.1% to
~0.5β0.9%, and the O3 contention governor makes the busy-poll spin
self-disengaging on contended cores. Efficiency that shows up on the
electricity bill, not just the benchmark.
The p=1 conquest β Redis's home turfΒΆ
Unpipelined, single-connection request/response was historically Redis's best
case and Moon's weakest. The --io-busy-poll-us poll-mode park closed it:
- 1.19β1.21Γ Redis on ARM (c4a Axion), 1.65β1.66Γ on x86 (c3) β same-instance A/B, n=3.
The honest catch was that the spin regressed on shared cores. v0.8.1's O3
governor removed the catch: each shard samples its own involuntary-preemption
rate and gates the spin automatically, so the win is safe on any host. One
flag β --profile standalone (or conf/moon-standalone.conf) β
now delivers the best single-shard tuning everywhere.
Latency: lower queue depth, lower tailΒΆ
| Metric | Redis | Moon | Improvement |
|---|---|---|---|
| p50 latency (8-shard) | 0.26β0.33 ms | 0.031 ms | 8β10Γ lower |
Multi-core parallelism reduces per-shard queue depth, so the median request waits less.
Memory: compact by constructionΒΆ
| Value size (1M keys) | Redis/key | Moon/key | Result |
|---|---|---|---|
| 1,024 B | 1,571 B | 1,153 B | 27% less |
| 1,024 B (63K keys) | 1,879 B | 1,207 B | 1.56Γ |
Per-key memory vs value size. Redis wins at tiny 32 B values; the lines cross near 256 B and Moon pulls steadily ahead from there.
CompactKey stores keys up to 23 bytes inline and CompactValue up to 12
bytes; larger payloads use HeapString(Vec<u8>) (48 B overhead) versus Redis's
robj + SDS chain (64β80 B). TTL packs as a 4-byte delta inside the entry at
zero extra cost, where Redis allocates a separate 24-byte dictEntry per
expiring key. (At tiny 32 B values Redis still wins on per-key overhead β Moon
is honest about that; the crossover is ~256 B.) Baseline RSS is identical
(7.0 MB empty), and tuned jemalloc decay (1 s dirty, background reclaim) returns
freed pages to the OS instead of hoarding them.
Durability: making it freeΒΆ
The per-shard WAL turns durable writes from a global serialization point into a parallel one. A 2026-07 write-path campaign closed the remaining gaps:
| Policy / workload | Before | After | vs Redis |
|---|---|---|---|
always SET p16 |
5.7K (0.12Γ) | 40.1K | 0.91Γ |
everysec SET p16 |
605K | 789K | 1.32Γ |
everysec SET p1 |
117K (0.80Γ) | 135K | 0.99Γ parity |
| Pub/sub fan-out | 438 msg/s (with drops) | 5.09M msg/s | 1.04Γ, zero drops |
The write-path campaign turned the fsync tax into near-parity: always p16
climbed 0.12Γβ0.91Γ of Redis, everysec p16 to a 1.32Γ win.
Per-batch group commit (one fsync per pipeline batch, not per command), a
single coalesced write_all, and a park-free AOF writer that deleted a
~150K/sec futex-wake storm. Durability semantics are unchanged (always = RPO
0, everysec = RPO β€ 1s) and verified by the SIGKILL crash-recovery matrix.
AI-native β in-core, not bolted onΒΆ
Vector search, full-text (BM25), and a property graph are compiled into the same binary, sharing the keyspace and durability β no module loader.
- Vector: 12.7K search QPS at 384d (HNSW + TurboQuant, COSINE), and bulk load β searchable at target recall beats Qdrant 1.6β2.3Γ via the parallel HNSW build. SQ8 quantization holds ~0.90 R@10 on real MiniLM 384-d embeddings across the full lifecycle (search / merge / persistence).
- Graph: 23Γ FalkorDB on bulk insert, ~2.4Γ Cypher QPS, and 2.78Γ on point-filter queries after the mutable-property fast path.
- Tiered engine offload: idle vector-index segments demote HOT β WARM (mmap) β COLD (unloaded stub, reload-on-search), giving memory back β β26% process RSS on a 40K Γ 768d corpus with identical search results after reload.
In-core vector and graph engines measured against the dedicated systems they replace β Moon's multiplier on each.
Production hardening β the part that makes it a databaseΒΆ
Efficiency means nothing if a kill -9 loses data. The v0.8.0 Storage
Kernel made durability a verifiable claim across every plane:
- Cross-plane kill-9 crash matrix: a 46-cell matrix (KV / vector /
graph / FTS / workspaces / MQ / temporal / txn Γ persistence mode Γ
disk-offload Γ shard count), all green β wired into scheduled CI (nightly
full matrix + weekly
ITERS=20soak). - 10Γ RAM datasets under disk offload: 2.6 GB of 10 KB values against a
256 MB cap β truthful
used_memoryat 1.00Γ cap steady-state, spill files cut from ~236,000 (one per key) to 840 (~280Γ fewer), worst cold-GET tail during an active spill flood 1,910 ms β 205 ms (9.3Γ better), and 500/500 kill-9 integrity. - Replication GA: real async replication with
WAIT/ACKacross all six data planes, validated by a 24-hour kill-9 soak β 114 alternating master/replica kills, 82,044WAIT-acked writes, zero lost.
The discipline: efficiency is what survives measurementΒΆ
The reason the numbers above hold up is a rule the project keeps: no measured win β no hot-path change. That rule is only credible because of what it rejected. A representative sample of ideas that sounded fast and were thrown out on the evidence:
- Value-line prefetch β measured "3β6% faster" until disassembly showed the
benchmark's
black_boxnever dereferenced the value; a forced-load harness showed it was 1β2% slower. Cut, and a cache-exceeding probe bench shipped so the trap can't recur. - A lock-free
Notifyreplacing theflumechannel β fully validated (loom lost-wake model, consistency suites, both-impl A/B) but neutral on the wall clock, because the wake is syscall-dominated. Theflumechannel stayed. - Transparent huge pages by default β real PMU gains (+12β24% GET), but a 45-minute RSS-drift soak showed idle khugepaged re-collapse drifts RSS +31%. Kept permanently opt-in, not the default.
- Per-segment
ef/βGbeam splitting for vector search β silently cost recall (0.9915 β 0.9295). Rejected; a recall-certified adaptive scheme shipped instead.
None of these are failures β they're the cost of knowing the wins are real. The efficiency Moon ships is the efficiency that survived.
Backing that discipline: ~4,980 tests, 11+ cargo-fuzz targets, loom
models for every lock-free structure, unsafe/unwrap ratchets that block new
unannotated uses, an RSS-regression CI gate, and the crash matrix above.
Where it's goingΒΆ
Moon is on a path to a v1.0 GA with every Production Contract box ticked. Next up: horizontal scale β cluster-on-monoio hardening and multi-shard replicas β then the enterprise foundation. The method won't change: measure, ship what wins, and say so honestly when it doesn't.
See the live numbers on the Benchmarks page, the full per-release history in RELEASES.md, and the GA scorecard in the Production Contract.






