The Moon JourneyΒΆ
The goal, from day one: a Redis-compatible server whose thread-per-core architecture measurably out-scales Redis on multi-core hardware β without sacrificing protocol compatibility or durability semantics.
Not "faster in a microbenchmark." Faster where it counts, proven on real hardware, with the durability and correctness a production datastore is required to keep. This page is the story of how Moon got there β and, just as importantly, the ideas it measured and threw away along the road, because "real efficiency" is what survives a benchmark, not what sounds fast.
Every number below is a measured result, and every one of them names the host
it was measured on. Figures marked Linux come from the production reference
(GCloud c3-standard-8 x86_64 and t2a-standard-8 ARM64); figures marked
macOS dev reference were measured on an Apple M4 Pro and are development
records only β CLAUDE.md forbids publishing them as production results, and
where Linux has since disagreed with one, this page says so instead of quietly
keeping the flattering number. The Benchmarks page and
BENCHMARK.md
carry the full methodology and raw data.
fsync=always durable
write throughput both landed with the v0.5.x busy-poll + group-commit work
and held through v0.8.1, where the O3 governor made the p=1 spin safe to
deploy on any host. Same-instance A/B vs Redis 8.6.1.
The thesis: efficiency is a shared-nothing problemΒΆ
Redis is single-threaded by design. On a modern many-core box that leaves throughput on the table, and its one global AOF file serializes every durable write. Moon's bet was that a thread-per-core, shared-nothing design β each shard owning its slice of the keyspace, its own event loop, and its own write-ahead log β could turn those cores into linear throughput and durable writes into a parallel, not a serial, operation.
The bet paid off, but only after the hard parts were solved: cross-shard dispatch overhead, the single-connection latency floor, allocator pressure, and making durability free instead of an 11Γ tax. The releases below are that work.
Milestone arcΒΆ
timeline
title Moon's road to a production-efficient database
v0.4.x : Shared-nothing core
: Lock-free cross-shard dispatch
: Zero-copy protocol + FTS/graph foundations
v0.5.x : KV correctness + WAL v3
: Won p=1 GET/SET vs Redis
: Durable write-path 2x campaign
v0.6.0 : Multi-db isolation
: Tiered engine offload (-26% RSS)
: --profile standalone preset
v0.7.x : Replication GA (6 planes)
: WAIT / ACK durability
: 24h kill-9 soak, zero loss
v0.8.0 : One Storage Kernel
: 46-cell kill-9 crash matrix
: 10x RAM via disk offload
v0.8.1 : Deploy-safe busy-poll (O3)
: p=1 win auto-gates on shared cores
: Single-shard tuning everywhere
| Release | Headline | The efficiency it bought |
|---|---|---|
| v0.4.x | Shared-nothing core + performance + FTS/graph foundations | Per-shard keyspace, lock-free cross-shard flume dispatch, zero-copy protocol parsing |
| v0.5.x | KV correctness + the p=1 win + WAL v3 | Beat Redis at single-connection GET/SET (its historical best case); durability write-path 2Γ campaign |
| v0.6.0 | Multi-db isolation + tiered engine offload | db-scoped indexes + per-db quotas; vector segments demote HOTβWARMβCOLD for β26% RSS; --profile standalone |
| v0.7.x | Replication GA for multi-shard masters | WAIT/ACK across all six data planes; 24 h kill-9 soak, zero acked-write loss |
| v0.8.0 | One Storage Kernel β kill-9-lossless + 10Γ RAM | Every plane crash-durable (46-cell matrix); datasets 10Γ RAM via disk offload with truthful accounting |
| v0.8.1 | Deploy-safe busy-poll + single-shard preset | The p=1 latency win auto-gates on shared cores β safe on any host, not just pinned ones |
Full detail per release: RELEASES.md Β· CHANGELOG.md.
Capabilities accreted without a module loader β every engine compiles into the one binary and shares the keyspace, durability, and crash matrix:
From a shared-nothing KV core with FTS/vector/graph foundations to a crash-lossless storage kernel with replication β one binary, growing.
What was achieved β by dimensionΒΆ
Throughput: turning cores into ops/secΒΆ
| Metric | Moon | vs Redis | Conditions |
|---|---|---|---|
| Peak GET (v0.1.6) | 5.11M ops/sec | 1.72Γ | Linux, GCloud c3-standard-8 x86_64, p=64. Redis io-threads setting and payload size not recorded β BENCHMARK.md Β§2.1 |
| Peak SET (v0.1.6) | 3.50M ops/sec | 1.92Γ | same run β Β§2.1 |
| Peak GET (ARM64, v0.1.6) | 3.47M ops/sec | 2.20Γ | Linux, GCloud t2a-standard-8 Neoverse-N1, p=64 β Β§2.1 |
| Peak GET (v0.8.7, current tree) | ratio only | 2.40Γ x86 / 2.29Γ ARM | Linux, GCE c3/t2a, Redis 7.0.15, --shards 1, c=50, p=64 β Β§2.12 |
| 8-shard SET p=16 | 2.52M ops/sec | 1.99Γ | macOS dev reference (Apple M4 Pro), c=50 β Β§4.2 |
Scope: this is the GET/SET inline path, not the whole command surface
On Linux, v0.8.7 measures every non-inlined command family β INCR, LPUSH,
SPOP, HSET β at 0.40β0.67Γ Redis at pβ₯8 (BENCHMARK.md Β§2.12). SET k v
runs 2.08Γ Redis; SET k v EX 100, the same work with one disqualifying
option, runs 0.87Γ. The boundary is the fast path, not the engine.
Peak throughput vs Redis 8.6.1 on the GCloud production reference (pipeline=64).
Sharding does scale, and an earlier version of this page said otherwise. The
claim that "1β8 shards is flat-to-slightly negative" was a harness artifact:
redis-benchmark -t lpush|sadd|hset|zadd drives ONE literal key and -r
randomises the element, not the key β so eight of twelve command families were
being asked to parallelise a single key, and 0.97Γ was the tautological answer.
Re-run on Linux with explicit __rand_int__ keys under a DBSIZE β₯ 50000 guard
proven to fire, moon's real 8-shard-over-1-shard scaling is 1.42Γ at p=1,
2.14Γ at p=8, and 3.79Γ at p=64 (BENCHMARK.md Β§2.14; Redis io-threads 8-over-1
over the same matrix is 1.19Γ / 1.23Γ / 1.08Γ). A genuinely single-hot-key
workload still cannot be sharded β that part was always true β and the
multi-shard win is amplified further by the per-shard WAL and independent event
loops parallelizing pipelined and durable work, plus hash-tag co-location:
At pipeline depth, the per-shard WAL parallelizes work that Redis's single event loop serializes β 1.48β1.99Γ across multi-shard configs. Apple M4 Pro development reference (BENCHMARK.md Β§4.2); not reproduced on Linux.
CPU efficiency: a tie on cost, a win on stabilityΒΆ
On Linux, CPU per operation is a tie. Moon --shards 8 costs 10.55 Β΅s/op
against Redis --io-threads 8's 11.33 Β΅s/op (GCE t2a-standard-8,
utime+stime from /proc/<pid>/stat, 5 reps, BENCHMARK.md Β§2.14). Moon's 6.9%
edge sits inside Redis's own 11.9% run-to-run spread, so it is not a win. What
is durable is stability: moon's CPU cost varies 2.0% run to run against
Redis's 11.9%, the latter being stopThreadedIOIfNeeded engaging and
disengaging io-threads under steady load.
This page previously claimed "23Γ less CPU" here. That figure has been
withdrawn: it came from an Apple M4 Pro table, its CPU column was sampled
with ps -o %cpu= (a process-lifetime average, not steady-state load), and its
CPU and RPS columns were not taken in the same run. Three independent reasons
not to publish it, and it was published under a Linux heading.
The engineering behind the cost that is measured is real, not incidental: lock-free oneshot channels
removed ~12% CPU of pthread_mutex contention, the cached shard clock removed
~4% of clock_gettime syscalls, and zero-copy parsing keeps the hot path
allocation-free. Idle cost was hunted just as hard β the adaptive idle-park and
an O(1) page-cache resident counter trimmed idle-shard CPU from 3.1% to
~0.5β0.9%, and the O3 contention governor makes the busy-poll spin
self-disengaging on contended cores. Efficiency that shows up on the
electricity bill, not just the benchmark. Those component figures are profile
deltas from the optimisation work, not a Moon-vs-Redis ratio.
The p=1 conquest β Redis's home turfΒΆ
Unpipelined, single-connection request/response was historically Redis's best
case and Moon's weakest. The --io-busy-poll-us poll-mode park closed it:
- 1.19β1.21Γ Redis on ARM (c4a Axion), 1.65β1.66Γ on x86 (c3) β Linux, same-instance A/B, n=3, on dedicated cores.
Re-measured on ordinary shared-tenant GCE instances (v0.8.7, BENCHMARK.md Β§2.12) the same flag yields 1.06β1.08Γ on x86 and nothing outside the noise floor on ARM, because the contention governor correctly self-gates there. Quote the dedicated-core figures only with that word attached.
The honest catch was that the spin regressed on shared cores. v0.8.1's O3
governor removed the catch: each shard samples its own involuntary-preemption
rate and gates the spin automatically, so the win is safe on any host. One
flag β --profile standalone (or conf/moon-standalone.conf) β
now delivers the best single-shard tuning everywhere.
Latency: lower queue depth, unmeasured on LinuxΒΆ
Multi-core parallelism reduces per-shard queue depth, so the median request
should wait less. That has not been measured on a Linux host, and the
"8β10Γ lower p50" figure this page used to carry has been withdrawn: it was an
Apple M4 Pro development reference published under a Linux heading, and it came
from redis-benchmark, a closed-loop tool that under-reports latency once the
server saturates. The raw macOS numbers are kept, correctly labelled, in
BENCHMARK.md Β§9.1.
Memory: a real win at large values, and half the number we publishedΒΆ
The shape held. The headline number did not. Re-measured on Linux
(BENCHMARK.md Β§3 β GCE c3-standard-8, x86_64, moon d5f3501b at
--shards 1, against Redis 7.4.2 built with jemalloc 5.3.0, fresh server per
point, redis-benchmark -r N):
| Value size | Keys | Redis/key | Moon/key | Result |
|---|---|---|---|---|
| 32 B | 63Kβ632K | 123β129 B | 143β186 B | Moon 11β51% worse |
| 256 B | 63Kβ632K | 408β410 B | 377β418 B | tie (Β±8%) |
| 1 KB | 63Kβ632K | 1,380β1,388 B | 1,155β1,256 B | Moon 9.5β16.4% less |
| 4 KB | 63Kβ632K | 5,259β5,266 B | 4,352β4,404 B | Moon 16.4β17.2% less |
This page used to say 27β35%. The number that moved was Redis's.
The old 1M Γ 1 KB row was Redis 1,571 B / Moon 1,153 B, measured on an Apple M4 Pro. Measured again on Linux against a jemalloc Redis: Redis 1,380 B, Moon 1,172 B. Moon's own number moved 1.6%; Redis's moved 12%. Both Redis binaries already on the benchmark host were libc-malloc builds β which inflate Redis RSS and would have biased the comparison in Moon's favour β so Redis was rebuilt against jemalloc for the run. The claim was never a Moon regression; it was an unrepresentative baseline. x86_64 only.
Per-key memory vs value size, measured on the Apple M4 Pro development rig: Redis wins at tiny 32 B values, the lines cross near 256 B, and Moon pulls ahead from there. The Linux measurement above reproduces that shape β loss at 32 B, tie at 256 B, win from 1 KB. What it does not reproduce is the size of the 1 KB win: 15% on Linux against the 27% plotted here. (At 4 KB the two agree, 17% vs 15%.) The plotted values are the superseded macOS ones.
CompactKey stores keys up to 23 bytes inline and CompactValue up to 12
bytes; larger payloads use HeapString(Vec<u8>) (48 B overhead) versus Redis's
robj + SDS chain (64β80 B). TTL packs as a 4-byte delta inside the entry at
zero extra cost, where Redis allocates a separate 24-byte dictEntry per
expiring key. That model explains the large-value win and not the 32 B loss,
where the same arithmetic predicts Moon should win; no cause is asserted.
Tuned jemalloc decay (1 s dirty, background reclaim) returns freed pages to the
OS instead of hoarding them.
Baseline RSS is the one that got worse under measurement. This page's ancestor
claimed 7.0 MB empty for both servers, measured on macOS. On Linux, across all 12
points of the run above: Redis 7.5β7.7 MB, Moon 12.6β12.9 MB β ~1.7Γ worse.
Nobody was watching that number; it is now tracked as
#821. On aarch64 at
--shards 8, BENCHMARK.md Β§2.14 separately measures idle RSS 1.26Γ above
Redis's, roughly 520 KB per extra shard, four threads per shard.
Durability: making it freeΒΆ
The per-shard WAL turns durable writes from a global serialization point into a parallel one. A 2026-07 write-path campaign closed the remaining gaps:
| Policy / workload | Before | After | vs Redis |
|---|---|---|---|
always SET p16 |
5.7K (0.12Γ) | 40.1K | 0.91Γ |
everysec SET p16 |
605K | 789K | 1.32Γ |
everysec SET p1 |
117K (0.80Γ) | 135K | 0.99Γ parity |
| Pub/sub fan-out | 438 msg/s (with drops) | 5.09M msg/s | 1.04Γ, zero drops |
The write-path campaign turned the fsync tax into near-parity: always p16
climbed 0.12Γβ0.91Γ of Redis, everysec p16 to a 1.32Γ win.
Per-batch group commit (one fsync per pipeline batch, not per command), a
single coalesced write_all, and a park-free AOF writer that deleted a
~150K/sec futex-wake storm. Durability semantics are unchanged (always = RPO
0, everysec = RPO β€ 1s) and verified by the SIGKILL crash-recovery matrix.
AI-native β in-core, not bolted onΒΆ
Vector search, full-text (BM25), and a property graph are compiled into the same binary, sharing the keyspace and durability β no module loader.
- Vector: 12.7K search QPS at 384d (HNSW + TurboQuant, COSINE), and bulk load β searchable at target recall beats Qdrant 1.6β2.3Γ via the parallel HNSW build. SQ8 quantization holds ~0.90 R@10 on real MiniLM 384-d embeddings across the full lifecycle (search / merge / persistence).
- Graph: 23Γ FalkorDB on bulk insert, ~2.4Γ Cypher QPS, and 2.78Γ on point-filter queries after the mutable-property fast path.
- Tiered engine offload: idle vector-index segments demote HOT β WARM (mmap) β COLD (unloaded stub, reload-on-search), giving memory back β β26% process RSS on a 40K Γ 768d corpus with identical search results after reload.
In-core vector and graph engines measured against the dedicated systems they replace β Moon's multiplier on each.
Production hardening β the part that makes it a databaseΒΆ
Efficiency means nothing if a kill -9 loses data. The v0.8.0 Storage
Kernel made durability a verifiable claim across every plane:
- Cross-plane kill-9 crash matrix: a 46-cell matrix (KV / vector /
graph / FTS / workspaces / MQ / temporal / txn Γ persistence mode Γ
disk-offload Γ shard count), all green β wired into scheduled CI (nightly
full matrix + weekly
ITERS=20soak). - 10Γ RAM datasets under disk offload: 2.6 GB of 10 KB values against a
256 MB cap β truthful
used_memoryat 1.00Γ cap steady-state, spill files cut from ~236,000 (one per key) to 840 (~280Γ fewer), worst cold-GET tail during an active spill flood 1,910 ms β 205 ms (9.3Γ better), and 500/500 kill-9 integrity. - Replication GA: real async replication with
WAIT/ACKacross all six data planes, validated by a 24-hour kill-9 soak β 114 alternating master/replica kills, 82,044WAIT-acked writes, zero lost.
The discipline: efficiency is what survives measurementΒΆ
The reason the numbers above hold up is a rule the project keeps: no measured win β no hot-path change. That rule is only credible because of what it rejected. A representative sample of ideas that sounded fast and were thrown out on the evidence:
- Value-line prefetch β measured "3β6% faster" until disassembly showed the
benchmark's
black_boxnever dereferenced the value; a forced-load harness showed it was 1β2% slower. Cut, and a cache-exceeding probe bench shipped so the trap can't recur. - A lock-free
Notifyreplacing theflumechannel β fully validated (loom lost-wake model, consistency suites, both-impl A/B) but neutral on the wall clock, because the wake is syscall-dominated. Theflumechannel stayed. - Transparent huge pages by default β real PMU gains (+12β24% GET), but a 45-minute RSS-drift soak showed idle khugepaged re-collapse drifts RSS +31%. Kept permanently opt-in, not the default.
- Per-segment
ef/βGbeam splitting for vector search β silently cost recall (0.9915 β 0.9295). Rejected; a recall-certified adaptive scheme shipped instead.
None of these are failures β they're the cost of knowing the wins are real. The efficiency Moon ships is the efficiency that survived.
Backing that discipline: ~4,980 tests, 11+ cargo-fuzz targets, loom
models for every lock-free structure, unsafe/unwrap ratchets that block new
unannotated uses, an RSS-regression CI gate, and the crash matrix above.
Where it's goingΒΆ
Moon is on a path to a v1.0 GA with every Production Contract box ticked. Next up: horizontal scale β cluster-on-monoio hardening and multi-shard replicas β then the enterprise foundation. The method won't change: measure, ship what wins, and say so honestly when it doesn't.
See the live numbers on the Benchmarks page, the full per-release history in RELEASES.md, and the GA scorecard in the Production Contract.





