Skip to content

The Moon JourneyΒΆ

The goal, from day one: a Redis-compatible server whose thread-per-core architecture measurably out-scales Redis on multi-core hardware β€” without sacrificing protocol compatibility or durability semantics.

Not "faster in a microbenchmark." Faster where it counts, proven on real hardware, with the durability and correctness a production datastore is required to keep. This page is the story of how Moon got there β€” and, just as importantly, the ideas it measured and threw away along the road, because "real efficiency" is what survives a benchmark, not what sounds fast.

Every number below is a measured result, and every one of them names the host it was measured on. Figures marked Linux come from the production reference (GCloud c3-standard-8 x86_64 and t2a-standard-8 ARM64); figures marked macOS dev reference were measured on an Apple M4 Pro and are development records only β€” CLAUDE.md forbids publishing them as production results, and where Linux has since disagreed with one, this page says so instead of quietly keeping the flattering number. The Benchmarks page and BENCHMARK.md carry the full methodology and raw data.

Moon vs Redis β€” the efficiency journey
Three efficiency curves that each crossed Redis parity, shown as a multiple of Redis by release. The single-connection (p=1) latency win β€” Redis's historical home turf β€” and the fsync=always durable write throughput both landed with the v0.5.x busy-poll + group-commit work and held through v0.8.1, where the O3 governor made the p=1 spin safe to deploy on any host. Same-instance A/B vs Redis 8.6.1.

The thesis: efficiency is a shared-nothing problemΒΆ

Redis is single-threaded by design. On a modern many-core box that leaves throughput on the table, and its one global AOF file serializes every durable write. Moon's bet was that a thread-per-core, shared-nothing design β€” each shard owning its slice of the keyspace, its own event loop, and its own write-ahead log β€” could turn those cores into linear throughput and durable writes into a parallel, not a serial, operation.

The bet paid off, but only after the hard parts were solved: cross-shard dispatch overhead, the single-connection latency floor, allocator pressure, and making durability free instead of an 11Γ— tax. The releases below are that work.


Milestone arcΒΆ

timeline
    title Moon's road to a production-efficient database
    v0.4.x : Shared-nothing core
           : Lock-free cross-shard dispatch
           : Zero-copy protocol + FTS/graph foundations
    v0.5.x : KV correctness + WAL v3
           : Won p=1 GET/SET vs Redis
           : Durable write-path 2x campaign
    v0.6.0 : Multi-db isolation
           : Tiered engine offload (-26% RSS)
           : --profile standalone preset
    v0.7.x : Replication GA (6 planes)
           : WAIT / ACK durability
           : 24h kill-9 soak, zero loss
    v0.8.0 : One Storage Kernel
           : 46-cell kill-9 crash matrix
           : 10x RAM via disk offload
    v0.8.1 : Deploy-safe busy-poll (O3)
           : p=1 win auto-gates on shared cores
           : Single-shard tuning everywhere
Release Headline The efficiency it bought
v0.4.x Shared-nothing core + performance + FTS/graph foundations Per-shard keyspace, lock-free cross-shard flume dispatch, zero-copy protocol parsing
v0.5.x KV correctness + the p=1 win + WAL v3 Beat Redis at single-connection GET/SET (its historical best case); durability write-path 2Γ— campaign
v0.6.0 Multi-db isolation + tiered engine offload db-scoped indexes + per-db quotas; vector segments demote HOTβ†’WARMβ†’COLD for βˆ’26% RSS; --profile standalone
v0.7.x Replication GA for multi-shard masters WAIT/ACK across all six data planes; 24 h kill-9 soak, zero acked-write loss
v0.8.0 One Storage Kernel β€” kill-9-lossless + 10Γ— RAM Every plane crash-durable (46-cell matrix); datasets 10Γ— RAM via disk offload with truthful accounting
v0.8.1 Deploy-safe busy-poll + single-shard preset The p=1 latency win auto-gates on shared cores β€” safe on any host, not just pinned ones

Full detail per release: RELEASES.md Β· CHANGELOG.md.

Capabilities accreted without a module loader β€” every engine compiles into the one binary and shares the keyspace, durability, and crash matrix:

Feature development β€” capability accretion by release

From a shared-nothing KV core with FTS/vector/graph foundations to a crash-lossless storage kernel with replication β€” one binary, growing.


What was achieved β€” by dimensionΒΆ

Throughput: turning cores into ops/secΒΆ

Metric Moon vs Redis Conditions
Peak GET (v0.1.6) 5.11M ops/sec 1.72Γ— Linux, GCloud c3-standard-8 x86_64, p=64. Redis io-threads setting and payload size not recorded β€” BENCHMARK.md Β§2.1
Peak SET (v0.1.6) 3.50M ops/sec 1.92Γ— same run β€” Β§2.1
Peak GET (ARM64, v0.1.6) 3.47M ops/sec 2.20Γ— Linux, GCloud t2a-standard-8 Neoverse-N1, p=64 β€” Β§2.1
Peak GET (v0.8.7, current tree) ratio only 2.40Γ— x86 / 2.29Γ— ARM Linux, GCE c3/t2a, Redis 7.0.15, --shards 1, c=50, p=64 β€” Β§2.12
8-shard SET p=16 2.52M ops/sec 1.99Γ— macOS dev reference (Apple M4 Pro), c=50 β€” Β§4.2

Scope: this is the GET/SET inline path, not the whole command surface

On Linux, v0.8.7 measures every non-inlined command family β€” INCR, LPUSH, SPOP, HSET β€” at 0.40–0.67Γ— Redis at pβ‰₯8 (BENCHMARK.md Β§2.12). SET k v runs 2.08Γ— Redis; SET k v EX 100, the same work with one disqualifying option, runs 0.87Γ—. The boundary is the fast path, not the engine.

Peak throughput vs Redis

Peak throughput vs Redis 8.6.1 on the GCloud production reference (pipeline=64).

Sharding does scale, and an earlier version of this page said otherwise. The claim that "1β†’8 shards is flat-to-slightly negative" was a harness artifact: redis-benchmark -t lpush|sadd|hset|zadd drives ONE literal key and -r randomises the element, not the key β€” so eight of twelve command families were being asked to parallelise a single key, and 0.97Γ— was the tautological answer. Re-run on Linux with explicit __rand_int__ keys under a DBSIZE β‰₯ 50000 guard proven to fire, moon's real 8-shard-over-1-shard scaling is 1.42Γ— at p=1, 2.14Γ— at p=8, and 3.79Γ— at p=64 (BENCHMARK.md Β§2.14; Redis io-threads 8-over-1 over the same matrix is 1.19Γ— / 1.23Γ— / 1.08Γ—). A genuinely single-hot-key workload still cannot be sharded β€” that part was always true β€” and the multi-shard win is amplified further by the per-shard WAL and independent event loops parallelizing pipelined and durable work, plus hash-tag co-location:

Multi-shard pipelined throughput vs Redis

At pipeline depth, the per-shard WAL parallelizes work that Redis's single event loop serializes β€” 1.48–1.99Γ— across multi-shard configs. Apple M4 Pro development reference (BENCHMARK.md Β§4.2); not reproduced on Linux.

CPU efficiency: a tie on cost, a win on stabilityΒΆ

On Linux, CPU per operation is a tie. Moon --shards 8 costs 10.55 Β΅s/op against Redis --io-threads 8's 11.33 Β΅s/op (GCE t2a-standard-8, utime+stime from /proc/<pid>/stat, 5 reps, BENCHMARK.md Β§2.14). Moon's 6.9% edge sits inside Redis's own 11.9% run-to-run spread, so it is not a win. What is durable is stability: moon's CPU cost varies 2.0% run to run against Redis's 11.9%, the latter being stopThreadedIOIfNeeded engaging and disengaging io-threads under steady load.

This page previously claimed "23Γ— less CPU" here. That figure has been withdrawn: it came from an Apple M4 Pro table, its CPU column was sampled with ps -o %cpu= (a process-lifetime average, not steady-state load), and its CPU and RPS columns were not taken in the same run. Three independent reasons not to publish it, and it was published under a Linux heading.

The engineering behind the cost that is measured is real, not incidental: lock-free oneshot channels removed ~12% CPU of pthread_mutex contention, the cached shard clock removed ~4% of clock_gettime syscalls, and zero-copy parsing keeps the hot path allocation-free. Idle cost was hunted just as hard β€” the adaptive idle-park and an O(1) page-cache resident counter trimmed idle-shard CPU from 3.1% to ~0.5–0.9%, and the O3 contention governor makes the busy-poll spin self-disengaging on contended cores. Efficiency that shows up on the electricity bill, not just the benchmark. Those component figures are profile deltas from the optimisation work, not a Moon-vs-Redis ratio.

The p=1 conquest β€” Redis's home turfΒΆ

Unpipelined, single-connection request/response was historically Redis's best case and Moon's weakest. The --io-busy-poll-us poll-mode park closed it:

  • 1.19–1.21Γ— Redis on ARM (c4a Axion), 1.65–1.66Γ— on x86 (c3) β€” Linux, same-instance A/B, n=3, on dedicated cores.

Re-measured on ordinary shared-tenant GCE instances (v0.8.7, BENCHMARK.md Β§2.12) the same flag yields 1.06–1.08Γ— on x86 and nothing outside the noise floor on ARM, because the contention governor correctly self-gates there. Quote the dedicated-core figures only with that word attached.

The honest catch was that the spin regressed on shared cores. v0.8.1's O3 governor removed the catch: each shard samples its own involuntary-preemption rate and gates the spin automatically, so the win is safe on any host. One flag β€” --profile standalone (or conf/moon-standalone.conf) β€” now delivers the best single-shard tuning everywhere.

Latency: lower queue depth, unmeasured on LinuxΒΆ

Multi-core parallelism reduces per-shard queue depth, so the median request should wait less. That has not been measured on a Linux host, and the "8–10Γ— lower p50" figure this page used to carry has been withdrawn: it was an Apple M4 Pro development reference published under a Linux heading, and it came from redis-benchmark, a closed-loop tool that under-reports latency once the server saturates. The raw macOS numbers are kept, correctly labelled, in BENCHMARK.md Β§9.1.

Memory: a real win at large values, and half the number we publishedΒΆ

The shape held. The headline number did not. Re-measured on Linux (BENCHMARK.md Β§3 β€” GCE c3-standard-8, x86_64, moon d5f3501b at --shards 1, against Redis 7.4.2 built with jemalloc 5.3.0, fresh server per point, redis-benchmark -r N):

Value size Keys Redis/key Moon/key Result
32 B 63K–632K 123–129 B 143–186 B Moon 11–51% worse
256 B 63K–632K 408–410 B 377–418 B tie (Β±8%)
1 KB 63K–632K 1,380–1,388 B 1,155–1,256 B Moon 9.5–16.4% less
4 KB 63K–632K 5,259–5,266 B 4,352–4,404 B Moon 16.4–17.2% less

This page used to say 27–35%. The number that moved was Redis's.

The old 1M Γ— 1 KB row was Redis 1,571 B / Moon 1,153 B, measured on an Apple M4 Pro. Measured again on Linux against a jemalloc Redis: Redis 1,380 B, Moon 1,172 B. Moon's own number moved 1.6%; Redis's moved 12%. Both Redis binaries already on the benchmark host were libc-malloc builds β€” which inflate Redis RSS and would have biased the comparison in Moon's favour β€” so Redis was rebuilt against jemalloc for the run. The claim was never a Moon regression; it was an unrepresentative baseline. x86_64 only.

Bytes per key vs value size β€” Moon vs Redis

Per-key memory vs value size, measured on the Apple M4 Pro development rig: Redis wins at tiny 32 B values, the lines cross near 256 B, and Moon pulls ahead from there. The Linux measurement above reproduces that shape β€” loss at 32 B, tie at 256 B, win from 1 KB. What it does not reproduce is the size of the 1 KB win: 15% on Linux against the 27% plotted here. (At 4 KB the two agree, 17% vs 15%.) The plotted values are the superseded macOS ones.

CompactKey stores keys up to 23 bytes inline and CompactValue up to 12 bytes; larger payloads use HeapString(Vec<u8>) (48 B overhead) versus Redis's robj + SDS chain (64–80 B). TTL packs as a 4-byte delta inside the entry at zero extra cost, where Redis allocates a separate 24-byte dictEntry per expiring key. That model explains the large-value win and not the 32 B loss, where the same arithmetic predicts Moon should win; no cause is asserted. Tuned jemalloc decay (1 s dirty, background reclaim) returns freed pages to the OS instead of hoarding them.

Baseline RSS is the one that got worse under measurement. This page's ancestor claimed 7.0 MB empty for both servers, measured on macOS. On Linux, across all 12 points of the run above: Redis 7.5–7.7 MB, Moon 12.6–12.9 MB β€” ~1.7Γ— worse. Nobody was watching that number; it is now tracked as #821. On aarch64 at --shards 8, BENCHMARK.md Β§2.14 separately measures idle RSS 1.26Γ— above Redis's, roughly 520 KB per extra shard, four threads per shard.

Durability: making it freeΒΆ

The per-shard WAL turns durable writes from a global serialization point into a parallel one. A 2026-07 write-path campaign closed the remaining gaps:

Policy / workload Before After vs Redis
always SET p16 5.7K (0.12Γ—) 40.1K 0.91Γ—
everysec SET p16 605K 789K 1.32Γ—
everysec SET p1 117K (0.80Γ—) 135K 0.99Γ— parity
Pub/sub fan-out 438 msg/s (with drops) 5.09M msg/s 1.04Γ—, zero drops

Durability write-path campaign β€” before vs after

The write-path campaign turned the fsync tax into near-parity: always p16 climbed 0.12Γ—β†’0.91Γ— of Redis, everysec p16 to a 1.32Γ— win.

Per-batch group commit (one fsync per pipeline batch, not per command), a single coalesced write_all, and a park-free AOF writer that deleted a ~150K/sec futex-wake storm. Durability semantics are unchanged (always = RPO 0, everysec = RPO ≀ 1s) and verified by the SIGKILL crash-recovery matrix.


AI-native β€” in-core, not bolted onΒΆ

Vector search, full-text (BM25), and a property graph are compiled into the same binary, sharing the keyspace and durability β€” no module loader.

  • Vector: 12.7K search QPS at 384d (HNSW + TurboQuant, COSINE), and bulk load β†’ searchable at target recall beats Qdrant 1.6–2.3Γ— via the parallel HNSW build. SQ8 quantization holds ~0.90 R@10 on real MiniLM 384-d embeddings across the full lifecycle (search / merge / persistence).
  • Graph: 23Γ— FalkorDB on bulk insert, ~2.4Γ— Cypher QPS, and 2.78Γ— on point-filter queries after the mutable-property fast path.
  • Tiered engine offload: idle vector-index segments demote HOT β†’ WARM (mmap) β†’ COLD (unloaded stub, reload-on-search), giving memory back β€” βˆ’26% process RSS on a 40K Γ— 768d corpus with identical search results after reload.

AI-native engines vs specialized systems

In-core vector and graph engines measured against the dedicated systems they replace β€” Moon's multiplier on each.


Production hardening β€” the part that makes it a databaseΒΆ

Efficiency means nothing if a kill -9 loses data. The v0.8.0 Storage Kernel made durability a verifiable claim across every plane:

  • Cross-plane kill-9 crash matrix: a 46-cell matrix (KV / vector / graph / FTS / workspaces / MQ / temporal / txn Γ— persistence mode Γ— disk-offload Γ— shard count), all green β€” wired into scheduled CI (nightly full matrix + weekly ITERS=20 soak).
  • 10Γ— RAM datasets under disk offload: 2.6 GB of 10 KB values against a 256 MB cap β€” truthful used_memory at 1.00Γ— cap steady-state, spill files cut from ~236,000 (one per key) to 840 (~280Γ— fewer), worst cold-GET tail during an active spill flood 1,910 ms β†’ 205 ms (9.3Γ— better), and 500/500 kill-9 integrity.
  • Replication GA: real async replication with WAIT/ACK across all six data planes, validated by a 24-hour kill-9 soak β€” 114 alternating master/replica kills, 82,044 WAIT-acked writes, zero lost.

The discipline: efficiency is what survives measurementΒΆ

The reason the numbers above hold up is a rule the project keeps: no measured win β†’ no hot-path change. That rule is only credible because of what it rejected. A representative sample of ideas that sounded fast and were thrown out on the evidence:

  • Value-line prefetch β€” measured "3–6% faster" until disassembly showed the benchmark's black_box never dereferenced the value; a forced-load harness showed it was 1–2% slower. Cut, and a cache-exceeding probe bench shipped so the trap can't recur.
  • A lock-free Notify replacing the flume channel β€” fully validated (loom lost-wake model, consistency suites, both-impl A/B) but neutral on the wall clock, because the wake is syscall-dominated. The flume channel stayed.
  • Transparent huge pages by default β€” real PMU gains (+12–24% GET), but a 45-minute RSS-drift soak showed idle khugepaged re-collapse drifts RSS +31%. Kept permanently opt-in, not the default.
  • Per-segment ef/√G beam splitting for vector search β€” silently cost recall (0.9915 β†’ 0.9295). Rejected; a recall-certified adaptive scheme shipped instead.

None of these are failures β€” they're the cost of knowing the wins are real. The efficiency Moon ships is the efficiency that survived.

Backing that discipline: ~4,980 tests, 11+ cargo-fuzz targets, loom models for every lock-free structure, unsafe/unwrap ratchets that block new unannotated uses, an RSS-regression CI gate, and the crash matrix above.


Where it's goingΒΆ

Moon is on a path to a v1.0 GA with every Production Contract box ticked. Next up: horizontal scale β€” cluster-on-monoio hardening and multi-shard replicas β€” then the enterprise foundation. The method won't change: measure, ship what wins, and say so honestly when it doesn't.

See the live numbers on the Benchmarks page, the full per-release history in RELEASES.md, and the GA scorecard in the Production Contract.