BenchmarksΒΆ
Read the provenance label on every table
CLAUDE.md requires every published benchmark number to come from a Linux
host. Only the tables on this page explicitly marked Linux satisfy that.
Every table marked macOS dev reference was measured on an Apple M4 Pro
(12 cores, 24 GB) and is kept as a development record only β it must not be
quoted as a production result, and its ratios are not assumed to carry
over to Linux. Two of them demonstrably did not: the macOS per-key memory
tables claimed 27β35% less at β₯1 KB values, and the Linux re-measurement of
2026-09-08 (BENCHMARK.md Β§3)
puts it at 16% at 1 KB; and every non-GET/SET command family runs
0.45β0.87Γ Redis at pβ₯8
(BENCHMARK.md Β§2).
All runs co-locate client and server using redis-benchmark β a closed-loop
tool β with a fresh server instance per memory data point. The canonical
report is BENCHMARK.md,
re-measured 2026-09-08; every run it supersedes is preserved verbatim in the
benchmark archive,
which is where the Β§2.x / Β§3.x / Β§7.x / Β§10.x section numbers now live.
Executive summary β Linux onlyΒΆ
Every row below was measured on Linux. Rows dated 2026-09-08 come from the current report; rows with an earlier date are the most recent measurement of that subsystem and were not re-measured on that date. Where a condition is not recorded in the source report, this table says so rather than filling it in.
| Metric | Moon vs Redis | Conditions |
|---|---|---|
| GET, p=64 | 1.62Γ x86 / 1.60Γ ARM | 2026-09-08, moon ae6cd003 (v0.8.9), GCE c3-standard-8 x86_64 + t2a-standard-8 aarch64, Redis 7.0.15, --shards 1, c=50, keys over -r 100000, n=5 interleaved β BENCHMARK.md Β§2 |
| SET, p=64 | 1.32Γ x86 / 1.55Γ ARM | same run β Β§2 |
| GET / SET, p=8 | 1.20Γ / 1.20Γ x86 Β· 1.18Γ / 1.19Γ ARM | same run β Β§2 |
| Every other command family, p=8 | 0.62β0.81Γ x86 / 0.64β0.77Γ ARM | INCR, LPUSH, SADD, SPOP, HSET, ZADD β same run β Β§2 |
| Every other command family, p=64 | 0.45β0.75Γ x86 / 0.53β0.87Γ ARM | same run β Β§2 |
| All families, p=1 | 0.89β1.08Γ x86 / 0.77β1.05Γ ARM | roughly a tie; ARM container families are the weak end β same run β Β§2 |
| Peak GET, absolute | 2.97M ops/s x86 vs Redis 1.83M Β· 1.53M ARM vs 0.95M | same run β Β§2 |
| Peak SET, absolute | 2.01M ops/s x86 vs Redis 1.52M Β· 1.32M ARM vs 0.85M | same run β Β§2 |
| vs v0.8.7, non-inlined families | +5 to +23% x86 / +5 to +18% ARM (raw ops/s) | v0.8.7 d63ffcd8 rebuilt on the same hosts, same day, same harness; Redis control drifted a median β0.4% (x86) / +0.3% (ARM) β Β§2 |
| Memory per key, 8 B β 1 KB | 0.78β0.92Γ x86 / 0.84β0.92Γ ARM (Moon uses 8β22% less on x86, 8β16% on ARM) | 2026-09-08, --shards 1, Redis 7.0.15/jemalloc, -r 200000, fresh server per point, n=3 (n=2 at 16/32/48/128 B), arithmetic-floor guard β Β§3 |
Idle RSS, --shards 1 |
tie β 0.99Γ x86, 0.96Γ ARM | 13.01 vs 13.19 MB (x86), 11.43 vs 11.93 MB (ARM) β same run β Β§3 |
| Peak GET (v0.1.6, absolute) | 5.11M ops/sec (1.72Γ) | GCloud c3-standard-8 x86_64, p=64. Redis io-threads and payload size not recorded, and the run drove a single hot key rather than -r N β archive Β§2.1, superseded |
Memory, 64 B values (aarch64, --shards 8) |
1.16Γ worse than Redis | 2026-09-01, GCE t2a-standard-8, Redis 7.0.15 --io-threads 8 --io-threads-do-reads yes, -r 200000 β archive Β§2.14; a different configuration from the --shards 1 rows above |
| Idle RSS, 8 shards (aarch64) | 1.26Γ worse than Redis | same run β archive Β§2.14 |
| CPU per operation | tie | 10.55 Β΅s vs 11.33 Β΅s, inside Redis's own 11.9% spread β 2026-09-01, archive Β§2.14 |
| Shard scaling (s8/s1) | 1.42Γ / 2.14Γ / 3.79Γ | p=1 / p=8 / p=64, explicitly-keyed families β 2026-09-01, archive Β§2.14 |
moon s8 vs Redis --io-threads 8 |
1.26β1.32Γ (p=8), 2.74β2.91Γ (p=64) | 2026-09-01, both arches β archive Β§2.14 |
AOF everysec SET p=16 |
1.32Γ Redis | 2026-07-08, GCE c3-standard-8, Redis 7.0.15, --shards 2 β archive Β§7.3 |
AOF always SET p=16 |
0.91Γ Redis | same run β archive Β§7.3 |
| Vector search (384d) | 12.7K QPS | 2026-04-15, GCloud c3-standard-8 x86_64, HNSW + TurboQuant 8-bit, COSINE, 50K vectors, K=10 β archive Β§10.1, and superseded there by Β§10.5's concurrent-client reframing. The current vs-competitor figures are the 2026-07-08 Qdrant runs (archive Β§10.9). |
| Data correctness | 132/132 tests | All types, 1/4/12 shards |
Scope of the pipelined GET/SET win
GET and plain SET key value are the two commands Moon serves from an inline
byte path that bypasses frame construction and the dispatch table. Those two
win at depth β 1.62Γ / 1.32Γ on x86 and 1.60Γ / 1.55Γ on ARM at p=64.
Every other family measured (INCR, LPUSH, SADD, SPOP, HSET, ZADD) loses at
pβ₯8, by 13β55%, on both architectures. The boundary is the fast path, not
the engine: measured on v0.8.7, SET k v ran 2.08Γ Redis while
SET k v EX 100 β same work, one disqualifying option β ran 0.87Γ
(archive Β§2.12). Any summary quoting the GET number alone describes a
two-command fast path, not the server. Full matrix:
BENCHMARK.md Β§2.
Memory efficiencyΒΆ
Per-key memory (Linux, --shards 1, both arches) β measured 2026-09-08ΒΆ
Hosts: GCE c3-standard-8 (Xeon 8481C, x86_64) and t2a-standard-8
(Neoverse-N1, aarch64), 8 vCPU, Ubuntu, kernel 6.17.
Moon: ae6cd003 (v0.8.9), --shards 1 --appendonly no --disk-offload disable.
Oracle: Redis 7.0.15 (jemalloc), --save "" --appendonly no.
Method: scripts/bench-ab-memory.sh, fresh server instance per data point,
-r 200000 distinct keys (DBSIZE β 173,000), per-key =
(loaded RSS β idle RSS) / DBSIZE, RSS from /proc/<pid>/status, n=3. Every row
is checked against its arithmetic floor (key(16) + value + 24) and a leg
reporting below its own floor aborts, because that is physically impossible.
Full table and method: BENCHMARK.md Β§3.
| Value size | Moon x86 | Redis x86 | Moon / Redis | Moon ARM | Redis ARM | Moon / Redis |
|---|---|---|---|---|---|---|
| 8 B | 97.9 B | 125.1 B | 0.78Γ | 97.5 B | 113.0 B | 0.86Γ |
| 16 B β | 113.3 B | 136.1 B | 0.83Γ | 113.3 B | 128.7 B | 0.88Γ |
| 32 B β | 130.1 B | 158.3 B | 0.82Γ | 129.7 B | 145.6 B | 0.89Γ |
| 48 B β | 146.1 B | 174.9 B | 0.84Γ | 145.7 B | 161.8 B | 0.90Γ |
| 64 B | 163.9 B | 178.8 B | 0.92Γ | 162.8 B | 177.0 B | 0.92Γ |
| 128 B β | 230.0 B | 259.8 B | 0.89Γ | 228.9 B | 257.3 B | 0.89Γ |
| 256 B | 362.2 B | 421.7 B | 0.86Γ | 361.1 B | 419.0 B | 0.86Γ |
| 1,024 B | 1,164.6 B | 1,393.3 B | 0.84Γ | 1,161.9 B | 1,386.1 B | 0.84Γ |
| idle RSS | 13.01 MB | 13.19 MB | 0.99Γ (tie) | 11.43 MB | 11.93 MB | 0.96Γ |
β = n=2 repetitions; every other row is n=3.
Read it as: Moon uses less memory per key at every size measured, on both
architectures β 8β22% on x86 and 8β16% on ARM, weakest at 64 B on both;
strongest at 8 B on x86 and at 1 KB on ARM, where
CompactValue inlines the value into the entry (values β€12 B). The result does
not depend on the idle-RSS subtraction: Moon's absolute loaded RSS is lower at
every size too (1 KB values, x86: 210 MB vs 249 MB). Scope: --shards 1, string
values, one key shape; 4 KB was not part of this run.
Two claims from the 2026-09-04 run did not reproduce β both stay on the record
The 2026-09-04 Linux run reported Moon 11β51% worse at 32 B on x86
--shards 1 (archive Β§3.2), and an empty-server RSS of 12.6β12.9 MB against
Redis 7.5β7.7 MB β Moon 1.7Γ worse (archive Β§3.1). 16/32/48/128 B were
measured on 2026-09-08 specifically to test the first point: Moon is 0.82Γ
at 32 B on x86 and 0.89Γ on ARM β an 18% and 11% win. On idle RSS the two servers tie. Neither older figure is
being called wrong here β the newer measurement did not reproduce it, and the
conditions differ: the 2026-09-04 oracle was Redis 7.4.2, this one is
Redis 7.0.15. On idle RSS the disagreement is on the Redis side
(7.5β7.7 MB then vs 13.19 MB now, same host class), which is what
#821 tracks. Both runs are on
the record; the older one is preserved in the
archive.
Scope β every row above is --shards 1
At --shards 8 on aarch64, archive Β§2.14 (2026-09-01) separately measures a
16% loss at 64 B and a 26% idle-RSS loss, because the default config
spawns four threads per shard. Nothing above contradicts that; it is a
different configuration, and it has not been re-measured on this tree.
Tip
Moon's large-value advantage comes from HeapString(Vec<u8>) (48 bytes
overhead) vs Redis's robj + SDS chain (~64β80 bytes overhead); its
small-value advantage comes from CompactValue inlining values β€12 B into the
entry. The TTL-overhead claim that used to sit here is unverified: the
harness section that would measure it omits redis-benchmark -r, so it loads
one key.
Baseline RSS (empty server) β LinuxΒΆ
| Server | 2026-09-08, x86_64 | 2026-09-08, aarch64 | Previously published |
|---|---|---|---|
| Redis (jemalloc) | 13.19 MB (7.0.15) | 11.93 MB (7.0.15) | 7.5β7.7 MB (Redis 7.4.2, 2026-09-04) |
| Moon (1 shard) | 13.01 MB | 11.43 MB | 12.6β12.9 MB (2026-09-04) |
| Moon (12 shards) | not measured on Linux | not measured on Linux | 15.7 MB (unverified / stale, macOS) |
Moon's own idle figure is stable across the two Linux runs (12.6β12.9 MB then, 13.01 MB now); Redis's is not, and that is the open question in #821. The ARM pair independently reproduces the archive's Β§2.14 sidebar (11.43 vs 11.93 MB here; 11.68 vs 11.96 MB there, a different date). Do not quote the retired "identical 7.0 MB" row β it was an Apple M4 Pro development reference.
Superseded: per-key memory (Linux x86_64, 2026-09-04)ΒΆ
Kept as a record β superseded by the 2026-09-08 run above
Measured against Redis 7.4.2 / jemalloc 5.3.0 on GCE c3-standard-8, moon
d5f3501b, --shards 1, scripts/bench-resources.sh, redis-benchmark -r N,
one point per cell (no repetitions, so no run-to-run spread is reported). Its
32 B and idle-RSS rows did not reproduce on 2026-09-08; its 4 KB row was not
re-measured. Full twelve-point table:
benchmark archive Β§3.2.
| Value size | Keys | Redis 7.4.2/key | Moon/key | Moon / Redis |
|---|---|---|---|---|
| 32 B | 63Kβ632K | 123β129 B | 143β186 B | 1.11β1.51Γ |
| 256 B | 63Kβ632K | 408β410 B | 377β418 B | 0.92β1.02Γ |
| 1 KB | 63Kβ632K | 1,380β1,388 B | 1,155β1,256 B | 0.84β0.90Γ |
| 4 KB | 63Kβ632K | 5,259β5,266 B | 4,352β4,404 B | 0.83β0.84Γ |
| Empty-server RSS | β | 7.5β7.7 MB | 12.6β12.9 MB | 1.7Γ |
The retired '27β35% less memory' claim β and Moon is not what changed
The old published 1M Γ 1 KB row was Redis 1,571 B / Moon 1,153 B. Re-measured on Linux on 2026-09-04: Redis 1,380 B / Moon 1,172 B. Moon's own figure moved 1.6%; the oracle moved 12%. The old claim was inflated by a Redis baseline measured on macOS and/or without jemalloc, not by anything Moon did β both Redis builds already on that host were libc-malloc, which inflates Redis RSS, so Redis was rebuilt against jemalloc for the run.
Superseded: per-key memory (macOS dev reference, 1-shard)ΒΆ
Kept as a development record only β do not quote
Measured on an Apple M4 Pro (12 cores, 24 GB). The Linux tables above supersede it. Its Redis column in particular is the source of the retired 27β35% claim.
| Value size | Keys | Redis/key | Moon/key | Winner | Ratio |
|---|---|---|---|---|---|
| 32 B | ~63K | 118 B | 147 B | Redis | 0.80x |
| 256 B | ~63K | 412 B | 407 B | Tied | 1.01x |
| 1,024 B | ~63K | 1,879 B | 1,207 B | Moon | 1.56x |
| 4,096 B | ~63K | 5,131 B | 4,352 B | Moon | 1.18x |
At 1M keys:
| Value size | Redis RSS | Moon RSS | Redis/key | Moon/key | Winner |
|---|---|---|---|---|---|
| 32 B | 78.2 MB | 95.8 MB | 118 B | 147 B | Redis |
| 256 B | 231.5 MB | 234.4 MB | 372 B | 376 B | Tied |
| 1,024 B | 954.2 MB | 703.0 MB | 1,571 B | 1,153 B | Moon |
ThroughputΒΆ
macOS dev reference
Both tables in this subsection were measured on an Apple M4 Pro, not on
Linux. The Linux throughput matrix is the executive summary above: the
current --shards 1 matrix is BENCHMARK.md Β§2, and the multi-shard /
io-threads rows are the archive's Β§2.14 (2026-09-01).
Single-shard SET (macOS dev reference, pipeline=16, 50 clients)ΒΆ
| Value size | Redis SET/s | Moon SET/s | Ratio |
|---|---|---|---|
| 32 B | 1,298,701 | 1,754,386 | 1.35x |
| 256 B | 1,219,512 | 1,639,344 | 1.34x |
| 1,024 B | 1,010,101 | 1,030,928 | 1.02x |
| 4,096 B | 540,541 | 571,429 | 1.06x |
Multi-shard peak throughput (macOS dev reference)ΒΆ
| Config | Moon | Redis | Ratio |
|---|---|---|---|
| 8-shard GET p=16 c=50 | 2.60M | 1.41M | 1.84x |
| 8-shard SET p=16 c=50 | 2.52M | 1.27M | 1.99x |
| 4-shard GET p=64 c=50 | 3.79M | 2.41M | 1.57x |
CPU efficiencyΒΆ
On Linux, CPU per operation is a tie. The archive's Β§2.14 (measured
2026-09-01, and not re-measured on 2026-09-08) puts moon --shards 8 at
10.55 Β΅s/op against 11.33 Β΅s/op for Redis --io-threads 8
(GCE t2a-standard-8, utime+stime from /proc/<pid>/stat, 5 reps). Moon's 6.9%
edge sits inside Redis's own 11.9% run-to-run spread, so it is not a win. The
durable difference is stability: moon's CPU cost varies 2.0% run to run against
Redis's 11.9%.
The "45Γ / 23Γ better CPU" figure this page used to publish has been removed.
It was derived from an Apple M4 Pro table whose CPU column was sampled with
ps -o %cpu= β a process-lifetime average, not steady-state load β and whose CPU
and RPS columns were not taken in the same run. The underlying macOS table is kept
in the benchmark archive Β§5.1 as a development record, with no ratio
derived from it.
Persistence (AOF) performanceΒΆ
macOS dev reference β the Linux figures are the campaign table below
This first table was measured on an Apple M4 Pro. The Linux write-path
measurement (GCE c3-standard-8, Redis 7.0.15, --shards 2, 3 alternated
reps) is the max-durability table that follows it: 1.32Γ for everysec
SET p=16 and 0.91Γ for always SET p=16.
| Pipeline | Moon SET/s | vs Redis (no AOF) | vs Redis (AOF everysec) |
|---|---|---|---|
| p=1 | 146K | 0.95x | 0.95x |
| p=8 | 1,117K | 1.68x | 1.68x |
| p=16 | 1,887K | 1.90x | 2.21x |
| p=64 | 2,778K | 1.80x | 2.75x |
Note
Moon's per-shard WAL avoids the global serialization point that Redis's single AOF file introduces. The advantage grows with pipeline depth because per-shard WAL scales linearly with shards.
Max-durability (appendfsync always) and everysec write pathΒΆ
A 2026-07 write-path campaign (GCE c3-standard-8, Redis 7.0.15, --shards 2, 3 alternated reps) closed Moon's remaining AOF-on deficits:
| Policy / workload | Before | After | vs Redis |
|---|---|---|---|
always SET p16 |
5.7K (0.12x) | 40.1K | 0.91x |
always SET p1 |
~3.2K | 3.1K | parity (fsync-device-bound) |
everysec SET p16 |
605K | 789K | 1.32x |
everysec SET p1 |
117K (0.80x) | 135K | 0.99x parity |
| Pub/sub fan-out delivery | 438 msg/s (drops) | 5.09M msg/s | 1.04x, 0 drops |
The wins come from per-batch group commit (one fsync per pipeline batch, not per command), a single coalesced write_all per batch, a park-free AOF-writer poll that removes a ~150K/sec futex-wake storm on the shard thread under everysec, and coalesced pub/sub delivery writes. Durability is unchanged (always = RPO 0, everysec = RPO β€ 1s), verified by the SIGKILL crash-recovery matrix. Measured 2026-07-08 and not re-measured on 2026-09-08 β it is carried forward with its date in BENCHMARK.md Β§4, and the full detail is in the benchmark archive Β§7.3.
LatencyΒΆ
No Moon-vs-Redis latency comparison has been measured on Linux. The 8-shard p50
figure this page used to publish was an Apple M4 Pro development reference, taken
with redis-benchmark β a closed-loop tool, which under-reports latency once the
server saturates (see Coordinated Omission). It is retained with
its provenance in the benchmark archive Β§9.1 and is not republished
here as a result.
Architecturally, multi-core parallelism reduces per-shard queue depth, so the median request should wait less. That expectation has not been confirmed on a Linux host.
Production workload patterns (macOS dev reference)ΒΆ
Warning
Measured on an Apple M4 Pro, not on Linux, and not reproduced there. On Linux, the 2026-09-08 matrix (BENCHMARK.md Β§2) measures several of these command families (INCR, LPUSH, SADD, HSET, ZADD) at 0.45β0.87Γ Redis at pβ₯8, so these ratios should not be read as production expectations.
| Scenario | Description | Moon vs Redis |
|---|---|---|
| Session store | 80% GET / 15% SET, 512B values | 1.24x |
| Rate limiting | INCR with 100-200 clients | 1.15x |
| Leaderboard | ZADD + ZRANGEBYSCORE | 1.06-1.25x |
| App caching | 1KB-4KB values, MSET batch | 1.10-1.27x |
| Job queue | LPUSH/RPOP producer-consumer | 1.06x |
| User profiles | HSET, HGET | 1.10x |
How to reproduceΒΆ
# Build with native CPU optimizations
RUSTFLAGS="-C target-cpu=native" cargo build --release
# Memory and CPU benchmark
./scripts/bench-resources.sh --shards 1
# Production workload scenarios
./scripts/bench-production.sh --shards 1
# Multi-shard scaling
./scripts/bench-production.sh --shards 4
./scripts/bench-production.sh --shards 8
# Data consistency tests
./scripts/test-consistency.sh --shards 1
./scripts/test-consistency.sh --shards 4
Warning
Co-located benchmarks (client and server on the same machine) are conservative. Separate-machine benchmarks with dedicated NICs show higher throughput. Always use redis-benchmark -r <num_keys> to generate unique keys. Use redis-benchmark 8.x which correctly handles \r in progress output.