Moon Storage Format v1ΒΆ
Audience: operators and tooling authors. This document is the public commitment Moon makes about on-disk formats. For exact byte layouts the authoritative source is the Rust constants and module-level rustdoc in
src/persistence/; this doc summarizes the guarantees, not the bits.
1. What "Storage Format v1" MeansΒΆ
Starting with v0.2.0, Moon's on-disk file formats are versioned as a single umbrella tag, storage format v1. The umbrella covers three sub-formats that together make a Moon shard durable:
| Sub-format | File(s) | Source of truth | Tag |
|---|---|---|---|
| WAL v3 | <persistence_dir>/wal/<shard>/<segment>.wal |
src/persistence/wal_v3/{record.rs,segment.rs} |
RRDWAL magic, version byte = 3 |
| RDB v2 (per-shard snapshot) | <persistence_dir>/<shard>.rdb |
src/persistence/snapshot.rs |
RRDSHARD magic, version byte = 2 |
| AOF multi-part | <persistence_dir>/appendonlydir/{moon.aof.manifest, *.aof, *.rdb} |
src/persistence/aof_manifest.rs, src/persistence/aof.rs |
manifest framing |
The version byte inside each file is the canonical machine-readable marker. The "storage format v1" umbrella is the human-readable, release- level commitment that all three of those byte-level versions are guaranteed to be supported.
2. The Compatibility GuaranteeΒΆ
For the duration of the Moon v0.2.x minor series (β₯ 18 months of LTS β see
docs/SUPPORT.md):
- Forward read. Every v0.2.x release reads files written by every earlier v0.2.x release.
- Reverse read on retirement. When v0.2 enters maintenance and v0.3 becomes the active line, v0.3 reads v0.2 files using v0.3's recovery path. v0.3 may, however, write files only v0.3+ can read.
- Crash recovery. Every file format includes per-record or per-page CRC32C / CRC32 checksums; corruption is detected and quarantined rather than silently propagated.
- No silent format bumps. Any change to the WAL v3 record byte layout, the RDB v2 preamble, or the AOF manifest framing requires the parent version byte to increment. The umbrella tag becomes "storage format v2" simultaneously.
- Migration. If a future Moon release retires a format inside the v1
umbrella, an automatic in-place rewrite (or an offline
moon migratetool) is guaranteed before the format is dropped.
3. Sub-format SummaryΒΆ
3.1 WAL v3 β Per-Shard Write-Ahead LogΒΆ
Authoritative source: src/persistence/wal_v3/.
- Segment header (64 bytes):
RRDWALmagic + version=3 + flags + shard id + epoch +redo_lsn+base_lsn+ segment size + reserved (byte layout insegment.rs).base_lsnis the LSN of the segment's first record (the next LSN to be assigned, while the segment holds none). Readers β the WAL recyclers and CDC.READ's seek β rely only on the bound it implies: every record of every earlier segment has an LSN< base_lsn. A header that overstates it therefore makes them conservative (a segment recycled later, a seek that starts one segment early), never wrong. Such headers exist only in WAL directories written by the unreleased moon#1188 build before the moon#1221 review fix, which stamped a segment opened after an off-loop rotation fsync with the LSN after the records appended while that fsync was in flight. - Record (variable): little-endian, self-describing.
Offset Size Field 0 4 record_len (u32 LE) 4 8 lsn (u64 LE) β monotonic log sequence number 12 1 record_type (u8) β Command / FullPageImage / Checkpoint / VectorUpsert / VectorDelete / β¦ 13 1 flags (u8) β bit 0 = LZ4-compressed payload (FPI) 14 2 reserved (zeroes) 16 N payload 16+N 4 crc32c (u32 LE) β over bytes [4 .. 16+N] - FPI compression: payloads β₯
FPI_COMPRESS_THRESHOLD(256 B) are LZ4-compressed; flag bit 0 set. - Atomicity: segment rotation is fsync-fenced; partial trailing records are detected by CRC and truncated on recovery.
3.2 RDB v2 β Per-Shard Point-in-Time SnapshotΒΆ
Authoritative source: src/persistence/snapshot.rs.
- Preamble (35 bytes):
RRDSHARDmagic + version=2 + shard_id (u16 LE) + epoch (u64 LE) + last_lsn (u64 LE) + created_at_unix_ms (u64 LE). - Body: value-encoded keys + entries (custom RDB-style, supports listpack / intset / hashtable / sorted-set / stream encodings).
- Per-field hash TTL trailer (v2-only): every
TYPE_HASHbody is followed by[ttl_count u32][field_len varint | field_bytes | ttl_ms u64]*.ttl_count = 0for plain hashes (no per-field TTL). v1 readers stop after the hash body and reconstruct a plainHash; v2 readers consume the trailer to rebuildHashWithTtl. Authoritative encoder/decoder:src/persistence/rdb.rsβ search forhas_hash_ttl_trailer. - Trailer:
0xFFEOF byte + CRC32 over the whole file. - Cold-graves trailer (optional, moon#1281): between the
0xFFEOF byte and the global CRC32 a snapshot taken WITHOUT an AOF may carry the spill-file slots that were already dead when it started:"MCGV" | ver u8 = 1 | file_count u32 | { file_id u64 | slot_count u32 | { page_idx u32 | slot_idx u16 }* }* | crc32 u32(all little-endian; the trailer's CRC covers"MCGV"through the last slot). Boot of a no-AOF process drops exactly those slots before rebuilding the cold index, so a cold key deleted before a successful snapshot stays deleted after a crash. Every reader stops at the EOF byte, so readers that predate the trailer load the file unchanged and ignore it (and resurrect those keys, as before) β the version byte is NOT bumped. Absent trailer = no graves. A trailer that fails its own checks is logged and ignored; the global CRC still guards the keys. Authoritative codec:src/persistence/snapshot/cold_graves.rs(fuzz targetsnapshot_cold_graves). - PITR: the embedded
last_lsnties each snapshot to the WAL position it shadows; replay resumes atlast_lsn + 1. - Forkless: snapshots are produced by cooperative segment iteration with per-snapshot overflow buffers β no
fork(), no COW RSS spike.
A v0.2.x reader still accepts the legacy v1 preamble (19 bytes, no
last_lsn / created_at_unix_ms) for snapshots produced by v0.1.x.
3.3 AOF Multi-Part β Append-Only Log BundleΒΆ
Authoritative source: src/persistence/aof.rs, src/persistence/aof_manifest.rs.
- Directory:
<persistence_dir>/appendonlydir/ - Manifest:
moon.aof.manifestβ versioned listing of (base RDB, sequence of incremental AOF segments). - Base file: optional RDB-format snapshot of state at last rewrite.
- Incremental files: RESP-encoded write commands appended since last rewrite.
- Legacy single-file
appendonly.aofis recognized on first boot, captured as seq-1 of the multi-part structure, and the legacy file is renamed toappendonly.aof.legacy. Older v0.1.x AOF files are read once, never written. - Per-shard incremental files (
--shardsβ₯ 2) frame every record as[u64 lsn LE][u32 len LE][RESP command]; the multi-part top-level incr and the flatappendonly.aofhold bare RESP.
Replay-only records (MOON.* pseudo-commands)ΒΆ
Besides client write commands, the writer interleaves a few records that only
a replay acts on. Each is an ordinary RESP array (never a #β¦ annotation
line, which moon's parser reads as a malformed RESP3 boolean, and never a new
WAL record type byte), written with lsn = 0 in the framed layout so it never
moves the replication offset a replay recovers. A client that sends one gets
ERR unknown command. Authoritative source:
src/persistence/replay/pseudo.rs and src/persistence/cold_records.rs.
| Record | Written | Replay effect |
|---|---|---|
SELECT <db> |
before a record whose execution db differs from the stream's | switches the db the following records apply to |
MOON.COLDCUT <watermark> |
first record of every generation (boot head, rewrite head) | opens the cold-tier replay gate: cold files with file_id < watermark are a valid base for what follows |
MOON.TS <ms> (moon#1283) |
right after MOON.COLDCUT in every generation head, before the first record a writer appends to a file it reopened, then before any record whose shard clock differs from the last MOON.TS in the stream |
sets the expiry-judgment clock (see below) |
MOON.TS <ms> CLOSE (moon#1283) |
last record a writer appends when it stops in order (SHUTDOWN, SIGTERM) with the file still the live incr, made durable by its final sync | marks a clean close: what follows it up to the next stamp is another binary's (see below) |
MOON.SPILLED <file_id> key⦠|
when a spill publishes keys into the cold index | demotes replay-built hot copies of those keys to their cold entries |
MOON.TS <ms> carries the shard's cached clock in unix milliseconds β the
clock the command after it judged key expiry with β read in the same
synchronous section as the mutation. It is at most one record per 1 ms clock
tick in which the shard logged a write (β€ 44 bytes each), plus one per record
whose producer parked between its mutation and its enqueue. It is AOF-only:
the replication stream never carries it.
A replay judges every record by the last MOON.TS read in the current
file (not a running maximum: a parked producer's record carries an older
stamp than the records it lands after), except a foreign segment (below).
Until a file's first MOON.TS β an older binary's log, or the stamp-less
prefix of a file an older binary started β it falls back to the file's
modification time capped at the wall clock (moon#1277), exactly as before. A
MOON.TS of 0, past the year 9999, or malformed is skipped. Clock stamps are
observations, not data: a replay that skips a block of data records still
applies the stamps inside it.
Foreign segments (positional). A writer that reopens a file writes a
MOON.TS before its first append, and one that stops in order ends the file
with MOON.TS <ms> CLOSE. So the records after a CLOSE and before the next
stamp (plain or CLOSE) were not written by a binary that knows this rule:
they are a foreign segment β an older binary's appends after a downgrade.
The replay judges them by that next stamp (the later session's first stamp,
never earlier than their own write time; late is how the older binary itself
judges them, by the mtime), or, when the segment runs to the end of the file,
by max(<ms> of the CLOSE, the file's mtime capped at the wall clock). The
boot that meets such a segment at the end of the file writes that judgment as
the first stamp of its own appends, so every later boot judges the segment
exactly as it did. The rule depends only on record positions: no rewrite is
needed, a file's mtime moved forward (touch, a cp restore) re-judges
nothing, and a CLOSE immediately followed by a stamp (a clean restart) is
an empty segment.
Compatibility. Adding MOON.TS changes no byte layout covered by Β§2
rule 4 (WAL v3 records, the RDB v2 preamble, the manifest framing), so it is
not a storage-format bump:
- Upgrade: a log without stamps replays exactly as before (mtime judgment).
- Downgrade: a binary that predates MOON.TS sends it β and MOON.TS <ms>
CLOSE β to command dispatch, gets "unknown command", counts it as an
unhandled record, and replays every data record around it with its own
(mtime) expiry judgment. The boot does not fail. (A build that knows only
the one-argument MOON.TS skips the CLOSE form as a malformed stamp.)
An older binary's own BGREWRITEAOF drops every marker, which is fine.
- Downgrade, then re-upgrade: the older binary appends to the same file
with no stamps. After a CLEAN stop of the newer binary those records form
a foreign segment and are judged as above; keys the older binary saw expire
and restarted (INCR onto an expired counter, with no DEL logged first)
come back as they were live, on the re-upgrade boot and on every boot after
it, however the re-upgraded server is stopped later.
Not protected: a downgrade after an UNCLEAN stop of the newer binary
(kill -9, OOM kill, crash, power loss) β there is no CLOSE, so the older
binary's records follow the newer binary's last stamp, hours or days stale,
and those keys replay onto their old value and deadline and are lost.
Procedure: stop the newer binary cleanly (SHUTDOWN, SIGTERM) before
downgrading; after it crashed, start it once and stop it cleanly first.
Otherwise, as the older binary's last action, run BGREWRITEAOF and stop
it once the rewrite completed: its history is then in the new base, an
image that judges nothing, and only what it appended after the rewrite
replays as a stamp-less prefix (mtime judgment, as on any upgrade). Also
not protected:
a newer binary whose writer could not finish its stop (abandoned after the
stop timeout, a torn write it latched, a failed marker append, a panic) β
its log shows no "AOF writers drained and synced" line: it shows
"AOF writers stopped ... WITHOUT their clean-close marker" (or the
abandoned-writer error) instead.
- Redis: a redis server cannot load a moon AOF regardless (MOON.COLDCUT
is already an unknown command there).
The per-shard WAL v3 KV log (--appendonly no) does not carry MOON.TS
yet: its replay keeps the mtime judgment of its newest segment.
4. Configuration SurfaceΒΆ
A --storage-format <v1> CLI flag is reserved and will be introduced in a
follow-up PR (issue #103 second checkbox). It accepts only v1 today; the
flag exists so that future Moon releases can offer v2-or-newer write
opt-in while still defaulting to v1 for compatibility.
Operators do not need to set this flag in v0.2.x. It is documented here so that automation written today against v0.2 continues to be correct when storage format v2 is introduced.
5. Migration PolicyΒΆ
When the umbrella tag bumps to storage format v2 (no earlier than v0.3), the release notes will contain:
- The list of byte-level format changes (which sub-formats changed; which stayed identical).
- An automatic in-place rewrite path β Moon detects v1 files at startup,
rewrites them to v2 atomically with a
*.v1.bakretained, and proceeds. - A
moon migrate --storage-format v2 --dry-runoffline tool for operators who prefer to validate before bumping production. - A minimum-required Moon version for the rewrite to be safe.
No v1 file will ever be silently rewritten to v2 without operator opt-in once v2 is the default.
6. Backwards-Compatibility Promise (Summary)ΒΆ
| Promise | Scope |
|---|---|
| v0.2.x reads files written by all v0.2.x releases | hard guarantee |
| v0.3.x reads v0.2.x files (one-way) | hard guarantee |
| v0.2.x reads v0.1.x files (one-way, best-effort) | currently implemented for RDB v1 + legacy appendonly.aof; not a hard guarantee |
| Major version bump never requires manual format conversion | hard guarantee β in-place rewrite is automated |
| Crash recovery from any v1 file at any LSN | hard guarantee β every record / page is CRC-protected |
7. Cross-ReferencesΒΆ
- WAL v3 record layout:
src/persistence/wal_v3/record.rs - WAL v3 segment header:
src/persistence/wal_v3/segment.rs - RDB v2 preamble:
src/persistence/snapshot.rs - AOF manifest framing:
src/persistence/aof_manifest.rs - Support / LTS policy:
docs/SUPPORT.md(to be added in #104) - Security disclosure:
SECURITY.md - Operator capacity planning:
docs/OPERATOR-GUIDE.md - v0.3.0 roadmap (this file is a v0.2.0 deliverable):
.planning/milestones/v0.3.0-ROADMAP.md