Skip to content

Changelog

All notable changes to this project will be documented in this file. The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.

[Unreleased]

Added

  • CLIENT TRACKINGINFO and CLIENT GETREDIR (refs moon#632), matching redis 8.6.1 byte for byte: flags (on/off, bcast, optin, caching-yes, optout, caching-no, noloop, broken_redirect), the redirect id (0 for none, -1 with tracking off) and the sorted BCAST prefixes; a map with a set of flags under RESP3. Both, and CLIENT CACHING (moon#1049), are published in COMMAND and CLIENT HELP and queue inside MULTI.

  • Six sorted-set commands that were unknown command, and ZADD ... INCR (moon#959). ZRANGEBYLEX, ZREVRANGEBYLEX, ZREMRANGEBYRANK, ZREMRANGEBYSCORE, ZREMRANGEBYLEX and ZDIFFSTORE are implemented, wired into every dispatch path, registered as @sortedset, and covered by rows in both parity harnesses; docs/commands.md had advertised ZRANGEBYLEX while dispatch rejected it. ZADD ... INCR — which redis-py's zadd(..., incr=True) sends — replies the new score as a bulk string, or nil when NX/XX/GT/LT refuse, in Redis's decision order. Every reply, error surface included, was read off redis-server 8.6.1 before the code was written: the range grammar is checked before the key is consulted, a ZREMRANGEBY* that drains a key deletes it, a listpack zset is trimmed in place and never converted, and the used_memory ledger stays exact on both encodings. ZDIFFSTORE joins the ZUNIONSTORE family's numkeys and option rules, refusing WEIGHTS/AGGREGATE as syntax error, and — because it writes a destination it is not routed on — it also joins the moon#592 cross-shard WRITE guard, so ZDIFFSTORE across shards is CROSSSLOT rather than an ack whose destination lands nowhere. (It shares the guard's (10, b'z') match arm with moon#962's ZINTERCARD; both spellings are named there.)

Changed

  • BEHAVIOUR CHANGE — SWAPDB is refused while either database has cold-tier data (moon#1237). This covers spilled keys, and spill files not yet reclaimed, on any shard. Error: ERR SWAPDB is not allowed while either database has keys or unreclaimed spill files in the disk-offload cold tier; after deleting or reading back its cold keys, run BGREWRITEAOF (appendonly yes) or BGSAVE (appendonly no) and retry.
  • If another shard holds a database through the whole bounded check (a few ms), SWAPDB answers ERR SWAPDB could not check the disk-offload cold tier of every shard, try again and swaps nothing.
  • Redis never refuses SWAPDB. The alternative, re-tagging spill files at swap time, cannot be made crash-consistent with the logged SWAPDB record without a format change.
  • A replica whose own cold tier holds either database answers its master's SWAPDB with a full resync (moon#1278).
  • BEHAVIOUR CHANGE — with --appendonly no, boot loads the snapshot and replays no KV log over it (moon#1267, redis parity).
  • It no longer replays a legacy appendonly.aof in --dir, nor WAL v3 records left by an earlier --appendonly yes run. Nothing writes those logs in that mode, so they were older than the snapshot, and replaying them reverted keys (50 of 50 in a reproduction).
  • This applies with or without --save, and in both --disk-offload modes. The cold tier and its manifest still recover.
  • Switching appendonly yes → no: run BGSAVE (or SHUTDOWN SAVE) under yes first. Without a snapshot, the server boots empty, as redis does, and logs a WARN naming the AOF it did not load (appendonlydir/ or appendonly.aof).
  • BEHAVIOUR CHANGE — snapshots always go to and load from --dir, whatever the save rules (moon#1267). BGSAVE, SHUTDOWN SAVE and boot recovery now work with --appendonly no and no --save. Before, they were refused and an existing snapshot was ignored. As in redis, save rules only schedule automatic saves.
  • BEHAVIOUR CHANGE — SIGTERM and SIGINT save first when save points are set (moon#1263), as a plain SHUTDOWN does.
  • They wait for a save that is already running.
  • With --appendonly yes plus save points, this also writes a final snapshot, which takes time proportional to the dataset.
  • There is no deadline. While the save makes no progress for 20 s, the server logs redis's shutdown error lines and keeps the stop armed; it exits once the save is on disk.
  • If the save fails, the server logs it and keeps running, as redis 7 does. A second SIGINT exits with status 1, skipping the save but still writing the acknowledged AOF records. SHUTDOWN ABORT cancels.
  • Size the supervisor's stop timeout to the final save plus 60 s for the AOF drain: systemd TimeoutStopSec (default 90 s), Docker --stop-timeout, Kubernetes terminationGracePeriodSeconds. See docs/production-guide.md, "Graceful shutdown and stop timeouts".
  • BEHAVIOUR CHANGE — SHUTDOWN and FLUSHALL wait for their save while it makes progress (moon#1263). They give up only after 20 s without progress: SHUTDOWN then answers -ERR SHUTDOWN failed: background save made no progress for 20 s, check logs and keeps running, and FLUSHALL logs the failed save and still answers +OK, as redis does. Before, one fixed 20 s deadline failed any longer save.
  • BEHAVIOUR CHANGE — FLUSHALL with save points saves the empty dataset before it replies (moon#1264, redis parity). This also covers ASYNC, FLUSHALL in MULTI, and FLUSHALL from a script.
  • It blocks the calling connection and its pipeline for one save: about 17–24 ms on a test host. Other connections are not blocked.
  • The save points used are the current ones, so CONFIG SET save applies here, as it does for SHUTDOWN and signals.
  • A replica applying its master's FLUSHALL, and the admin console's flush, do not save.
  • BEHAVIOUR CHANGE — a graceful exit waits for the AOF writers to write and fsync their queues (moon#1274). This covers SIGTERM, SIGINT and SHUTDOWN. A BGREWRITEAOF still in flight at the stop is aborted, and the old generation stays authoritative. The wait is bounded at 60 s; a writer still running after that is named in the log, and the exit status is non-zero.
  • BEHAVIOUR CHANGE — AOF replay judges key expiry by when the log was last written (moon#1277): its newest file's mtime, capped at the wall clock. It no longer uses the clock at restart.
  • If the log's mtime is earlier than its last write (the clock stepped back during a run, or a network filesystem whose server clock runs behind), replay can keep keys that had expired, and a key written again after a lazy expiry can come back with a wrong value. A logged time record is tracked in moon#1283.
  • BEHAVIOUR CHANGE — an AOF base is loaded with its expired keys (moon#1236), as redis loads an AOF preamble. So is a snapshot when logs are replayed over it (moon#1277), so the replayed writes see the keyspace they were logged against. Active expiry then reaps them, logging their DELs. Until then they count in DBSIZE and in memory.
  • Measured: 200k expired keys in a base showed DBSIZE 199k and 49 MB right after boot. They were reaped in about 95 s and added 5.29 MB of DEL records to the AOF.
  • Under --maxmemory with noeviction, writes can answer -OOM until the reap finishes.
  • BEHAVIOUR CHANGE — with --appendonly no, an unreferenced spill file is removed only after the next successful snapshot that started after it emptied (moon#1260). It is removed right after that snapshot, and counted in cold_files_pending_unlink until then.
  • Without save rules, only a manual BGSAVE or SHUTDOWN SAVE releases such files, and until then SWAPDB is refused on their databases.
  • As with any snapshot-only setup, a restart returns to the last snapshot. So with --appendonly no and no save rules, a DEL or FLUSHALL of cold keys inherited from an earlier run is undone by any restart, a clean SHUTDOWN included, until a BGSAVE succeeds.
  • SHUTDOWN ABORT (moon#1264):
  • It now cancels a SHUTDOWN, SIGTERM or SIGINT that is still saving. It answers +OK, the shutdown does not happen, and the waiting client gets -ERR Errors trying to SHUTDOWN. Check logs..
  • With nothing in progress it answers -ERR No shutdown in progress.; the trailing period is new.
  • ABORT combined with another modifier is a syntax error. A repeated modifier, such as NOSAVE NOSAVE, is accepted.

  • BEHAVIOUR CHANGE — FLUSHDB and SWAPDB during a BGSAVE no longer fail the save (moon#1228). It completes with the keyspace as it was when the save started, as redis's forked child does, so a workload that flushes more often than a save takes can now save at all.

  • The save holds a flushed database's start-of-save rows until it writes them; INFO current_cow_size reports them. Rows written after the save began are trimmed away over the following ticks, at most 512 row operations per tick (a collection costs one more per 64 elements), so the memory held is bounded by that database's size when the save began. What the save lets go of is freed off the shard thread. Release build: a 2M-row database grown by 6M rows and flushed during a save kept PING at p99.9 0.9–1.5 ms and max 9–14 ms over the whole save (main's plain FLUSHDB of the same table: 174–193 ms).
  • FLUSHALL still fails an in-flight save, as redis's does, and so does a replica full resync.
  • So does a FLUSHDB of a database that grew during the save while another grown database is still waiting or being trimmed, once together they hold more than 8 MiB of rows written since the save began (two such flushes within one tick — a pipeline, MULTI, a script or two clients — or during the trim). A database that did not grow never fails the save.
  • BEHAVIOUR CHANGE — a BGSAVE that writes are outpacing does more work per tick (moon#1228): up to 16× the entries and segments and 4× the bytes, scaled by how far writers are ahead of it. Saves under a write flood finish sooner and hold less copy-on-write memory, and write tail latency is higher while the save runs.
  • Release build, --shards 1, SET flood: the save finished 2.6–2.9× sooner (2.26 s → 0.88 / 0.77 s).
  • Write p99 during the save roughly doubled (804 → 1,818 µs light, 2,383 → 4,294 µs heavy), and p99.9 rose 25–80%. Max latency did not change.
  • Read-only load and --shards 4 show no regression.
  • New INFO persistence field current_cow_size: the memory a running save holds for itself. As in redis, it is not part of used_memory.
  • BEHAVIOUR CHANGE — MQ CREATE and MQ PUSH answer -OOM over maxmemory under noeviction, or evict under an evicting policy, as XADD does (moon#1250). POP and ACK are never refused.
  • BEHAVIOUR CHANGE — --save rules fire in the sharded server (moon#1232). --save "<seconds> <changes>" never fired there before; deployments that pass --save now get the periodic snapshots the rules describe.
  • Rules trigger on rdb_changes_since_last_save and on the time since the last successful save; a failed save is retried after 5 s whatever the rule's seconds.
  • rdb_changes_since_last_save now counts what redis 7.0.15 counts: collection writes by their redis rule (HSET field-value pairs, LPUSH elements, SADD members added, ZADD added plus rescored, pops by the elements popped, geo stores by the members stored), blocking commands served at once as their non-blocking twins (also at --shards > 1), SWAPDB as one change (also when replicated or replayed from the AOF), and writes made while a save runs.
  • It no longer counts a DEL or EXPIRE of a missing key, key or hash-field expiry, eviction, a read that brings a cold key back into memory, RENAME k k, booting from a snapshot, or a replica's full sync (which made --save "3 100" rewrite the whole snapshot 3 s after every boot).
  • The remaining known differences are listed in src/command/keyspace_changes.rs.
  • BEHAVIOUR CHANGE — SHUTDOWN during a running save waits for it, then saves (moon#1232), within one 20 s deadline, instead of failing with "Background save already in progress". SHUTDOWN NOSAVE and SIGTERM still exit at once.
  • BEHAVIOUR CHANGE — a spill file below the latest AOF rewrite's cut stays on disk until the next committed rewrite and sweep (moon#1231), and shows in INFO cold_files_pending_unlink. While such files keep the dead-slot ledger over its threshold they ask the auto-rewrite monitor for a rewrite, also with auto-aof-rewrite-percentage 0.
  • New INFO field spill_thread_alive (refs moon#1265): 0 once any shard's spill thread was found dead. A dead spill thread is logged once at error and its in-flight compactions are abandoned so they no longer block the shard's reclaim; the thread is not yet respawned.
  • BEHAVIOUR CHANGE — DEL/UNLINK of a key whose TTL has passed answers 0 and publishes expired, not del (moon#1234 review), on the plain, spanning, MULTI and Lua paths, as redis 7.0.15 does. The deletion still reaches the AOF and replicas.
  • BEHAVIOUR CHANGE — a KNN prefilter accepts up to 128 conditions. More answers ERR invalid FILTER expression (inline …=>[KNN …] and FILTER, at every shard count).
  • BEHAVIOUR CHANGE — XSETID below the stream's top entry is refused with redis's error (moon#1249). Before, a later XADD * could overwrite existing entries. Replaying an old AOF that contains such an XSETID now keeps the original entries.
  • BEHAVIOUR CHANGE — multi-shard FT.SEARCH honours an inline KNN prefilter (moon#1238).
  • At --shards > 1 the scatter used to drop the @field:{v} prefix of @field:{v}=>[KNN …] and return the nearest documents unfiltered. These queries now return what --shards 1 returns.
  • An unparseable prefilter now answers ERR invalid FILTER expression instead of rows.
  • A shard that cannot evaluate the prefilter fails the query rather than being skipped.
  • An explicit FILTER clause is still refused at --shards > 1.
  • BEHAVIOUR CHANGE — a full-text KNN prefilter that cannot be evaluated is refused, not answered empty (moon#1226).
  • This applies with MOON_VECTOR_PAYLOAD_TEXT=off, in a build without text-index, and for a text match on a field the declared schema does not index. Before, an @field:{multi word} prefilter silently matched nothing.
  • It now returns ERR full-text KNN filter … on every search path.
  • FT.INFO reports payload_text_index on|off.
  • BEHAVIOUR CHANGE — a SCRIPT LOAD that cannot reach every shard answers -MOONERR partialfanout … (moon#567). Before, it returned a sha that some shards would answer NOSCRIPT for.
  • BEHAVIOUR CHANGE — list and set moves refuse with -IOERR before popping when an endpoint's cold-tier copy cannot be read (moon#1225). This covers LMOVE, RPOPLPUSH, BLMOVE, BRPOPLPUSH, LMPOP, BLMPOP and SMOVE. Before, the element was popped and returned, but its push onto the unreadable destination was dropped, so it was gone from both keys.
  • BEHAVIOUR CHANGE — a BGSAVE crossed by FLUSHDB / FLUSHALL / SWAPDB on a database it has not finished now fails (moon#1224) — error log, rdb_last_bgsave_status:err, the previous snapshot file kept — instead of panicking the shard (FLUSH*) or writing a mixed file (SWAPDB). A replica full resync during a BGSAVE fails it the same way. Snapshot segment blocks are now written in hash order (the loader ignores the segment index field).
  • BEHAVIOUR CHANGE — data-type edge cases now match redis 7 (moon#1168–#1172): SINTERCARD answers WRONGTYPE even when an earlier key is missing; BITFIELD grows the string under OVERFLOW FAIL; ZRANDMEMBER with count ≥ size answers highest score first; GEOSEARCH keeps redis's unsorted order, honours ANY and uses redis's error texts; APPEND / SETRANGE / SETBIT / BITFIELD keep the key's LFU counter and record one access. GEOSEARCH / GEORADIUS* parse like redis's georadiusGeneric: a negative or non-numeric radius, width or height is an error with redis's text, clauses may come in any order, a missing key still validates every option, FROMMEMBER of an absent member is an error, and positions decoded from out-of-range scores are clamped to WGS84. ZRANDMEMBER / HRANDFIELD parse the count before the key with redis's strict integer rules (+1, 01, -0 and an extra argument are errors).
  • Vector: WARM segments rank like their HOT source (moon#1213) — they carry real sub-centroid signs and use the 32-level LUT. The signs cost ⌈padded dim / 8⌉ bytes per vector (+128 B at 768d, padded to 1024) and count in resident_bytes, so the WARM mmap budget demotes more segments. New opt-in MOON_VECTOR_PAYLOAD_SCHEMA=declared indexes only declared TAG/NUMERIC payload fields and routes KNN TEXT filters to the BM25 plane (filters on undeclared fields then match nothing). index_persist sidecars stay v5, which the previous release reads, unless MOON_VECTOR_PAYLOAD_SCHEMA=declared is on and an index declares a payload field: only then is v6 written (older binaries refuse a v6 sidecar; replicas still receive v5 definitions). A v5 sidecar carries no payload schema, so an index created while the mode was off keeps the index-every-field payload policy after a restart with it on — re-create it to opt in.

  • BEHAVIOUR CHANGE — a command the user's ACL denies inside MULTI now aborts the whole transaction (moon#1035). EXEC answers -EXECABORT Transaction discarded because of previous errors. and applies nothing, where it used to apply every command except the denied one. Measured against redis-server 8.6.1 with a +@all -flushall user: MULTI / SET mx 1 / FLUSHALL / EXEC answered *1 +OK on moon (mx set) and -EXECABORT on redis (mx unset); moon now matches. The denied command still gets its -NOPERM reply at queue time, with the same text and the same ACL LOG behaviour as outside a transaction. This covers every ACL refusal — command, key pattern, and channel pattern: PUBLISH/SPUBLISH to a denied channel inside MULTI used to answer +QUEUED and was refused only inside EXEC's reply after the rest had applied; it is now refused at queue time. Holds on both runtimes, at --shards 1 and --shards 4 (including a body routed to its owner shard, moon#247), pipelined or not; DISCARD clears the poison and the aborted EXEC clears WATCHes. Each handler's ACL gate calls one ConnectionState::flag_transaction on the verdict of check_command_permission(user, cmd, args), so per-subcommand rules (-config|set) poison the transaction as soon as that check enforces them. A client that relied on the partial commit — treating the NOPERM as a per-command failure and the rest as applied — now sees nothing applied.

  • Check (macOS) and Check (Windows) run their tests in three shards, cutting the critical path of a workflow_dispatch roughly in half. Measured on a real run before changing anything: the macOS job spent 146s compiling and 527s RUNNING 5,909 tests — 85% of 793s — and Windows 739s of 854s, with clippy only 15-40s because sccache and rust-cache already make the build cheap. So the cost was never compilation, and caching it harder would have bought nothing. cargo nextest --partition count:N/3 splits the RUN across three machines; each still pays the ~146s compile, trading 2x146s of CPU for ~350s of wall clock per job. Whole-tree audits (both clippy invocations, the x86_64-apple-darwin cross-build) are pinned to shard 1 rather than repeated three times. Shards are separate machines, so the fixed-port server suites cannot collide, and fail-fast: false keeps a failure in one shard from hiding the others. Free on a public repo; the one real limit is the 5-concurrent-macOS-job ceiling, which a single dispatch stays under.

  • The AOF RDB-preamble load no longer wipes the cold plane on restart (moon#1007). Recovery Phase 3 rebuilds the cold index from the shard manifest; Phase 4b's replay_aof then loaded the MOON preamble that BGREWRITEAOF writes, and rdb::load_from_bytes swaps fresh Database temporaries over the live ones (*live = temp) — dropping cold_index and cold_shard_dir, live-tier topology the hot snapshot does not carry. The server then came up with a wired-but-EMPTY cold plane and every spilled key read as an ABSENT key, with DBSIZE agreeing. Measured at --shards 1 after any BGREWRITEAOF: 28,868 cold keys gone, and gone again on every later boot. The damaged-file scenario that surfaced it is a red herring — an undamaged run loses just as much. Only tokio --shards 1 takes this path (--shards >= 2 uses the PerShard manifest, whose shard_replay already brackets the same swap via take_cold_wiring); monoio is exposed for exactly one boot when upgrading from a legacy AOF, during which an INCR/APPEND against a vanished key mints from zero and corrupts it permanently. The preamble load is now bracketed the same way, restored BEFORE the RESP tail so replayed DEL/FLUSH* still tombstone cold. Fixed in replay_aof rather than rdb::load_from_bytes on purpose: the generic loader also serves replica full-sync and DEBUG RELOAD with a FOREIGN dataset, where preserving this node's index would surface stale reads.

  • BEHAVIOUR CHANGE — a disk-offload server that cannot prove where its cold file ids resume now refuses to start (moon#997). If a shard's data/ or vectors/ directory exists but cannot be listed, an entry in it cannot be read, or its shard manifest is full-length but cannot be opened, startup exits with refusing to start: cannot prove shard N's cold file_id seed …, the OS error, and what is safe to do: for a permission or I/O error nothing needs removing; for a manifest whose two root pages are both corrupt the message says NOT to delete it (the next boot would delete every heap file beside it as an orphan). It used to log a warning and restart the counter at 1, after which the next spill renamed its batch onto the live heap-000001.mpf (reproduced on both runtimes at --shards 1 and 4 with a -wx data/ directory). A manifest shorter than its two root pages is NOT refused: only an interrupted create produces one, it holds no entry, and it is re-created empty with a WARN naming the file. Nothing on disk is changed by a refusal.

  • BEHAVIOUR CHANGE — ZADD ... GT LT and a NaN WEIGHTS value now error where they previously succeeded (moon#969). ZADD k GT LT 1 m used to reply (integer) 1 and, on an existing member, (integer) 0 with the score left alone; it is now ERR GT, LT, and/or NX options at the same time are not compatible, as on Redis. ZUNIONSTORE/ZINTERSTORE/ZUNION/ZINTER with WEIGHTS nan used to be accepted and poison every aggregated score; it is now ERR weight value is not a float. Infinite weights remain legal. A client relying on either form silently doing nothing will now see an error.

  • BEHAVIOUR CHANGE — BLMPOP/BZMPOP whose keys span shards are refused with CROSSSLOT at --shards > 1 (moon#989), the rule moon#962 already applies to LMPOP/ZMPOP. They used to answer, and measured at --shards 4 over 16 three-shard placements, 20 of 32 probes popped a key the reply did not name: either the WRONG key (a later local key served over an earlier remote one) or a second key whose element no client ever received. "Pop from the first non-empty key in argument order, exactly once" is a property of the whole key vector that no single shard can see. The refusal is decided from the key names before anything is touched, so the keyspace is unchanged. Keys under one {hash} tag, and every placement at --shards 1, are unaffected.

  • BEHAVIOUR CHANGE — ACL SETUSER now REJECTS rule tokens it used to answer +OK for and silently drop (moon#970, moon#979). The parser ended in _ => {}, so a token it did not recognise was ignored while the call succeeded. Each of these now errors with redis's own text and applies NOTHING -- the rule list runs on a copy of the user, committed only if every rule applied, so a rejected call neither creates the user nor keeps the prefix that parsed (on >pw ~* +@all bogus used to create the user with +@all): an unknown token (totalnonsense, nocommand, @read, ' on') -> Syntax error; a malformed % selector (%X~k, %RR~k, %) -> Syntax error; an unknown command (+bogus, a typo'd -flushal, +, +config|bogus) -> Unknown command or category name in ACL; a #/! hash that is not 64 lowercase hex -> redis's hash error (an uppercase hash used to be stored and then never authenticate); <pw / !hash for a password the user does not hold -> The password you are trying to remove from the user does not exist; a (...) selector -> ACL selectors are not supported (redis accepts selectors; moon does not implement them and refuses rather than drop the grant); a non-UTF-8 rule argument -> ACL rules must be valid UTF-8 (it used to be dropped before the parser saw it). An aclfile line carrying any of these is no longer loaded -- the user is absent and a WARN names the rule -- where it used to load with the token dropped; fix the line (usually a typo'd command or an uppercase hash) before upgrading. Two further visible changes: a user with no key patterns can now run keyless commands such as PING (redis gates only keyed commands on key patterns; keyed commands are still denied), and allkeys / ~* and allchannels / &* now REPLACE the pattern list as on redis, so ~a allkeys is reported and saved as ~* rather than ~a ~* -- the same permission either way. Not changed: a pattern added AFTER ~* is still accepted (redis rejects it), so an existing aclfile holding ~* ~x keeps loading.

  • BEHAVIOUR CHANGE — REPLICAOF/SLAVEOF host port and CLUSTER REPLICATE on a --shards > 1 node now error instead of replying +OK (moon#1015). Streaming replication applies into one shard only (multi-shard replicas are moon#406), so the replica task already refused such a node — but only in the server log, AFTER the handler had acked +OK, flipped the node to a read-only replica and killed any running replica task. The node then refused every write while holding none of the master's data. The handlers now refuse first, with ERR replica mode requires --shards 1: this node runs more than one shard and multi-shard replicas are not supported yet (moon#406), and leave the role, the running replica task and the cluster view untouched. REPLICAOF NO ONE is unaffected, and so is a --shards 1 node. One shared predicate gates all four call sites (monoio and tokio REPLICAOF, monoio and tokio CLUSTER REPLICATE) and the replica task's own guard, so they cannot drift apart. An admin script that retried until +OK will now see the error.

Performance

  • An overwrite no longer hashes the key into an empty cold-tier index. With disk offload enabled (the default) every hot-key overwrite asked the cold index to drop a spilled copy; with nothing spilled that still cost a 64-bit key hash and a map probe. The index now answers at once when it holds nothing: 7.2% fewer instructions per SET on the write path (callgrind, 300k SETs, P16, --shards 1), about 2.5% less server CPU per op natively. The same review measured wave 2a's own write-path additions (the moon#1299 hold check and the moon#1286 overwrite check) at +11 instructions per SET (+0.4%).

  • A write that evicts against a backlogged AOF writer no longer stalls its shard for k × 500 ms (moon#1294). Every connection-path eviction gate (monoio write gate, tokio per-command and MQ gates, script bridge, inline SET) shares one reason-DEL backpressure bound per eviction run; past it the remaining victims' DELs fail fast into aof_reason_del_dropped (counted, logged, aof_last_append_status:err). One evicting write against a stalled writer: 21.5–45.1 s → 0.52 s, both runtimes.

  • Active expiry adapts to an expired backlog (moon#1288). When a cycle runs out of time with keys still due, the 1 ms tick drains the backlog in slices capped at 25% of the shard and 1 ms per tick (a token bucket; a saturated loop gets fewer slices). 1.84M keys expired during a 6 s stall clear in ~5–6 s instead of ~4 min (7–8K → 300–370K keys/s; redis 7.0.15: 3.8 s), at about +250 µs PING/GET p99 while it drains. Not on replicas; stands down when the AOF channel is nearly full. New INFO expired_time_cap_reached_count, expire_cycle_cpu_milliseconds.

  • SSCAN pages a large set by position: O(COUNT) per call (moon#1287). It materialized and sorted the whole set on every call (a full scan of 1M members ~9–10 min; now 2.1–2.7 s, redis 3.0–3.3 s), and a member removed before the cursor could make it skip members present for the whole scan. Small (intset/listpack) sets answer in one call with cursor 0, as redis does; an intset answers in numeric order and SSCAN key -1 is accepted as in redis. A set rebuilt mid-scan (SUNIONSTORE/SINTERSTORE/SDIFFSTORE, RESTORE, the cold tier, or another key's set moved in by RENAME / COPY REPLACE) is detected at zero bytes per set and the scan continues on a hash-ordered cursor that always terminates, so no member present for the whole scan is skipped. One case is not detected: a copy of the scanned set itself, modified and moved back onto it mid-scan, keeps its tag. HSCAN keeps the old path until moon#1171's IndexMap.

  • A stalled shard no longer replays every missed periodic tick (moon#1280). Every interval now skips missed ticks (both runtimes; clippy.toml forbids the raw constructors), the monoio chores are due by elapsed time instead of tick counts, and active expiry scales its one catch-up sweep (capped at 4×). After a 3 s SIGSTOP with a queued lazy-free backlog the first PING waits ~0.7 ms instead of 177-207 ms (monoio) or 23-40 ms (tokio); a 3 s stall fires 1-2 back-to-back ticks instead of ~3,000. The per-tick duties (the lazy-free drain, the snapshot walk) owe a late tick the milliseconds it missed (at most 32 ticks' budget in one tick), so a loop saturated by long commands frees and saves at its idle pace instead of one slice per round. New INFO shard_tick_late_total and shard_tick_burst_max. instantaneous_ops_per_sec now reports a real rate (it read 0).

  • DEL / UNLINK of a large value during a save no longer copy it first (moon#1269). The copy-on-write capture deep-cloned the value of every key a write touched before the command ran — for a removal, a copy of something about to be freed, O(elements) on the shard thread. The removed entry is now the snapshot's pre-image, moved, and its size rides along so the save does not walk it again. UNLINK of a 5M-field hash mid-save: 586-676 ms → 69-83 ms on monoio, 852-1624 ms → 65-80 ms on tokio (worst PING gap on another connection 807-898 ms → 70-83 ms on monoio); DEL 876-2392 ms → 75-101 ms. An in-place write to a large collection (HSET of one field) still copies it.

  • Cold-tier reclaim runs off the shard thread (moon#1240): it no longer reads, writes or fsyncs spill files, or waits for manifest fsyncs, on the shard thread. Same-host A/B: PING p99 during cold-delete churn −23%, p99.9 about 3× lower. Holding the spill files of a large cold tier after a FLUSHALL no longer costs O(N²) on the shard thread, and the per-tick "does a held file need a rewrite" check is O(1).

  • The moon#1232 change counting costs no measurable throughput: release A/B at --shards 1, SET/HSET/LPUSH/SADD/ZADD at pipeline 1 and 16, every row within ±3% of main or faster, CPU per op unchanged. 2026-09 performance review fix wave, part 3b (index moon#1199). Evidence is in .add/milestones/v0-9-2-perf-review/plans/{WS7,WS10,WS17,WS18,FIX-1238}-*/SUMMARY.md. The A/Bs are relative, from a shared 4-vCPU Linux container; re-measure on the GCE rig before quoting them.

  • The connection hot path takes no process-global lock per batch (moon#1175, moon#1165, moon#1176, moon#1178, moon#1198).

  • CLIENT PAUSE costs one relaxed load per batch unless a pause may be in force, and query-buffer ceilings are read once per connection.
  • After an ACL change, existing connections re-resolve their cached ACL at the top of the next batch. Before, one ACL SETUSER cost every existing connection its inline GET/SET path for life: GET went from −36…−39 % to −1…+7 %. Revocations still apply on the next command.
  • Once a replica attaches, writes issue LSNs and record replication without the ReplicationState lock.
  • Prometheus counters live in per-thread slots and are published at scrape time, so --admin-port now costs GET about 0 % instead of about −11 %. That delta is within this box's noise: the exporter leg's own repetitions spread +1/−1/+29 %. Histogram upkeep runs every 5 s, so an unscraped exporter no longer grows without bound.
  • Read-only peeks on the inline paths take the shared db guard.
  • One idle CLIENT TRACKING client no longer slows every writer (moon#1166, partial).
  • Writes to keys nobody tracks skip the tracking lock.
  • The inline SET path stays enabled under tracking.
  • SET with an idle tracker went from −39 % to −11 % (redis: −9 %). Table striping remains open.
  • Lists (moon#1212, moon#1226).
  • LPUSHX, RPUSHX and the blocking serve path keep small lists in listpack form, and SORT reads listpack elements borrowed.
  • Listpack buffers grow to the allocator size class of the new length instead of doubling: a capped list went from 3,791 to 2,257 B/key (redis 2,224).
  • RPOP key n and LMPOP … RIGHT COUNT n cut the tail once, and SRANDMEMBER key -N on a listpack set copies each member once.
  • Vector (moon#1226, moon#1228).
  • The mutable segment is exact-reranked from its f16 rows: far-query R@10 went from 0.70 to 0.99 on the test fixture, at 1.6–1.7× the mutable-scan CPU (measured at K=10 only). Under a broad filter (HnswPostFilter), the rerank depth is the same k on every search path, so a plain FT.SEARCH returns the same documents as the same query with RANGE or SESSION.
  • PreparedTqQuery is built only for queries that span two or more graph segments.
  • A declared-TEXT KNN filter resolves membership by bitmap, 1.9× faster.
  • The HNSW prefetch covers misaligned code rows.
  • Text (moon#1226): term-at-a-time scoring uses a bounded scratch, query leaves are found by binary search, TopK sorts once when k covers every match, and posting runs no longer keep twice their memory.
  • Shard upkeep (moon#1198, moon#1226).
  • Hot-key sampling uses a per-thread tick, stream IDs render without format!, and KEYS no longer clones or probes per key.
  • The 1 ms lazy-free tick is gated per shard.
  • Each tokio AOF writer's buffer is 2 MiB instead of 8 MiB.

2026-09 performance review fix wave, part 3a (index moon#1199; evidence in .add/milestones/v0-9-2-perf-review/plans/{WS8,WS15}-*/SUMMARY.md). Relative A/Bs on a shared 4-vCPU Linux container. No throughput change is claimed; only the deterministic counts below are.

  • Spanning MSET/DEL/UNLINK send one sub-command per owner shard (moon#1184) and log one AOF record per owner. A DEL/UNLINK that removes nothing is no longer logged: DEL of 10 absent keys went from 51.7 MB of AOF to about 0.3 KB, and MSET of 10 pairs is 23% smaller.
  • Cross-shard writes do less redundant work (moon#1177): request frames are moved instead of deep-cloned, AOF records are passed on without a copy, and the replication backlog lock is skipped when there are no replicas.
  • WATCH over remote keys (moon#1183) reads versions through the foreign-read fast path and asks every remaining owner at once, under shared guards. Owner-loop wakes for 200 WATCHes of 3 remote owners went from 603 to 3.
  • Multi-shard FT.SEARCH (moon#1182) sends its remote legs before searching locally, and every KNN leg runs on the cooperative yielding path.
  • SPSC batch arms take one exclusive db guard per command (moon#1198). Three shard message variants that nothing produced are removed, along with about 690 lines of their handler arms.
  • moon#1214 item 1 (the SPSC Notify lock) was profiled and deferred. The removable share is at most 0.83%, which is below the noise floor on this host; re-profile at --shards 8 on the GCE rig.

2026-09 performance review fix wave, part 2 (index moon#1199; evidence in .add/milestones/v0-9-2-perf-review/plans/{WS2,WS4,WS9,WS11,WS12,WS13}-*/SUMMARY.md). Relative A/Bs on a shared 4-vCPU Linux container — re-measure on the GCE rig before quoting.

  • Large multibulk uploads parse in O(n) (moon#1164): a resumable tri-state scan keeps a per-connection parse state across reads — one RPUSH of 1M elements in 64 KiB writes 15.7–22.7 s → 121–126 ms (redis 112–126 ms). Wire path (moon#1179): Frame 72 → 40 bytes, direct reads into the read buffer sized from the parse hint, lazy batch scratch and response slots — 64 KiB SET +60%, SET … EX P16 +8%, server CPU per request −6 to −11%.
  • Sorted sets, sets, strings, geo (moon#1168–#1172, #1174 §4, #1189 B+tree): ZRANGEBYSCORE LIMIT 48.7 ms → 118 µs and ZCOUNT 32.4 ms → 55 µs via order statistics; ZRANDMEMBER 20.3 ms → 56 µs; SINTER small ∩ big 232 ms → 85 µs; SETBIT on 12.5 MB 8.75 ms → 59 µs and APPEND 40K × 100 B 12.0 s → 0.2 s (in place); GEOSEARCH over 200K points 22.4 ms → 84 µs; a 1M rising zset's B+tree 363 → 180 MB.
  • EVAL/EVALSHA reuse compiled functions (moon#1167): EVALSHA 2.28× (1-line) and 5.02× (1.4 KB script). PUBLISH snapshots subscribers with one Arc clone (moon#1180). Keyspace events with no __key* subscriber allocate nothing (moon#1214): SET with notify-keyspace-events KEA 1.53×.
  • BGSAVE converges under an insert flood (moon#1216): a per-tick budget over the hash-space walk — 1M keys: 35 s → 1.5 s, max PING 40–420 ms → 1–3 ms, peak RSS 950 → 200 MiB. The budget also caps a tick's output (1 MiB, one segment past at most) and pauses while the snapshot writer is backlogged.
  • Vector (moon#1213): immutable segments drop never-read QJL data and collections no longer hold QJL matrices (4 shards × 2 EXACT 768d indexes +145.5 → +1.4 MB RSS); EXACT compaction at 768d 12.3 → 2.5 s; HNSW prefetch widened (−4 to −9% per query, identical results).
  • Text / graph (moon#1220): wide prefix/fuzzy expansions scored term-at-a-time (ka* over 200K matches 4.5×); posting positions stored contiguously per run (RSS −37% at 200K docs); Cypher RETURN … ORDER BY … LIMIT projects only the kept rows (2.1–3.0×) and range IndexScans stream (9.4×).

2026-09 performance review fix wave, part 1 (index moon#1199; per-workstream evidence in .add/milestones/v0-9-2-perf-review/plans/*/SUMMARY.md). Numbers are relative A/Bs on a shared 4-vCPU Linux container against HEAD 935c555 — re-measure on the GCE rig before quoting.

  • DashTable probes compare ~1 key instead of ~7 (moon#1159). The H2 fingerprint came from the same top hash bits the segment directory routes on, so past ~5K keys every key in a segment had the same fingerprint. H2 now uses bits 32–38: key compares per hit 6.7 → 1.0, per miss 20.5 → 0.16; GET +11–17%, SET +36% at 1M keys. Segment splits move only the upper half in place, and a SET overwrite of a >23-byte key no longer builds and drops an owned key.
  • Removing a large collection no longer stalls the shard (moon#1190, partial). UNLINK and active expiry of values with ≥ 65,536 elements hand them to a per-db lazy-free queue drained on the tick and a lazily started moon-lazyfree thread (both runtimes): UNLINK of a 1M-field hash replies in ≤ 0.5 ms (was 105–173 ms). The freed bytes stay in used_memory until the drain releases them.
  • Expiry sweep: one table probe per expired key (moon#1189, expiry-index part), no key clone, ~65% more keys expired per sweep budget.
  • LREM is one pass, LPOS scans in place (moon#1173): LREM on a 200K list 13.2 s → 2.5 ms, LPOS … MAXLEN 10 on 1M elements 39 ms → 71 µs.
  • Small lists stay listpack (moon#1174 §1–§3): LTRIM/LREM/LINSERT/LMOVE/RPOPLPUSH/LMPOP no longer convert a listpack list to linkedlist (10K capped lists: RSS 201 → 62 MB); listpack range reads seek once (LRANGE ~2.3×); owned listpack reads allocate once; HKEYS/HVALS/HEXISTS/HSTRLEN read fields in place.
  • Vector search (moon#1192, #1193, #1194 part, #1196): the budgeted 16-level TQ-ADC kernel is 8-accumulator and safe (2.6–3.4× per candidate); sub-centroid signs are persisted (sub_signs.bin, segment format v2 — additive, v1 directories load unchanged) so restarted segments keep the 32-level LUT; BUILD_MODE EXACT computes QJL data on the compaction worker (compaction-submit stall 3.2 s → 12 ms) and scans the mutable segment with TQ-ADC + FastScan (EXACT rankings change: recall vs exact rises in-distribution 0.71 → 0.83 and for non-unit L2 0.12 → 0.83, falls for far queries 0.83 → 0.73); FT.SEARCH builds its rotation/LUT once per query, takes the tombstone guard once, and no longer clones the SESSION member map; the f16 rerank sidecar is served from its mapped file once persisted; payload-index keys no longer pin request buffers. New opt-out MOON_VECTOR_PAYLOAD_TEXT=off skips payload full-text indexing. Format v2 compatibility: older binaries ignore sub_signs.bin and load v2 segments with the 16-level LUT; a segment directory of a newer format than the binary supports is refused rather than misread (the loader leaves it untouched; its keys are re-indexed from the keyspace).
  • Full-text search scores each match once (moon#1191): bitmap membership, one BM25 pass with doc-ordered cursors, a bounded top-k, keys cloned for the returned page only — broad-term FT.SEARCH … LIMIT 0 10 70–103× faster at shards 1 (25× at shards 4), scores bit-identical.
  • Text upserts no longer memmove whole postings (moon#1195): re-indexing the oldest of 200K docs 15.5 → 0.42 ms. Text side tables are dense columns (moon#1194, text part): RSS for 100K tagged docs −55%, and their used_memory billing now matches real size.
  • Cypher LIMIT streams and ORDER BY … LIMIT keeps only the page (moon#1197): MATCH (n:L) RETURN n.x LIMIT 10 on 200K nodes 143 → 0.18 ms.
  • Persistence off the event loop (moon#1181, #1185 part, #1186, #1187 part, #1188): BGSAVE streams to a writer thread and finalizes (fsync/rename) off the shard (max stall on 1.5M keys 1.3 s → 12 ms, peak RSS growth +257 → +2 MiB); AOF rewrite streams its base image instead of deep-copying the keyspace (peak RSS growth +567 → +6 MiB; the fold is still O(dataset) on the shard — moon#1185 stays open — and the image travels in 1 MiB chunks over an unbounded channel the AOF writer reads only after its phase-3 drain and fsync, so on a slow disk up to one serialized image per shard can wait in memory used_memory does not count; the +6 MiB was measured where fsync is nearly free); WAL segment-rotation fsync goes through the sync agent (while it is in flight new WAL records wait in process memory — at most 4 MiB or 1 s, then the rotation completes inline — so a SIGKILL there can lose up to that much of the WAL-only planes, workspace / MQ / temporal, still within everysec's second); each AOF record costs exactly one allocation; CDC.READ seeks by segment header and reads in chunks on a CDC read pool (a poll at a ~1M-record tail 5.8 s → 0.17 ms).

Fixed

  • A TXN ABORT could overwrite another client's acknowledged write to a key the transaction had touched (moon#1299). A KV key written inside an open cross-store TXN is now held until TXN COMMIT / TXN ABORT: another client's write to it answers -TXNCONFLICT key held by an open transaction on every write path (plain, inline, MULTI/EXEC, scripts, blocking pops, MOVE/COPY ... DB, MQ, routed), and FLUSHDB / FLUSHALL / SWAPDB on a database with held keys answer -TXNCONFLICT database has keys held by an open transaction. A cross-shard MSET / DEL / UNLINK refused on one shard while others applied says so: -TXNCONFLICT key held by an open transaction: command partially executed; .... Eviction and active expiry skip held keys; blocked clients are served once the key is released. A replica applies its master's stream unconditionally. There is no idle timeout: a transaction holds its keys until it ends or its connection closes, and every connection exit — a protocol error, a blocked pop whose client vanished, an output-buffer-limit disconnect, QUIT or an error in subscriber mode, PSYNC — now rolls an open transaction back and releases its keys before the socket closes, so a client that has seen its QUIT reply finds the keys free (CLIENT KILL still replies first, moon#1312). RESET ends an open TXN the same way, as redis RESET discards MULTI state. New INFO stats fields: txn_open, txn_oldest_age_ms, txn_held_keys, txn_conflicts_refused. Graph writes are not isolated yet (moon#1307).

  • A TXN connection write that answered an error kept its undo capture (moon#1303). SET k v BADOPT, INCR of a non-number or a WRONGTYPE inside a TXN no longer leaves a write intent, key hold or pre-image behind, so TXN ABORT can no longer restore a stale value over another client's write.

  • TXN.COMMIT refused with snapshot too old (after KILL SNAPSHOT) kept the transaction's writes. It now rolls them back, as its reply says the commit failed.

  • Graph, MQ, workspace and temporal WAL records past the shard's 4096-slot append channel were acknowledged but dropped (moon#1302). A 6000-node Cypher CREATE kept 4096 nodes after kill -9, and a TXN COMMIT of 6000 MQ PUBLISH kept 4096. Records past the channel now wait in an in-memory queue on the shard's own thread, which the next 1 ms tick appends in order; a graceful shutdown appends them before its final WAL flush. The on-disk format is unchanged. New INFO field: reclamation_wal_append_overflow_total. TXN ABORT of a large graph rollback no longer answers MOONERR WAL backpressure: the rollback is written to the WAL before it is replicated, and only what the WAL accepted is replicated, so a master and its replica agree after a restart.

  • After a restart, writes to db 0 could replay into another database. A reopened AOF writer assumed its stream was at db 0, but the file ended at the previous run's last SELECT, so SELECT 3; SET a 1, restart, SET b 1, restart put b in db 3. A reopened writer now starts at an unknown db, as redis does (aof_selected_db = -1), so its first record always carries a SELECT. Every AOF layout, graceful and kill -9 restarts alike, was affected.

  • AOF replay judged key expiry by the log file's mtime (moon#1283). Keys that expired while the server ran, and were rewritten before their DEL was logged, came back with old values when a log's mtime was earlier than its last write (a clock stepped back, a lagging network filesystem, a touch -d restore). The writer now emits a MOON.TS <ms> record whenever the shard clock changes and at every generation head, and replay pins its clock to it. Logs written before this change replay as before; an older binary skips MOON.TS as an unknown command. A clean stop (SHUTDOWN, SIGTERM) ends each incr with a close marker MOON.TS <ms> CLOSE and a restarted writer stamps its first write, so records an older binary appends after a downgrade are recognised by their position and judged as the older binary judged them, on the re-upgrade boot and every boot after it, with no rewrite; a touch or cp of the AOF never re-judges its last writes. Not protected: a downgrade after an unclean stop of the newer binary — stop it cleanly first, or run BGREWRITEAOF as the older binary's last action (docs/STORAGE-FORMAT-V1.md §3.3). size_of::<AofMessage>() is 80 bytes (was 72). New fuzz target aof_incr_replay.

  • appendfsync everysec lost most acknowledged writes to a process crash (moon#1266, Option 3). The once-a-second fsync now runs on a per-writer agent thread (aof-fsync-<n>), so a slow disk no longer stops the AOF writer from draining acknowledged writes into the file. The tokio writer flushes every batch to the kernel (the 8 KiB user-space tail is gone); the monoio writer polls every 500 µs while writes stream in and parks otherwise, so the first write after idle is picked up at once instead of after up to 50 ms. With SIGKILL 1 ms after the last acknowledgement, 20 reps per cell: losses in 9 of 240 reps (16 of 360 across both tokio runs), down from 226 of 240 (4-vCPU Linux container). The remaining sub-millisecond window is moon#1266 Option 1A. appendfsync always is unchanged. A stalled fsync is reported as redis does: the log line "Asynchronous AOF fsync is taking too long (disk is busy?)", new INFO fields aof_pending_bio_fsync (writers with an fsync in flight: 0..N at --shards N) and aof_fsync_in_flight_ms, and aof_delayed_fsync counted once per 2 s of an ongoing postpone (unlike redis, the write itself is never postponed). A failed post-rewrite fsync is recorded and retried within about 100 ms. Diagnostic knob: MOON_AOF_WARM_POLL_US.

  • CONFIG SET appendfsync answered OK but the writers kept their startup policy, so writes were acknowledged without the fsync always promises. The change now takes effect at once, as in redis; leaving everysec first waits for an in-flight background fsync, and values other than always|everysec|no are rejected.

  • After a failed everysec fsync, aof_last_fsync_status returned to ok on a retry with nothing new written. It now clears only once a write issued after the failure has been fsynced. (Refusing writes meanwhile, as redis's -MISCONF does, is moon#1309.)

  • INFO stats expired_keys counted nothing (moon#1286). It now counts every expiry-driven whole-key removal: the active cycle (including moon#1288's fast slices), the lazy-reap drain, a DEL / UNLINK or write that lands on an expired key (SET, SETNX, GETSET, APPEND, INCR*, SETBIT, PFADD, MSET, a COPY / RENAME / *STORE destination, a key a read had hidden, a key only the cold tier held), and the cold-tier TTL sweep and on-read reclaim. AOF replay counts nothing, a replica applying its master's DEL does not count it, and CONFIG RESETSTAT resets it, as redis 7.2.7 does; hash-field expiry is not counted. An absolute deadline already in the past (EXPIREAT / PEXPIREAT, GETEX ... EXAT/PXAT, RESTORE ... ABSTTL) now deletes the key at once and publishes del rather than counting an expiry, and the master propagates that delete as DEL; a replica applying its master's stream stores the deadline instead. EXPIRE k -1 publishes del, and GETEX k EX -1 answers ERR invalid expire time in 'getex' command.

  • ACL GETUSER, ACL LIST and ACL SAVE sorted command rules alphabetically (moon#1296). They now render them in the order they were applied, as redis 7.2+ does, and a command grant under +@all (+@all +get -set) is kept, so ACL SAVE / ACL LOAD reproduce the rules exactly. Existing ACL files re-save with a different rule order; permissions are unchanged, but diff-based config management sees a one-time change. Category tokens (+@read) are still expanded into their commands, unlike redis (moon#1306).

  • Held cold spill files were released only by a manual BGREWRITEAOF or BGSAVE (moon#1289). After three orphan sweeps (about two minutes by default) with no committed fold, moon releases them itself: with an AOF the auto-rewrite monitor folds, and without one a rate-limited snapshot is requested, at most one per 10 sweep intervals (10 minutes by default). Without an AOF this automatic snapshot runs even with save "" and overwrites the dump file like any BGSAVE. It never contains a TXN's uncommitted writes: it waits while any TXN is open, and it is abandoned whole — no shard file replaced, LASTSAVE unmoved, no held file released — and retried at a later sweep if any shard holds an uncommitted TXN write when that shard starts its part; TXN writes after a shard's start are saved at their pre-transaction value. An open TXN, or unbroken TXN traffic, therefore keeps held files on disk (and SWAPDB refused) until a snapshot can run. BGSAVE, SAVE and the save rules are not covered yet (moon#1300). New INFO fields: cold_held_files_stale_databases, cold_held_release_folds_requested, cold_held_release_snapshots_requested, cold_held_release_snapshots_deferred_txn, cold_held_release_snapshots_abandoned_txn.

  • volatile-lru, volatile-lfu and volatile-random could answer OOM while a key with a TTL existed. Victim sampling draws random table segments and gives up after 8 × maxmemory-samples draws, so a database with a few TTL keys among many could miss them all (1 TTL key in 301: 14 misses in 3,000 runs). When sampling finds no candidate, these policies now fall back to the key with the nearest deadline, as redis's sampling of its expires dict cannot miss one.

  • A crash while a boot opened a fresh AOF generation could bring back cold keys deleted before the switch to --appendonly yes (moon#1293). The manifest committed before each incr file's MOON.COLDCUT head (and its cold DELs) was written, so a crash between the two left a headless generation that the next boot replayed ungated: 88–167 of 88–167 deleted probes came back across both runtimes at --shards 1 and 4. Every head is now written and fsynced (with its directory entry) before the manifest commits; a crash before the commit leaves no manifest and the next boot redoes the initialization. Every newly created data, offload or AOF directory (binary and embedded entry) has its entry fsynced too. On a filesystem without directory fsync (squashfs, vboxsf, WSL1 drvfs, macOS exFAT/SMB: EINVAL, EBADF, ENOTSUP, ENOTTY, ...) the fsync is skipped at every level, new or pre-existing, with one warning; a pre-existing ancestor the process may not open is skipped too. Neither fails boot; EIO does, including on an auto-resolved user-data directory. An auto-resolved user-data directory that cannot be created at all still falls back to the current directory.

  • Cold-key graves survive a smaller --databases and a failed spill commit (moon#1291). Restarting a no-AOF server with fewer databases aborted the snapshot load before its graves trailer, so every deleted cold key of the dropped databases came back at the next restart with the original count (129–173 probes per run); out-of-range databases are now skipped and their graves carried forward. A failed manifest commit during a no-AOF durable spill no longer leaves the file listed with slots neither indexed nor graved.

  • With --appendonly no --disk-offload enable, eviction lost keys instead of tiering them (moon#1290). The connection write gates and the cross-shard gate had no manifest to spill durably with and plain-dropped victims: 3.3–5.4K of 16.2K keys were gone in the repro, with no error to the client. They now tier every victim durably (0 lost, including after BGSAVE + kill -9), and so does a script routed to another shard at --shards ≥ 2 (EVAL, FCALL, or a script inside a routed MULTI: 1.5–2.2K of 16K keys were lost). The async-spill route is chosen by whether an AOF writer exists, not by the CONFIG SET appendonly string. Cost: without an AOF each spill batch is fsynced and committed on the shard thread, so a write flood over maxmemory runs at the durable spill rate (monoio --shards 4: ~150–160K rps with the 1024-entry / 1 MiB batches, where it used to run 315–559K rps while dropping keys); AOF-backed tiering is unaffected (~170–195K rps).

  • TXN.ABORT is durable and replicated on both runtimes (moon#1285, moon#1185 option b). An aborted write used to come back after a kill -9 restart (both runtimes) and stayed on replicas: the forward writes reached the AOF and the replica stream, the rollback only memory. The abort now logs compensating records stamped with the fold epoch — DEL for an undone insert, RESTORE … REPLACE ABSTTL (plus HPEXPIREAT per field deadline) for an undone update or delete — and moves each replaced value into an in-flight BGSAVE's pre-image, so the image stays point-in-time. A SELECT inside a transaction no longer makes the abort restore into the wrong database. Vector / text documents of restored keys are rebuilt live and on replicas (RESTORE now updates FT indexes, keeping FT.SEARCH … AS_OF history). Graph rollbacks were never logged at all; they are now WAL-logged and replicated, with three new WAL records (GRAPH.DELPROP, GRAPH.UNDELETENODE, GRAPH.UNDELETEEDGE; fuzz target graph_wal_replay). TXN.ABORT answers the AOF's refusal instead of +OK when its records cannot be queued (the rollback is applied either way); a refused rollback record on the disconnect / dirty-COMMIT path is counted (aof_backpressure_dropped) and logged, never silent. Cost: the abort DUMPs each restored value, O(value) on the shard thread (a 200k-field hash: ~11 ms more).

  • TXN.ABORT no longer answers +OK when its graph rollback records were dropped (moon#1285, PR #1301 review). The rollback's WAL records went through an unchecked try_send into the shard's 4096-slot append channel: a 6000-node transaction answered +OK, showed 0 nodes live, and brought 1904 aborted nodes back after kill -9, on the local leg and on a remote shard's leg. The records are now appended checked and in order (the WAL holds a prefix of the rollback, never a gap). A refusal answers MOONERR WAL backpressure: TXN rolled back in memory, but its graph rollback records were not all queued for persistence; ..., is counted in INFO persistence txn_rollback_wal_dropped, and a remote shard's rollback that was not delivered or not acknowledged fails the abort too. Durability matches forward graph writes (no reply waits for a WAL fsync), so abort latency is unchanged. A rollback larger than the free channel capacity (about 4096 records per shard) now always answers this error.

  • Writes made by EVAL, EVALSHA or FCALL inside an open TXN are rolled back by TXN.ABORT (moon#1285, PR #1301 review) — live, after a restart and on replicas. They used to bypass the transaction's undo log, so the abort left them in place. Each key a script writes is pre-image captured once, before its first write, and joins the transaction's undo log when the script returns. A write a script cannot run — one only a connection-level handler serves (blocking pops, FCALL, MQ, WS, FT.*, GRAPH.*, FUNCTION, TEMPORAL.*), an argv its arity rejects (too short, or longer than an exact arity: SETNX k v extra), or a write that answers any other error (SET k v BADOPT, MOVE k <same db>, WRONGTYPE) — captures nothing, so the abort never restores over another client's write. Nor does a write-flagged command whose argv the key walker reads as naming no written key (SORT src or GEORADIUS src ... without STORE), on the connection or from a script: it captured src, and the abort restored src over concurrent writes, logged to the AOF and replicas. Over-capture remains in these cases, each of which lets an abort restore a key over another client's write until moon#1299 keeps other writers off the keys an open TXN holds: LMPOP / ZMPOP capture every candidate key, not only the one popped; two walker corner cases (GEORADIUSBYMEMBER g STORE 100 km whose member is literally STORE captures 100; XGROUP HELP <x> captures <x>); a write that succeeds as a no-op keeps its capture (SETNX or SET … NX on an existing key, COPY without REPLACE onto an existing key, RENAMENX, LPUSHX on a missing key, SMOVE of a missing member); and a write on the connection itself that answers an error keeps its capture (moon#1303; scripts drop it). Behaviour change: inside a TXN, a script's FLUSHDB, FLUSHALL, SWAPDB, MOVE or COPY … DB, or an arity-valid write whose keys cannot be enumerated (a malformed LMPOP, ZMPOP or XREADGROUP), is refused with ERR TXN cannot roll back this command from a script …, and a read-write script whose keys live on another shard is refused with the TXN cross-shard error (EVAL_RO / EVALSHA_RO / FCALL_RO and an FCALL of a function registered no-writes still route). Both refusals poison the TXN, each refused command counting once. A function registered no-writes now runs read-only under plain FCALL as well, as in Redis: a write from it answers ERR Write commands are not allowed from read-only scripts. — an ordinary command error, as in Redis 7: redis.call raises it and redis.pcall returns it as {err = ...} (also for EVAL_RO / FCALL_RO, where it used to be raised even from redis.pcall). Cost: none outside a TXN; ~0.7 µs per captured write inside one.

  • FLUSHDB / FLUSHALL inside a TXN no longer wipe data TXN.ABORT cannot bring back (moon#1285, PR #1301 review). Sent on the connection inside an open TXN they ran, and the abort answered +OK restoring nothing (SET k v; TXN BEGIN; FLUSHDB; TXN ABORT; GET k → nil, both runtimes, any shard count). Behaviour change: they are now refused before they run with ERR TXN cannot roll back this command (whole-database write) … and poison the TXN (TXN.COMMIT answers EXECABORT), as a script's already were. SWAPDB, MOVE and COPY … DB keep their existing TXN refusal.

  • Without an AOF, a deleted cold key came back after BGSAVE + kill -9 (moon#1281). The durable state is the last snapshot plus every listed spill file, and a spill file is unlinked only when its last live key leaves it, so a cold key deleted (or overwritten, flushed, read back and deleted) before a successful save was re-indexed from its slot at boot: 100% of the deleted cold probes came back on both runtimes at --shards 1 and 4, and the moon#1236 second-crash variant returned 61-96 of 100 OLD values. Each no-AOF snapshot now records the spill slots that were already dead when it started, as a trailer after its EOF marker (older readers stop at EOF and load the file unchanged; the format version is not bumped), and the boot drops exactly those slots before it resolves each key's newest copy — a dead newer slot can no longer shadow a live older one — then carries them into the next snapshot, also for a listed file this boot could not read. The graves apply whenever the snapshot is the KV base, so a no-AOF directory booted with --appendonly yes keeps its deletes too. A kill -9 matrix (before, during and after the save, across the post-save sweep's unlinks and manifest commit, and a second generation; 24 runs × both runtimes × --shards 1 and 4) brings back none. New INFO cold_grave_slots, cold_grave_bytes; new fuzz target snapshot_cold_graves. A deletion is durable no later than the next successful snapshot.

  • monoio never reclaimed spill files emptied in the first minute after a boot (moon#1279). The unlink hold takes its baseline from the first view it observes; only the orphan sweep took views, and monoio's first sweep runs one interval after boot (tokio's at t=0), so every file spilled in that interval was held as if inherited, waiting for an AOF fold that never came (cold_files_pending_unlink stuck, disk growing under delete churn). The shard now observes the boot view before its loop on both runtimes. Also: a spill every key of which was superseded while in flight is no longer listed in the manifest with only ghost slots (it stayed cold_files_dead forever); its file is unlinked at once.

  • A write refused because the AOF writer was backlogged reported a disk fault (moon#1272). Under appendfsync everysec/no, a write that waits --aof-fsync-timeout-ms for room in its shard's writer queue (or finds a rewrite's overflow full) is refused, and answered -ERR AOF fsync failed; write not durable though no fsync ran or failed. It now answers -MOONERR AOF backpressure: write applied in memory but not queued for persistence; the AOF writer is backlogged, is counted in INFO aof_append_backpressure_refusals and /metrics moon_aof_append_backpressure_refusals_total, and is logged once per stall; aof_fsync_failures and aof_last_fsync_status no longer move for it. The policy is unchanged and documented in docs/guides/persistence.md: moon refuses rather than acknowledging a record its writer never received; the write stands in memory, so retry only idempotent writes. Under appendfsync always, writes whose batch fsync barrier cannot be queued for the same backlog answer -MOONERR AOF backpressure: write applied in memory and queued, but not confirmed durable; the AOF writer is backlogged and are counted the same way. A real write or fsync failure still answers -ERR AOF fsync failed; write not durable.

  • scripts/test-consistency.sh and scripts/test-commands.sh died silently with redis-cli older than 7.4 and leaked servers that corrupted the next run (moon#1276). redis-cli -t exists only in 7.4+; on 7.0.x the run ended under set -e with no summary and left its auxiliary moon running, which the next run then shared through SO_REUSEPORT. Both scripts now bound every probe with timeout/gtimeout (a whole-command bound; -t bounds only the connect and is the fallback, or they warn loudly), track and kill every auxiliary server on exit, refuse a port that is already taken, name the line of a set -e death, and no longer pkill their own command line. A header names the oracle version the expected values assume (redis 7.2+/8.x).

  • Under the tokio runtime, keys evicted by maxmemory came back after an AOF restart. The tokio write gates (the per-command gate, the MQ write gate and the script bridge's gate for redis.call writes) dropped eviction victims without appending a DEL, so the AOF replayed them: 29,172 of 29,175 evicted keys returned in the new plain_evictions_are_not_resurrected_by_the_aof test (8 MB allkeys-lru, --appendonly yes). They now log a DEL per victim, as the monoio gates always did; a degraded spill thread (moon#1265) drops through the same sink.

  • A panicked cold-tier spill thread was never restarted (moon#1265): one panic (a corrupt spill file read by the cold reclaim, say) left writes under memory pressure answering -OOM and in-flight payloads pinned in RAM until a restart. The thread now runs under catch_unwind and the shard respawns it on its tick with bounded exponential backoff (100 ms → 30 s) on the same channels; queued requests keep their file ids, payloads the dead thread held go back to RAM and are re-spilled, and the moon#1253 superseded sets are rebuilt from what is still queued, so no in-flight key is lost or resurrected (checked with the panic injected into the moon#1253 crash suite on both runtimes). After 5 respawns in 10 minutes the shard is declared degraded: it stops spilling (evicting policies drop, noeviction answers OOM), keeps serving RAM and existing cold files, and warns. New INFO spill_thread_restarts, spill_thread_degraded, spill_thread_rehydrated (spill_thread_alive now means "no shard's thread is down right now"); new metrics moon_spill_thread_deaths_total, moon_spill_thread_restarts_total, moon_spill_threads_degraded. A cold-reclaim file that kills the thread twice is given up (left on disk and still readable, not compacted again), so one corrupt file cannot use up the restart budget; INFO cold_reclaim_files_given_up counts every reclaim give-up. The restart budget runs on a monotonic clock: a wall-clock step no longer refills it. A death during a cold-reclaim job no longer spends that budget: it respawns after 100 ms uncharged, against a reclaim budget of its own (8 per 10 minutes) whose exhaustion disables cold reclaim on that shard until restart (one WARN, INFO cold_reclaim_disabled, gauge moon_cold_reclaim_disabled_shards) while the shard keeps spilling — so neither a few corrupt spill files nor a systematic reclaim bug degrade it.

  • Data loss on a graceful exit with the default --appendonly yes (moon#1274). SIGTERM (systemctl stop), SIGINT or SHUTDOWN lost acknowledged writes still queued for the AOF writers: all 300 of 300 at --shards 1. Shutdown now stops the shards, then lets every AOF writer write its queue and fsync before exiting, as redis's prepareForShutdown does.

  • Data loss on SIGTERM/SIGINT with save points and --appendonly no (moon#1263). Every write since the last automatic save was lost. The server now saves first; see Changed.
  • Data loss: a per-shard AOF writer that started late deleted other shards' in-progress rewrite files (moon#1271). The rewrite then aborted, or it committed a manifest that named deleted files, and those shards' keys were gone after the next restart: 165 of 200 in a forced reproduction. Loading the AOF manifest no longer deletes anything; stale rewrite files are swept at boot only.
  • Data loss: MOVE of a cold-tier key (moon#1254). A key that was spilled, or whose spill was in flight, was deleted from both databases, and MOVE answered :0. Reproduced on 69–95 of 200 cold keys. The key is now read back and moved. A MOVE or COPY … DB n of a cold key whose data cannot be read answers -IOERR.
  • Data resurrection: a cold key overwritten with a TTL came back with its old value after a restart past the TTL (moon#1236):
  • after an AOF rewrite: 46–89 of 100 keys;
  • with --appendonly no after a snapshot: 77–95 of 100.

A snapshot boot now drops the cold copy of every key it holds as expired, and an AOF base keeps its expired keys (see Changed). With --appendonly no, the drop is not yet durable: after a later snapshot and a crash, the old value can return. That is the no-AOF cold-deletion gap tracked in moon#1281. - Data resurrection: AOF replay after a downtime longer than a key's TTL (moon#1277). A TTL-preserving write logged while the key was alive (APPEND, INCR, HSET, SETRANGE, …) was replayed onto an absent key. The key came back with a wrong value and no TTL: 4 of 4 keys, on both runtimes, at --shards 1 and 4. - SWAPDB with cold keys (moon#1237): - After an AOF rewrite and a restart, cold keys reappeared in their old database, and keys deleted after the swap came back. - A restart without a rewrite moved every key spilled after a SWAPDB into the other database. - A replayed SWAPDB now moves only the cold data that existed at that point in the log. - SWAPDB on a replica with its own cold tier (moon#1278): after a failover and a restart, its cold keys came back in the old database. - Data loss on tokio --shards 1 with --appendonly yes (moon#1275). Any SWAPDB followed by kill -9 lost every key: the swap's WAL record made recovery skip appendonly.aof. - Data loss with --appendonly no (moon#1260): - A cold key inherited from an AOF run was lost after it was read, the orphan sweep ran and the process was killed: 179 of 179. - Cold keys deleted before a snapshot came back after a kill -9 that followed it (a regression inside this release, caught in review). - A key evicted while a BGSAVE ran was missing from the snapshot (moon#1257). All of the 9,952 (--shards 1) and 13,404 (--shards 4) keys evicted during a held save were missing. - Eviction now hands the removed value itself to the save, as the key's pre-save state. It is not copied, and it is freed off the shard thread once written. - With no save running, the cost is one flag check per victim. - During a save, the evicted value's memory is released once the save has written it; current_cow_size reports it until then. - SHUTDOWN ABORT could answer +OK while the server still exited (moon#1264). The abort and the commit are now decided under one lock. - Tests: - Six integration tests no longer fail on a slow or stalled host, as seen on hosted Windows runners (moon#1065, moon#1273). They poll for the condition they need instead of relying on fixed sleeps, round counts, or 50–300 ms windows. Cold-tier fillers retry an AOF backpressure refusal. - scripts/test-consistency.sh no longer hangs in the moon#1235 race rows: a bare wait also waited for the script's own servers. - perf_ws21_flushall_save: the SIGTERM leg of the CONFIG SET save test is now its own Unix-only test, like the other signal suites. Windows has no SIGTERM, and its kill cannot see a native pid, so the Windows leg of main's CI went red after part 5.

  • A BGSAVE records each key as it was when the save started, even when that key changes before the save writes it (moon#1228).
  • Before, a key changed by any of these could be saved in its later state:
    • MOVE and COPY … DB n (the key could come back in both databases or in neither);
    • WS DROP, locally and on a replica;
    • the MQ subcommands, TXN MQ.PUBLISH and replicated MQ records;
    • a stream waker's group read.
  • Queues and streams could come back with messages and pending entries from after the save started.
  • Each of these now saves the key's start-of-save state first.
  • MQ writes are charged to used_memory (moon#1250, moon#1261), as XADD is: MQ CREATE, PUSH, POP, ACK, TXN MQ.PUBLISH, stream-wake group reads, replicated MQ records and WAL replay.
  • Before, 20,000 pushes of 100 B added about 24 KB to used_memory for a 5.7 MB queue.
  • MQ POP/ACK churn keeps the charge exact: a POP's released surplus used to stay charged, so used_memory drifted up.
  • After a restart's WAL replay and on a replica, queues are billed as on the master. Replayed or replicated POPs never charged their pending entries while ACKs credited them, so a churned queue's bill drained toward 0 (28,421 B against 124,097 B live).
  • MQ and stream edge cases (part-4 review fixes):
  • MQ POP … COUNT near 2^64 no longer overflows (a debug-build crash) and delivers every queued message.
  • MQ POP of an empty queue no longer creates the internal consumer on the master only; master, replicas and a restart agree on XINFO CONSUMERS and MEMORY USAGE.
  • At the last possible stream ID, XADD * and MQ PUSH answer an error instead of crashing a debug build; a TXN MQ.PUBLISH the stream refuses is no longer written to the WAL, and a dead letter whose dead-letter stream is full stays pending in its queue instead of being lost.
  • XADD key <ms>-* when the top ID already has the last possible sequence for <ms> answers "ERR The ID specified in XADD is equal or smaller than the target stream top item" (a debug build crashed; a release build wrapped).
  • current_cow_size no longer counts a queued lazy-free value twice when a FLUSHDB lands during a BGSAVE (moon#1228).
  • Data loss: a cold key could be lost at the next restart, or come back with the wrong value (moon#1231), when it was cold at an AOF rewrite, was then read back into memory or read-modify-written, and the orphan sweep later removed its spill file (reproduced on 143–179 of 200 keys). A spill file that a replayable AOF generation may still read is now kept until a later rewrite has committed, also when it is only briefly unreachable while the sweep runs. The cold-reclaim adoption no longer removes a compacted file whose survivors changed after the rewrite.
  • Scripts over maxmemory (moon#1241), as in redis 7.0.15:
  • Inside EVAL/EVALSHA, commands that can only free memory (DEL, UNLINK, HDEL, LPOP, EXPIRE, …) are no longer refused with -OOM; growing commands still are.
  • A function registered with allow-oom runs any command over maxmemory; eviction still runs first. A per-database db_maxmemory quota (a moon extension) still refuses its growing writes.
  • A function without allow-oom still has every write refused.
  • Cold-tier housekeeping (moon#1231, moon#1240, refs moon#1253):
  • A compacted output whose keys all changed while its listing was committing is reclaimed, instead of staying on disk until a restart.
  • A listed spill file found missing at boot is retired by the first sweep, so the "cold index rebuild DEGRADED" alarm appears on one boot, not on every boot until a rewrite.
  • The set of in-flight spills superseded by a write (moon#1253) is bounded even when a completion never arrives.
  • Tests: the fan-out probe source scan (perf_ws7_repl_offsets) normalizes CRLF, so it passes on a Windows checkout.
  • P0: a key deleted, flushed or overwritten while its disk-offload spill was still in flight came back after BGREWRITEAOF + restart (moon#1253). The moon#1215 fix covered keys that were already cold. It missed a DEL that lands while the spill request is in flight when a fold runs before that spill completes.
  • The rewrite now writes a head DEL for every spill request a write retired in flight.
  • Every completion outcome, including one after SWAPDB, settles its request.
  • FLUSHALL still leaves used_memory at 0.
  • A real-server suite, crash_recovery_cold_del_inflight_1253, now runs nightly. Before the fix, 1–12 probes per case came back.
  • An expired in-flight spill no longer drops a collection write that re-creates its key (moon#1255). Before, an acknowledged RPUSH/HSET/SADD/ZADD could be lost on restart.
  • Keyspace events name the right database after SWAPDB, and after a restart from an AOF with an RDB base (moon#1234 review). An RDB load used to reset every database's number to 0.
  • Lua scripts no longer read an expired key as live. Routed scripts, tokio's local scripts and MULTI-queued scripts now refresh the database clock.
  • Script bodies and pub/sub channel names no longer pin the connection's read buffer (moon#1160).
  • FT.SEARCH SESSION on the hybrid and sparse paths filters in place, and L2 vectors whose components are below f16 precision keep their quantized distance in the rerank, on HOT and WARM segments.
  • The test suite passes on macOS and Windows again:
  • the vector checksum test pins libm-dependent goldens only on Linux x86_64 (the product issue is moon#1256);
  • the APPEND amortisation bound is per platform;
  • the WAL poison, LFU, BGSAVE-abort and MSET-during-BGSAVE tests are deterministic.
  • Collection elements no longer pin the connection read buffer (moon#1160).
  • Hash fields and values, set and sorted-set members, list elements, and stream fields and names are now stored as exact-size copies.
  • AOF replay streams through a bounded 1 MiB buffer.
  • 100K small set members held +561 MiB RSS; they now hold +10 MiB.
  • Streams are charged to used_memory (moon#1163), about 206 B per one-field entry, so maxmemory binds on stream workloads. DEL credits exactly what was charged, and MEMORY USAGE of a stream reports its measured size, now in O(1). Durable MQ queues are billed separately in moon#1250.
  • Listpack backlen bytes use redis's byte order and widths (moon#1206). Backward walks over entries of 128 B or more are now correct.
  • DEL, UNLINK and GETDEL publish the del keyspace event (moon#1234), once per key removed. This holds on every path: plain, spanning shards, MULTI/EXEC and Lua. GETDEL of a missing key publishes keymiss.
  • SCRIPT LOAD racing SCRIPT FLUSH, and FUNCTION LOAD racing FUNCTION FLUSH, no longer leave shards disagreeing (moon#1235). At --shards 4, mixed trials went from 173/200 (SCRIPT) and 167/200 (FUNCTION) disagreeing to 0.
  • An LREM that removes nothing, and an LINSERT with a missing pivot, no longer abort a WATCHing EXEC (moon#1226).
  • In a pipeline with a protocol error, the commands before the error are answered first, as in redis (moon#1226). This covers commands deferred behind a blocking pop, SUBSCRIBE or the cross-shard ordering guard, and the fast-path GET/SET replies. RESP2 subscriber mode reports the error instead of closing silently.
  • A prepared vector query no longer returns empty results when its collection lacks the sub-centroid table (moon#1226).
  • A .tpost file with doc-id holes is no longer refused at boot (moon#1220). Refusing it forced a full text-index rebuild.
  • P0: a deleted cold-tier key came back after an AOF rewrite and any restart (moon#1215, pre-existing in v0.8.9, default config). The cold index now remembers every dead slot still on disk, and every rewrite opens its new generation with plain DEL records for the keys that are dead at the fold instant. The real-server crash suite went from 44–157 resurrected keys per case to 8/8 cases passing at --shards 1 and --shards 4 on both runtimes. Every moon version replays DEL, so a downgrade keeps the deletes until the older binary runs its own rewrite. The dead-slot ledger is bounded:
  • only slots that can come back are recorded (not slots whose own TTL has passed, and nothing without an AOF);
  • maxmemory binds on the ledger at write admission (deletes are never refused), without making eviction push live keys out to pay for it;
  • the rewrite streams its head DELs in bounded chunks;
  • mostly-dead spill files are compacted and reclaimed once a committed rewrite makes that safe.

INFO reports the ledger as cold_dead_slots / cold_dead_slot_bytes. Under a 16 MB maxmemory cold-delete churn it settles at 0.15× maxmemory, peaking at about 1.3–1.4× right after a burst of deletes, where writes answer OOM until reclaim runs. Without the bound it reached 5×. - P1: a key whose spill was in flight when a BGREWRITEAOF fold cut its base could be lost (moon#1223) if the spill then did not publish (the marker was refused under AOF backpressure, the pwrite failed, or the file id was re-issued). The fold base image now includes in-flight spill payloads. On ae21476, a kill -9 run lost 1,067 of 63,049 acknowledged keys. A payload that cannot be re-encoded now fails the rewrite, which keeps the previous generation, instead of committing a generation without the key. - rdb_last_bgsave_status now reports the last save (moon#1230). Before, one failed sharded BGSAVE latched err forever and every later SHUTDOWN SAVE was refused. LASTSAVE/rdb_last_save_time and rdb_changes_since_last_save now move only on a successful save. A shard that cannot write a snapshot (lost data dir) now fails the save instead of leaving rdb_bgsave_in_progress:1 forever, and sharded auto-saves start as counted saves. With no persistence directory at all, BGSAVE and SHUTDOWN SAVE now answer an immediate -ERR instead of starting a save that can only fail. - A spanning DEL/UNLINK, a fanned-out FLUSHDB/FLUSHALL, or an HSET/HDEL executed on another shard now updates vector and text indexes and drops durable queues on every owner shard (moon#1162). Before, all 20 deleted docs were still returned by FT.SEARCH at --shards 4. - SCRIPT FLUSH now empties every shard's script cache before replying (moon#1229). Before, EVALSHA still ran for up to 36 of 48 keys at --shards 4. An invalid mode now returns redis's error. - The MSET coordinator's local leg now captures BGSAVE pre-images (moon#1228, MSET item), so a snapshot taken during a spanning MSET stays point-in-time. - MSET and MSETNX emit a keyspace set event per pair, as redis does. This applies on every path: a single shard, spanning owner legs, MULTI, Lua, replica apply and AOF replay. Before, a spanning MSET notified only some keys, and the local leg and --shards 1 notified none. - A delete routed to another shard is never refused for memory. DEL/UNLINK sent to an owner shard over its budget answered -OOM under noeviction (144 of 200 at --shards 4); the connection's own shard already let deletes through. - P0: BGSAVE lost pre-snapshot keys when a DashTable segment split during the save (moon#1216): 1,744 of 2,000 keys in the repro, 698K–726K of 1M under an insert flood. The epoch now walks hash space, so a split cannot move keys out of the walk. Copy-on-write also captures every key a multi-key write modifies, not only the first (moon#1217; MOVE / COPY … DB n remain), including a woken BLMOVE's keys. - FLUSHDB / FLUSHALL / SWAPDB during a BGSAVE no longer panics the shard (moon#1224; see Changed), and an aborted save's writer is cancelled and joined before the next save can reuse its temp file. Blocking commands served immediately (BLMOVE, BLPOP, BZPOP, XREADGROUP, also inside MULTI) capture their keys' pre-images too (moon#1217). - Sorted-set B+tree corruption (moon#1205): an internal split lost a separator and a subtree and a right-borrow lost a subtree count, so ZRANK answered nil for existing members and ZREM'd members came back in ZRANGE for any zset over 128 members built from non-monotone scores. Restart after upgrading: persisted zsets rebuild a correct tree. - Vector index sidecars lost RERANK_MULT and EXACT_BEAM on every restart (moon#1194): the writer wrote v4. - Cross-shard TAG/NUMERIC consistency suites run again* (moon#1219) — on shards 1 and 4 against an in-test oracle, no longer #[ignore]d.

  • A failed WAL v3 fsync is never followed by a durability claim (moon#1221 review, refs moon#1188). Once the off-loop sync agent's fsync of a segment failed, the writer completed the pending rotation by fsyncing the same file inline — which succeeds after the kernel has reported the error to the agent's descriptor — and published the watermark over it, so the checkpoint's log-before-data wait accepted pages whose WAL fsync had failed (a retried inline fsync did the same without an agent). Any failed fsync now poisons the WAL for good: every durability wait checks the poison before the watermark and fails, no fsync is retried, a pending rotation fails loudly and past its memory bound opens the next segment with no durability claim, a graceful shutdown still writes buffered records to the page cache, and an inline fsync publishes only once the agent fsyncs of the same file that could have consumed the error have settled. The decisions are loom-checked in tests/loom_wal_sync_agent.rs against the shipping code.
  • A spill's MOON.SPILLED cut record is never dropped under AOF backpressure (moon#1202). When the AOF writer's channel stayed full past the 500 ms bound, the record was dropped while its keys stayed published to the cold tier. Recovery was still value-correct, because the log rebuilt those keys in RAM, but the drop was reported as a lost acknowledged append (aof_last_append_status:err) and the keys came back in RAM instead of cold. A later retry was not an option: a record logged after a newer write to the key makes replay lose that write. The record is now logged before the keys are published. If the writer cannot take it, the spill is withdrawn: the keys go back to RAM, the file stays out of the manifest (the startup orphan sweep removes it), and the next eviction pass spills them again. New INFO persistence field: spill_completion_marker_withdrawn. One backpressure budget now covers every record in a completion drain, so a saturated writer blocks the shard thread for at most 500 ms per drain instead of 500 ms per spill file.
  • A spilled value larger than ~4 MB is readable again (moon#1201). The cold-tier overflow-chain reader capped a chain at a fixed 1000 pages (~4.03 MB), while the spill writer has no size cap, so every larger value — in practice a consumer-group stream whose PEL grows without XACK — was written intact but refused on read as OverflowBroken: the key stayed indexed and answered IOERR to XADD, XREADGROUP, GET and every other reader until overwritten (in v0.8.9 and earlier the same read was a silent miss, so XADD started a fresh stream and the old one was lost). The cycle guard is now the file's own page count, which no acyclic chain can exceed. No on-disk format change; files written by any earlier version read back.

  • An AOF rewrite that is started while the previous one is still draining is no longer lost (moon#1158). A per-shard rewrite released the in-progress flag when its manifest committed, before every writer had written out the appends that spilled during the fold. On a real disk under sustained writes that drain takes seconds, so the auto-rewrite monitor could start the next rewrite during it. The drain then consumed and dropped the new rewrite request. Its countdown never finished, the flag stayed set, and no rewrite ran again: the incr AOF grew until appends were dropped and the disk filled. The flag is now released when the last writer finishes its drain. The same ordering is fixed on the --shards 1 tokio writer, and a per-shard rewrite whose fan-out fails part-way no longer clears the flag while the writers that received it are still working. A rewrite request that reaches a drain anyway aborts loudly instead of disappearing, and the monitor logs an error when a rewrite has been running for more than 5 minutes.

  • Reads never refreshed LRU/LFU metadata, so allkeys-lru/allkeys-lfu evicted the most-read keys first (moon#1161): hot-key retention in the review scenario 0.6% → 99.5% (LRU) / 100% (LFU). OBJECT IDLETIME resets on read, OBJECT FREQ grows under LFU, TOUCH touches, OBJECT on a missing key answers nil. lfu-log-factor / lfu-decay-time now take effect, and MEMORY USAGE, DEBUG OBJECT and DEBUG DIGEST no longer count as access (moon#1211).

  • LPOS / LMOVE Redis parity (moon#1209): option errors are checked before values (ERR syntax error), the RANK 0 and COUNT/MAXLEN error texts match redis, LMOVE k k keeps the key's TTL, and a refused LMOVE no longer converts its source list.
  • FT.SEARCH panicked on a LIGHT TQ4A2 index over a non-empty mutable segment (moon#1207), and EXACT compaction gave rows after a deleted vector their neighbour's QJL signs (moon#1208).
  • FT.SEARCH prefix/fuzzy results differed between processes (moon#1218): the 50-term expansion cap broke document-frequency ties in HashMap order (and, with an FST plus newer terms, by an unstable re-sort). One capped selection now spans the FST and the newer terms and keeps the top 50 by document frequency, ties by term bytes — a function of the matching terms and their frequencies only, so it no longer depends on term-id assignment (which a rebuild or a replica makes differently).

  • The boot crash-orphan sweep no longer lets a cold file id be issued twice in one AOF generation (moon#1114). A spill's MOON.SPILLED <N> marker reaches the AOF before its deferred manifest commit; a crash in between leaves heap-<N>.mpf on disk with no manifest entry. The next boot deleted it as a crash orphan, and nothing recorded N once it was gone, so the boot after that resumed the id counter at or below N and re-issued it. The old marker then authorised the new file early on replay: RPUSH X a once read back a a. Recovery now commits the highest orphan id to the manifest as a page-less Tombstone (an id reservation, no format change) before the sweep deletes anything, and deletes nothing if that commit fails. At most one entry per boot whatever the orphan backlog; a crash between the commit and the unlink re-sweeps the file on the next boot instead of leaking it.

  • A document written to a shard while that shard runs its boot index rescan stays searchable (moon#1124). Writes routed from another shard are applied during the walk (the -LOADING gate guards only the connection path), and the rescan's deletion probe tombstoned every recovered document whose key the walk had not observed — including a key that was deleted while the server was down and re-created by a routed HSET mid-walk. The key existed and FT.SEARCH never returned it. The live auto-index hook now records (db, key) in a shard-local ledger while recovery runs, a live DEL/UNLINK forgets it again, and both the vector and the text probe skip a recorded key. Outside recovery the ledger is None: one branch per indexed write. A 2-shard kill-9 restart that lands the write mid-walk (tests/vector_rescan_live_write.rs) lost the document on every in-window run before the fix (4/4) and keeps it after (6/6).
  • A plain XREADGROUP, and XCLAIM / XAUTOCLAIM, are logged as the effect they had, so consumer-group state survives a restart and reaches replicas intact (moon#1130). All three were written to the AOF and the replication stream verbatim, and their outcome depends on the clock: the replayed read stamped every pending entry with the REPLAY time (idle times restarted from zero after a kill -9), and a replayed XCLAIM ... 1000 then found the entry a few ms idle and claimed nothing, so XPENDING named the old owner again. They now propagate as redis does: a > read as XCLAIM key group consumer 0 <ids> TIME <delivery-ms> RETRYCOUNT 1 FORCE JUSTID LASTID <last> per stream served (XGROUP CREATECONSUMER + XGROUP SETID under NOACK), a claim as XCLAIM ... 0 <claimed ids> TIME <ms> ... FORCE with its own RETRYCOUNT/JUSTID/LASTID, and an XAUTOCLAIM as one XCLAIM of the entries it claimed and the deleted ones it dropped. Reads that deliver nothing, and history reads, log nothing. The same records are written from every executor — connection, cross-shard, MULTI/EXEC and Lua. A read over several streams writes one record per stream, each its own AOF record.
  • The library entry point logs a write before another connection can log a later one (moon#1099). handler_single, which listener::run_with_shutdown and moon::server::handle_connection drive, collected a pipelined batch's AOF records and sent them after the batch: after its WAIT / CLIENT PAUSE awaits, after the reply writes under everysec, and always after the db lock was released. A write from another connection that was applied later could be logged first, and replay restored the older value (MULTI / SET k 1 / EXEC / WAIT with a SET k 2 during the WAIT replayed 1; eight connections pipelining APPENDs to shared keys replayed a different order on every run). Each record is now enqueued while the guard it was applied under is held — inside EXEC's body, inside MOVE / COPY ... DB n's two-db section, and with every db held across FLUSHALL's clear. Room in the writer is awaited before the guard is taken; the appendfsync always fsync barrier stays one per batch. A record that cannot be enqueued now turns its reply into WRITEFAIL under everysec / no too, where it used to be dropped after +OK. The shipped binary (run_sharded) was not affected.
  • A script's plain COPY to an undeclared key on another shard is refused with CROSSSLOT instead of acking a copy nobody can read (moon#1133). On a connection a DB-less COPY is coordinator-routed and correct across shards; a script has no coordinator, and route_script_keys sees only the declared keys. So EVAL "return redis.call('COPY', KEYS[1], ARGV[1])" 1 src dst with dst owned by another shard wrote the copy into the source's shard under a name normal routing never looks for there — at --shards 4, :1 for 8 of 8 destinations, 1 of 8 readable, DBSIZE counting all 8. The scripting bridge's cross-shard guard now covers plain COPY (with or without REPLACE) like the rest of the two-key write family, for EVAL, EVALSHA and FCALL. Same-shard and {hash}-tagged pairs, --shards 1, and COPY sent on a connection are unchanged.
  • test: the ACL CAT diff in scripts/test-consistency.sh no longer truncates the suite under a non-C locale. Its sort -u and comm ran under the caller's collation; under en_US.UTF-8 GNU comm rejected the sorted files as out of order, set -e ended the run, and the consistency gate reported a TRUNCATED RUN on a freshly provisioned Ubuntu 26.04 VM (reproduced on main b65a73aa). Both now run with LC_ALL=C; the gate completes, tolerating only the documented moon#536 row.

  • A key spilled again after a BGREWRITEAOF keeps its pre-rewrite value across a kill -9, even when the respill's MOON.SPILLED marker never reached the AOF (moon#1140). The marker is emitted into the AOF writer's channel while the spill file's manifest entry is committed by the manifest-sync thread, so a process death (or AOF backpressure) can leave the file named by the manifest with nothing in the log that authorizes it. The rebuilt cold index points the key at that newest file, the MOON.COLDCUT gate hides it, and the copy below the cut — the only base the generation's records were written against — was shadowed by it. Every post-rewrite write then replayed onto an empty value and the end-of-replay resolution kept that truncated copy: a list acknowledged as a,b,c,d,e,f came back as f, an INCRBY counter as 1 instead of 8. A gated replay now falls back to the key's newest AUTHORIZED copy, which is what the newest-wins index would hold had the hidden file never been written; the rebuild keeps the copies it superseded for the length of the generation and releases them when it closes. Reproduced on both runtimes at --shards 1 (a 200 ms kill -9 after the respill failed 2 of 4 runs; the crash image fails recovery every time, and appending only the missing marker to it recovers correctly), and present on 3b596be0 too — it predates moon#1085, moon#1075 and moon#1118.

  • Waiters a refused AOF record left parked are served once the writer has room (moon#1111). When the AOF writer could not take a served pop's record within the backpressure bound, the wake stopped (every further serve would have waited out another bound on the shard thread) and the waiters behind it stayed parked beside data until another write to the key or their own timeout — never, for BLPOP k 0 on a queue whose producer went quiet. The key is now remembered and retried by the shard's 10 ms blocking tick as soon as the writer has room again (never while it is still full, so the retry spends no bound on the shard thread), each serve still logged in its own synchronous stretch; stream group readers skipped for the same reason are retried the same way.

  • Becoming a replica releases every blocked client. A client parked on a node that then ran REPLICAOF host port stayed parked. Before the fix above, nothing woke it until its timeout. After it, the first replicated push served it: a BLPOP took the element out of the replica's copy only. Every parked client is now answered -UNBLOCKED force unblock from blocking operation, instance state changed (master -> replica?) and its connection closed, dropping anything pipelined behind the blocking command. That is redis's disconnectAllBlockedClients (text and close checked against redis-server 8.6.1, for BLPOP, XREADGROUP and XREAD).

  • A replica wakes the clients blocked on what replication writes (moon#1096). Every write a replica holds arrives through replication::apply, and nothing there reached the ready-key hook, so an XREAD BLOCK parked on a replica answered nil at its own timeout while its stream filled. Each applied command now serves the clients blocked on the keys it wrote, through the master's own hook — its written keys, both databases of a SWAPDB, the destination of a MOVE/COPY ... DB n — after the write is applied, as redis does (signalKeyAsReady from the keyspace write; measured against redis-server 8.6.1: woken 0.41 s after a master XADD issued 0.4 s into the wait, also through MULTI, MOVE and SWAPDB). Blocking pops and XREADGROUP stay refused with -READONLY on a replica, in both servers.

  • A parked XREADGROUP is answered at once when its stream or group goes away (moon#1086). Redis unblocks a group reader when any write deletes or retypes its stream or destroys its group; moon left it parked until its own timeout and then answered nil. Every such change now answers it immediately, with redis-server 8.6.1's text: -NOGROUP No such key 'k' or consumer group 'g' in XREADGROUP with GROUP option after DEL, UNLINK, RENAME away, MOVE, SWAPDB, FLUSHDB, FLUSHALL, XGROUP DESTROY or expiry, and -WRONGTYPE after SET or a RENAME of another type onto it — on the connection, inside MULTI, from a script, and across shards. A write that names the key signals it directly (while a group reader is parked, the ready-key gate admits every write, not only list/zset/stream writers, and an inline SET takes the generic path); a flush, the source side of a MOVE and an expiry mark the shard for a recheck its 10 ms blocking tick runs. XREAD waiters stay parked, as in redis. The same text now answers an XREADGROUP issued against a missing key or group (it said ERR The XREADGROUP subcommand requires the key to exist. or NOGROUP No such consumer group for key name). A multi-stream XREADGROUP now checks every stream, group and id before it reads any stream. It used to move the first stream's entries into the PEL and then fail on the second. Ids are checked with redis's texts:

  • $ and + get ERR The $ ID is meaningless in the context of XREADGROUP: ... (and the + form of it) instead of being parsed.
  • A malformed id (-, bad, 1-x, an out-of-range number) gets ERR Invalid stream ID specified as stream command argument. + and - used to be accepted as ids.

  • A blocked XREADGROUP that times out as it is served no longer strands entries in its PEL (moon#1047). The owner's stream waker ran the read — moving the entries into the consumer's PEL and advancing the cursor — before claiming a remote waiter; a waiter whose timeout settled its claim in between answered nil while the entries sat undelivered in its PEL, and a > loop never saw them again. The waker now claims first and reads only on a won claim, the claim-token order moon#1045 gave the list and zset wakers. The readiness check before the claim is read-only. It used to take the stream mutably, and a wake that served nothing then aborted every EXEC watching the stream. A consumer the check must create does not signal watchers, as in redis.

  • XCLAIM honours its options (moon#1104). Every argument that parsed as a stream id was claimed, so RETRYCOUNT 1 claimed entry 1-0 and a TIME value claimed another, and IDLE, TIME, RETRYCOUNT, FORCE, JUSTID and LASTID were ignored — which also meant the records redis propagates for a group read (and moon now logs) could not be replayed. The ids now run until the first non-id argument and every option is applied as in redis 8.6.1, with its error texts: NOGROUP No such key 'k' or consumer group 'g' for a missing key or group (it answered an empty array), Unrecognized XCLAIM option, and the Invalid ... argument for XCLAIM family.

  • A blocking XREADGROUP that delivers is logged, and survives kill -9 (moon#1104). A consumer-group read moves what it delivers into the PEL and advances the group's last-delivered id, but through BLOCK it went through the blocking intercept and nothing reached the AOF or the replication stream: after a restart the PEL was empty, the cursor was back, and the next > reader was handed the same entries again (every row of the new kill -9 test, --shards 1 and --shards 4, immediate and parked, with and without NOACK). The shard that serves the read now logs, in the read's own synchronous stretch and before the reply leaves, what redis-server 8.6.1 propagates for it: one XCLAIM key group consumer 0 id TIME t RETRYCOUNT n FORCE JUSTID LASTID id per delivered entry, XGROUP SETID for the cursor, and XGROUP CREATECONSUMER for a NOACK read that created its consumer (ENTRIESREAD is omitted because moon keeps no such counter, and the records are not wrapped in MULTI, which moon's log never carries). Under appendfsync always the reader confirms the fsync on the stream owner's writer, as a blocking pop does; a record the writer refuses answers the reader with the standard MOONERR AOF backpressure error.

  • A cross-shard COPY keeps the source's absolute expiry (moon#1095). The TTL crossed as a relative duration — PTTL read on the source shard's cached clock, PEXPIRE re-anchored on the destination's — so the copy's deadline moved by the drift between the two (measured: PTTL 500004 after PEXPIRE 500000; 5 to 32 of 64 copies moved per run). The copy now carries the source's PEXPIRETIME and is written as one SET dst v PXAT <deadline> [NX], which also closes the window in which a crash between the old SET and PEXPIRE left a copy that never expired. Every other cross-shard hop that could carry a TTL was audited: RENAME/RENAMENX, SMOVE, LMOVE/BLMOVE, SORT ... STORE and the *STORE family are refused across shards (CROSSSLOT, moon#592/#570), and MOVE / COPY ... DB n stay on one shard, carrying the stored entry and its absolute deadline.

  • MOVE and COPY ... DB n work inside scripts, and land where the effect record says (moon#1068). A script reaches the keyspace through the one database it runs in, so redis.call('MOVE', k, n) answered ERR MOVE requires handler-level dispatch, and redis.call('COPY', a, b, 'DB', n) answered :1 while writing b into the SCRIPT's database — and logged COPY a b DB n to the AOF and the replication stream, which a replica and a restart apply into db n. Both now run from EVAL, EVALSHA, FCALL and a script queued in MULTI exactly as on the connection, as redis 8.6.1 does: the reply, the landing database, the key's absolute deadline, REPLACE, redis's errors for the script's own db and an out-of-range db, and one verbatim effect record on :1 only. A client blocked on the key in the destination database is served once the script returns, after the record is logged (the moon#1056 rule), so a restart neither loses nor resurrects the element it popped.

  • scripts/test-consistency.sh: six rows no longer fail with CROSSSLOT at --shards 4 (moon#1106). ZRANGESTORE still-negative stop and five CLIENT TRACKING destination controls named keys on different shards, so they tested the routing refusal instead of the command. Their keys (and their sibling rows') now share a {tag}. The redirect-transcript drain used read -t 0.3, which macOS /bin/bash 3.2 rejects ("invalid timeout specification"), ending the drain at once; it now uses -t 1.

  • CLIENT INFO / CLIENT LIST report a subscriber's flag, counts and protocol, and RESET from RESP2 subscriber mode resets everything (moon#1105), matching redis 8.6.1 on both runtimes. A subscribed client was listed as flags=S (redis's REPLICA flag) with sub=0 psub=0 ssub=0 and resp=2 whatever it held; it is now P with its real channel, pattern and shard-channel counts, every client's resp follows HELLO, and flag characters combine in redis's order (Px, Pb, then t/R/B) instead of keeping only the first. RESET sent from the RESP2 subscriber loop only unsubscribed, so the connection kept its db, CLIENT TRACKING, name and authentication; both handlers now run the same RESET as everywhere else. That shared RESET also tore down only channels and patterns: a RESP3 client's SSUBSCRIBE survived it, and SPUBLISH still counted it as a receiver. It now clears all three namespaces and the remote shard maps.

  • A restart no longer drops vector documents whose hash is in the cold tier or carries a field TTL (moon#1074). At boot, index recovery walks the keyspace, and any recovered document whose key the walk did not see is deleted as "removed while the server was down". The walk read only the hot table, and within it skipped the per-field-TTL hash encoding. So every document whose HASH allkeys-lru eviction had spilled to the cold tier, and every document that had had HEXPIRE applied to one of its fields, dropped out of FT.SEARCH after a restart, while EXISTS/HGETALL still returned the key. Measured on a 300-document index pushed to the cold tier: 299-300 of 300 documents were unfindable after a clean restart, and 294-295 of 300 after kill -9 (the rest were hot at boot). A cold document that was still in the mutable segment at shutdown was never re-indexed at all. The walk now reads cold-tier hashes from disk (one page read per key, in file order) and reads field-TTL hashes without their expired fields. A cold entry that cannot be read keeps its recovered document: a read fault is not a delete. The walk lists keys up front but takes each key's payload only when it reconciles it: the keyspace is not frozen while the walk runs (writes routed from another shard are applied, and eviction spills), so a payload captured at the start could be stale by the time it was reconciled. Text indexes loaded from .tpost are reconciled by the same walk.

  • A re-written or deleted vector document stops matching in every tier, at runtime and across a restart, and FT.INFO num_docs counts each live key once (moon#1066, moon#1073). Measured on a real server, 1000 keys, before the fix:

  • an HSET re-write of a key whose old copy was WARM or COLD left that copy live: KNN for the OVERWRITTEN vector still returned the key top-1, and num_docs read 1001. The update path tombstoned the mutable and HOT segments only; DEL went through every tier. Both now share one sweep (SegmentList::tombstone_key).
  • across a restart, a HOT segment's copy of a re-written key came back live: Stack B writes a segment once, so a tombstone applied in memory is gone after the reload, and recovery admitted a key when ANY loaded row had its key_hash. After a kill -9 with the re-write still in the mutable segment, the key answered ONLY to its overwritten vector and its current one was lost. A DEL'd key came back as vec:<id>. Recovery now keeps a loaded row only if it is the copy the persisted keymap names (key_hash AND global_id), once, and tombstones every other row in its own segment -- no key_hash-wide tombstone, no on-disk format change.
  • a row killed at install time (the key was deleted or re-written while its background build ran) came back live after HOT -> WARM: mvcc.mpf carried the delete_lsn, and the warm reader ignored it. It is now a dead row in WARM and COLD, by position, so a live sibling copy of the key stays.
  • num_docs was wrong both ways: a HOT segment never counted a steady-state tombstone (1000 after a DEL of 1000), and every WARM/COLD segment counted every tombstone whether it held the key or not (1998 after one DEL across two segments). A tombstone is now recorded and counted only by the segment holding a live row for the key, through a per-segment key_hash index (4 bytes per row in HOT and WARM segments; a COLD stub keeps 8 bytes per live row). Tombstone sets no longer grow with every DEL in every segment.
  • FT.COMPACT draining an in-flight background merge replayed the sources' tombstones key_hash-wide, killing the NEW copy of a key re-written while the merge ran; it now uses the origin-gated replay the background install already used.
  • a synchronous merge (VACUUM VECTOR, and the autovacuum pass) installed its output without replaying the sources' steady-state tombstones at all: after a DEL, its row matched again top-1 as vec:<id> and was counted (num_docs 2000 instead of 1999); after a re-write whose new copy was still in the mutable segment, the old copy matched beside it. It now replays them like the background installs.
  • at boot, a WARM row killed at install time counted as evidence that its segment served the key. When the keymap named that row's copy (recovery kills a duplicate of the current copy that way), a segment holding the dead duplicate could claim the key first and the live copy in another segment was then tombstoned, losing the document. Only live rows count now.

  • A slow everysec fsync no longer makes a multi-shard server answer "write applied in memory but not queued for persistence" (moon#769).

  • Before. At --shards > 1 a write to a key another shard owns runs on that shard, and the shard applied it first and only then waited 5 ms for room in its AOF writer's 10k channel. A writer stalled on a slow fsync fills that channel quickly under pipelined load, so such writes were refused after they were already in memory, and their records never reached the AOF. Redis and Valkey complete the same redis-benchmark -P 16 workload. The write leg on the connection's own shard already waited up to --aof-fsync-timeout-ms (2 s by default).
  • Now. Before a routed write runs, its shard checks that the AOF writer has room for the write's records. The shard thread never waits for it: a write that finds no room stays queued, unapplied, at the head of the queue from the shard that sent it, and is retried on every loop tick, while reads, local connections and the other shards' queues keep flowing. Later commands from the same sending shard wait behind it, so nothing overtakes an earlier write. A stall shorter than --aof-fsync-timeout-ms (2 s by default; 0 means 10 s here, and the wait never exceeds 10 s) is absorbed. A write still without room after that is refused without being applied, with -MOONERR AOF backpressure: command not executed, the AOF writer is stalled; retry, and while the writer stays stalled the next routed writes are refused at once instead of each waiting again. The keyspace, the AOF and the replicas keep agreeing, and the client can retry.
  • A multi-shard MSET, DEL or UNLINK whose part on a stalled shard is refused while its other parts ran answers -MOONERR AOF backpressure: command partially executed; ... instead, because it was not "not executed". A multi-shard FLUSHALL/FLUSHDB keeps its existing MOONERR FLUSH partial reply.
  • The check applies under every appendfsync policy (always, everysec, no): it is about room in the writer's channel, which all three use. Admission reserves room for every write it lets through in the same loop pass, plus 256 spare records for what it cannot count in advance (eviction deletes, a script's extra writes, other shards' fsync barriers). A single write that needs more than that spare can still meet a full channel after it applied and get the existing fail-loud error.
  • While a rewrite is folding, writes spill to the rewrite overflow instead of waiting. The script-load fan-out budget grows by the admission wait, so a slow disk is not reported as a divergent shard.
  • INFO persistence gains aof_backpressure_stalls (routed writes that waited) and aof_backpressure_refused (commands refused unapplied).
  • The write leg on the connection's own shard is unchanged. It still applies the write, waits up to the bound on its connection task, and then reports ERR AOF fsync failed; write not durable.

  • ZRANGESTORE checks its range before it looks up the source (moon#1102), the ZRANGESTORE sibling of moon#1060. A malformed rank, BYSCORE or BYLEX bound against a missing source answered :0 and deleted the destination; against a source of another type it answered WRONGTYPE. Redis parses the range first, so both now answer its parse error (min or max is not a float, min or max not valid string range item or value is not an integer or out of range), with or without REV and LIMIT, and the destination is left alone.

  • A WAIT queued inside MULTI no longer holds EXEC for its timeout (moon#1098). Redis runs a transaction body with CLIENT_DENY_BLOCKING, so a queued WAIT answers the current replica ack count at once. Moon filled it with the live WAIT, which polled until its deadline: MULTI / INCR k / WAIT 1 1500 / EXEC answered after 1.5 s instead of at once, and WAIT 1 0 inside MULTI never answered. The reply is still the number of replicas that have acknowledged the master's current offset, now sampled once, on both runtimes and on the local and owner-routed EXEC paths. A WAIT outside MULTI still blocks as before. (WAITAOF is not implemented, so it has no queued form to fix.)

  • The disk-offload checkpoint only publishes a redo point over heap pages it has made durable (moon#452 item 3). The fuzzy checkpoint pwrote dirty heap pages and never fsynced a heap file, then published redo_lsn in the control file and recycled every WAL segment below it — FullPageImages included. A power loss after that recycle could roll a page back with nothing left in the WAL to redo it. Four changes close the protocol:

  • Finalize fsyncs every heap file the flush wrote before it writes the WAL checkpoint record. The fsync runs on a helper thread and the shard thread never waits for it: each tick polls the helper without blocking, and a shard has at most one helper outstanding, so a hung disk leaves one stuck thread rather than one per retry. Only the shutdown checkpoint, which serves no clients any more, waits for that one helper, for at most WAIT_DURABLE_TIMEOUT counted from the helper's start. The WAL-ceiling checkpoint runs while the shard serves clients, so it never waits: it leaves a pending fsync to the periodic tick, and its emergency recycle cuts only below the redo point already published. A failed control-file write no longer leaves the unpublished redo point in memory, where that recycle would have cut below it. If a file cannot be opened, Finalize retries with backoff. If an fsync returns an error, the shard never publishes a redo point again, because a retried fsync can report success for pages the kernel already dropped. The WAL above the last good redo point is kept, and recovery replays from there.
  • A heap file that no longer exists does not hold the redo point back. The cold-tier GC unlinks a file once its last live key is gone, and only then; nothing references it and no fsync can reach the unlinked inode. So the checkpoint stops waiting on it, drops its pages from the dirty set, counts it and logs it at debug. If the manifest still lists the file as live, something other than the GC removed it: that is counted separately and logged as an error. Without this, every Finalize failed on NotFound and the WAL grew until the disk was full.
  • Log before data, one WAL wait per batch: the flush appends the FullPageImage of every page in the batch, waits once for the WAL to be durable through the batch's highest LSN, and only then overwrites the pages in place. The flush used to collect every image and append them only after the whole batch had been pwritten, into the WAL writer's memory buffer. A crash during the pwrite left a torn page and no image of it anywhere.
  • A page that fails any flush step (its image, the WAL wait, or its write) makes the checkpoint flush again, keeping its redo point. Before, the page was counted as flushed and the checkpoint finalized over a change that was on disk in neither the heap nor the WAL.

No production path marks a heap page dirty today (PageCache::mark_dirty has no callers outside tests), so this closes the hazard before the first in-place cold-page write can reach it. Pinned by the tests in shard::persistence_tick::checkpoint_tick_tests, persistence::data_file_sync and persistence::page_cache.

  • aof_last_append_status and aof_last_write_status return to ok after a rewrite that covers the dropped append (moon#1094). The status latched to err at the first dropped acked append and stayed there for the life of the process, even after BGREWRITEAOF folded the live keyspace, the dropped write included, into a fresh base. Each AOF writer now tags a drop with its fold epoch, and clears it when a fold COMMITS whose snapshot was taken after that drop (the same folded_below rule moon#455 uses to drop records the new base already holds). A drop after the snapshot, an aborted fold, or a drop on another writer keeps err; with the PerShard layout the status is the AND across writers. Redis clears aof_last_write_status on the next successful write because it keeps the failed data in aof_buf and retries it; moon drops the record, so only a covering rewrite makes the log whole. A reason-DEL drop no longer sets a second, never-clearing latch of its own.

  • An AOF rewrite no longer replays a write twice, and a rewrite that fails late no longer leaves the writer appending to a deleted file (moon#455).

  • Double apply. A rewrite split the append stream by position: whatever reached the writer after its snapshot went into the new incr. A write whose record arrives after the snapshot while its effect is already in the snapshot was therefore replayed on top of the new base after a restart. An INCR came back incremented twice, an LPUSH pushed twice. Records reach the writer late when the producer awaits between the mutation and the enqueue: an EXEC whose body holds WAIT, CONFIG or another connection intercept, the local half of a multi-shard FLUSHDB/FLUSHALL (which could then wipe writes the base had taken after it), the local slice of a multi-shard MSET, or any producer parked on a full AOF channel. Every record now carries the rewrite epoch in force when its mutation ran. Once a rewrite takes effect, the writer drops each record stamped before that rewrite's snapshot, wherever the record surfaces. An aborted rewrite drops nothing. INFO persistence counts the dropped records as aof_rewrite_late_records_folded.
  • Late failure. The writer opened the new incr only after the manifest had switched to it. If that open failed, the rewrite was reported aborted while the manifest named the new generation, and the writer kept appending, and fsyncing, into the old incr the switch had deleted. Nothing it wrote after that could be recovered. The new incr is now opened before the manifest switches, so every failure leaves the old generation committed and the writer on it. A failed manifest write also no longer advances the in-memory sequence past the one on disk.
  • SWAPDB during a rewrite. SWAPDB logged its record, awaited, and only then swapped. A rewrite that snapshotted in that gap had a base without the swap and dropped (or folded into the old incr) the record, so the acknowledged SWAPDB was gone after a restart. The record is now enqueued, the swap applied and the replication record emitted in one synchronous step; while the AOF channel is full, SWAPDB waits for room without holding its record, and is refused unapplied after --aof-fsync-timeout-ms. Under appendfsync always the fsync is now confirmed after the swap, so an fsync failure is reported on a swap that stays applied on every shard, like any other always write.
  • Directory fsyncs. A per-shard rewrite created the new incr after its last fsync of the shard directory, and the tokio --shards 1 rewrite renamed its new file into place without a directory fsync. After a power loss the committed manifest could name a missing incr, or the old appendonly.aof could return, losing every record written after the rewrite. Both directories are now fsynced before the new file is used.

  • The console lints, and CI runs it (moon#1082). pnpm run lint exited with a missing-config error because ESLint 9 reads only flat config and the console had none. console/eslint.config.js now wires the plugins that were already installed (@eslint/js and typescript-eslint recommended, react-hooks, react-refresh), and the unit job of console-integration.yml runs pnpm run lint before the tests. Its first run found a real bug: GraphCosmos returned its Canvas2D fallback before its hooks, so the re-render after a failed WebGL init called fewer hooks than the first render and React threw; only the parent's error boundary hid it. The fallback now returns after the hooks, with a unit test that failed before the change. badge.tsx stops exporting the unused badgeVariants, and no-unused-vars accepts a leading underscore, the convention tsc already applies under noUnusedParameters.

  • A COLD vector segment leaves unloaded with the search that reloads it, and a delete that lands while the reload is waiting to install is no longer lost (moon#1070). Since the off-loop reload pool (prod-hardening #18) a search only SUBMITTED the reload and answered from it; the reloaded segment sat in the pool until some later search installed it, while the COLD stub stayed the index's segment and the only place a DEL could be recorded. The install then threw the stub away without replaying it: a document deleted after the first search came back, as vec:<id>, on the next one (reproduced on a real server). The install now replays the stub's tombstones, and the yielding FT.SEARCH handlers install finished reloads as soon as the query that awaited them is back on the shard, so FT.INFO unloaded_segments drops to 0 with that search. Four tests/vector_idle_unload.rs tests that were #[ignore]d -- and silently red on main -- now run in CI; only the ps-based RSS measurement stays ignored.

  • A restart no longer re-issues a cold-tier file id, which could apply a write twice (moon#1067). A restart resumes the shard's cold file-id counter at one past the highest id the manifest or the disk still holds. Manifest tombstone GC could prune the entry that held the highest id once its file was reclaimed, and the counter then moved backwards. The AOF appends to the same generation across restarts, so the generation still held a MOON.SPILLED <id> record for the old file. On replay, that record made the re-issued id readable early, and a write logged before its key was spilled into the new file was applied on top of the value it had already produced. Measured with --appendonly yes --disk-offload enable and the tombstone retention at zero: RPUSH X a once, then spill, restart, spill again, restart, and LRANGE X read a a, after both kill -9 and SHUTDOWN, on both runtimes. The same sequence with the default retention (tombstones outlive the restart) read a. GC now keeps the tombstone that holds the highest file id until a higher id is in the manifest. That pins at most one manifest entry per shard. No on-disk format change.

  • CLIENT TRACKING never drops an invalidation silently, scripts are tracked, and a RESP2 subscriber's pipeline follows its live subscription count (moon#1088, moon#1089, moon#1090, refs moon#1078). Every wire reply was read off redis-server 8.6.1 over a raw socket first.

  • moon#1088: a tracking connection's invalidations went through a 256-slot channel with try_send, so one 400-key MSET delivered 256 pushes and the client kept serving the other 144 keys stale. The queue is now unbounded in slots and bounded in bytes by --client-output-buffer-limit-normal; past it the connection is closed through the CLIENT KILL path, which is what redis does at its output-buffer limit (measured: normal 8192 0 0 closes the tracker on that MSET). The limit bounds what is QUEUED, as redis's bounds its output buffer. A tracker that stops reading is always closed. A tracker whose connection drains the queue on another shard thread while the burst arrives can receive every push instead. A burst is written in socket writes that never exceed that limit, and a RESP2 connection, which redis writes nothing for its own invalidations, queues nothing. The RESP2/RESP3 switch takes effect at the HELLO itself, so a pipelined HELLO 3 / GET k / BLPOP still gets the push for k. A RESP3 tracker that runs MONITOR keeps receiving its pushes, as on redis; on monoio it used to receive none. A REDIRECT target that is subscribed receives through its pub/sub channel: one command's invalidations, or a whole script's or EXEC body's, now take one slot there instead of one per key, and a target whose channel is still full is disconnected rather than silently shorted.
  • moon#1089: a write made through redis.call invalidated nothing, and a read made inside a script was never tracked. The scripting bridge now applies tracking to every command a script runs, as redis does inside call(): writes invalidate (NOLOOP judged against the caller), reads register for the client that ran the script under its OPTIN/OPTOUT/CACHING state. EVAL, EVALSHA, EVAL_RO, FCALL, FCALL_RO, scripts routed to another shard and scripts queued inside MULTI are all covered, and a FCALL of a function that only reads no longer counts as a write of its keys. A script's tracking effects are recorded without a lock as it runs and applied in order under one tracking lock when it ends, so a script's redis.calls do not each contend for the process-wide tracking mutex.
  • moon#1090: commands a RESP2 client pipelined after its last UNSUBSCRIBE (or RESET) were refused with the subscriber-context error on monoio and left unanswered on tokio. The subscriber gate now judges each command by the subscription count it runs under, and hands the rest of the batch back to the normal path.
  • moon#1078: CLIENT LIST and CLIENT INFO report tracking — flags=t, R for a broken redirect, B for BCAST, and redir= (0 with no redirect, -1 with tracking off). A RESP3 REDIRECT target that neither subscribed nor enabled tracking still receives nothing: reaching it needs either a channel on every RESP3 connection, which keeps each one out of idle task parking (measured on monoio, macOS, --conn-park-secs 2: 2000 idle HELLO 3 connections report parked_clients:2000, the same connections holding a channel parked_clients:0), or a cross-thread wake of a parked connection, which the idle-park machinery does not have. That part of moon#1078 stays open.
  • A served blocking pop is logged by the shard that popped it, at the moment it popped (moon#1056, moon#1097). Since moon#827 a BLPOP/BRPOP/ BLMOVE/BRPOPLPUSH/BLMPOP/BZPOPMIN/BZPOPMAX/BZMPOP that popped propagated a synthesised record, but the WAITER's connection wrote it, after the reply had reached it. At --shards > 1 that is usually not the shard owning the key, so the record sat in the wrong shard's AOF and replay dropped it: 33 of 48 probes in the new kill -9 test (four owners, one waiter connection, appendfsync always) came back with the popped element restored. At any shard count the record could also land behind a later write to the same key ([a,b], pop a, LPUSH x recovered as [a,b] instead of [x,b], on disk and on a replica), and an AOF rewrite snapshot could fall between the pop and its record, applying the pop twice. The owner now appends the record to its own AOF and replication stream in the pop's synchronous stretch, before the reply leaves (wake-served, claim-won and immediately-served pops alike); the waiter only confirms the fsync on the owner's writer under appendfsync always. A write that wakes a waiter (a plain write, EXEC, MOVE/COPY ... DB n on the connection and on every cross-shard SPSC arm) now logs itself before it serves the waiter, so the pop always follows the push that fed it. A record the AOF writer cannot take within its backpressure bound answers the waiter with the same MOONERR AOF backpressure error as every other synchronous write, instead of the element; that error, like AOF fsync failed, means the element may have been consumed. One wake pass shares one backpressure bound across all the pops it logs, so a saturated writer stalls the shard thread once, not once per served waiter.

  • A write is logged in the order it was applied, even when it waits after applying (moon#1084). Three paths applied a write, awaited something, and only then appended it to the AOF (and, on monoio, to the replication stream), so a write another client made in between was logged first and replay applied the two in the wrong order:

  • EXEC with a queued connection intercept (WAIT, CLIENT, CONFIG, SCRIPT, FUNCTION, ...) logged its body after filling the intercept replies. SET k 0, then MULTI / INCR k / WAIT 1 1500 / EXEC with another client's SET k 5 landing while EXEC was parked in WAIT: the server acknowledged 5, and after kill -9 it recovered 6; a replica settled on 6 too. Seen on both runtimes at --shards 1 and --shards 4. The body is now logged and replicated right after it runs, before any intercept is filled; the intercepts do not change the keyspace. If the append fails, EXEC reports the error without running the intercepts, as the owner-routed path already did.
  • A typed FLUSHDB/FLUSHALL at --shards > 1 logged this shard's flush after broadcasting it to the other shards, so a key written to this shard during the broadcast was replayed BEFORE the flush and vanished after a restart. The flush is now logged before the broadcast, which also means a broadcast that fails part-way no longer leaves this shard's flush out of the log.
  • A scattered MSET logged its local slice after awaiting the remote legs, so a newer write to one of its local keys was replayed under the MSET value. The slice is now logged right after it is applied.

  • Three wire-parity gaps found probing redis-server 8.6.1 raw sockets (moon#1060, moon#1076, moon#1077).

  • ZRANGEBYSCORE, ZRANGE ... BYSCORE/BYLEX and ZREVRANGEBYSCORE (moon#1060) looked a key up before validating the min/max grammar, so a malformed bound against a MISSING key answered an empty array where Redis parses the grammar unconditionally and answers a parse error. ZRANGEBYLEX/ZREVRANGEBYLEX (moon#959) already had the order right.
  • Inside MULTI, a container subcommand with the wrong number of arguments (CLIENT CACHING, CLIENT SETNAME, CONFIG GET, CONFIG SET, ...) was queued instead of aborting the transaction (moon#1076). The queue-time gate checked the container's own arity and, since moon#670, whether the subcommand was known, but never the subcommand's own arity even though SUBCOMMAND_META records it. EXEC now answers EXECABORT the same way it does for an unknown subcommand.
  • The unknown-command error always appended , with args beginning with: and never listed the arguments (moon#1077). Redis appends that clause only when there is at least one argument, then lists each one quoted and space-separated with no commas, truncated at a combined 128-byte budget (and, C's %.*s being what it is, at an argument's first embedded NUL). One builder (command::helpers::err_unknown_command) now serves command::dispatch, command::dispatch_read and the MULTI queue-time gate, so the three cannot drift the way they had — the live paths never listed arguments and the queue-time gate listed them with ', ' where Redis has no comma. The builder writes raw bytes rather than a lossy String, so a non-UTF8 argument reaches the wire unchanged; CR and LF are still mapped to a space so a client-chosen argument cannot split this reply into a second, forged one on a pipelined connection (the same class of bug as moon#1031, whose general fix — the RESP2/RESP3 line writers — remains open).

  • Restart no longer deletes warm vector segments it is serving, and a warm segment is superseded per key instead of as a whole directory (moon#893). Boot recovery judged a warm segment "already covered" when ANY one of its key_hashes was already indexed, and then ran remove_dir_all on its directory. Two ways in, both measured on a real server with 1000 warm keys: re-inserting ONE key and compacting it into a HOT segment retired the whole directory on the next boot, and the other 999 vectors were re-encoded; and a manifest written by an older build that lists one segment id twice (its id counter re-issued live ids) had the directory attached on the first entry and deleted on the second. Worse, when the re-inserted key's own segment had also gone warm, recovery registered the key under its new global_id, attached the older segment, and deleted the newer one: after the restart the key answered to its OVERWRITTEN vector and its current one was gone. Now each key is decided on its own. The persisted keymap names a key's current copy by global_id. A copy that is not that one, or that another live segment already serves, is tombstoned in that segment only. The rest stay served from it. A directory is retired only when none of its keys is live, never one recovery attached. The directory is renamed out of discovery's reach before it is deleted, so a crash mid-delete cannot leave a half-deleted segment. Each segment id is handled once. Duplicate manifest entries are collapsed at boot and healed on disk, keeping the last. For a spill file, that last entry is the one that describes the file. ShardManifest::add_file now refuses a second entry for an (id, type) that is already listed (it returns DuplicateFileEntry), and a warm transition onto an id whose directory or entry already exists is refused before anything is written, where it used to commit the entry and then fail the rename with ENOTEMPTY. A spill batch whose file id the manifest already lists puts its keys back in RAM from their in-flight payloads, never publishes them cold, and is counted in the new INFO field spill_completion_id_rejected. A reattached warm segment raises the vector global_id allocator above its own ids, so a key written after the restart can never be given an id a warm row already holds. Covered by tests/warm_segment_restart_893.rs, which restarts twice (clean and kill -9) and checks that the directories persist, nothing is re-encoded, and FT.SEARCH is identical.

  • A write that lands data on a key wakes the clients blocked on it, whatever command wrote it (moon#1059, moon#1069). Only six "producer" commands (LPUSH, RPUSH, LMOVE, RPOPLPUSH, ZADD, XADD) used to wake a blocked client, so a BLPOP/BZPOPMIN/XREAD BLOCK stayed parked until its own timeout — or forever with timeout 0 — beside data it could pop when the key was written by RENAME, RENAMENX, COPY, MOVE, COPY ... DB n, SORT ... STORE, ZUNIONSTORE/ZINTERSTORE/ZDIFFSTORE/ZRANGESTORE, ZINCRBY, GEOADD, RESTORE, SWAPDB, a script (EVAL/FCALL), or any of those inside MULTI. A BLMOVE/BRPOPLPUSH served by a wake, or served immediately, pushed onto its destination without waking the BLPOP parked there, so move chains stalled at the first hop. The wake is now keyed on the keys a command WRITES (the shared key walker's write positions), not on command names; MOVE/COPY ... DB n wake the destination database and SWAPDB every key parked in either database. A wake-served move feeds its destination's waiters in the same pass, over a worklist bounded by the waiters parked when it began — chains and cycles are served as redis's handleClientsBlockedOnKeys serves them. All the keys one command, one EXEC or one script made ready form a single batch that is served before the keys the moves it serves push onto. So, with BLMOVE a c, BLMOVE b c and BRPOP c parked, MULTI; RPUSH a x; RPUSH b y; EXEC hands the BRPOP y and leaves c = [x], as redis does. A key that becomes the wrong type for its waiter still leaves the waiter parked, as in redis. The wake costs nothing measurable while a client is parked on the shard: a write that can only produce a string, hash, set, bitmap or HyperLogLog skips the key walk entirely, as redis's signalKeyAsReady returns early on type, and the registry is probed with borrowed keys. A 10-key MSET with one BLPOP parked runs within noise of the build before the wake existed. Every case was measured against redis-server 8.6.1 first (served within about 0.3 s) and now matches it at --shards 1 and --shards 4 on both runtimes.

  • The last-resort WAL v3 replay now restores KV writes instead of none (moon#1026). When appendonly.aof is missing and the WAL carries KV records (--wal-kv-log on), boot falls back to replaying the WAL. That fallback passed each record's raw RESP payload to dispatch as the command name with no arguments. Every record came back "unknown command" and nothing reached the keyspace, yet the boot log said replayed N WAL v3 records. Measured end to end at --shards 2: 0 of 64 keys came back and the log claimed 35 records. The fallback now uses the same payload decoder as the Phase 4 WAL pass (replay::replay_resp_payload), so the two cannot drift. After the fix the same run restores 29 of 64 keys; the other 35 were connection-local writes, which the WAL does not log by design. Both fallback sites (shard/mod.rs and recovery Phase 4b) now log the KV commands they applied, alongside records read, non-KV commands and undecodable records. The legacy-dir fallback also closes the replay generation, as its AOF sibling does. replay_wal_auto had the same defect and uses the same per-record replay now. Each record replays into the db it was written in (moon#1039), so a MOVE or COPY ... DB n takes its source from that db (moon#1046). A record for a db beyond --databases is skipped with a warning, not folded into db 0. The WAL remains a partial source: it is never the recovery authority.

  • REPLICAOF <hostname> <port> no longer aborts the server (moon#1034). On the default monoio runtime, the replica task parsed host:port as a socket address with .expect, and that parse accepts IP literals only. REPLICAOF localhost 6379, or any DNS name (the normal way to name a master in Kubernetes), panicked on the shard thread, and the panic hook aborted the whole process (SIGABRT). This happened after the command had already replied +OK. The host is now resolved as redis does it: inside the reconnect loop, with backoff. An IP literal (including ::1 and [::1]) needs no lookup. A name is resolved with the system resolver on a helper thread, never on the shard thread, and the wait is bounded at 5 s. At most one lookup is in flight per replica task, even when the resolver hangs. Every resolved address is tried in order, each connect bounded at 5 s: localhost gives ::1 first, which is refused when moon binds 127.0.0.1, so the task falls through to 127.0.0.1. While a host does not resolve, the node stays up and reports master_link_status:down, and it picks the master up once DNS recovers. REPLICAOF NO ONE and re-pointing still supersede the task. The tokio build formatted host:port into one string, which cannot express an IPv6 literal. It now shares the same resolver.

  • One GRAPH.* write no longer makes a restart discard every acknowledged KV write (moon#1018). This hits runtime-tokio with --shards 1 and the graph feature. That configuration has no AofManifest, so recovery picks the KV authority itself: the shard's WAL v3 if replaying it produced KV history, otherwise appendonly.aof. Every graph write lands in that WAL as a Command record. Replay hands it to the graph engine, never to the keyspace, yet it was counted as KV history. The AOF was skipped, and the keys lived only there. Measured with --appendfsync always and kill -9: DBSIZE 3 → 0, and the same again on every later boot. Recovery now counts a record only when replay applied it to the keyspace. The replay engine reports where each record went (keyspace, cold plane, graph, or unhandled), so this does not depend on an exclusion list: moon#914 and this issue each found a record class such a list missed. A record this build cannot handle (for example, GRAPH.* without the graph feature) no longer counts either. monoio, and tokio with --shards ≥ 2, have a manifest and never reach this decision. CI's per-PR Test (graph) step now also runs the recovery/replay unit tests and tests/graph_wal_kv_authority_1018.rs under runtime-tokio + graph, a combination no per-PR leg built before. It adds ~19 s.

  • CLIENT TRACKING ... REDIRECT <id> now reaches its target, and CLIENT CACHING works (moon#1048, moon#1049). Every wire reply below was read off redis-server 8.6.1 over a raw socket first.

  • REDIRECT (the RESP2 client-side-caching pattern): a target subscribed to __redis__:invalidate received nothing, for a write, an expiry, a BCAST match or a FLUSHALL. It now gets *3 message __redis__:invalidate *1 <key> (a flush sends a null payload); a RESP3 target gets the invalidate push. REDIRECT to an id that does not exist (or 0, or a negative id) is refused with -ERR The client ID you want redirect to does not exist, and when the target disconnects a RESP3 source is told tracking-redir-broken <id> and CLIENT TRACKINGINFO reports broken_redirect. Invalidations reach a target on any shard.
  • CLIENT CACHING yes|no answered "unknown subcommand", so OPTIN tracked every read and OPTOUT could opt out of none. It is implemented with redis's errors and its one-command lifetime: the flag covers the NEXT command (or the whole next transaction), survives other CLIENT subcommands and an open MULTI, and is consumed by anything else — an unknown command included.
  • Reads inside MULTI/EXEC are now tracked, as redis tracks them; before this only a transaction's writes invalidated. Each read is tracked under the modes in force at its own position in the body: a CLIENT CACHING queued mid-body covers only the commands after it, and a CLIENT TRACKING on queued in the body tracks the reads after it — on the local and the routed EXEC path alike.
  • A REDIRECT target's invalidations are framed for the protocol it speaks now (a HELLO or RESET since it first subscribed is honoured), and a target that has unsubscribed from everything gets nothing, as in redis, instead of collecting invalidations that its next SUBSCRIBE replayed.
  • CLIENT TRACKING option errors now follow redis: an unknown option is syntax error, OPTIN with OPTOUT and either with BCAST are refused, switching OPTIN/OPTOUT or BCAST without turning tracking off first is refused, a second REDIRECT is refused, overlapping BCAST prefixes are refused, and re-enabling tracking replaces the redirect and flags instead of being silently ignored.
  • A RESP2 connection with tracking on (and no redirect) no longer has a RESP3 > push frame written into its reply stream; redis sends it nothing.
  • A RESP3 connection that tracks and is also subscribed now receives its invalidations while idle; on the monoio runtime they waited until it unsubscribed.
  • Known gap: a RESP3 redirect target that neither subscribed nor enabled tracking itself cannot be reached (redis pushes to it); giving every connection a delivery channel would cost every idle connection its park (moon#1078). Also still open: a burst of more than 256 invalidations to one connection drops the excess silently (moon#1088), and writes and reads made by scripts are invisible to tracking (moon#1089).

  • blocking_spanning_claim bsc8 no longer fails its precondition on Linux when a waiter lands on its key's owner (moon#1083). The test raced one waiter connection per owner shard and required at least SHARDS - 1 of the four races to happen. Only a waiter on a REMOTE owner can race, because a waiter on its own owner has no claim token and unregisters as soon as it sees the disconnect. That requirement holds on macOS, where every connection lands on one shard. On Linux each connection is placed on its own (the kernel's SO_REUSEPORT hash, or the central listener's round-robin), so two local owners in one sequence failed the run. Measured in a Linux container: 6/10 failures on both origin/main and #1045's own tested head 80885575. #1045's own dispatch showed it too: TRY 1 FAIL, then FLAKY 2/3. It was not the merge. The test now tries up to 8 fresh connections per owner until one lands on another shard, and keeps the same SHARDS - 1 precondition. It passed 20/20 on Linux against the same binary. It still fails when the settle window is not held open ("0 of 4 owners … 8 connections each"). Against the old restore server (ffacf2ef) it fails with 8/8 SERVE UNDONE: the correctness assertion now runs before the precondition, where the old test could report "only 2 of 4" instead. Test only; the server is unchanged.

  • test/ci: cargo test --release --lib is green on main again — five CONFIG SET tests were permanently leaking a published maxmemory into every test that ran after them, and the CI waiver hiding it is retired (moon#856). Not a race, as it had been read for months, but a PERMANENT leak through PRODUCTION code: config_set publishes to five process-global atomics on the CONFIG SET path (MAXMEMORY_GLOBAL, MAXMEMORY_HINT, MAXMEMORY_PER_SHARD_HINT, MAXMEMORY_POLICY_GLOBAL, DB_MAXMEMORY_ANY_SET) and the tests driving it never put them back. Under cargo test --lib — one process for the whole suite, unlike nextest's process-per-test — whichever test later READ that state failed, which is why the victim (scripting::bridge::tests::gate_is_skipped_with_spill_sender_when_no_limit_is_configured) looked order-dependent while the cause was not. Measured on a8eb2efc: 5558 passed / 1 failed before, 5559 / 0 after; each of the five writers was proven individually sufficient by running it alone with the victim under --test-threads=1, and -- --skip command::config passed 5541 / 0 as the attribution control. A #[cfg(test)] PublishedLimits RAII guard in storage::eviction now snapshots all five and restores them on Drop — including the UNWIND path, which the three tests that restored manually at the end of their bodies did not cover, so a mid-body panic there leaked worse than never restoring at all. The command::config tests call through a config_set_scoped helper, and a module-local shadow of the glob-imported config_set makes reaching the unguarded publisher from that module impossible rather than merely discouraged. scripts/libtest-singleproc-gate.sh no longer waives anything (LIBTEST_KNOWN_FAILURES=0, LIBTEST_KNOWN_FAILURE empty); its self-test gains "the retired waiver no longer waives moon#856's victim" plus three cases that re-arm the waiver mechanism through the environment and prove it still grants and still refuses. The victim's 100-attempt retry loop — what let a deterministic leak read as a flake — is deleted in favour of a single assertion that names the state it observed.

  • MOVE and COPY ... DB n queued inside MULTI now run at EXEC, into the database they name (moon#1062). The transaction executors sent every queued command to the single-db dispatch, which cannot reach a second database. MOVE answered -ERR MOVE requires handler-level dispatch in its slot. COPY a b DB 4 was worse: it answered :1, wrote b into the SOURCE db, and logged the command verbatim, so a replica (and AOF replay, once moon#1046 lands) put b in db 4 while the master had it in db 0. Every executor now runs both commands against both databases with the same helpers as the live paths: the sharded one on monoio and tokio, the owner-routed TxnExecute, and the embedded handler_single one. A SELECT queued earlier in the body chooses the source db. Only a :1 is logged, verbatim, under the source db, which is the record the live paths already write.

In the same family, redis 8.6.1's replies now come back on every path. MOVE key <current db> answers ERR source and destination objects are the same (it answered :0). A negative or too-large db index answers ERR DB index is out of range for both MOVE and COPY (they answered ERR value is not an integer or out of range and ERR invalid DB index). COPY k k gives the same-object error even when k is missing.

Cross-shard COPY ... DB n is refused. At --shards > 1, COPY src dst DB n whose two keys hash to different shards was routed by src alone, and dst was written into src's shard. There, no normally routed read could see it: 24 of 24 constructed split placements acked :1 and read back nil, and the AOF recorded them on the wrong shard. It now answers the two-key-write CROSSSLOT error, as RENAME does. A {hash} tag co-locates the keys, and COPY without a DB clause still works across shards.

  • The cold-index rebuild no longer drops entries silently, and an indexed-but-unreadable cold entry is no longer a "miss" in code (moon#875). ColdIndex::rebuild_from_manifest_per_db skipped a heap file that failed to read (any io::Error), a page that failed its magic/type/CRC check, and a trailing partial page (chunks_exact discards the remainder) with no log line, no counter and no error — and every entry lost that way then read as an ABSENT key, indistinguishable to a client from one that was never written. Two more silent paths were found on the way: a slot inside a CRC-valid page that does not decode (an unknown ValueType, i.e. a downgrade), and a file truncated on a page boundary, which no per-page check can see. Every one is now counted by cause (INFO → reclamation_cold_recovery_{files_missing,files_unreadable,files_short,pages_rejected,partial_page_bytes,entries_rejected}_total), logged with the file_id and page, and rolled into one per-shard summary that says cold index rebuild clean or cold index rebuild DEGRADED. A valid KvOverflow page — which KvLeafPage::from_bytes also rejects — is classified from its header first and is NOT counted as loss. The decision per class: a NotFound file (the orphan sweep's unlink-before-commit crash window, or external removal) is warned about and queued so the sweep retires its manifest entry instead of re-warning on every boot; any other read error is logged at error and the file is skipped, never tombstoned, so a restart after the operator fixes it recovers the keys (a recovery Err today falls back to v2 recovery, which would discard the whole v3 replay — refusing to boot needs a path shard init does not have yet, left as a follow-up); corrupt pages, partial pages and undecodable slots are bytes that are gone, so they are counted and logged. On the read side, ColdReadOutcome gained Unreadable(ColdReadFault): Miss now means only "no index entry", and a read whose index entry points at bytes that cannot be produced is counted (reclamation_cold_read_unreadable_total), logged with its location, and leaves the index entry in place so a later read heals. At the wire such a key now answers -IOERR cold tier: key is indexed but its data could not be read (see server log) instead of nil, on both dispatch paths: the fabricating accessors (get_or_create* and their compact siblings), INCR*/INCRBYFLOAT, APPEND, SETRANGE, GETSET and SET … GET|KEEPTTL refuse BEFORE mutating — previously INCR minted a counter from zero and HSET/LPUSH/… fabricated a fresh value that shadowed the cold copy until the orphan sweep reclaimed it for good — and a dispatch-boundary gate turns every remaining Database::get- shaped reply into the error (one relaxed load per command). EXISTS, DBSIZE and TYPE keep reporting the key present, plain SET still overwrites (it never reads the old value) and DEL discards: the two escape hatches. Proved two ways against the pre-fix binary at --shards 1 and 4: (a) a real spill → BGREWRITEAOF → SIGKILL → on-disk damage → restart lifecycle, where the four damaged keys answered nil with nothing in INFO or the log and now come with the counters and the file ids; (b) a file removed and another chmod 000 under the RUNNING server, where GET answered $-1 for an indexed key and now answers -IOERR, APPEND and INCR are refused, and after chmod 644 the ORIGINAL value is served.

  • An acknowledged MOVE or COPY src dst DB n survives kill -9 and restart (moon#1046). Both commands need two databases, so the live paths run them through a two-db intercept, but AOF/WAL replay handed the logged record to the single-db dispatch. MOVE hit its "requires handler-level dispatch" error and replay dropped the error, so the key came back in its source db. COPY ... DB n copied within the source db instead: the destination copy was lost and a stray key appeared in the source db. The replay engine now intercepts both, the way it already did for SWAPDB and FLUSHALL. It applies them with the same parsers and core helpers as the live intercept, using the record's SELECT context as the source db. MOVE onto an existing key, COPY without REPLACE onto an existing key, and a missing source all stay no-ops. The TTL moves with the entry, and replaying a record twice does nothing the second time. A record naming a db the server no longer has is skipped with a warning. On both runtimes, at --shards 1 and --shards 4, with --appendfsync always, and with or without a BGREWRITEAOF in the middle, the restarted keyspace now matches the acknowledged one exactly (per-db DBSIZE, values and TTLs). Before the fix, 117 of those checks failed without a rewrite and 61 with one.

  • ZRANGE, ZREVRANGE and ZRANGESTORE no longer clamp a still-negative STOP to 0 (moon#1001). Redis's rank-window rule only clamps a still-negative START; a STOP still negative after len + stop is left negative, so start > stop reports the window empty. Both range helpers (zrange_by_rank against the B+tree, zrange_from_entries against a listpack) instead did (len + stop).max(0), turning -1 into rank 0. Verified against redis-server 8.6.1 on a five-member zset: ZRANGE z -10 -6 answered [a] where redis answers [] (also wrong on ZREVRANGE, ZRANGE ... REV and ZRANGESTORE, which shares the same helper). Both helpers now call the rank_window helper moon#959 added for ZREMRANGEBYRANK, so the four commands cannot drift apart again.

  • maxmemory with an evicting policy no longer refuses every write once the cold tier holds most keys — the moon#1036 change is reverted (PR #1055 reverted). #1055 made every eviction gate compare maxmemory against hot bytes PLUS the cold index's RAM. Under allkeys-* with disk offload, eviction SPILLS victims instead of dropping them, and each spill turns hot bytes into a cold-index entry, so once the cold index dominates, spilling cannot lower the figure: eviction stalls and every write answers -OOM, although an evicting policy must never refuse a write (moon#273). Measured on the v0.8.10 release candidate (--shards 4, AOF everysec, maxmemory 512mb allkeys-lru, 1M SET preload then 150 s of HSET -r 3M): used_memory 234 MB, 2,472,426 cold keys, 724 of 730 HSET runs ended in -OOM and a plain SET was refused afterwards; v0.8.9 on the same load accepted the SET. The revert restores v0.8.9's per-write gate (hot bytes only); moon#1036's between-tick overshoot is reopened and tracked with the spill-cannot-progress fix.

  • A tokio --shards 1 AOF written before MOON.COLDCUT existed no longer compounds its damage on every restart (moon#914 (b)). Such a file has no cut, so replay reads every cold file ungated and re-applies a write on top of the spilled copy of its own result. Nothing rewrote the file, so each boot replayed the same log over the last boot's re-spilled values. Measured on a synthesized legacy file: all 96 non-idempotent probes grew again on every boot (a five-element list read 5 or 10, then 10 or 15, then 15 or 20), and the 24 SET controls stayed put. When boot replays a cut-less appendonly.aof while cold files exist, moon now logs a WARN and runs ONE background AOF rewrite through the #433 auto-rewrite monitor. The rewritten file opens with its cut, so boots 2..N serve exactly what boot 1 served. The first boot's damage can't be undone: the log doesn't record which writes preceded which spill. The rewrite is dispatched after the shards start and doesn't block accept. A crash before its atomic rename leaves the old file authoritative, and the next boot retries. A failed rewrite retries after 60 s. tests/legacy_aof_rewrite_on_boot_914.rs is RED 3/3 without the fix (96 probes changed boot to boot) and GREEN 3/3 with it. The tokio TopLevel rewrite now also sets INFO aof_last_bgrewrite_status, which it never did.

BEHAVIOUR CHANGE: the first boot after upgrading a tokio --shards 1 deployment that uses disk offload may run one AOF rewrite, even with auto-aof-rewrite-percentage 0. It costs one snapshot of the hot dataset written and fsynced as a new RDB-preamble AOF. Measured on macOS for local iteration (not a Linux number), with 100k keys and 46 cold files, on a loaded host: boot-to-accept was unchanged within noise (48-72 ms with the trigger, 58-97 ms without). The rewrite took 17-30 ms, wrote 8.0 MB (the 10.2 MB legacy AOF shrank to 8.0 MB), and finished about 1.1 s after spawn, one monitor tick. With --appendonly no, or an --appendfilename other than appendonly.aof, moon can't rewrite the file that was replayed. It logs a WARN naming the manual remedy instead.

  • WAL v3 KV records now replay into the database they were written in (moon#1039, P0). Under --wal-kv-log on, a write that executes on a shard thread (a pipelined cross-shard write, an active-expiry reason-DEL, a MOON.SPILLED cold marker) was logged as a bare command with no SELECT, and recovery replayed every shard's WAL starting from db 0. When the WAL was the KV authority, a restart put every db 1-15 write into db 0, and a DEL logged for db 3 deleted db 0's key of the same name. On the tokio runtime at --shards 1 the WAL is the authority by default, so a plain SIGKILL and restart with the AOF untouched was enough: an expiry in db 3 erased db 0's k. On monoio, and at --shards 4, it took a missing multi-part AOF manifest (k3_1 answered from db 0, db 3 empty). The db now travels in the record header: a new flag bit plus the two header bytes that were always-zero padding. Old WALs still replay (a record with no db context goes to db 0, as before, and never inherits the previous record's db), and an older binary still reads a new WAL without error. A record for a db beyond the configured --databases count is dropped with a warning instead of being folded into db 0.

  • SPUBLISH queued inside MULTI is now delivered at EXEC (moon#1043). The command was answered +QUEUED, then EXEC answered -ERR unknown command for that slot while the rest of the transaction committed, so the shard-channel message was silently dropped. The transaction executors intercepted PUBLISH for their deferred post-body fan-out but not SPUBLISH, which fell through to the keyspace dispatch table. Now both are recorded with their namespace, an enum and not a bool, since the two namespaces can share a channel name without sharing subscribers. After the body, each fans out through the same registry, remote-subscriber map, and SPSC message as its immediate form. This holds on every handler, including an owner-routed EXEC at --shards > 1. The channel ACL is checked at fan-out exactly as for PUBLISH, so a denied channel answers NOPERM in its slot and is never delivered. Measured against redis 8.6.1 at --shards 1 and --shards 4: the EXEC reply (*2 +OK :3) and delivery to every subscriber now match.

  • runtime-tokio with --shards 1 now opens every AOF generation with a MOON.COLDCUT, so a kill -9 no longer double-applies writes to spilled keys or drops acknowledged post-rewrite writes (moon#914). This is the one configuration with no AofManifest (creating one there wipes state on the next boot, #96), so neither seed_cold_cut nor the rewrite's head ever ran. Its replay read every cold file ungated, and moon#902 and moon#912 were both still live with moon#965's fix applied: 79–82 of 216 probes double-applied on the first restart, and 13–14 of 24 acknowledged post-BGREWRITEAOF SETs came back holding the pre-rewrite value. The legacy appendonly.aof now carries the head itself. Boot writes it when the file holds no record yet, and BGREWRITEAOF writes it right after the RDB preamble, before the file is renamed into place. The record and its meaning are the same as in the monoio incr head, and no manifest is created. Two fixes ride along. With --wal-kv-log on, a single MOON.SPILLED marker mirrored into the WAL was counted as KV history: recovery took the WAL as the authority, skipped the AOF, and lost its entire history (DBSIZE 248 → 103 after one restart). Cold-plane records no longer count. The end-of-replay reconcile closed only db 0, so a gate left on SELECT 1..N would have hidden later spills. Every AOF replay path now closes the generation on every database, and main.rs warns and closes any gate a missed path leaves open. Not covered: an AOF written before this change has no head, and replays ungated until its first rewrite. Run BGREWRITEAOF once after upgrading a tokio --shards 1 deployment that uses disk offload.

  • A multi-key BLPOP/BRPOP/BZPOPMIN/BZPOPMAX whose keys span shards pops exactly one element, from the first non-empty key in argument order (moon#1019). At --shards > 1 such a command could pop the WRONG key (the client's shard served a later key it owned over an earlier non-empty key another shard owned) or pop on two or three shards at once and deliver one reply, destroying the other elements. These commands keep working across shards — an untagged BLPOP q1 q2 q3 0 worker loop is not refused — and now answer as standalone Redis does, including -WRONGTYPE for the first existing key of the wrong type. A waiter registered on several shards carries one claim token: a shard pops, must then win the token before it answers, and puts the element back in the same step if it loses. The keys are registered in argument order, one run of same-shard keys at a time, and each run is acknowledged before the next, so a later key never answers ahead of an earlier non-empty one. This costs one extra round trip per remote run that is not the last. Co-located keys and single-key waits pay nothing extra, and a waiter whose keys all live on its own shard carries no token. Measured at --shards 4 (tests/blocking_spanning_claim.rs, macOS, both runtimes, before → after): immediate pops differing from Redis 91/192 (monoio) and 95/192 (tokio) → 0/192; the -WRONGTYPE ladder skipped 14/16 and 10/16 → 0/16; concurrent pushes to three owners destroyed elements in 39/80 and 64/80 waits → 0/80.
  • A blocking pop no longer loses an element that is served as its timeout, disconnect or shutdown fires (moon#1023). The wait path sent BlockCancel and dropped its receivers. A serve that landed between the two went into a still-live receiver, so the owner's undo, which runs only when its send fails, never ran, and the element was dropped with the receiver. The waiter now closes its claim token first. If no shard has claimed it, none can any more, and there is nothing to drain. If one has, its reply is taken and delivered on a timeout or shutdown: the client was served before the end of the wait was observed. If the client is gone, the serve stands, as it does in Redis: its effect is recorded through the same path as a delivered reply (tracking invalidation included), and the reply is dropped with the socket. That path has two known gaps, which this change inherits and does not close. The record goes to the connection's shard AOF, so for a key owned by another shard, replay drops it (moon#1056). On runtime-tokio, the record is appended to the AOF only, never to the replication stream. The element is not put back. A put-back would land after whatever other clients wrote to the key in the meantime, on top of a pop that was never logged. After RPUSH k b; LPOP k the master would then hold [a] while its AOF replays [b], and a DEL k would be undone. Measured with a test-only window (MOON_TEST_BLOCK_SETTLE_DELAY_MS) at --shards 4 on both runtimes:
  • a push racing a timeout destroyed its element 13/16 → 0/16;
  • a served-then-disconnected waiter's element came back over later writes 6/6 → 0/6 (monoio).
  • A spanning blocking pop's timeout is the client's timeout. Each cross-shard registration step waits for its owner's acknowledgement. That wait was bounded only by the 30 s internal reply timeout, not by the client's deadline. So BLPOP q1 q2 q3 1 with a stalled owner answered after the stall, or with a MOONERR after 30 s, and a shutdown during the wait also answered MOONERR. The acknowledgement now races the client's deadline and shutdown, and each ends the wait with its normal reply (nil, or the shutdown error). With a 3 s test-only owner stall (MOON_TEST_BLOCK_ACK_STALL_MS), BLPOP k1 k2 k3 0.5 answered after 3.0 s and 6.0 s (both runtimes) → nil in under 1.5 s.
  • A blocking pop's element that must go back (the waiter was won by another shard, or its reply could not be sent) keeps the key's TTL. When that pop had emptied the key, the put-back recreated it with no TTL, so the master kept a key that every replica expired. The encoding is not preserved: the key is recreated in the natural encoding for its size. A wake on a key that holds nothing now answers nobody, instead of answering a parked BLPOP k 0 with nil.
  • A blocking pop parked on a key that now holds another type stays parked. Example: BZPOPMIN k 0 parks, then RPUSH k x y creates k as a list, then a remote BLPOP k 0 registers. That registration runs every waker on k. The zset waker took the BZPOPMIN waiter, found nothing to pop, and answered it nil, which Redis never sends a timeout-0 waiter. A waker with nothing to pop now puts the waiter back at the front of its queue, unanswered.
  • A timed-out spanning wait no longer leaves a registration behind. A later local run of a spanning wait (BLPOP a b c 1, with a and c local and b remote) was registered without the client's deadline. The timeout sweep removed the waiter's deadlined entries and forgot its id, which orphaned that entry until its key was next pushed. The sweep now removes every registration of a timed-out waiter.
  • One push that carries several elements serves every parked waiter those elements cover, as Redis does. Two clients in BLPOP k 2 and one RPUSH k a b used to answer one waiter and leave the other parked next to b until its timeout; now both are answered at once.

  • CLIENT TRACKING now invalidates a key that expires or is evicted (moon#1013). Only command writes used to push invalidate, so a client-side cache served an expired or evicted value forever, with no signal. redis 8.6.1 pushes in both cases, and moon now does too. The fix covers the active expiry sweep, the lazy-expiry drain, hash-field expiry (the hash is invalidated even when it survives, as in redis), cold-tier expiry (on read and in the periodic sweep), and plain-drop eviction. A victim spilled to the cold tier is not invalidated, because it stays readable with the same value. All of these go through one hook on the existing process-global tracking table, so BCAST prefixes and NOLOOP behave as they do for writes. An expiry is nobody's own write, so NOLOOP does not suppress it, matching redis. The owner shard pushes straight into a reader on any other shard. With no tracking client the hook costs one relaxed atomic load per removed key. Measured red/green against redis at --shards 1 and --shards 4: before the fix every case got zero pushes, and after it every case got a push.

Also fixed here: a hash-field TTL on an idle database was never reaped. The field reaper runs against the database's cached clock, and only commands advance that clock, so a due field sat in memory (and its invalidation stayed unsent) until unrelated traffic touched the database. The expiry tick now moves that clock forward, never backwards, and only when a hash-field TTL exists.

Found alongside, filed separately, and not changed here: REDIRECT (the RESP2 mode) delivers no invalidation at all (moon#1048), and CLIENT CACHING is unimplemented, so OPTIN/OPTOUT track every read (moon#1049). Both affect writes and expiry alike.

  • Scripts queued inside MULTI now run at EXEC (moon#894). EVAL, EVALSHA, EVAL_RO, EVALSHA_RO, FCALL and FCALL_RO were answered +QUEUED and then -ERR unknown command at EXEC, while the rest of the transaction committed. That one step was silently dropped from an otherwise successful transaction, and the reply array was still full length. The transaction executor now runs a queued script in place, in body order, on whichever shard the body runs. The moon#247 locality rules apply unchanged: a script whose declared keys live on another shard routes the whole body to that owner, and a body spanning shards is refused with CROSSSLOT before anything runs. Every inner redis.call is authorized as the calling user, on the owner too.

The script's effect records are captured and spliced into the body's own record list, so they reach the AOF, the WAL and the replication stream in body order and exactly once. Letting the script emit them directly, as it does outside MULTI, would log SET k 1; EVAL "APPEND k x" as APPEND k x; SET k 1, and a restart would read 1 where the master answered 1x. A mutation run that disabled the capture reproduced exactly that divergence after SIGKILL.

Semantics match redis-server 8.6.1: - an unknown sha is NOSCRIPT in its slot; - EVAL_RO refuses a write; - a script that errors is one error element, its earlier writes stay, and the rest of the body runs; - a WATCH conflict aborts before any script runs.

  • A write replayed after its key's spill marker is no longer discarded on restart (moon#965). On runtime-tokio with --shards 1 no AofManifest exists, so the AOF never carries a MOON.COLDCUT head — but MOON.SPILLED markers are emitted unconditionally, so that configuration ran moon#902 half-armed. A replayed marker drops the key's hot copy; its next write then replays through Database::set's Inserted arm, which by design leaves the cold shadow standing; and with no cut the end-of-replay reconcile took the legacy cold-wins branch and threw the newer write away — logging it as 1 hot shadow(s) demoted to cold stubs. The issue's "AOF tail loss" title named the wrong mechanism: both writes were on disk and both replayed. A replayed marker now proves the generation is #902-era, and reconcile resolves hot-wins for it whether or not the head carries the cut. That is value-correct in every way a key can end hot and cold there: written after its marker, a stale entry in a still-listed file, or a marker lost under backpressure (both planes then hold the same value). Logs with neither record keep the task #56 path unchanged. tests/cold_shadow_single_shard_tokio.rs, written for exactly this configuration, was RED on main for weeks and never ran: it was #[ignore]d because it shelled out to redis-cli. It now speaks RESP through common::Conn and runs in every runtime-tokio leg; with the fix inert it fails 3/3 on 36-47 stale probes.

  • Cluster mode no longer refuses MSET/MSETNX with CROSSSLOT when a VALUE hashes to another slot (moon#1012). The cluster pre-check slot-hashed every argument after the routing key as if it were a key, so MSET {t}a x {t}b y — two keys in one slot — was refused because x hashed elsewhere; only a user who hash-tagged their values got through. It now reads key positions from the shared key walker (acl::keyspec::command_key_positions, the one ACL, the moon#592 cross-shard write guard and cache invalidation already use): MSET's first/last/step of 1, -1, 2. Measured against redis-server 8.6.1 on one node holding all 16384 slots, 15 rows (MSET/MSETNX/MGET/DEL/BITOP/COPY, same-slot and spanning) are now byte-identical on both runtimes at --shards 1 and --shards 4; before, 4 accepted writes were refused. Keys genuinely in two slots are still CROSSSLOT. COPY's REPLACE literal and BITOP's operation token now fall out of their key specs instead of a special case.

  • The moon#507 pipeline wait set is derived from COMMAND_META instead of a hand-written list, and WATCH inside its own pipeline no longer aborts the transaction (moon#937, moon#946). must_wait_for_pending_remote decides whether a pipelined command may run while its batch's earlier cross-shard writes are still undispatched; its list of inline-intercepted commands was a (len, first byte) table that failed OPEN and had drifted — TXN, TEMPORAL.INVALIDATE (moon#937), the blocking family (moon#946), and, found on the way, WATCH, SPUBLISH/SSUBSCRIBE/SUNSUBSCRIBE, MQ and WS, each intercepted through a predicate the drift scanner could not see. The predicate now reads the registry's NO_INTERCEPT bit — the same bit the connection handler uses to skip its gate chain — so a command waits unless the registry proves no interceptor can claim it, and both consumers fail safe. To keep that free, the 98 routable commands no gate claims (HGETALL, LRANGE, ZRANGE, XADD, SMEMBERS, …) are now marked NO_INTERCEPT (67 → 165), which also lets the handler skip its gate chain for them; a coupling test fails in either direction — a claimed command that is marked, or a routable unclaimed one that is not. Behaviour change, measured at --shards 4 with remote writes pending in the same batch: TXN BEGIN, TEMPORAL.INVALIDATE, SPUBLISH and WATCH now wait for those writes to land (they ran inline before); GET/HMSET/LRANGE/ZRANGE/XADD still do not wait. The user-visible one is WATCH: SET k; WATCH k; MULTI; GET k; EXEC in one write aborted the transaction 18 of 24 times because WATCH captured the key's version before its own batch's SET landed (redis 8.6.1: 0 of 24). TXN and TEMPORAL.* also join extract_primary_key's keyless table, so the cluster slot router no longer hashes "BEGIN" or an entity id into one fixed slot. moon#946 is closed as a documented non-bug: the blocking family is deferred by the #438 early-flush guard one statement above this predicate in both handlers whenever remote work is pending, and the derived predicate now says wait for them regardless.

  • Multi-key blocking pops no longer destroy an element from a key they did not answer with (moon#989). BLMPOP, BZMPOP, BLPOP, BRPOP, BZPOPMIN and BZPOPMAX over keys co-located under one {hash} tag replied correctly while a SECOND non-empty key silently lost its head element: BLMPOP 0.3 3 {t}a {t}b {t}c LEFT answered {t}b B1 and left {t}c at [C2], with C1 delivered to nobody. It happened whenever the client's connection lived on a different shard than the keys: the client's shard could not see them, so it sent one registration per key to their owner, and the owner served the same waiter once per non-empty key — the client kept the first reply and dropped the rest. The same fan-out popped a key named twice (BLMPOP 0 2 k k LEFT) twice, and skipped Redis's -WRONGTYPE for an earlier key of the wrong type. Measured at --shards 4 (BLMPOP, BZMPOP, BLPOP, BZPOPMIN × 16 tags × 3 server instances, against redis 8.6.1): 139 of 192 probes destroyed an element before, 0 after; --shards 1 was and is 0. The client now sends ONE registration per owner shard carrying every key it owns, and the owner registers, type-checks and serves them in one synchronous stretch, so a waiter is served at most once there.

  • ACL SETUSER implements allkeys, allcommands, allchannels and every spelling of the %R~ / %W~ / %RW~ key selectors (moon#970). All were dropped with +OK, so a user provisioned with the standard redis idioms came out inert: ACL SETUSER w1 on >pw allkeys allcommands reported -@all and every command was NOPERM; %RW~cache:* +@read left the user with no key access at all; %r~ / %wr~ (lowercase flags) vanished. ACL LIST, ACL GETUSER and ACL SAVE now render these exactly as redis 8.6.1 does (%RW~p as ~p, one-sided grants as %R~p / %W~p), through one key-pattern renderer -- ACL GETUSER carried a second copy that would have rendered a no-access pattern as a write grant. Every touched permission set, including moon#981's +@all -x, is verified to round-trip SETUSER -> LIST/GETUSER -> SAVE -> restart with LOAD unchanged, token order included.

  • A restart no longer re-issues a warm vector segment's id to a KV spill file, and retiring a segment entry no longer tombstones a spill file (moon#893, moon#997). One per-shard counter names both KV spill files and warm vector segments, but its restart seed scanned data/heap-*.mpf only. When the highest id in use belonged to a vector segment, the next spill file took the same id; when that segment's directory later vanished, recovery retired its manifest entry with remove_file(id), which tombstoned every entry with that id — the live spill file's included — so its keys read as absent after the restart. Measured before the fix (durable cold keys absent after one restart): monoio 256/894 at --shards 1 and 540/901 at --shards 4 after a BGREWRITEAOF, 2/4 and 4/226 under --appendonly no; tokio 514/902 at --shards 4, 2/4 and 209/219 under --appendonly no. The seed is now one authority (storage::tiered::file_id_seed): the maximum over every manifest entry of every type and status, every heap-*.{mpf,tmp} and every segment-* / .segment-*.staging directory, computed once after recovery and shared by the spill counter and the MOON.COLDCUT watermark. ShardManifest::remove_file now matches (file_id, file_type), so a data dir a pre-fix build already wrote the collision into keeps its spill file. Both spill writers refuse to replace an existing heap-*.mpf; a re-issued id becomes a failed spill that keeps the values hot instead of overwriting live cold data. The seed scan also no longer skips a directory entry it cannot read.

  • A shard's cold file ids come from one counter that never moves backwards (moon#893, moon#997). The event loop kept a second copy of the counter, re-synced once per tick. On tokio the cross-shard SPSC drain ran after that sync and advanced the shared counter; the eviction tick and warm vector transitions then allocated from the stale copy, re-issuing the drain's ids, and wrote it back with a plain set that moved the shared counter backwards. The copy is gone: every consumer allocates through file_id_seed::allocate_from, and every write-back is monotonic.

  • ShardManifest::create is atomic (temp file, fsync, rename, directory fsync). A crash part-way through it used to leave a manifest shorter than its two root pages at the real path, which every later open rejected. Tombstones are also aged per (file_id, file_type), not per id.

  • ZUNIONSTORE/ZINTERSTORE report WRONGTYPE before an option error, and no longer flatten a listpack source (moon#959). Redis looks every source up before it parses WEIGHTS/AGGREGATE, so ZUNIONSTORE d 1 <string-key> BOGUS is WRONGTYPE on redis 8.6.1; moon answered syntax error. The store family also read its sources through the promoting accessor, converting a listpack source to skiplist as a side effect of reading it — the moon#928 defect the read-only set operations were already cured of. Both fixes came with the shared implementation ZDIFFSTORE now uses.

  • Commands routed to another shard are counted and timed (moon#982). At --shards > 1 a command whose key lives on a shard other than the connection's went through no telemetry probe at all — neither the cross-shard read fast path (executed on the origin thread) nor any of the six SPSC execute arms on the owner (Execute, ExecuteSlotted, PipelineBatch, PipelineBatchSlotted, MultiExecute, MultiExecuteSlotted) observed what they ran. INFO total_commands_processed and moon_command_duration_microseconds therefore counted only connection-shard-local commands: measured on one connection, 400 SMEMBERS over 16 untagged keys counted 400 / 150 / 100 / 50 at 1 / 2 / 4 / 8 shards, and 0 at 8 shards with the fast path off when the connection's shard owned none of the keys. Every ops/sec dashboard derived from that counter was under by 1/shards, silently, in the reassuring direction. Each SPSC drain cycle now builds one LatencyProbe from a per-shard sampler and metric-handle cache and every execute arm observes its cmd_dispatch through it — the same observe the connection handlers use, so there is still one spelling of "time this command" on both sides of the SPSC boundary; the fast-path read and the local part of a spanning read (moon#768 fan-out) are observed on the origin thread by the connection's own probe. Duration for a routed command is its execution on the owning shard — the same quantity as for a local command and what Redis reports — not the queue and reply wait, which the cross_spsc / remote_awaits_parked counters already expose. A spanning multi-key read now counts once per shard it touches; commands that take the multi-key coordinator (MSET, spanning DEL/EXISTS, KEYS/SCAN/ DBSIZE aggregation) count their remote legs but not yet their local leg. A slowlog entry for a routed command carries an empty client address and name (the message does not carry the client). Cost: one probe construction and drop per drain cycle, and per routed command the same counter increment and 1-in-16 Instant the local paths already pay; no allocation.
  • Sorted-set argument validation reports the error CLASS Redis reports (moon#969). Nine forms answered the wrong class, which matters beyond wording: redis-py raises a distinct exception type per class, so a client branching on the exception took the wrong branch and retried a request that could never succeed. ZPOPMIN k notanint/k -1 now say value is out of range, must be positive; ZINTERCARD 0 k and ZUNIONSTORE d 0 k now say at least 1 input key is needed for '<cmd>' command; ZMPOP 0 k MIN says numkeys should be greater than 0; and a short WEIGHTS list, a dangling AGGREGATE/LIMIT/ COUNT, a numkeys overrunning the key list, and ZADD k 1 a 2 are all syntax error rather than arity errors. The set-operation family SPLITS into two classes exactly as Redis does — not-a-number is the generic integer error, a number below 1 names the command — while ZMPOP does not split, and arity is checked first so ZUNION 0 stays an arity error. Also fixed while reproducing: ZINTERCARD k LIMIT -1 (LIMIT can't be negative), ZMPOP ... COUNT 0 (count should be greater than 0), and one moon#967 leftover where ZUNIONSTORE d 1 k BOGUS stepped over the unknown token and answered a different, successful command. Four ZRANGE-family sites the issue also cites were verified against a redis 8.6.1 oracle to be ALREADY correct and are deliberately unchanged, with harness rows pinning them.
  • ZADD ... CH counts a rescore exactly instead of against an epsilon window (moon#792). Both mutation loops decided changed with an ABSOLUTE f64::EPSILON, where Redis's zsetAdd compares exactly. f64::EPSILON is the gap between 1.0 and the next double — a RELATIVE quantity — so as a fixed tolerance it swallowed real moves at every magnitude below 1: rescoring 0.0000000001 to 0.00000000010000001, six significant figures, replied 0 while ZSCORE showed the new value. The window also disagreed with zset_update_existing, which moves the member on to_bits() inequality, so the write happened and only the tally pretended otherwise. Any client using CH as a did-anything-change signal silently skipped those updates. Fixed on both the listpack and B+tree arms, which carried separate copies.

Security

  • fuzz/Cargo.lock is audited and clean (moon#1092). The fuzz workspace has its own lockfile, which no gate checked, and it carried RUSTSEC-2026-0204 (crossbeam-epoch 0.9.18) and RUSTSEC-2026-0258 (h2 0.4.13), plus the unsound anyhow 1.0.102 and memmap2 0.9.10 and the yanked spin 0.9.8. Each is bumped to the version the root Cargo.lock already ships (0.9.20, 0.4.18, 1.0.104, 0.9.11, 0.9.9), and nothing else in the lockfile moves. supply-chain.yml now triggers on fuzz/Cargo.toml and fuzz/Cargo.lock and runs cargo audit --file fuzz/Cargo.lock next to the root audit, so the fuzz lockfile cannot drift behind again unnoticed.

  • Error and status replies can no longer be split by client input (moon#1031). Error text quotes client input (an unknown command name, an ACL SETUSER rule), and CR/LF bytes in it were written raw, so one command could produce several RESP replies. That desynchronises any client or proxy that pipelines on a shared connection. Every line-framed reply (+, -, RESP3 () now goes through one writer that maps CR and LF to spaces, as redis does, and so do the inline quota error and the protocol-error echo. It is allocation-free, and a reply with no CR/LF costs the same as before.

  • rustls bumped past RUSTSEC-2026-0285 / GHSA-2mjx-qc3c-rqvc ("TLS 1.3 handshake messages incorrectly accepted across encryption level boundaries"), 0.23.44 → 0.23.45 in Cargo.lock (and 0.23.37 → 0.23.45 in fuzz/Cargo.lock, which was independently behind). moon's TLS path uses rustls directly and through the vendored vendor/monoio-rustls, whose Cargo.toml already declared rustls = "~0.23.4" — wide enough to admit the patched release without a manifest change. cargo audit and cargo deny check advisories licenses bans sources are clean on Cargo.lock, and both TLS integration suites (tests/tls_idle_downshift_parity.rs, tests/tls_park_keyupdate.rs — 4 tests) pass against a fresh binary on both the monoio and runtime-tokio,jemalloc builds. cargo audit's own .cargo/audit.toml ignore for RUSTSEC-2026-0097 (rand < 0.9.3, aka GHSA-cq8v-f236-94qc — the same advisory as the rand Dependabot alert below) is now moot since rand is 0.9.3+ in Cargo.lock, so it was removed rather than left as stale documentation. Also picked up while re-locking: the yanked chacha20 0.10.0 (pulled in transitively via rand) to 0.10.2, in both Cargo.lock and fuzz/Cargo.lock.

  • Dependabot alerts on console/pnpm-lock.yaml, Cargo.lock and fuzz/Cargo.lock cleared with minimal, scoped bumps. js-yaml (4.3.0 → 4.3.2, GHSA-2883-xcg3-v3hh / GHSA-5p4m-2wfm-xmqj), baseline-browser-mapping (2.10.43 → 2.11.25, GHSA-w5vr-8v7q-w6rv), fflate (0.8.2 → 0.8.3 and 0.6.10 → 0.6.11, GHSA-px8p-9vwx-vf98), @humanfs/node (0.16.7 → 0.16.8, GHSA-p498-v437-472g), browserslist (4.28.6 → 4.29.0, GHSA-73wf-gq98-2v4g), react-router (7.18.1 → 7.18.4, GHSA-qwww-vcr4-c8h2), dompurify (3.4.12 → 3.4.15, GHSA-55q2-fjhq-7xh7) and postcss (8.5.15 → 8.5.28, GHSA-fxqj-rqcc-2cmp / GHSA-r28c-9q8g-f849) are all transitive dev/runtime deps several levels deep behind eslint, vite, @react-three/drei, @vitejs/plugin-react, react-router-dom and @cosmos.gl/graph; each is pinned via a version-bounded pnpm.overrides entry rather than an unconstrained bump, so no direct dependency's declared range and no major version moved (an unbounded override on react-router first resolved to the incompatible 8.x line and was re-bounded to <8). rand (0.9.2 → 0.9.3 in Cargo.lock, pulled in via metrics-util; 0.10.0 → 0.10.1 in fuzz/Cargo.lock, GHSA-cq8v-f236-94qc) via cargo update -p rand --precise. nltk (sdk/python/uv.lock, GHSA-8mgp-746c-j5xp, no patched version exists upstream) is left as-is: it reaches the SDK only transitively through the optional moondb[llamaindex] extra's llama-index-core dependency, which imports only nltk.tokenize.PunktSentenceTokenizer, nltk.corpus.stopwords, nltk.data.find and nltk.download — never the vulnerable model-artifact APIs (TransitionParser.train/parse, AveragedPerceptron.save/load, PerceptronTagger.save_to_json, save_maxent_params), and the advisory's own PoC additionally requires the consumer to opt into nltk.pathsec.ENFORCE=True, which neither moondb nor llama-index-core sets.

  • Subcommand ACL rules are now enforced (moon#1030). +@all -config|set was accepted, listed and saved, but the permission check only ever looked up the bare command name, so the user could still run CONFIG SET. The same gap meant +config|get or +select|0 grants never took effect. The check now consults cmd|<first arg> first, with redis's last-rule-wins ordering: a bare rule or category clears that command's subcommand rules. NOPERM text and ACL LOG name the subcommand (config|set). ACL LIST, GETUSER and SAVE emit bare rules before subcommand rules, so a saved file reloads to the same permissions; a line with no subcommand rules is byte-identical.

  • Setting a password on a nopass user now actually requires it, and nopass now revokes the old passwords (moon#999). Two credential fail-opens, both answering +OK. (a) >pw / #hash did not clear the nopass flag: ACL SETUSER fa on nopass ~* +@all then ACL SETUSER fa >realpw left an account that accepted ANY password -- measured against redis-server 8.6.1, AUTH fa totallywrong answered OK on moon and WRONGPASS on redis, over AUTH, HELLO 3 AUTH and inline AUTH alike, while ACL GETUSER showed a password set. (b) nopass did not drop the stored hashes, so rotating a credential through nopass (>oldpw, nopass, >newpw) left the compromised oldpw valid indefinitely. > and # now clear nopass, and nopass clears the password list, as redis does; both hold across ACL SAVE / ACL LOAD and a restart.

  • The ACL keywords an operator uses to contain a compromised account now take effect (moon#979). ACL SETUSER svc nocommands (the arm did not exist) and ACL SETUSER svc OFF / RESET / RESETKEYS / RESETCHANNELS / RESETPASS / NOPASS (every non-lowercase spelling -- redis compares keywords case-insensitively) answered OK and changed nothing: the account stayed live with +@all and the operator was told the lockdown worked. Keywords are now matched case-insensitively in exactly one table, and every redis keyword has an arm.
  • An ACL category Moon does not implement is an error, not a grant of every command (moon#978). get_category_commands ended in _ => &[], so an unknown category resolved to an EMPTY command list rather than failing. deny_command walked that empty list, inserted nothing, and then unconditionally rebuilt the permission set as Specific { base_allow: true, allowed: {}, denied: {} } — base-allow with an empty deny set, i.e. every command permitted. is_command_allowed fell through to base_allow, while user_to_acl_line (which discards base_allow) printed the user as -@all, so introspection actively concealed the state: measured against redis-server 8.6.1, ACL SETUSER v on >pw ~* &* +@all then ACL SETUSER v -@bitmap answered OK on both, after which moon's ACL LIST reported user v on #… ~* &* -@all while that user ran SETBIT, BITCOUNT and FLUSHALL — redis reported +@all -@bitmap and answered NOPERM. Six real redis categories reached that arm (bitmap, hyperloglog, geo, fast, slow, blocking), as did every non-lowercase spelling of a category Moon did have — the old match compared lowercase literals while redis category names are case-insensitive, so -@DANGEROUS was a full grant. Category lookup is now case-insensitive, an unresolvable name returns redis's ERR Error in ACL SETUSER modifier '<rule>': Unknown command or category name in ACL, and the rejected ACL SETUSER mutates nothing — neither creating the user nor applying the prefix of the rule list that parsed. This is a second, distinct path into the state disclosed as GHSA-9x86-7597-5wwj, and worse in one respect: there, stored and reported state agreed.
  • +@read no longer grants commands that mutate, and -@dangerous now revokes what redis revokes (moon#980). @read contained getdel, getex and sort, so a -@all +@read user could run GETDEL vic — measured returning "hello" and leaving EXISTS vic at 0, where redis answers NOPERM — and SORT … STORE wrote a new key. @dangerous was missing SWAPDB, INFO, CLIENT, ROLE, SHUTDOWN and RESTORE, so +@all -@dangerous left all of them runnable. Every category's membership is now derived from a live redis-server 8.6.1 ACL CAT, restricted to the commands Moon implements, with redis's per-subcommand classification collapsed onto the bare container name Moon's permission check actually sees — in the direction that makes -@dangerous deny the whole container. The six missing categories are implemented, and Moon-only families (FT.*, GRAPH.*, TXN, TEMPORAL.*, MQ, WS, CDC.READ, VACUUM) are classified explicitly instead of falling into the hole this fixes. @transaction listed a bare temporal, which Moon does not dispatch at all — the real names are TEMPORAL.SNAPSHOT_AT and TEMPORAL.INVALIDATE, so that carve-out had never covered either, and the unit test asserting it did was probing an un-dispatchable name and could not fail.
  • ACL CAT publishes exactly the categories +@/-@ resolves. The name list existed in three hand-maintained copies — two in src/command/acl.rs (one published by the no-argument form, one gating the single-category form) and the match arms in src/acl/rules.rs. A name could be published without resolving, or resolve to nothing while ACL SETUSER still accepted it. There is now one table with three consumers.
  • +CONFIG|GET is stored lowercased. The subcommand branch of allow_command stored the rule verbatim while is_command_allowed lowercases the incoming name before probing, so a mixed-case subcommand grant could never match anything.
  • The consistency suite has ACL rows. scripts/test-consistency.sh had none — the only ACL mention in either harness was a container-subcommand list — so both bugs above shipped unnoticed. It now carries the three measured escalations, the case-insensitivity case, moon#971's base_allow polarity pair, and an ACL CAT diff of all 21 redis categories against the live oracle that fails on an escalation in a permissive category or a missing member of @admin/@dangerous.

Added

  • Console vitest unit suite (10 files / 56 tests) now runs in CI (Closes moon#964). Nothing in .github/workflows/ ran it — a dependency bump could break console/src with every check green, which is exactly what happened in moon#909's first commit (vitest bumped to ^5 while @vitest/coverage-v8 stayed on ^2.1.8, fixed in the same PR's second commit before merge). Added a unit job to console-integration.yml, which already carried the correct console/** path filter and installed Node/pnpm without ever using them. Runs pnpm test only — coverage stays unthresholded, since the current 6.21% figure is an artefact of the vitest config's include scope pulling in untested Three.js/graph UI, not a signal a threshold could usefully gate.

Fixed

  • ACL SAVE no longer inverts a +@all -<cmd> user into -@all -<cmd> (moon#981). CommandPermissions::Specific carries the base_allow polarity added by moon#971, but the serializer behind ACL SAVE, ACL LIST and ACL GETUSER discarded it and wrote -@all plus the sets for EVERY Specific user. A service account defined as "everything except FLUSHALL" was written to disk as "nothing, and also not FLUSHALL", and came back from ACL LOAD or a restart with --aclfile unable to run a single command — GET k answered NOPERM while SAVE and LOAD had both answered +OK. It failed closed, so it was an outage rather than an escalation, but a silent one. The writer now emits +@all followed by the revocations (then any re-grants) when the base is allow, and -@all followed by the grants when it is deny — the same line redis 8.6.1 writes, verified against a real aclfile on both engines. The reader was already correct; only the writer changed. ACL GETUSER's commands field, a second copy of the same serializer, now shares the one implementation and reports +@all -flushall as redis does.
  • Twelve multi-key commands now return CROSSSLOT at --shards >= 2 instead of answering — and, for LMPOP/ZMPOP, MUTATING — from one shard's slice (moon#962). This is a behaviour change. Routing picks a command's FIRST key and ships the whole command to that key's owner, which then executes it against its own keyspace slice; every other key reads as ABSENT rather than erroring. SINTER, SUNION, SDIFF, SINTERCARD, ZDIFF, ZINTER, ZUNION, ZINTERCARD, LCS, PFCOUNT, LMPOP and ZMPOP therefore answered confidently wrong: measured against redis 8.6.1 at --shards 4, SDIFF/ZDIFF returned EXTRA members the remote operand should have subtracted, SINTER/ZINTER/*CARD empty or 0, SUNION/ZUNION a short set (and ZUNION wrong SCORES), LCS empty, PFCOUNT an undercount — 156 of 156 constructed cross-shard placements. LMPOP and ZMPOP are flags: W and were worse than a wrong answer: they POPPED a key the command is defined never to reach and acked it (LMPOP 3 {t1}a {t2}b {t7}c LEFT answered {t7}c C1 where redis answered {t2}b B1, 24 of 24). All twelve now fail closed before anything is read or written. Co-located and --shards 1 usage is unaffected, and {hash} tags are the remedy — the same trade moon already made for the *STORE family in moon#592. TOUCH is the one member that does NOT error: it is per-key decomposable, so it fans out and sums exactly like EXISTS, and now answers the correct total where it previously undercounted.

  • Writes report their real latency on the shipped runtime, and GET/SET are in the histogram at all (moon#941, moon#963). The monoio write path constructed its 1-in-16 latency timer AFTER the with_shard closure that ran the command had returned, so every write on the runtime that ships reported 0 µs and SLOWLOG was structurally unable to fire for a write — measured on one connection: 20 sampled SADDs of 3000 members summed to 0 µs while the SMEMBERS control over the same members summed to 2407. Separately, try_inline_dispatch (plain GET/SET, the hottest path) recorded nothing, so moon_command_duration_microseconds{cmd="get"|"set"} did not exist as a series. Both were instances of the three-dispatch-paths trap. All five timing sites (monoio write/read, tokio sharded write/read, tokio single) and the inline path now go through ONE LatencyProbe::observe that takes the command as a closure and brackets exactly it, so the timer cannot be placed on the wrong side of the work again; the inline path takes the probe as a mandatory parameter so an uninstrumented arm cannot be spelled. SlowlogArgv lets the inline path offer its raw argv to the slowlog without building a Frame. The tokio single-shard handler, which timed every command unconditionally, now samples 1-in-16 like the others. Cost is unchanged on the generic paths (same counter, same branch, same Instant cadence); the inline path gains the same per-command counter increment and branch, with total_commands_processed still flushed once per batch.

  • A re-spilled key recovers to its NEWEST on-disk copy, not whichever file the manifest happened to list last (moon#983). The cold-index rebuild resolved a key present in two Active heap files by "last one seen wins", where "last" was ShardManifest::files() order — add_file push order, which is not recency order: the async spill path pushes a file when its background completion is applied, the durable-batch path pushes at eviction time, so a CONFIG SET appendonly flip with a completion still in flight registers a higher file_id ahead of a lower one, and every restart preserves that order. The rebuild then served the superseded value after recovery with no error and no log line. Duplicates are now resolved by ColdLocation::recency_key() — (file_id, page_idx, slot_idx), the spill allocation sequence — regardless of manifest order. Reproduced on the pre-fix binary by booting it on a newest-first manifest built with the spill thread's own writers: GET answered the stale copy; after the fix, the fresh one. The real spill → overwrite → re-spill → SIGKILL → restart lifecycle is pinned at --shards 1 and 4.
  • LMOVE/RPOPLPUSH/BLPOP no longer strand 56 B every time they drain a list to empty (moon#949). Database::list_pop_front/list_pop_back credited the popped element back to used_memory on the non-empty branch but not on the empty one, on the theory that whole-key removal recovers it via entry_overhead. It does not: entry_overhead is computed from the CURRENT value, which by then no longer holds the element, so the push-time charge was never given back. The drift is UPWARD and unbounded on an EMPTY keyspace — --maxmemory and eviction firing on a server holding nothing — and it reaches exactly the reliable-queue pattern that drains a list over and over. Measured at 56.00 B/cycle before and 0.00 after on all three paths, with LPOP/RPOP (which route through pop_eager and always credited) and a SET/DEL pair as controls at 0 both times. STRANDED_BY_LIST_POP in the list tests goes from 56 to 0.
  • A type-refused write no longer aborts a watching transaction (moon#940). Acquiring a mutable handle IS the WATCH version bump (moon#926), and it fired before the arm that answers WRONGTYPE — so SADD/HSET/ LPUSH/ZADD against a key of the wrong type dirtied that key and aborted a concurrent EXEC that redis lets through. The command mutated nothing, so a watcher had nothing to miss: measured against redis 8.6.1 over a raw socket, EXEC answered nil on moon and the queued replies on redis, for all four families. Every accessor now decides the type before it stamps, and the positive control (a real SET) still aborts on both engines. The accepted paths stamp exactly as before, so the moon#942 probe budgets and the "exactly once per command" version counts are unchanged; only the refusal differs.
  • HINCRBY refuses to wrap at the i64 boundary (moon#952). The increment used a plain +, which wraps in release builds, so HINCRBY h f 1 on a field holding i64::MAX replied (integer) -9223372036854775808 and STORED it — silent corruption reported as success, written to the AOF and shipped to every replica. INCR on a string already refused the identical case with ERR increment or decrement would overflow; the hash family now applies the same rule, on both the listpack and the owned-HashMap arm, leaving the field untouched when it declines.
  • Sorted-set range bounds are compared exactly, and an infinite bound now has a direction (moon#966, moon#961). ScoreBound::includes/includes_upper tested an inclusive bound with score >= v || (score - v).abs() < f64::EPSILON. f64::EPSILON is the gap between 1.0 and the next double — a RELATIVE quantity used here as an ABSOLUTE tolerance, which below magnitude 1 spans many ulps: at a bound of 0.5 one ulp is 1.11e-16, so ZCOUNT k -inf 0.5 answered 3 where redis answered 2, returning a member strictly outside the range. The same two methods answered true for NegInf and PosInf in BOTH directions, so ZRANGEBYSCORE k +inf -inf returned the whole set instead of nothing. ZRANGE also handed argv straight to the range helpers, whose contract is (min, max) in semantic order, so the redis spelling ZRANGE k 3 1 BYSCORE REV arrived inverted and matched nothing.
  • ZINCRBY refuses a NaN result and aggregates clamp one to zero (moon#960). The guard checked is_nan() on the increment but never on current + increment, so inf + -inf stored the string NaN; redis answers ERR resulting score is not a number (NaN) and leaves the score untouched. The listpack arm already declined the case but fell through to the eager get_or_create_sorted_set, permanently flattening the encoding from a command that wrote nothing. ZUNIONSTORE/ZINTERSTORE/ZUNION/ZINTER take the opposite rule — redis clamps a NaN aggregate to 0.0 — applied at both the SUM and the weight multiply, since inf * 0 is NaN before any aggregation.
  • Unknown sorted-set option tokens are rejected instead of skipped (moon#967). Every option loop ended in a bare } else { i += 1; }, so a mis-spelled option became a DIFFERENT, successful command: ZINTERCARD 1 k BOGUS 1 dropped the intended LIMIT and returned the unbounded cardinality, and ZMPOP 1 k MIN MAX popped on input redis rejects outright. Six sites. ZRANGE now also rejects BYLEX with WITHSCORES, and a negative LIMIT offset returns nothing rather than being clamped to zero.
  • HINCRBYFLOAT no longer flattens a small hash (moon#958). It was the eighth secondary writer, and the one moon#897 missed: it reached straight for the eager get_or_create_hash, which upgrades on ACCESS, so one HINCRBYFLOAT converted a small hash to hashtable permanently (nothing demotes). It now increments inside the listpack like redis and converts only when a threshold is genuinely crossed — including when the RENDERED value outgrows hash-max-listpack-value, which the count-based listpack_fits check cannot see and a rendered f64 has no useful constant bound for (format_float never uses exponent form, so 1e300 + 1 renders 301 chars).
  • HDEL no longer leaves an empty hash behind when the emptying field is not the last argument (moon#942). hdel tracked emptiness in a last_was_empty variable reassigned on EVERY iteration, including the ones that removed nothing, so the emptiness the real removal reported was overwritten by the false a later absent field produced. HDEL h only absent therefore answered 1, emptied the hash and kept the key: EXISTS h answered 1 and HLEN h answered 0 on a hash redis had already deleted, and the empty container was written to the AOF and shipped to every replica. HDEL h absent only — the same two fields, reversed — deleted correctly, which is why this survived. Emptiness is now a property of the hash after the whole batch.
  • ZADD <missing-key> XX <score> <member> no longer creates an empty zset (moon#942). moon reaches the keyspace through get_or_create_*, which fabricates the container before the mutation loop can discover that XX refuses every member of the batch; Redis short-circuits first (if (zobj == NULL) { if (xx) goto reply_to_client; }) and creates nothing. Both encodings leaked, because a member past zset-max-listpack-value skips the listpack entry gate and lands on get_or_create_sorted_set, which fabricates too. ZADD now drops an empty container the way ZREM always has.

Verified against a live redis 8.6.1 — after one ZADD ghost XX 1 m:

redis moon (before)
EXISTS ghost 0 1
TYPE ghost none zset
DBSIZE 0 1
KEYS * (empty) ghost
DEBUG DIGEST 000…000 01158ee1…
ZCARD ghost 0 0 (agrees — which is why nothing caught it)

This was unbounded keyspace growth on a path any unprivileged client can drive — an entry plus a fabricated container per call, none of which ever appears to hold anything — and it moved DEBUG DIGEST, so a replica or a reloaded RDB disagreed with its master about the keyspace. The fix is pinned on both encodings, together with the ledger (a refused ZADD charges nothing) and the case it must not break (a refused XX on a zset that DOES exist leaves it and its members alone).

Performance

  • ZADD and ZINCRBY on a skiplist zset hash the member once, not three times (moon#942). Redis's zsetAdd does ONE dictFind and writes the new score through the entry it found. moon did three — members.get for the flag decision, then members.remove and members.insert inside zadd_member — and cloned the Bytes twice on a path src/command/ forbids cloning on at all. A new zset_update_existing looks the member up once with get_mut and writes through the slot; the flag decision rides inside its closure, so it has one spelling and the write and the CH tally can never disagree about it.

The B+tree is now touched only when the score actually MOVES, which is Redis's own if (score != curscore) and matters more here than it does there: BPTree::remove builds a Bytes::copy_from_slice(member) to form its lookup key, so an idempotent re-post used to cost an allocation as well as a tree delete and a tree insert.

Measured with cfg(test) counters:

on a 200-member (skiplist) zset before after
members hashes, ZADD onto an existing member 3 1
members hashes, ZADD of a new member 3 2
members hashes, ZINCRBY onto an existing member 3 1
members hashes, ZREM 1 1 (control)
B+tree re-insertions, same score re-posted 1 0
B+tree re-insertions, score genuinely moved 1 1 (control)

Counts, not a throughput claim (PERF-08, moon#789).

The comparison is to_bits(), not ==, for the same reason the listpack arm compares bytes: -0.0 == 0.0 is true while -0 and 0 render differently, and moon has always stored whichever spelling the client sent. the_bptree_identical_score_skip_is_decided_on_bits pins it.

The ledger is guarded, not argued. zset_member_cost moved from is_new at the end of zadd_member into the ABSENT arm of a match, and a charge that moves is a charge that can be dropped or doubled (moon#814/#788). every_restructured_arm_keeps_the_ledger_exact walks both encodings through create, the skipped write, a widened score, a narrowed score, NX and XX refusals, ZINCRBY both ways and ZREM, asserting the running ledger against a from-scratch recalculate_memory at every rung — and returning the listpack ladder to its exact floor. Proven able to fail: deleting the charge line makes it report a 41,265 B ledger against a 42,865 B recompute.

  • A ZADD that consults nothing stops decoding the stored score, and a score already in place is no longer written back over itself (moon#942). A zset listpack keeps a score as canonical decimal text, so reading one is a real str::parse::<f64> — and the plain ZADD z <score> <member>, which is the benchmark's shape and most applications', consulted it for nothing: should_update was unconditionally true and the changed tally it fed is not what the command replies. It is now decoded only for CH, GT and LT (NX refuses on presence, which locating the pair already established). On top of that, Redis's zsetAdd re-inserts only if (score != curscore) while moon spliced the rendering back over itself every time — so the idempotent re-post a leaderboard client makes paid an encode_entry and a write_entry to change nothing.

Measured with cfg(test) counters, on the benchmark's own zadd z:<n> 1 m:<n> shape:

ZADD onto an existing listpack member before after
stored-score decodes, no flag and no CH 1 0
stored-score decodes with GT / CH 1 1 (control)
score entries written, same score re-posted 1 0
score entries written, score genuinely moved 1 1 (control)

A call count is not a throughput claim and this entry makes none (PERF-08, moon#789).

The skip is decided on the RENDERED BYTES, not on old == score, and that is deliberate: -0.0 == 0.0 is true while -0 and 0 are different bytes, so comparing doubles the way Redis does would have silently started answering 0 to a client that wrote -0. moon has always stored whichever spelling the client sent, and the_listpack_identical_score_skip_is_decided_on_bytes pins that. Writing bytes that are already there changes nothing, so declining is observationally identical everywhere else.

  • ZADD parses each score argument once instead of twice (moon#942). moon#814's validation pre-pass — which must keep proving every pair before the keyspace is touched, because the mutation loop returns from inside the table_before … charge_memory() window — threw every decoded f64 away, and the loop then re-ran str::parse::<f64> over the same bytes. Redis's zaddGenericCommand parses once into its own scores array. The pre-pass now keeps what it decodes, in a fixed 32-entry stack array: src/command/ forbids the heap allocation a SmallVec spill would make, so a batch larger than that re-parses in the loop exactly as before.

Measured with a cfg(test) counter on the parse itself: a one-pair ZADD went from 2 parses to 1, a four-pair batch from 8 to 4. A call count is not a throughput claim and this entry makes none (PERF-08, moon#789).

The all-or-nothing contract is unchanged and guarded: a_rejected_batch_still_validates_every_pair_before_the_keyspace asserts that a bad pair anywhere still errors and still leaves the key uncreated — the one thing caching the pre-pass's output could have broken.

  • An integral sorted-set score is no longer rendered by core::fmt's f64 Display (moon#942). Every score moon stored in a listpack, replied to a client, or wrote to an RDB went through write!("{score}") — the shortest-round-trip (Grisu/Dragon) formatter plus the whole Formatter machinery — including the literal 1 the ZADD benchmark writes on every call and every integral leaderboard score ever posted. Redis's d2string has always forked here: double2ll first, ll2string when it succeeds, and fpconv_dtoa only when it does not. zset_score::render_score now takes the same fork through a new integral_score, and format_score / format_score_bytes — which were an independent second transcription of the same three rules — delegate to it, so ZSCORE, ZRANGE … WITHSCORES, ZINCRBY, ZPOPMIN/ZPOPMAX, ZMSCORE, the blocking BZPOPMIN wakeups and DEBUG DIGEST all take it too.

Measured with a cfg(test) counter on the slow arm, on the benchmark's own zadd z:<n> 1 m:<n> shape: a ZADD with an integral score reached core::fmt's float formatter 1 time before, 0 after; a ZADD with 1.5 still reaches it exactly once, which is the control that proves the counter is live.

A call count is a count. It is not a throughput claim and this entry makes none — PERF-08 (moon#789) measured +11% on aarch64 and −17% on x86_64 for one work reduction in this repo, so the wall-clock question belongs to a Linux bench host.

The bytes are identical wherever the fast path is taken, and that is asserted against core::fmt DIRECTLY rather than against another moon function: the_integer_fast_path_is_byte_identical_to_core_fmt sweeps the hand-written cases, both sides of the 2^53 cutoff, every integer in ±1100, and 20,000 deterministic f64 bit patterns plus their truncations. Three exclusions carry the proof — -0.0 (which prints -0, not 0), non-integers, and anything past 2^53 (which catches both infinities and NaN, whose fract() is NaN). Both were proven able to fail: widening the cutoff to f64::MAX makes 1e21 report 1000000000000000000000 against -9223372036854775808, and dropping the -0.0 arm makes -0.0 report 0 against -0. A listpack or RDB written before this change still reads back byte-for-byte the same, which is what parse_score's round-trip exactness and every reader of a stored score depend on.

Performance

  • HDEL costs ONE key lookup for the whole command, not two per field: a three-field HDEL went 6 → 1 (moon#942). hdel called Database::hash_delete_field once per argument, and that method costs two DashTable probes — its own data.get_mut(key) and, on a real removal, stamp_hash_field_mutation's, which re-finds the same entry to stamp it. Redis pays one dictFind for the command and then looks each field up inside the hash; Database::hash_delete_fields now does the same, walking the batch inside the handle the single lookup already holds and stamping that same handle. hash_delete_field is kept as a delegator to it, so the one-field and many-field spellings cannot drift.

Measured with the cfg(test) DashTable key-lookup counter:

HDEL arm before after
one field, present 2 1
three fields, all present 6 1
three fields, none present 3 1

It also makes one command ONE WATCH version bump (moon#926). HDEL h f1 f2 moved the version by two, because the per-field loop stamped once per removed field — the same observable shadow of a duplicated accessor that sadd_bumps_the_watch_version_exactly_once_on_every_encoding pins for SADD. A batch that removes nothing still leaves a watcher alone, as before.

  • HINCRBY walks a hash listpack ONCE, not twice (moon#942). Its listpack arm called Listpack::pair_value — a borrowed scan that located the field — and then Listpack::replace_pair_value, which located the SAME field all over again from the head, so a 128-field hash was walked up to 256 entries to change one value. That is the defect moon#799 removed for HSET and moon#942 removed for ZADD/ZINCRBY; Listpack::update_pair_value is the one-scan primitive built for callers whose replacement depends on the value currently stored, and HINCRBY was the last hash caller still walking twice. Structural, not counted: this repo has no listpack entry-decode counter, and one is proposed rather than asserted here.

  • HSET and HMSET walk their argv once, not twice (moon#942). The moon#823 all_args_are_bytes validation pass matched every frame and the moon#896 entry gate's max chain then matched every frame again to read its length. One loop now does both, and still refuses a non-argument-shaped frame strictly before get_or_create_hash_listpack opens the mutation window.

  • HSET's probe budget is now ratcheted (moon#942). Measured on the harness's own shape — scripts/bench-ab-matrix.sh:84 is hset hash:__rand_int__ f __rand_int__ over a 100 000-key keyspace, so the benchmarked HSET touches hashes holding exactly ONE field named f:

HSET arm probes notes
absent key (create) 3 hot_state + insert + get_mut
listpack steady state 2 the benchmarked arm; at this accessor family's floor
full-HashMap steady state 5 open — 2 thrown away + 2 + 1, see below

The full-HashMap arm's five is the hash equivalent of the SADD defect fixed above and is NOT fixed here: get_or_create_hash_listpack answers Ok(None) for a Hash or HashWithTtl (2 probes, and a WATCH bump, discarded), get_or_create_hash then re-runs the whole classification (2 more, and a second bump), and hash_clear_field_ttls re-finds the entry a third time (1) only to discover a plain Hash carries no sidecar. The fix is a HashHandle mirroring SetHandle in storage/db/accessors.rs, which this change does not own.

None of these numbers is a throughput claim and this entry makes none. The PERF-08 precedent (moon#789) measured +11% on aarch64 and −17% on x86_64 for a probe reduction — the two architectures disagreed in sign — so the wall-clock question belongs to a Linux bench host and to nothing else.

The ledger is byte-identical across the HDEL move. The per-field credits a batch books are the same hash_field_cost / hash_ttl_field_cost figures summed into one credit_memory, and the listpack arm's one before/after estimate_memory pair reports the same capacity-based delta the per-field snapshots did (Vec::drain does not shrink capacity). A new guard asserts the running ledger against a from-scratch recalculate_memory after a mixed hit/miss batch on a 4-field listpack and a 400-field HashMap.

Every new guard was proven able to fail, by mutation: restoring the per-field stamp makes the probe budget report (2, 4, 1) and the WATCH guard report two bumps; forcing empty = false resurrects the empty-hash leak; dropping the refusal from the fused argv walk fails the moon#823 pin; taking the LAST element's length instead of the longest fails the moon#896 gate pin; and making HINCRBY's one-scan closure refuse a Str entry fails the round-trip pin on a non-canonical stored spelling.

Performance

  • Every integer argument the string, bitmap and list commands take is read in one pass (moon#942). command::string::parse_i64 — the parser behind INCRBY/DECRBY's delta and behind every offset, index and count in SETRANGE, GETRANGE, SETBIT, GETBIT, BITCOUNT, BITPOS, LRANGE, LINDEX and LPOS — walked its argument twice, once for from_utf8 and once for str::parse. It now calls storage::numeric::parse_i64_bytes, and the accepted set is byte-for-byte unchanged (see the entry below for how that equivalence is pinned).

  • INCRBYFLOAT allocates twice per call, not three times (moon#942). format_float built a String with format!, trimmed it, and then built a second String with to_string() to hold a prefix of the one already in hand; the handler then clone()d the result so the stored Entry and the reply could have one each. format!, to_string() and clone() are all three banned in src/command/ by CLAUDE.md, and this one command paid all three per call.

The trim is now a truncate — a length store, no copy — and the Entry is built from the rendered bytes rather than from a second copy of them, which for any result inside CompactValue's 12-byte SSO window means it needs no buffer at all. Rendered output is unchanged, pinned by a differential against the exact pre-#942 body over ~2,400 values.

Measured with a counting GlobalAlloc, per call, after warm-up:

INCRBYFLOAT arm before after
integral result 3 2
fractional result 3 2
negative fractional 3 2
wider than the SSO window 2 2

The remaining two are format!'s String growth and Bytes::from(String) reallocating in into_boxed_slice. Reaching the true floor (1, or 2 outside the SSO window) needs an f64 Display renderer that writes into a stack buffer, which is not done here.

  • INCR stops walking its counter twice to read one integer (moon#942). The path parsed the stored value as std::str::from_utf8(bytes) followed by str::parse::<i64>(). The first walk proves the slice is UTF-8; the second walks the same bytes again rejecting everything that is not an ASCII digit. The first is redundant by construction — every byte string the i64 grammar admits is [+-]?[0-9]+, which is pure ASCII and therefore always valid UTF-8 — so from_utf8 can only ever reject inputs the digit scan was going to reject anyway, and its verdict is never the deciding one. Redis reaches the same answer in one pass, in string2ll.

storage::numeric::parse_i64_bytes reads that grammar straight off the bytes. It also hoists the range question out of the per-digit loop: once leading zeros are skipped, more than 19 significant digits cannot fit an i64 at all, and 19 digits of 9 is comfortably inside u64, so the accumulator provably cannot wrap and the two checked_* operations and char::to_digit's radix handling go with it.

It is exactly from_utf8(b).ok().and_then(|s| s.parse::<i64>().ok()), including every permissive spelling str::parse accepts and canonical_i64 does not ("007", "+5", "-0"). Changing which spellings a counter accepts would be a client-visible behaviour change, so the equivalence is pinned rather than described: differentials over every 1-byte input, every 2- and 3-byte word from a discriminating alphabet, both i64 boundaries digit by digit, zero-padded and over-long forms and 200,000 randomised inputs; an end-to-end INCR/DECR differential against the pre-#942 from_utf8 + str::parse reference over 25 accept/reject shapes; and a new parse_i64_bytes_differential fuzz target registered in both matrices of .github/workflows/fuzz.yml.

No throughput claim. Removing work is a hypothesis about wall clock, not a measurement of it.

  • Creating a counter with INCR costs 2 key lookups, not 5 (moon#942). The in-place fast path added for INCR/INCRBY/DECR/DECRBY made the hot counter cost one DashTable probe — Redis's single lookupKeyWrite — but made the absent one cost five, up from four: its declined get_mut was spent before incrby_general started over with Database::get (classify + cold guard + re-probe) and Database::set. The module docs called that a deliberate trade on the theory that a counter is created once and incremented many times. True in production; false in the instrument this work is measured on — scripts/bench-ab-matrix.sh seeds only key:* and set:*, so all 100k of its ctr:* keys are created by the benchmark itself, which is essentially every INCR at p=1 and about a quarter of them at p=8.

Database::incr_string (was incr_hot_string_in_place) now owns the create as well. It fabricates nothing until promote_cold_known_absent has ruled out both the in-flight spill plane and the cold tier — the same method the container accessors use, whose precondition is discharged by the get_mut immediately above it — and hands a key that really was promoted straight back to the general path, which re-reads it.

Measured with the cfg(test) DashTable key-lookup counter, over the whole INCR command rather than the storage method (a decline whose fallback costs four more probes is invisible at method level — that is how this one got in):

INCR arm before after
hot, live, integer 1 1
absent from every plane 5 2
present but TTL-expired 3 3

A probe count is a count, not a throughput claim, and this entry makes none: PERF-08 (moon#789) measured a probe reduction at +11% on aarch64 and −17% on x86_64 — the two architectures disagreed in sign. The wall-clock question belongs to a Linux bench host. - The list pops stop paying for work they already had in hand (moon#942). Three reductions on the LPUSH/RPUSH/LPOP/RPOP family, all inside src/command/list/:

path before after unit
LPOP/RPOP off a listpack 2 1 heap allocations per element
LPOP/RPOP off a linkedlist 4 3 DashTable key lookups
LPUSHX/RPUSHX, hit 4 3 DashTable key lookups
LPUSHX/RPUSHX, miss or WRONGTYPE 2 1 DashTable key lookups
  1. listpack_pop_end decoded the popped element through the OWNING ListpackEntry: Listpack::get_at allocates a Vec and ListpackEntry::to_bytes then goes through as_bytes, whose string arm CLONES it — two heap allocations, the first dropped having been copied and never read. It now decodes through the borrowed ListpackRef that Listpack::iter_refs already hands out. One is the floor, not zero: the reply owns its bytes and remove_at mutates the buffer on the next line. src/command/ is forbidden from allocating on the hot path at all, and moon#897 made the listpack encoding survive a pop — so every small work queue in the tree now takes this path on every drain, where before #897 it flattened once and paid zero thereafter.
  2. pop_eager popped through get_or_create_list, dropped the borrow, and then asked get_list_ref_if_alive a FOURTH time whether the list was now empty — a question the &mut VecDeque it had just been holding answers for free.
  3. LPUSHX/RPUSHX's exists-and-is-a-list gate was db.get_list(key) = get_promoted: two lookups, and a &mut self accessor whose ListKind::upgrade flattens the compact encoding just to answer a yes/no. It is now the &self router LPOP already uses (moon#897), whose receiver makes the rewrite unrepresentable. The mutable accessor behind it is unchanged, so the encoding OUTCOME is unchanged too.

    The one trade: for a COLD-SPILLED list the &self router decodes a throwaway copy that get_promoted would have promoted outright, so a spilled key now costs one extra decode on these two commands — the same trade pop_generic already makes, and hot keys pay nothing. What it must NOT do is answer 0, which for LPUSHX is a silently dropped write wearing a success-shaped reply; pushx_sees_a_cold_spilled_list_and_appends_to_it asserts the cold read-through rather than trusting the accessor's comment, because moon#610 is exactly the class of a read-only path that forgets the cold tier.

Which arm does the benchmark exercise? The listpack one, and only that. scripts/bench-ab-matrix.sh's row is LPUSH list:__rand_int__ xxxxxxxx at -r 100000, so ~2.0M pushes spread over 100,000 keys leave a list: key holding roughly 5 elements when the p=64 point starts and 20 when it ends, against a list-max-listpack-size of 128. No timed LPUSH in the matrix ever reaches the quicklist arm, and LPUSH on the listpack arm was already at the accessor skeleton's 2-lookup floor before this change. None of the three reductions above can move the benchmarked LPUSH number, and this entry does not claim they do. They are worth having for real lists, which are longer than a benchmark's, and for the queue drain, which a benchmark of pushes does not measure at all. No throughput number was measured and none is claimed; PERF-08 (moon#789) is why — a probe reduction in this repo measured +11% on aarch64 and −17% on x86_64.

Still open, and now pinned rather than described: LPUSH onto a linkedlist costs FOUR lookups and moves a watched key's version by TWO, because get_or_create_list_listpack answers Ok(None) for a key already holding the full VecDeque and lpush answers that by re-running the whole skeleton through get_or_create_list. That is exactly the shape moon#942 closed for SADD, and closing it needs the same SetHandle-shaped change to src/storage/db/accessors.rs. lpush_end_to_end_probe_budget and list_writes_bump_the_watch_version_exactly_once assert both at their CURRENT values, in both directions, so the fix has to update them deliberately.

Every new guard was proven able to fail by mutation, and the mutants are named in the tests. One claim did not survive that check and was corrected rather than kept: deleting pop_listpack's db.adjust_memory(before, after) leaves the ledger tests green, because Listpack::estimate_memory bills the size class of the buffer's CAPACITY and nothing shrinks a listpack's buffer on removal — so before == after on every pop and the call is a genuine no-op there.

A bug this work FOUND and did not fix. LMOVE, RPOPLPUSH and the whole BLPOP/BRPOP/BLMOVE/BRPOPLPUSH family strand used_memory every time they drain a list to empty, without bound, on a keyspace that ends up empty. Database::list_pop_front and list_pop_back (src/storage/db/accessors.rs:1192-1222) credit list_elem_cost(&val) on the else branch and not on the if empty branch, on the stated theory that "whole-key removal recomputes the (now-empty) entry cost via entry_overhead" — but entry_overhead is computed from the CURRENT value, which no longer holds the element, so the push-time charge is never given back.

Measured on an otherwise empty Database: one RPUSH k e followed by one list_pop_front leaves used_memory at 56 B against a from-scratch recalculate_memory of 0 B, on both encodings; ten create/drain cycles leave 560 B. The drift is UPWARD, so the consequence is --maxmemory and eviction firing on a server that is actually empty — the opposite direction from moon#814, and reached by precisely the reliable-queue pattern that drains a list over and over. The LPOP/RPOP command paths are exact and are unaffected.

The fix is one line in each accessor — credit the element unconditionally, as pop_eager already does — but accessors.rs was out of scope here. every_list_writer_that_empties_a_list_removes_the_key enumerates all nine list state writers that can remove the last element, asserts all nine do remove the KEY (they do), and pins this drift at its measured 56 B so the fix shows up as a test that needs updating rather than as silence.

  • SADD on a hashtable set stops paying for a second accessor: 4 key lookups → 2 (moon#942). get_or_create_set_listpack answered Ok(None) for a set that was already an IndexSet — or a SetIntset the moon#899 absorb refused — and every caller then called get_or_create_set, which re-ran the entire accessor skeleton (hot_state, settle_not_live, get_mut, stamp_mutation, SetKind::upgrade) against the key the first call had already classified and was still holding. The accessor now returns a SetHandle { Listpack(&mut Listpack), Full(&mut SetValue) } and runs the upgrade on that handle, so the second accessor is gone.

This lands on the hashtable regime, not the miss regime — 122 of 200 keys at the benchmark's own p=64 point — which is why it is the largest remaining item in the SADD residue the previous entry left open.

Measured with the cfg(test) DashTable key-lookup counter, on the benchmark's own SADD set:<12-digit> <12-digit> shape:

SADD arm before after
absent key (create) 3 3
listpack steady state 2 2
hashtable steady state 4 2
refused-intset promotion 4 2

A probe count is a count. It is not a throughput claim, and this entry makes none: the PERF-08 precedent (moon#789) measured +11% on aarch64 and −17% on x86_64 for a probe reduction — the two architectures disagreed in sign — so the wall-clock question belongs to a Linux bench host and to nothing else.

The ledger is byte-identical across the move. The used_memory delta that moved is the same SetKind::upgrade call returning the same isize, applied to the same counter one accessor earlier; nothing is charged twice and nothing is dropped. Two new guards in ledger_consistency_788 assert the running ledger against a from-scratch recalculate_memory at every rung of one key's ladder — empty → intset → the moon#899 absorb → listpack → the 64/65-byte set-max-listpack-value boundary in both directions → hashtable → duplicate → DEL back to the floor — and again across the intset → IndexSet promotion. Both were proven able to fail: deleting the delta line from the new Full arm makes the refused-intset guard report a 729 B ledger against a 14,665 B recompute. moon#814 is why this matters — a charge stranded on a branch drives used_memory monotonically DOWN, without bound, on a path any unprivileged client can drive, until --maxmemory can never fire.

Side effect on WATCH, in the correct direction. moon#926's rule is that acquiring a mutable handle IS the version bump, and the hashtable arm was acquiring two, so one SADD moved a watched key's version by two. It now moves it by one. A watcher aborted either way, so this is not a behaviour fix; it is the observable tell that the duplicate accessor is gone, and a new test pins it on all three encodings. moon#940 was untouched by this change and was still open when it landed: stamp_mutation fired before the arm that answers Err(WRONGTYPE), so a rejected SADD still dirtied the key, and a test pinned that at exactly one bump so this change was provably neutral on it. moon#940 is fixed separately in this same release — see the entry above; that pin has been replaced by one asserting no bump at all.

Encoding behaviour is unchanged and was checked against a live redis 8.6.1 on every rung of the ladder above, plus the entry-count boundary, SREM on both compact forms, and the moon#795 byte-transparency cases (007, +7, -0, both i64 limits) re-asked after each promotion: every OBJECT ENCODING, reply and SISMEMBER matched, and the whole-dataset DEBUG DIGEST was identical (verified discriminating — one extra member on moon alone changes it).

  • The accessor miss path drops its fourth probe: cold promotion stops re-asking whether the key is hot (moon#942). promote_cold_if_present opens with a contains_key, and accessors::settle_not_live — the only caller on the get_or_create* preamble — reaches it exclusively from HotState::Absent (hot_state's get has just answered None) or from HotState::Expired after remove_hot, which removes unconditionally. Both arms leave the key provably absent, so that contains_key could not do anything but re-answer a question one probe old. It is now promote_cold_known_absent, the same method without the re-ask, and settle_not_live calls that instead.

Measured with the cfg(test) DashTable key-lookup counter added by the preceding entry, which is also where the pre-change numbers below come from — this change moves only the miss column, and no hit path moves at all:

accessor miss before miss after
get_or_create 4 3
get_mut_if_present 3 2
get_promoted 3 2
get_or_create_intset / _hash_listpack / _list_listpack / _zset_listpack / _set_listpack 4 3

End to end, SADD on an absent key goes 4 → 3. Its hashtable regime stays at 4 and is untouched: that one is two accessors' worth (get_or_create_set_listpack answers Ok(None), get_or_create_set then repeats the skeleton) and collapsing it needs the accessor skeleton itself to hand back the encoding it already holds, which this change does not do.

The skipped contains_key is load-bearing for data integrity, not just for speed, which is why this is a second entry point and not a deletion. promote_inflight_if_present does not re-check residency — it calls Database::set unconditionally — so on a key that is hot AND still carries an in-flight spill record, that probe is the only thing between a live value and the older spilled body overwriting it. (promote_cold_outcome, the on-disk arm, carries its own guard and is safe either way.) promote_cold_known_absent is therefore pub(super), documents the precondition, and has exactly one caller. A new test builds that exact state and pins it: removing the guard from promote_cold_if_present makes it report the 3-member spilled body where the 2-member live value should be. A second new test pins that a key DEL'd mid-spill does not resurrect through the accessor (moon#459); dropping spill_inflight_forget from remove_cold_only makes that one report 3 members instead of 0.

No throughput number is claimed, and none was measured. The same discipline as the entry below applies and applies harder here, because the quantity is smaller: moon#789 measured the previous probe-count reduction in this repo at +11% on aarch64 and −17% on x86_64, disagreeing in sign, and b3083c5a found the wall-clock net guarding it was timing a page-fault artifact. This removes one probe from a miss, so at the SADD benchmark point (75 of 200 keys absent) it is a fraction of a single probe per operation and is expected to sit below the ~1.5% noise floor on its own — it is committed separately so it can be A/B'd and dropped on its own evidence. Fewer probes is a structural fact; whether it is faster is a question only a Linux benchmark host may answer, on both arches.

Behaviour is unchanged. Beyond the two new tests, every guard the preceding entry installed still passes unmodified — expired-reads-as-absent through every accessor with no expiry-index leak (moon#541), cold promotion before fabrication on both the absent and the expired arm (moon#459), WRONGTYPE on a live key of the wrong type, one WATCH version bump per mutable handle (moon#926; moon#940 is neither fixed nor worsened, and stamp_mutation did not move), and used_memory agreeing with an independent whole-keyspace recount. Separately verified against a live redis 8.6.1 oracle: 33 of 33 rows agree, covering every set/zset encoding transition through the create path this change rewrote — intset → listpack (moon#899), listpack → hashtable at the entry and the 64/65-byte value boundary, the zero-padded benchmark shape staying a listpack with its bytes intact (moon#795), zset listpack → skiplist, and SADD/ZADD creating a fresh container after the key expired rather than resurrecting the old members — with a byte-identical DEBUG DIGEST over the whole dataset at the end.

  • Writing a listpack entry no longer touches the allocator, and ZADD no longer walks the listpack twice (moon#942). Two changes in src/storage/listpack.rs, both on the write path HSET, LPUSH, SADD and ZADD share.

encode_entry built a Vec<u8> per entry written, copied it into the listpack and dropped it. Measured with a counting allocator over 100 same-width replacements — a width that neither grows nor shrinks the buffer, so the only correct answer is zero — it was four allocations per string entry and two per integer entry: Vec::new() starts at capacity 0, so the encoding head and the backlen each grew it, and encode_backlen allocated a second Vec of its own. The encoding head now goes into [u8; LP_MAX_ENTRY_HEAD] and the backlen into [u8; LP_MAX_BACKLEN], both bounds derived from the encoder's own arms rather than picked; the payload is borrowed and copied straight into the listpack, because no constant can bound it (hash-max-listpack-value and friends are runtime config, and the RDB/AOF loaders rebuild listpacks with no element-size limit at all). The four Vec::splice call sites became one write_entry that moves the tail once with copy_within — and not at all when the replacement is the same width, which the common HSET/ZADD update is. Same shape as Redis's lpInsert. The counter now reads 0.

ZADD and ZINCRBY scanned a zset listpack twice to change one score: a borrowed walk down to a pair ORDINAL, then replace_at walking back to that ordinal from the head. That is the defect moon#799 fixed for HSET and left standing for the sorted set. Listpack::update_pair_value is HSET's locate_pair/replace_pair_value generalised to a caller whose replacement depends on the old value, so the scan that finds the member writes the new score where it stopped; replace_pair_value is now that method with a constant decision. A test-only seek counter pins it — the shape this replaces seeks from the head once, the new one never does.

No throughput number is claimed here. Benchmarks are Linux-only and these were written on macOS; the allocation and walk counts are counts, and the ops/s effect is unmeasured. The encoding itself is unchanged, byte for byte, which is the part that matters for a format DUMP/RESTORE and the RDB both write out: goldens captured from the old encoder — every integer width at both signs of every boundary, every string width at its seam, non-UTF8 payloads, the moon#795 leading-zero and leading-+ families, and widening/narrowing/equal-width mutation sequences — were committed before the rewrite and still pass unedited.

Documentation

  • BENCHMARK.md: re-measured the eight command families on GCE against 5bf716a9, and settled moon#923. Ten PRs had landed since the last run, two on the measured hot path. #938's stamp_mutation (+8 instructions per container write, measured by disassembly and never throughput-tested) and #929's keyless-table arms are both below the noise floor: fifteen of sixteen container-write cells are ties across two architectures at n=10, the architectures disagree in sign on six of eight rows, and the one non-tie is a gain. Recorded as an unresolvable effect rather than a measured zero — 8 instructions is ~0.1% of a ~3.2 µs/op budget and needs perf stat, not ops/s.

moon#923 is attributed: the whole ae6cd003..1a7bd83f window is #861 (box RedisValue's fat variants). A four-point, two-pass bisect at n=10 puts SADD p=64 at -6.4%, SPOP at -4.5% and ZADD at -4.7% in that single commit, with the two flanking windows tied on all 48 of their cells — overturning the issue's "spread across three windows" framing. The mechanism is boxed-variant indirection, and it correlates without exception: the three families that regressed hold a variant #861 boxed at their own measurement point (set hashtable, zset skiplist), and the two that tie hold 24-byte listpack variants it did not box. A symbol-resolved perf A/B localises the cost to the handler bodies (set_write::sadd 0.80% → 6.45%, spop 6.10% → 10.15%) and rules out the allocator (1.94% → 2.02%). Establishing this needed the harness leg replayed in order and the encoding population tallied — the families mutate each other's keyspace, and the regime oscillates per key. The standing refutation of #861 was right about HSET, which the bisect confirms is a tie, and wrong only to generalise from it.

Also measured for the first time: container memory. §3 had covered string values only, so the encoding campaign's payoff was measured nowhere. Two figures per type, and the pair is the point — before #920/#921/#922 one secondary write per key cost a hash +344% and a zset +1959% (217 B → 4,478 B); at HEAD it costs +0.1-0.5%. Against Redis 7.0.15 moon is 0.54-0.56x on sets and 0.89-0.91x on lists, but 1.21-1.29x on hashes and 1.29-1.38x on sorted sets — and only hash and zset are like-for-like encoding comparisons, so §1's "moon uses less memory per key" is now marked as string-only.

Confirmed ZADD's +38% held (+39.8% at p=64) and now survives a read: before #932 one ZSCORE on the mutable dispatch path flipped a listpack zset to a skiplist for good, at +1935% RSS (218 B → 4,433 B per key).

Added

  • A test that enumerates COMMAND_META and drives the real wait decision, intercepted_commands_wait_for_pending_remote_writes_moon937 (src/server/conn/shared.rs). tests/intercept_flag_drift.rs guards one direction — marked NO_INTERCEPT implies no gate claims it — whose flag fails safe. The other direction, a gate claims it implies must_wait_for_pending_remote says wait, was unguarded, and is_inline_intercepted fails open: a missing entry reads as permission, which is moon#507's shape and how moon#937 happened. The new test drives the real decision function over the registry rather than a hand list, probes each command twice (a key-shaped modifier and a numkeys form — the one-probe version stayed green under a deliberate deletion), and probes the gates whose name predicate lives outside the five scanned files explicitly, since no text scan can see those. Six commands answer "safe" today; each is waived by name with a reason, and the waiver is pinned to exactly the set that still offends, so a new undeclared interceptor fails one assertion and fixing one of the waived commands fails the other. Verified by mutation: five separate breakages — dropping EVALSHA, dropping EVAL, declaring TXN, breaking the gate scan, breaking a delegated probe — each turned it red.
  • Fuzz target canonical_i64_differential — pins is_canonical_i64(d) == canonical_i64(d).is_some() for every byte string, plus the moon#795 property itself (an accepted value must render back to the caller's exact bytes) and a suffix-extension check that a length-confused recognizer would fail. Two hand-written recognizers of one grammar stay honest only by agreeing, and a divergence here is a data-corruption bug on a hot path, not a cosmetic one. Registered in fuzz/Cargo.toml and in both matrices in .github/workflows/fuzz.yml — an unlisted target never runs.

The target was proved to find its own bug before being trusted. Against the correct implementation it ran 16,290,362 executions in 46s with no finding; against an implementation with the -0 rule removed it crashed on [45, 48] — "-0", one of the original moon#795 vectors — reached by the suffix-extension check from the one-byte input "-".

  • storage::numeric::is_canonical_i64 — canonical_i64's verdict without its value, for the callers that only need to ROUTE. canonical_i64 costs a UTF-8 validation, an i64 parse, an itoa render and a memcmp; that is the right price when the value is wanted, and pure waste when the answer is a yes/no. SADD's all_integers pre-pass is the first caller: it walks the batch to decide whether it belongs in an intset and then discards every number it parsed, because the push loop re-derives the ones it needs. is_canonical_i64 decides the same question from the bytes — sign, digits, no leading zero, no -0, and one slice compare against the i64 boundary at 19 magnitude digits.

Verdict-identity is the whole contract, and moon#795 is why: a non-canonical spelling that slips into an integer encoding destroys the caller's bytes (SADD s 000000012345 came back as 12345). So the equivalence is pinned DIFFERENTIALLY against canonical_i64 itself, never against hand-written expectations — exhaustively over every 1-byte input and over a discriminating alphabet at widths 2 and 3, across both i64 boundaries digit by digit, and over 200,000 randomised digit-heavy inputs. All four differential tests were confirmed to FAIL against a deliberately mutated implementation before being trusted (dropping the -0 rule; dropping the range check), and a command-level test pins the observable consequence: each of the moon#795 vectors — 007, +7, -0, " 7", "7 ", the empty string, a 20-digit number, and i64::MIN/i64::MAX on both sides of their exact boundaries — must still route to the same encoding and come back byte for byte.

No throughput number is claimed here. The candidate came from a 75ad520c SADD-only profile showing from_utf8 1.37% + itoa 0.90% + canonical_i64 0.55%, but those symbols have other callers on the same leg (the listpack encode path among them), so their attribution to this pre-pass is unverified — benchmarks are Linux-only and this landed from macOS. The change is justified by doing strictly less work for a provably identical verdict, not by a measurement. - test: a DashTable key-lookup counter and a probe budget for the storage accessors (moon#942). DashTable's six lookup entry points (get, get_mut, insert, insert_or_update, remove, remove_entry) now record a per-thread count under cfg(test), behind the same #[cfg(not(test))] #[inline(always)] no-op that moon#789 established for note_simd_probe — zero production cost, in the hottest lookups in the codebase. It is deliberately COARSER than segment::take_simd_probes: that one counts control-byte group scans and varies with a segment's fill, this one counts how many times a caller hashes a key and walks a segment at all, which is a property of the accessor rather than of the table.

storage::db::probe_budget pins the measured budget of every get_or_create* / get_promoted / get_mut_if_present accessor and of SADD end to end, on a hit and on a miss. The baseline it records is measured, not read off the moon#942 audit — which undercounted get_or_create's miss path by one (promote_cold_if_present opens with its own contains_key). These are the PRE-reduction counts, which is the point of recording them; the reduced ones are under Performance below and are what the module asserts at HEAD:

accessor hit miss
get_or_create 3 6
get_mut_if_present 3 4
get_promoted 4 5
get_or_create_intset / _hash_listpack / _list_listpack / _zset_listpack 3 6
get_or_create_set_listpack 4 7
SADD end to end — absent key / listpack regime / hashtable regime [7, 4, 7]

Redis reaches all of these with one dictFind. No throughput claim is made or implied: the counter exists precisely because moon#789 measured the last probe-count change at +11% on aarch64 and −17% on x86_64 — the two architectures disagreed in sign — and because b3083c5a found the wall-clock net for it was timing a page-fault artifact. A probe count is a structural fact; only a Linux benchmark host may speak about time.

  • scripts/bench-ab-delta.py — compares moon against moon across two matrix runs, which bench-ab-report.py cannot do. Redis is the control: the tool picks the raw or the ratio column per row from the control's own measured movement, and publishes neither when the control's CV exceeds 5%. That guard exists because it fired: Redis's own ZADD p=8 series went bimodal within one session, which turns a flat moon row into a false +13%.
  • scripts/bench-ab-memory-containers.sh — per-key RSS for hash/list/set/zset, reported both untouched and after one secondary write, with a per-row arithmetic floor.
  • scripts/bench-zset-read-encoding.sh — checks whether a read destroys a small zset's encoding, sweeping the plain, MULTI/EXEC and Lua dispatch paths and reading both OBJECT ENCODING and /proc RSS. The plain leg is a deliberate negative control: the defect it checks for is unreachable from the read-only path, so a bare ZSCORE runs clean against a buggy binary.

Security

  • ACL SETUSER: a -rule after a +rule no longer flips the base policy (GHSA-9x86-7597-5wwj). CommandPermissions::Specific carried no base polarity — is_command_allowed inferred it from whether allowed happened to be empty. Both mutators prune the opposite set (allow_command does denied.remove, deny_command does allowed.remove), so the last rule applied could reverse the meaning of every rule before it:

  • -@all +get then -get emptied allowed and flipped the base to allow-all, escalating a restricted user to the entire server — SET, CONFIG GET, ACL WHOAMI and INFO all succeeded where Redis denies them. Reachable through --aclfile, so importing a Redis users.acl could silently grant full access to a locked-down user.

  • +@all -get then +get emptied denied and flipped the base the other way, leaving the user holding only get where Redis restores full access.

Specific now stores base_allow explicitly, set once at the transition that establishes it and never re-derived from set contents. Verified against a live redis 8.0.5 oracle: every row of the escalation/demotion probe now matches, with -@all (deny) and +@all (allow) as controls firing in both directions.

Not covered by this change and tracked separately: ACL SETUSER still accepts unknown rules and replies +OK, so nocommands, allkeys, allcommands, allchannels, clearselectors and sanitize-payload are silent no-ops.

  • ACL ~pattern now applies to MQ, and is no longer silently skipped for FT.* and CDC.READ (moon#927). ACL key enforcement derives key positions from COMMAND_META, where first_key: 0 meant two different things: "this command names no keyspace key" and "nobody filled the key positions in". MQ declared the second while meaning the first, so a user limited to ~cache:* was correctly refused HGET secret:doc1 and then created, filled and drained secretq through MQ CREATE/PUSH/POP — a real keyspace key of type stream, outside its pattern. The same user read secret:doc1 and its BM25 score out of FT.SEARCH, and the values of every key in the instance out of CDC.READ <wal_dir> 0.

Fixed in three parts. MQ's queue key is now enumerated per subcommand (CREATE, PUSH, POP, ACK, DLQLEN, TRIGGER, PUBLISH; an unrecognised subcommand fails closed). The registry now carries KeySpecClass, so every first_key: 0 entry must be classified as keyless-by-design, movable-key, or pattern-unscopable — an unclassified one is DENIED at runtime and refused by a registration test, which is what makes the ambiguity unrepresentable rather than merely commented. And the FT./GRAPH. skip in acl::keyspec, whose comment deferred to an issue that never existed, is now a decision that is written down: GRAPH.*, WS and TEMPORAL.* address namespaces with no keyspace reach and stay outside ~pattern (gated by command/category permissions, as pub/sub channels are gated by &pattern), while FT.* and CDC.READ hand back keyspace data and are refused.

Upgrade note. Only users with a key pattern OTHER than ~* are affected; unrestricted and ~* users short-circuit before key extraction and are unchanged. Such a user now needs ~<queue-pattern> to use MQ, and can no longer run FT.* or CDC.READ at all. Per-index scoping for FT.* — deciding whether the caller's patterns cover the index's PREFIX — is filed as the follow-up; grant ~* in the meantime if a restricted user must search.

Performance

  • INCR/INCRBY/DECR/DECRBY mutate the stored integer in place (moon#942). Adding one to a hot counter cost three independent DashTable probes — two in Database::get (the second is a documented NLL re-probe, kv_ops.rs:29/:36) and a third in Database::set — plus a CompactKey::from(key) re-copy of the key bytes, a whole new Entry, and two entry_overhead recomputations. Redis's incrDecrCommand does one lookupKeyWrite and rewrites o->ptr. Database::incr_hot_string_in_place (src/storage/db/incr.rs) is the equivalent: one get_mut probe, one CompactValue assignment, and nothing re-hashed or rebuilt.

Counted in the disassembly of the release binary (aarch64-apple-darwin, fat LTO, strip=false), by DashTable entry points reached on the live path — each entry point computes exactly one hash_key:

DashTable entry points key hashes
Database::get (live path) 2 x DashTable::get 2
Database::set 1 x insert_or_update 1
pre-#942 INCR hot path 3 3
incr_hot_string_in_place 1 x DashTable::get_mut 1

That also settles the open question about Database::get's NLL re-probe (kv_ops.rs:29/:36): LLVM does not eliminate it. It tail-merges the live-path re-probe with the post-cold-promotion probe into one branch target, but the live path still executes bl DashTable::get and then tail- calls DashTable::get a second time. DashTable::get is not inlined into Database::get even with hash_key fully inlinable, so no CSE is possible across the two. Fixing that re-probe remains open (it needs polonius, a RawEntry-style DashTable API, or unsafe); INCR no longer pays it because its hot path does not call Database::get at all.

No throughput number is claimed. Benchmarks are Linux-only and this change has not been run on the instrument; the probe counts above are static instruction-level facts, not a wall-clock measurement.

The risk in a fast path around Database::set is the side effects it quietly stops doing, so all eleven are enumerated in the module docs with a per-item decision, and each preserved one has a named test. The two cases the fast path refuses — absent key, TTL-expired key — fall back to the original get + set pair unchanged, because those are exactly where cold-tier promotion, in-flight-spill rehydration and lazy-expiry bookkeeping live; fabricating a 0 for a spilled counter would have been a silent data loss. Error replies (WRONGTYPE, non-integer, overflow) leave the keyspace bit-for-bit unchanged — no WATCH-version bump (moon#926/#940), no dirty-counter charge, no ledger movement, no keyspace notification.

The trade is stated in the module docs rather than left to be discovered: a refused call has already spent its probe, so INCR on an absent key now costs 5 probes where it cost 4. A counter is created once and incremented many times.

21 mutation-injected defects — one per preserved side effect, in both directions — were each confirmed to turn a named test red before the suite was trusted (11 against the unit suite, 10 against the live-server one). Two guards did not catch their mutation on the first pass and were rewritten until they did.

Fixed

  • scripts/test-consistency.sh: the WATCH/CAS rows no longer race moon's shard dispatch (moon#953). watch_cas_outcome and watch_cas_aba_outcome pipelined WATCH with MULTI/SET and then wrote from a second connection without ever reading WATCH's +OK, so on a thread-per-core server the interloper's write could reach the key's shard before the watch was registered there — and EXEC committed. Measured at --shards 4 against a redis 8.6.1 oracle: 2/10 wrong at delay 0, 0/10 once the client waits. Redis passes the pipelined form only by being single-threaded, so the row asserted something stronger than the contract its own comment documents. The fix reads WATCH's reply before MULTI is sent; an ECHO barrier cannot be used there, because after MULTI every command replies +QUEUED and the ECHO never echoes (exec_abort_reply_type already drained its acks and is the in-file precedent). Proved non-vacuous by mutation: swapping the setup command for PING — same round trips, nothing armed — flips the conflicting case from abort to commit, and the outcome still varies with conflict itself, 10/10 unanimous per cell on both engines. Failure sentinels now interpolate the port, so a double failure can no longer pass assert_eq vacuously.
  • SADD under-reported its reply when one batch crossed set-max-intset-entries (moon#944). The intset push loop breaks the instant the ceiling is crossed, and the upgrade path then re-inserted the whole argv into the new IndexSet while DISCARDING insert's "was this member new" bool — so added stopped at the crossing and every member positioned after it was stored but never counted. Measured against redis 7.4.0 with set-max-intset-entries 512: SADD on a 510-member intset with a 24-member batch adding 22 new members replied 3 where redis replied 22. The data was never wrong — SCARD and SMEMBERS agreed all along — which is why no harness row caught it: the reply is the only thing that diverged, and clients build dedup accounting and "did I win the insert" logic on exactly that number.

The upgrade path now counts IndexSet::insert's bool, and walks only the unabsorbed tail rather than the whole argv. The prefix is provably already in the set: Intset::to_set_value renders each value with its decimal spelling, and every value entered the intset through canonical_i64, so that rendering is the caller's exact bytes (moon#795) — re-walking it was an O(batch) no-op that also spent one Bytes clone per member on a src/command/ path, which CLAUDE.md bans. No unit test covered this because every existing encoding row crosses a threshold with a batch of one, and a one-member batch has no tail past the crossing. New rows in scripts/test-consistency.sh compare the reply against the real redis oracle for a straddle by twenty, a straddle by one, and a control wholly below the ceiling; the unit tests assert the reply against the SCARD delta rather than a hardcoded count, so neither the buggy answer nor an over-counting fix can pass them.

  • HINCRBY and HSETNX no longer flatten a small hash (moon#897). Both reached for Database::get_or_create_hash, whose contract is an EAGER upgrade to the full HashMap — so a single HINCRBY on a three-field hash promoted it from listpack to hashtable, and because nothing demotes (moon#832) it stayed there for the key's lifetime. Counter hashes, which are touched by HINCRBY in a loop, never got the compact encoding at all. HDEL was the one hash secondary write that already did the right thing; both commands now take the same in-place listpack route it does, gated by the single EncodingLimits authority (moon#896) so a hash that genuinely exceeds hash-max-listpack-entries or hash-max-listpack-value still promotes. Measured against redis 8.6.1 with the policies aligned, one shard: 24 divergences before, 0 after — including field ORDER on HGETALL/HKEYS, which only differed because the container had been flattened into an unordered map. Reply values, error strings, HSETNX's no-op-on-existing contract and the per-field TTL sidecar are unchanged; the listpack integer encoding still refuses non-canonical spellings, so +5 and 007 survive byte-for-byte (moon#795).
  • LSET, LPOP and RPOP no longer flatten a small list's compact encoding (moon#897). A three-element list built by RPUSH reported listpack and then linkedlist after ONE LSET, LPOP or RPOP, where redis 8.6.1 reports listpack. Nothing demotes (moon#832), so a small work queue lost the compact form on its first pop — LPOP is the queue primitive — and never got it back. Two mechanisms were at fault and both are gone from the compact path: get_or_create_list, whose ListKind::upgrade materialises the VecDeque unconditionally, and a get_list (= get_promoted) EXISTENCE probe that the pops ran purely to decide "does the key exist?" — so the miss check flattened the list before the pop had even started. All three now mutate the listpack in place, routed on a &self probe, and promote only when the ONE EncodingLimits authority (moon#896) says the container no longer fits: for LSET, that is a replacement element past the value threshold, the one list secondary write that can legitimately cross one. Replies (including LPOP's optional count, its *-1-vs-*0 miss, and RPOP's back-to-front order), the ERR index out of range / ERR no such key / count-range error strings and delete-when-empty are unchanged, verified row by row against a live redis oracle at, below and one past the threshold.

  • LSET on a missing key no longer creates it (moon#830). LSET ghost 0 v replied ERR no such key and left an empty list behind — charged to the memory ledger, visible to DBSIZE, and never propagated to the AOF or a replica, since propagation is gated on the reply not being an error. The cause was the same eager get_or_create_list call moon#897 removes: the create half fired before the error could be returned. The non-creating &self router answers "no such key" without it. Fixed as a consequence of moon#897, not as a separate change; EXISTS, TYPE and DBSIZE now match redis 8.6.1.

  • SREM no longer flattens a small set's compact encoding (moon#897). A three-member set built by SADD reported listpack (or intset) and then hashtable after ONE SREM of one member, where redis 8.6.1 reports listpack/intset. Nothing demotes (moon#832), so the promotion was permanent for the key's lifetime — a session set lost the compact form on its first update and never got it back. SREM now removes in place from BOTH compact forms (they are different code paths and both were wrong), routing on a &self probe that cannot itself rewrite the value, and consulting the ONE EncodingLimits authority (moon#896) afterwards so a container that genuinely exceeds the policy still promotes. The empty-set cleanup moved off get_set (= get_promoted), which meant the emptiness PROBE was a flattener in its own right. Replies, delete-when-empty, WRONGTYPE and byte transparency for numeric-looking members (+5, 000000012345, -0 — moon#795/#903) are unchanged, verified row by row against a live redis oracle at, below and one past every set threshold.

  • ZINCRBY and ZREM no longer flatten a small sorted set (moon#897). Both reached for the eager get_or_create_sorted_set, which upgrades a SortedSetListpack on ACCESS — so one ZINCRBY or ZREM turned a three-member zset into a skiplist, and because nothing demotes (moon#832) it stayed that way for the key's lifetime. Measured against redis 8.6.1 on one shard: zadd z 1 a 2 b 3 c was listpack on both, and one ZINCRBY z 5 b left moon on skiplist where redis stays listpack. Since moon#787 measured the full zset form at 20.5× Redis's, that made the compact encoding reachable only for a key nobody ever updates — which is the opposite of the leaderboard and counter workloads the encoding exists for. Both commands now mutate the listpack in place and consult the one EncodingLimits authority (moon#896) for the entry gate and the upgrade check, so a zset that genuinely crosses zset-max-listpack-entries (128) or zset-max-listpack-value (64) still promotes, and one already promoted is never demoted. Replies, score rendering, inf/-inf handling and the delete-on-empty rule are byte-identical to before: an A/B of the two binaries over 56 scripted cases diverges on OBJECT ENCODING and nothing else. A ZINCRBY whose result would be NaN (inf + -inf) is deliberately routed to the B+tree arm — render_score(NaN) is not round-trippable through a listpack (the moon#863 shape), and moon's pre-existing reply for that case is unchanged.
  • FLUSHALL ASYNC / FLUSHDB ASYNC cleared one shard and answered +OK (moon#925). extract_primary_key, the function that decides shard routing, keeps a keyless fast-path table — and it had no f arm at all. Both flush commands fell through to "the routing key is args[0]". The BARE forms were correct only by accident of arity: an earlier args.is_empty() guard returned None before the tail ran. Give either command its optional modifier and args[0] is the literal ASYNC/SYNC, which was hashed as though it were a key — is_local went false and coordinate_flush_broadcast, which sits inside the is_local block, never ran. Measured against the unfixed build with 60 keys: 30/60 survived at --shards 2, 43/60 at 4, 53/60 at 8, and the survivors were readable values, not tombstones. Silent in both directions — an operator who believed the keyspace was empty could reuse key names on top of live data. SYNC failed at a different set of shard counts than ASYNC, each passing wherever key_to_shard(<modifier>) happened to land on the connection's own shard, which is why the regression tests sweep --shards 1,2,3,4,5,8 with the bare form as an in-run control.

The class, not just the symptom. Sixteen commands that COMMAND_META records as first_key == 0 were missing from the routing table and would each have hashed a modifier as a key: FLUSHALL, FLUSHDB, KILL, LOLWUT, MODULE, MONITOR, PUBSUB, RANDOMKEY, RESET, ROLE, SHUTDOWN, SLOWLOG, TIME, UNWATCH, VACUUM, and BGREWRITEAOF — whose arm was spelled at length 13 for a 12-byte name and so could never fire. All are now keyless regardless of arity, and a new test walks COMMAND_META so a future keyless command cannot repeat the omission. Only the two flush commands were observably broken; the rest were saved by an interceptor or by taking no arguments, both of which are accidents rather than guarantees. SPUBLISH/SSUBSCRIBE/SUNSUBSCRIBE are deliberately NOT keyless despite first_key == 0: redis hashes their shard channel for cluster slot routing, and in moon the cluster slot router is the only caller they reach — keyless-for-ACL is not keyless-for-slots.

A second class, from the same defect. BGREWRITEAOF's arm was not merely missing, it was dead: (13, b'b') for a twelve-byte name. These dispatch tables hand-write the length and first byte beside the name literal, and nothing checks that the two agree — a mismatch compiles clean, passes every test, and yields an arm that can never fire. A new test sweeps every (len, first_byte) arm in server/conn/shared.rs and command/mod.rs (223 arms, 373 compares) and fails on any arm whose pattern cannot match the name it compares. The class was otherwise clean; the test is there so it stays clean. - WATCH now sees an in-place container mutation (moon#926). The per-entry version a watching transaction re-checks at EXEC moved only when a whole Entry was replaced, so SET and DEL aborted a watcher and every container write — HSET, HDEL, HINCRBY, LPUSH, LPOP, LSET, LINSERT, LREM, LTRIM, SADD, SREM, SPOP, ZADD, ZINCRBY, ZREM, ZPOPMIN, XADD, XDEL, EXPIRE, PEXPIRE, PERSIST, HEXPIRE, HPERSIST, HGETDEL — did not. Database::increment_version had zero production callers: the mechanism was written and unit-tested and nothing was ever wired to it. A canonical CAS loop (WATCH inventory / HGET / MULTI / HSET / EXEC) therefore lost the update under concurrency while EXEC returned the result array, telling the client the compare-and-swap had held.

The bump now lives in the storage accessors that hand out a mutable handle on a stored value, not in the ~60 write handlers — acquiring the &mut IS the bump — so a 61st write command inherits it. Read accessors are deliberately excluded: Database::get_promoted and friends take &mut self to rewrite the value's encoding but return a shared reference, and a bump there would abort a transaction that nothing wrote.

Measured against redis 8.6.1 over 113 two-connection WATCH/MULTI/EXEC rows: 42 divergences before, 7 after. All 41 "moon commits where redis aborts" are gone (the 42nd, SETBIT, was already the other way round). Six of the seven remaining are no-op writes by a competing client (SADD of a member already present, SREM/ZREM/LREM of an absent element, ZADD with the score the member already has, HSETNX on an existing field) where moon now aborts and redis does not — a CAS retry rather than a lost update, recorded and pinned in tests/watch_container_mutation_926.rs. Where the writer already knows for free whether it changed anything (HDEL of an absent field, HPERSIST on a field with no TTL, PERSIST on a key with no TTL, a pop from a missing key) moon matches redis exactly. The seventh, SETBIT writing a bit that already holds that value, predates this change. - A sorted-set READ no longer flattens a listpack zset (moon#928). All fourteen get_sorted_set call sites in command/sorted_set/sorted_set_read.rs reached the zset through an accessor documented as read-only that takes &mut self and returns the FULL (&HashMap<Bytes, f64>, &BPTree) pair — a shape a SortedSetListpack cannot satisfy, so obtaining one ran SortedSetKind::upgrade. That conversion is one-way (nothing downgrades, moon#832), so ONE ZCARD taken on the mutable dispatch path — inside MULTI/EXEC, inside a Lua script, or through try_inline_dispatch — converted a small zset to a skiplist for the rest of its life. moon#853 fixed exactly this for the set and list families and predicted this one in writing; moon#878, which made ZADD produce listpacks, made the prediction live. Measured on the pre-fix binary (--shards 1, macOS host, used_memory ledger — accounting, not throughput): 1000 eight-member zsets went 293,055 -> 4,749,055 bytes (16.21x) after one ZCARD each through MULTI/EXEC; the same reads on a bare connection kept listpack, the negative control that proves the probe measures the dispatch path. That put a ceiling on moon#878's +38.0% ZADD win — a write-only benchmark measuring an encoding the first read destroyed. All seventeen affected handlers — ZSCORE, ZCARD, ZRANK, ZREVRANK, ZSCAN, ZRANGE, ZREVRANGE, ZRANGEBYSCORE, ZREVRANGEBYSCORE, ZCOUNT, ZLEXCOUNT, ZMSCORE, ZRANDMEMBER, ZDIFF, ZUNION, ZINTER, ZINTERCARD — now take their read through the &Database implementation that already backed dispatch_read. A shared borrow cannot reach K::upgrade, so the compiler enforces "reading does not rewrite", and the two implementations of each command collapse into one.

Deliberately not preserved: the mutable path used to reclaim an expired key and promote a cold-tier hit into hot RAM as side effects of a read. Neither is a correctness property — get_sorted_set_ref_if_alive still treats an expired key as absent and still reads the cold tier through — and the hash, list and set families have shipped this trade since moon#853. tests/zset_read_cold_tier_928.rs pins it: every read of a cold-spilled zset answers identically to a hot one holding the same members, and a WRITE still promotes (re-deriving the compact encoding, moon#898).

  • ZREVRANGE/ZRANGE … REV answered the wrong window on a listpack zset. zrange_from_entries' by-rank arm sliced the score-ASCENDING entries by [start..=stop] and then reversed, instead of counting ranks from the high-score end. Only the whole-range form came out right, which is why ZREVRANGE z 0 -1 looked fine: on {a:1, b:2, c:3}, ZREVRANGE z 0 1 answered [b, a] where redis 8.6.1 answers [c, b]. Reachable only through the compact-encoding branch, so it was live on the dispatch_read path before this change and would have spread to the mutable path with it. The index mapping now matches zrange_by_rank's against the B+tree exactly. A 648-row sweep against a live redis-server 8.6.1, rebuilding the fixture before every read so the pre-fix binary cannot self-flatten, goes from 73 divergences to 48 with zero new ones, and moon's own two dispatch paths go from 17 disagreements to 0.

  • --max-wal-size now reaches the WAL overflow ceiling (moon#916). The flag configured CheckpointTrigger but never WalWriterV3, whose bounds setter had no production caller at all — so the P6 check in maybe_force_checkpoint_on_wal_overflow, which compares against wal.max_wal_bytes(), enforced the 256 MiB DEFAULT_MAX_WAL_BYTES on every instance ever run. A live instance launched with --max-wal-size 1gb logged P6 WAL ceiling trigger — … > max 268435456 bytes. The bounds are now a required constructor argument (WalBounds) so a future construction site cannot omit them, and CheckpointTrigger, autovacuum Pass C and the writer all read the one ServerConfig::wal_bounds(). The recycler floor (min_wal_bytes) is unchanged at 48 MiB for any ceiling of 96 MiB or more.

Upgrade note — behaviour change for anyone who set --max-wal-size above 256 MiB. You now get what you asked for: each shard's WAL grows to the configured ceiling (so --shards N × --max-wal-size on disk) before the P6 trigger forces a checkpoint, restart replay is proportionally longer (WAL replay is a measured startup cost, moon#476), and --disk-free-min-pct is approached sooner. An instance on --max-wal-size 1gb goes from a 256 MiB WAL to a 1 GiB one per shard. Lower the flag if the old effective ceiling was what you were sized for.

Two configurations that were inert are now enforced, and one of them is refused. A ceiling below 96 MiB lowers the recycler floor to max / 2 so the regular recycler can still bring the WAL back under it (announced at startup). A ceiling below two WAL segments (--max-wal-size < 2 × --wal-segment-size, e.g. 8mb at the default 16 MiB segment) can never be enforced — the active segment is never recycled — and the server now refuses to start, naming both flags; --check-config reports it.

Performance

  • server/conn (monoio): a single-shard server stops building and polling the cross-shard coordinator's future once per command. try_handle_cross_shard_commands is an async fn taking six arguments whose very first statement is if ctx.num_shards <= 1 { return false } — but its call site was guarded only by !txn_multikey_write, so at --shards 1 every GET, INCR, SADD, LPUSH, HSET and ZADD constructed that future, awaited it, and was told there was nothing to coordinate. The same predicate now sits at the call site; the inner early-out stays, so the site is an optimisation rather than a correctness dependency.

It is deliberately not gated on skip_name_gates. MGET/MSET/UNLINK/EXISTS carry NO_INTERCEPT and are claimed by the gate's own multi-key arm, so the flag would silently route them by their FIRST key and execute the whole command against that one shard's slice. That failure was reproduced deliberately, by mutating the gate to ctx.num_shards > 8: MSET across four shards then loses three of its four values and DBSIZE answers 1 instead of 4.

Above one shard nothing changes — the predicate is the one the callee already applied. No throughput number is claimed: benchmarks are Linux-only and this was written and verified on macOS. tests/cross_shard_gate_shard_count.rs runs every command the gate claims at --shards 1 and --shards 4 with identical assertions, because this seam's defects (moon#507, moon#513, moon#592, moon#937) were all invisible at one shard. - server: the monoio write path stops taking a second exclusive database guard just to discover it has nothing to wake (moon#942). After dispatch, handler_monoio acquired s.databases.write(new_sel_db) — a guard_depth thread-local RMW plus a parking_lot write acquisition — on every successful write, solely to hand a &mut Database to wake_producer. That function's own first line is producer_family(cmd)?, which returns None for everything outside LPUSH/RPUSH/LMOVE/RPOPLPUSH/ZADD/XADD, so for INCR, SADD and HSET — three of the five families moon#942 targets — the guard was taken, nothing happened, and it was dropped. producer_family needs only the command NAME, so the question is now asked before the guard instead of inside it. The set of wake_producer calls is unchanged; only no-op acquisitions are gone, and the guard, when taken, is still taken on the POST-dispatch database index. No throughput number is claimed: the acquisition is uncontended at --shards 1 and this was not measured on a Linux host. The gate is producer_family itself and must never become a hand-written list of command names — a gate narrower than the mapping is a lost wakeup whose visibility depends on which shard owns the key, which is moon#595 exactly. New tests/wakeup_local_write_gate.rs pins BLPOP/LPUSH, XREAD/XADD and BZPOPMIN/ZADD at both 1 and 4 shards, plus an INCR/SADD/HSET control that must wake nothing and must not error; restoring the pre-#595 gate turns it red at 8/8 trials at one shard and 3/8 at four, which is that routing-dependence reproduced. handler_sharded (tokio) was checked and deliberately left alone: its single db_guard already spans dispatch and the wake, so it never paid for the no-op. - storage: the typed accessors probe the DashTable half as many times (moon#942). Every get_or_create* opened with a get (inside drop_if_expired) and a contains_key that answer the same question — "is there a live entry at this key?" — before the get_mut that hands the entry out. get_promoted paid a fourth lookup, a trailing get, purely to re-borrow immutably what the upgrade's &mut already had. get_or_create_set_listpack paid a fifth, because absorb_intset_into_listpack was a &mut self method keyed by &[u8] and took its own get_mut — on every SADD, including the overwhelming majority where the key is not an intset and the function does nothing. Each of those is a full xxh64 over the key plus a segment scan; nothing cached a probe anywhere.

One HotState classification now decides the whole preamble, the create arm reads promote_cold_if_present's own return value instead of re-issuing contains_key, and the intset edge runs on the handle the accessor was about to take anyway. Measured with a new cfg(test) DashTable key-lookup counter, per accessor, on a hit and on a miss:

accessor before (hit / miss) after (hit / miss)
get_or_create 3 / 6 2 / 4
get_mut_if_present 3 / 4 2 / 3
get_promoted 4 / 5 2 / 3
get_or_create_intset / _hash_listpack / _list_listpack / _zset_listpack 3 / 6 2 / 4
get_or_create_set_listpack 4 / 7 2 / 4

SADD end to end does not fit that table's hit/miss axis — its three regimes are an absent key, a listpack-encoded key and a hashtable-encoded one (122 of 200 keys at the benchmark's own p=64 point):

SADD regime before after
absent key (creates the container) 7 4
listpack-encoded key 4 2
hashtable-encoded key 7 4

Redis reaches the same key with one dictFind. The residual 4 on SADD's hashtable regime is two accessors' worth: get_or_create_set_listpack answers Ok(None) and get_or_create_set then repeats the skeleton. Collapsing that needs a change in src/command/set/.

No throughput number is claimed, and none was measured. That is deliberate, not an omission: moon#789 measured the previous probe-count reduction in this repo at +11% on aarch64 and −17% on x86_64 — the two architectures disagreed in sign — and b3083c5a found the wall-clock net guarding it was timing a page-fault artifact rather than the optimisation. Fewer probes is a structural fact; whether it is faster is a question only a Linux benchmark host may answer, and this change must be A/B'd on both arches before any performance claim is attached to it. What is verified beyond the counter is that the codegen moved: on aarch64 the emitted bodies of get_or_create_hash_listpack, _list_listpack, _zset_listpack, _intset and _set_listpack shrink by 45–48% (884→460, 880→456, 884→460, 984→556, 964→528 bytes). Code size is not speed either; it is evidence the compiler saw the change.

Behaviour is unchanged and pinned by tests written against the pre-change baseline: an expired key still reads as ABSENT (not WRONGTYPE) through every accessor and still leaves no expiry-index entry behind (moon#541); a cold-spilled value is still promoted before an empty container is fabricated over it, on the expired arm as well as the absent one (moon#459); a live key of the wrong type still errors; the WATCH version still bumps exactly once per mutable handle (moon#926 — moon#940, the WRONGTYPE path stamping too, is neither fixed nor worsened, and stamp_mutation did not move); and used_memory still agrees with an independent whole-keyspace recount.

That last one was first written as assert_eq!(run(), run()) over a pure deterministic closure, which is a tautology — a build that moved a charge across a branch would still agree with itself — and an adversarial review caught it. It now compares the incrementally maintained ledger against a fresh walk summing entry_overhead per entry, and a second test pins the one ledger write this change actually moved (the moon#899 intset→listpack swing, which used to be applied inside absorb_intset_into_listpack before stamp_mutation and is now applied by the accessor after it). Both were verified to FAIL against deliberately mutated builds — charge dropped, swing dropped — rather than assumed to be capable of failing.

  • storage: the listpack write path stops walking twice and stops allocating per entry it walks past (moon#799). get_at, remove_at and replace_at reached their index with the OWNED decoder, which copies every string entry into a fresh Vec — and remove_at/replace_at then threw that entry away. HSET onto an existing field also scanned the listpack TWICE: once allocation-free to find the field, then again from the head to reach the position the first scan had already arrived at. On a 128-field hash that was ~256 entries walked and ~260 malloc/free pairs to change one value, inside src/command/, which forbids allocating at all. The walk now uses the borrowed decoder (seek_to), and HSET/HMSET/HDEL/HGETDEL/ HGET/ZSCORE route through one-scan pair operations (replace_pair_value, remove_pair, take_pair_value, pair_value). Listpack::find — SISMEMBER against a set listpack — moved to the borrowed comparison too. Allocations per call, counted by a GlobalAlloc wrapper on an 8- vs 128-pair listpack: replace_at 20/260 -> 4/4, remove_at 16/256 -> 0/0, get_at 16/256 -> 1/1. Ratios only, macOS aarch64, --profile release-fast, interleaved legs — NOT Linux numbers. benches/listpack_write_path at 128 fields: HSET-existing 5.3x, HGET 5.5x, HDEL 8.8x, SISMEMBER 5.6x. End to end (redis-benchmark -t hset -P 16 -r N, --shards 1): 128 fields 3.39x, 64 fields 2.53x, 32 fields 2.10x. The internal control is the point — moon's listpack was 5.80x slower than moon's own hashtable on this host and is now 1.67x; the promoted hashtable rows are the negative control and did not move (0.97x, 1.01x).

  • text/vector: one tokenization per HASH field, shared by the BM25 text plane and the vector payload index (moon#885). An HSET into a key that matches both a text index and a vector index ran the whole analysis pipeline twice over the same bytes — NFKD, combining-mark strip, lowercase, UAX#29 segmentation and English Snowball — once in AnalyzerPipeline::tokenize_with_positions and again in filter::TextIndex::tokenize from index_payload_field. Measured on Linux/aarch64 as 48.7% of a no-.tpost boot reconcile. The two consumers are not interchangeable and were not made so: the text plane drops the 33 default stop words, honours NOSTEM per field and keeps word positions; the payload index keeps stop words (an @f:the filter still hits), always stems and has no positions. AnalyzedText now holds the normalized text and word spans once per value, with a per-word English stem filled by whichever consumer needs it first; an AnalysisCache scoped to one auto_index_hset call hands the second consumer the first one's work. A text-only NOSTEM field still never stems; a text-only stemmed field still never stems its stop words. Scope: this removes duplication only where both planes run — steady-state ingest, and recovery with no .tpost (first boot, corrupt or version-skewed file). With .tpost present the text plane is skipped on the reconcile path, so there is nothing to share and that boot time is unchanged; the payload index's own pass there is the only thing rebuilding it, because PayloadIndex has no durable form. Pinned by src/text/shared_analysis_tests.rs, including a thread-local segment_passes() probe asserting one pass per field value.

  • storage: the cold-index rebuild builds its ordered map in one bulk load instead of one tree descent per key. ColdIndex::rebuild_from_manifest_per_db fed every recovered key into a BTreeMap one at a time, in random hash order. Profiled over a real 466,912-entry / 1,816-file / 114 MiB spill corpus (generated by driving a live --shards 1 server past maxmemory with allkeys-lru), that insert loop was 75% of the whole function's wall time — file I/O was 13%, the per-page 4 KiB copy 2%, the CRC32C verify 3% and full entry decode 8%. The scan now collects ((scan_h48(key), key), ColdLocation) pairs into a flat per-db vector and hands them to BTreeMap's FromIterator, which stable-sorts once and bulk-loads bottom-up; file_refs, resident_bytes and pending_unlink are derived from the finished map. Measured, interleaved legs, fresh corpus copy per leg:

  • CPU only (in-process, warm page cache, median of 7): 0.200s -> 0.113s, 1.77x.
  • Whole restart phase at scale — a 1,629,070-entry / 6,490-file / 430 MiB corpus, cold cache, bracketed between the manifest recovered and rebuilt cold index log lines, median of 3 pairs: 3.10s -> 2.44s, 1.27x. Smaller than the CPU figure because at that size file I/O, not the map, is the larger term; every one of the three pairs favoured the new path.
  • Peak RSS over the same restart, /usr/bin/time -l, median of 3: 552 MiB -> 494 MiB (-10.5%). The transient pair vector is more than paid for by bulk_build's densely packed B-tree nodes; random insertion leaves them part-full.
  • Byte-heavy shape (2.1 GiB / 183,741 entries, file scan is 84% of the phase): 0.411s -> 0.369s.

The rebuilt index is byte-identical — same keys, same file_id/page_idx/slot_idx/ttl_ms/value_type, same resident_bytes, same referenced-file count — because FromIterator's stable sort keeps the last of every equal-key run, exactly as repeated insert did. Red/green: kv_spill::tests::test_rebuild_duplicate_key_across_active_files_keeps_the_later_file (a key live in two still-Active heap files must resolve to the later one; fails with Some(301) if the dedup is flipped to first-wins).

Fixed

  • persistence: a cold-tier key no longer comes back from a kill -9 with its non-idempotent writes applied twice (moon#902). Under --appendonly yes --disk-offload enable (the default) a spilled key was durable in BOTH planes — the shard manifest held the value, the AOF held the command that produced it — with no cut between them: recovery rebuilt the cold index first, then replayed the AOF, and every get_or_create_* a replayed write goes through promoted the cold copy before mutating it. RPUSH l a b c d e landed on the [a,b,c,d,e] it had already produced; INCRBY doubled, APPEND doubled, HINCRBY/ZINCRBY/BITFIELD INCRBY doubled — and it compounded by one per unclean restart, because the replay-driven eviction re-spilled the doubled value (measured k=3 after two crashes). Idempotent writes (SET, HSET, PFADD, SETRANGE, XADD with a resolved id) were unaffected, which is why it hid. The cut is now in-band in the AOF: every generation opens with MOON.COLDCUT <w> (cold files below w were sealed before the base was cut and are a valid base for every record that follows — the base RDB is hot-only, so they are the only copy), and every published spill batch appends MOON.SPILLED <file_id> key… (AOF only — file ids are shard-local and must never reach a replica). During an AOF-authority replay a cold entry is invisible to the value-giving read paths until its file is cut; the marker also drops the replay-built hot copy, which is exactly the restart-as-cold task #56 wanted. Tombstones (DEL/FLUSH*, moon#257) bypass the gate. A generation written before this fix has no MOON.COLDCUT and replays exactly as before. The same cut also fixes the OTHER direction, found by this fix's rewrite control: a SET issued to a cold key after a BGREWRITEAOF was silently dropped at the next crash (12/24 probes on the reference run) because the old end-of-replay demote treated the stale cold copy as authoritative. Reference run (macOS, --shards 1, release-fast): 96 corrupted probes in cycle 1 → 0; 112 in cycle 3 → 0. New crash suite tests/cold_tier_aof_double_apply_902.rs; the moon#898 suite's list tolerance is removed.
  • storage: the compact-encoding thresholds have ONE authority (moon#896). Three sites decided whether a container stays in its listpack/intset form — a write command's entry gate, its post-push upgrade check, and the restart-side re-derivation in compact_after_decode — each with its own copy of the arithmetic, and two of them disagreed about the UNIT: the entry gate compared args.len() - 1 (listpack ENTRIES, two per hash field or zset member) against the same 128 the upgrade check applied to lp.len() / 2 (ITEMS). All three now consult storage::encoding_limits::EncodingLimits, a Copy snapshot on Database, through a predicate that takes a Shape rather than a raw count, so the entries-per-item factor lives in one place and a caller cannot pass the wrong unit. The POLICY thresholds (*-max-listpack-*, set-max-intset-entries) and the SAFETY ceiling (LISTPACK_SAFE_BATCH_ENTRIES, which keeps a batch far below the listpack header's u16 range and closes SADD's O(n^2) scan window, moon#865) are now distinct named things; the raw constants are private to the module and scripts/audit-encoding-limits.sh keeps a hand-rolled threshold or unit conversion out of every consultation site. The authority carries moon's CURRENT values, and the refactor is proven behaviour-neutral: the encoding matrix (src/command/encoding_matrix.golden — every type at every size around the thresholds, built bulk, incrementally and via the restart path) is byte-identical to the one captured on the pre-authority tree, wrong cells included. The unit fix then lands on top of it: the HSET/HMSET and ZADD entry gates now convert their argv length through Shape::items_in, so a bulk HSET of 65..128 fields or ZADD of 65..128 pairs stays a listpack, as the same container built one item at a time always did and as redis 8.6.1 does; the 20 flipped golden cells are exactly that range at 8- and 64-byte elements, and a new agreement test asserts, for every shape and size, that the entry gate, the upgrade check, the restart path and the authority's own predicate reach one verdict. scripts/test-consistency.sh and scripts/test-commands.sh gain the bulk 64/65 (hash, zset) and 128/129 (all four types) boundary rows against a live redis oracle.

  • set: SADD of a string to a small intset reaches a listpack (moon#899). Redis's set state machine has three forward edges — intset -> listpack, listpack -> hashtable, intset -> hashtable — and moon had no intset -> listpack: SADD s 1 2 3 then SADD s abc was hashtable here and listpack on redis 8.6.1, and because nothing demotes (moon#832) the hashtable was permanent. get_or_create_set_listpack now takes the edge when the authority says the result fits — Redis's rule, intsetLen < set-max-listpack-entries and every member within set-max-listpack-value, with the intset's widest rendered integer measured too — and is self-accounting for the cost swing (ledger test). Measured at the boundary against the oracle: 127 ints + a string is a listpack on both, 128 + a string a hashtable on both; a fixture at 200 ints cannot see this defect. Byte transparency across the edge is guarded explicitly (moon#795): every intset value arrived through canonical_i64, so its itoa rendering is the client's exact bytes, and +5, 000000012345, -0 added in the same command stay distinct members. Found on the way: the restart-side re-derivation (compact_after_decode) decided "all integers" with a bare parse::<i64>(), so a set holding +5 reloaded as an intset and read back 5 — the moon#795 loss moon#802 had closed on the live path. It now uses canonical_i64; such a set reloads as a listpack, byte-exact.

  • replication/persistence: a full resync no longer flattens every compact encoding (moon#863). The Redis-wire RDB codec has no type tag for a compact encoding — SetIntset, SetListpack and the full Set all go out as RDB_TYPE_SET — so redis_rdb::read_rdb_entry rebuilt every container in its FULL form, and that codec is the one a FULLRESYNC payload travels through (replication::apply::load_snapshot -> load_rdb). A replica therefore held the master's data in the expensive form until each key was rewritten, and OBJECT ENCODING on the replica disagreed with the master's for all five affected types (set intset, set listpack, hash listpack, list listpack, zset listpack). moon#840 fixed the equivalent gap on the RESTART path only; this closes the replication one, reusing the same helper and the same thresholds (value_codec::compact_after_decode), so a value that arrives over the wire lands in exactly the encoding it would have had if the same commands had been replayed live. RESTORE (which shares the codec) benefits identically. Nothing about the bytes on the wire changes, so moon stays readable by — and able to read — real Redis in both directions; containers past the thresholds still load in their full form, complete. The RDB_TYPE_ZSET_2 arm now builds the canonical SortedSetBPTree instead of the legacy SortedSet, which this arm was the last producer of: the legacy form cost a SortedSetKind::upgrade on every zset command's first touch and billed used_memory with the retired per-member model. A NaN zset score (8 raw bytes on the wire; no writer produces one, ZADD rejects it) is now rejected rather than rendered into a listpack as the text NaN, which parse_score refuses and every later read would answer 0.0 for. Measured on macOS 15.7.9 / M4 Pro, release-fast (not a publishable benchmark profile), 3 interleaved A/B rounds over 100,000 keys across the five types: replica used_memory 46,628,615 B -> 24,228,615 B (-48.0%), which is now exactly the master's own 24,228,615 B; replica RSS 127.4 MB -> 92.0 MB (-27.7%). The master is unchanged.
  • storage: promoting a value back out of the cold tier no longer flattens its compact encoding (moon#898). The spill body format is canonical per LOGICAL type — Set, SetListpack and SetIntset all encode as ValueType::Set, and the body carries elements, never an encoding tag — so the cold decoder could only rebuild the FULL form, and every container that came back from disk landed as hashtable/linkedlist/skiplist. Nothing demotes (moon#832), so that was permanent for the key's lifetime, and --disk-offload is enabled by DEFAULT: the keys most likely to travel this path are exactly the long-lived, rarely-touched small collections the compact encodings exist for. moon#840 closed the same gap on the restart path and moon#863 on the replication one; this is the third and last instance of that shape. The re-derivation runs at the two boundaries where a cold value re-enters the HOT keyspace — Database::promote_cold_outcome (the on-disk plane) and eviction::rehydrate_spill_payload (the in-flight spill plane) — and deliberately NOT one layer down in kv_serde::deserialize_collection, which the NON-promoting cold read-through shares: ValueKind::classify_cold accepts only the full forms, and a compacted value there answers WRONGTYPE for a valid set (measured, not feared — there is a guard). Same helper and same thresholds as the other two paths, so a promoted value lands in exactly the encoding a live write would have produced; containers past the thresholds still come back full and complete, and nothing about the bytes on disk changes. Measured on macOS 15.7.9 / Apple M4 Pro, release-fast (not a publishable benchmark profile; a macOS number is not a Linux number), 3 interleaved A/B rounds over 10,000 5-field hashes with ~3,700-6,400 of them cold per round: used_memory per cold-promoted key 744.0 B -> 184.0 B (-75.3%), identical to the byte in all three rounds. Wire-level, 120 probes across all five affected encodings: 120/120 came back compact, against 0-6 per type before. Because it shares compact_after_decode with the restart and replication paths, this path also inherits moon#903: that helper's intset arm decided "all integers" with a bare parse::<i64>() rather than numeric::canonical_i64, admitting +5, 000000012345 and -0 — which an intset both rewrites (moon#795) and collapses into fewer members. A byte-transparency guard over exactly those spellings ships here, and the cold path is correct only with moon#903 fixed.
  • ci: the merge bar can now see process-global state leaks between tests, and lints console test code (moon#904, moon#905). Two proven holes in moon's own gates. (1) Every configured test gate — both VM suites and both --native suites in scripts/ci-local.sh, plus the hosted Check and check-monoio jobs — runs cargo nextest, which forks a process per test. A bug whose mechanism is process-global state leaking from one test into another therefore never gets a second test to reach, and moon#856 has been read as an environment-specific flake for months because of it. Measured on the macOS host at 65fa069e, same target dir and same test binary: cargo nextest run --lib 5338/5338 green, cargo test --lib 5337 passed / 1 failed. nextest is KEPT for the other ~6000 tests; a single-process scripts/libtest-singleproc-gate.sh stage now runs beside it (36s on macOS, ~70s on Linux). It ASSERTS four things — the suite reached its summary, it ran at least MIN_LIBTEST_TESTS, there is at most one failure, and that failure is moon#856 by name — and REPORTS the passed count against a platform-keyed baseline without gating on it, because Linux-only cfg code compiles ~21 Linux-only tests that macOS never runs and a single hardcoded figure would read as a phantom regression on the other platform. The gate ships a --self-test (13 synthetic transcripts) that ci-local and the hosted Lint job run first, so its waiver is proven able to refuse before any run is trusted. (2) cargo clippy --features console --all-targets failed on main: two deny-by-default clippy::approx_constant errors in console_gateway.rs's own #[cfg(test)] block. The lint needs --all-targets AND console, and no gate had both — ci-local's console leg omitted --all-targets, every --all-targets leg omitted console. The intersection was empty for as long as the module has existed. The test's 3.14 becomes 3.5 (the assertion never cared which float), and both the ci-local and check-console legs gain --all-targets.

  • text: the TMX3 TAG/NUMERIC sidecar block is now proven compatible in both directions, and the format no longer depends on the build's feature set (#880). The block itself landed with #879; this finishes it. TagFieldDef and NumericFieldDef lose their text-index gate (they are plain data with no dependencies), so text-indexes.meta is written and read identically on every build and the serializer carries no #[cfg] — the dead-code hazard on the tokio leg that failed #879's first CI run cannot recur. The reader is split into deserialize_body (behaviourally the pre-block reader: reads count indexes and returns) and deserialize_ext_block, so compatibility is unit-tested against the real code path rather than argued: an old reader given a new file restores every TEXT field and parks exactly on the block; a new reader given a v2 or v1 file without the block loads with empty TAG/NUMERIC; a torn block fails closed at every cut point; foreign trailing bytes are ignored. Measured over the wire as well, with the pre-#879 binary as the old side: FT.SEARCH ix hello finds doc:1 after a cross-version restart in both directions, and only the direction where the file never carried the fields answers unknown_field for @cat:{a}. New fuzz target text_index_meta covers the decoder (both fuzz.yml matrices), and tests/ft_text_meta_tag_numeric_restart.rs is the SIGKILL restart round-trip that fails on a pre-#879 binary.

  • storage: one large RPUSH/LPUSH/HSET silently dropped 65,536 elements, and the server acknowledged the write (#865). The listpack header counts its elements in a u16 and update_header advanced that count with wrapping_add, while the command layer only compared the count against LISTPACK_MAX_ENTRIES after the whole argument list had been pushed. Nothing between the client and the wrap enforced a bound. Measured over the wire against a v0.8.9 binary, with redis-server as the oracle:

    RPUSH lpbig e0000000 .. e0069999 moon replied 4464 Redis 70000 LLEN lpbig moon replied 4464 Redis 70000 HSET hbig (40,000 fields) moon replied 40000 Redis 40000 HLEN hbig moon replied 7232 Redis 40000

70000 - 65536 = 4464; 80000 - 65536 = 14464, halved for the field count = 7232. The bytes were still in the buffer -- only the header count wrapped -- so this was not a crash or a corrupt file but something worse: an acknowledged write whose length reply is neither the truth nor an error, and every later read iterates the wrapped count. HSET returned the correct count and then reported a different one from HLEN.

The batch is now bounded before the listpack path is entered (listpack_batch_fits, src/storage/db/mod.rs): a call carrying more than LISTPACK_MAX_ENTRIES new entries skips the listpack encoding entirely, exactly as an oversized element already did. Such a container was going to be upgraded on the very next line anyway, so this costs nothing and caps the count at 128 existing + 128 new -- three orders of magnitude below the wrap. It also closes a quadratic window that came with it: the membership scan inside the push loop is linear, so an unbounded batch was O(n^2) inside a single command on a single shard thread (7.1s for 70k members; 0.01s now). As a backstop, update_header saturates instead of wrapping and carries a debug_assert!, so a future caller that forgets the bound corrupts nothing worse than its own count and says so in a debug build.

Five lib tests plus two scripts/test-consistency.sh rows, all verified to fail against the pre-fix binary.

  • persistence: the WAL ceiling-trigger no longer re-reads every sealed segment every ~10 s to free nothing (#870). With the WAL past --max-wal-size, at least one completed checkpoint, and a workspace/MQ/temporal record in each sealed segment, the overflow pass content-scanned every immutable sealed file on every pass and the plane guard (correctly) refused them all — 7.96 TB read in 6.8 days on a live instance writing 30 KB/s, growing with the WAL. WalWriterV3 now memoizes the plane-guard verdict per sealed segment (in memory; a restart costs one scan per segment, once), and the --wal-max-checkpoint-lag-ms guard backs off ×2 per pass that frees nothing, capped at 64× (10 min 40 s at defaults), resetting the moment any recycler frees a segment. The fail-closed plane guard is unchanged and now pinned by a test on the memoized path. The issue's "break at the first blocker" was assessed and rejected: a pure-KV segment sealed after a plane-blocked one is freeable today, and break would pin the whole WAL behind the first MQ record. New # Reclamation fields reclamation_wal_plane_scan_total / reclamation_wal_plane_scan_bytes_total make the loop, and its absence, visible on a running server. Red/green in src/persistence/wal_v3/segment.rs and src/shard/persistence_tick.rs (test_870_*), asserting on scan counts, never wall time.
  • persistence: AOF auto-rewrite seeds its base size from the manifest's base RDB, not the whole appendonlydir (#868). auto_rewrite::init() recorded base + uncompacted incr as the growth baseline, so every restart raised the rewrite trigger by however much incr had accumulated — for good. On a host that restarts more often than the AOF doubles the monitor never fired: a live instance reported aof_base_size 3.41 GB against a 1.08 GB base file and sat at 4.8 GB on disk for 17 days. Boot now seeds base from the base file(s) the committed moon.aof.manifest names (one per shard in the PerShard layout) plus the legacy flat appendonly.aof, and current from the directory as before; the post-rewrite baseline is unchanged. Trigger percentage, min-size and cadence are untouched.

Added

  • text: text indexes load their postings on restart instead of re-tokenising every hash under a text prefix. A text index had no durable form — text-indexes.meta carried definitions and .tfst (written only by FT.COMPACT) the term dictionaries — so every boot rebuilt every posting from the keyspace: ~0.22 µs per token, half tokenising and half building postings, growing with the corpus forever. Each index now persists its complete state to {persist_dir}/{xxh64(name)}.tpost (postings, dictionaries, FSTs, doc maps, TAG/NUMERIC entries) plus one content checksum per document, and a boot loads the file, then walks the keyspace comparing the stored stamp with one computed over the live hash: equal docs are skipped, changed docs re-indexed, docs whose key is gone removed — so a stale or crash-torn file still yields exactly a rebuild's result. Any missing/corrupt/version-mismatched/schema-mismatched file means "rebuild this index"; nothing partial is ever installed. Encoding runs on the shard thread under a 2 ms-per-second budget and a 1 % duty cycle per index; every syscall (tmp → fsync → rename → dir fsync) runs on one moon-text-persist thread; graceful shutdown flushes and waits. Design and contract: docs/internal/text-postings-persistence.md. Fuzz target text_postings_file. Measured on a generated corpus (macOS, 1,000 indexes, same binary both sides): 49.5k 40-120-word docs, the index phase 3.22 s rebuilt → 0.95 s loaded (0.73 s decode of 93 MB + 0.17 s checksum walk); 19.8k 300-800-word docs, 6.40 s → 1.29 s. 1,100 queries (term, AND, phrase, prefix, stopword-only, TAG, NUMERIC, match-all) return identical keys and scores loaded vs rebuilt; flipping one byte in every file makes every index fall back and answer the same. Also fixed a pre-existing restart data loss: TAG and NUMERIC field definitions were dropped by the meta sidecar (see Fixed below).

Fixed

  • text: TAG and NUMERIC fields no longer vanish on restart (or on a replica). The text-indexes.meta sidecar only ever carried TEXT field definitions, so after every reboot FT.SEARCH ix '@cat:{a}' answered unknown_field and numeric ranges likewise; the replication snapshot streams the same bytes and lost them the same way. The definitions now ride a TMX3 block appended after the v2 body (the file keeps version = 2; an older binary stops at count and keeps restoring the TEXT fields it understands), and every restore path builds the index through one constructor, TextIndex::from_meta.

  • recovery: startup no longer spends minutes rebuilding index definitions it already has, re-loading a keyspace it then wipes, or sitting silent on an accepted socket. A live 2.2M-key instance with 4,191 text + 2,793 vector indexes took 6+ minutes to answer, and its log named four structural wastes, all fixed here: (1) every restored index definition rewrote the whole metadata sidecar (two fsyncs, O(n²) bytes) — TextStore::restore_index and VectorStore::create_index_unsaved register without the write and the sidecar is written once per store after the loop; (2) the boot rescan tested every prefix of every index against every key (O(keys × indexes)) and the HSET path did the same three times per key — both stores now keep a prefix -> index map (util::prefix_map) so the lookup is O(key length), on the boot path and on every steady-state HSET/DEL; (3) when the multi-part AOF is the KV authority, per-shard recovery still loaded the snapshot, applied WAL KV records, ran the "no appendonly.aof found — replayed N records from legacy-mode WAL v3" fallback over the SAME WAL directory and demoted 650k hot shadows, all of which db.clear() then discarded — restore_from_persistence now takes kv_authority_elsewhere and skips every hot-keyspace load while still recovering the cold index, warm segments, FPI repair and last_lsn; (4) warm-segment registration re-read every index's persisted keymap from disk for every segment — read once per index and indexed by key_hash. The per-segment INFO lines moved to DEBUG behind the existing summaries. Readiness: the definition-restore loops now yield to the event loop and the keyspace rescan slices are time-bounded (10 ms) instead of 1024-keys-bounded, so a client during recovery gets +PONG / -LOADING instead of a hung socket, and the reconcile progress line prints the cumulative and the recent keys/s, labelled, plus an ETA from the recent one (the cumulative alone read 21 keys/s on a live instance whose real rate was ~230/s). Measured on a generated 380k-key corpus (1,500 text + 400 vector indexes, 80k indexed hashes, 3 interleaved rounds, macOS): loading:0 at 22.0-25.1 s → 2.1-3.4 s; first +PONG at 17.7-21.9 s → 0.3 s with zero probe timeouts; state after restart identical (DEBUG DIGEST, per-index doc counts, text hits) before and after every round. Red/green: util::prefix_map tests, prefix_map_tracks_* in both stores against the brute-force walk, kv_authority_elsewhere_skips_every_hot_keyspace_load, and the_recent_rate_ignores_a_slow_start_and_drives_the_eta.

  • replication: coordinator local legs now replicate, and stop inflating master_repl_offset (#815). On a multi-shard master the in-process leg of a multi-key write (MSET/MSETNX co-located or scattered slices, DEL/UNLINK, COPY, BITOP at the connection's own shard) reached the AOF via persist_local_leg but never the replication backlog or a live replica, while issue_append_lsn still advanced the offset for it. With --appendonly yes the offset counted bytes no replica could ACK, so WAIT answered :0 forever after the first local leg; with --appendonly no the leg neither counted nor shipped and the replica silently diverged. persist_local_leg now takes the same two-plane contract as the handler's single-key write: record to the backlog + advance the offset synchronously when a replica is attached (AOF LSN 0), else advance through the AOF LSN. The multi-shard SWAPDB leg, which issued an AOF LSN and emitted a replication record, no longer counts its record twice. record_local_write/record_local_write_db and their ctx-free twins now share one implementation (replication::state::record_local_write_db_on). Red/green: tests/replication_local_leg_815.rs (#[ignore]d, real primary + replica, --shards 1 and --shards 4, both AOF modes).

  • bench-compare.sh reported the LPUSH seeding rate under the LRANGE label. redis-benchmark -t lrange_100 emits two requests per second lines — the LPUSH used to populate the list, then the LRANGE — and parse_rps took the first match, so all four LRANGE rows the script has ever printed were LPUSH numbers. The same parser scanned for the first numeric field, which reads the literal score out of zadd z:__rand_int__ 1 m:__rand_int__ and reports 1 rps. Both reproduced against live redis-benchmark 8.x output before and after the fix; the parser now anchors on the position of the words requests per second and takes the last match.
  • storage: used_memory bills the boxed payload block of every fat RedisValue variant. CompactValue::estimate_memory charged size_class(size_of::<RedisValue>()) for the outer Box<RedisValue> and nothing for the payload blocks inside it. Two consequences, both measured:

  • RedisValue::Stream has carried a Box<StreamData> since it was written, and its 112-byte block was never billed — a pre-existing under-count on every stream key, independent of the boxing above.

  • After boxing the fat variants the gap widens to all six: a Set key's ledger would fall 80 B while its RSS did not move at all, and a SortedSetBPTree key's would fall 128 B while its RSS rose 48 B. That is used_memory moving away from RSS — the exact failure mode moon#788 exists to prevent, because a ledger that drifts makes --maxmemory unable to bind.

estimate_memory now adds boxed_payload_bytes(), one size_class term per boxed field, derived with size_of_val of the pointee so it follows the declared field type rather than a copied constant.

Three encoding-upgrade sites in the command layer (HSET/HSETNX listpack -> Hash, SADD intset -> Set) re-derived the post-upgrade cost by hand and so missed the new block; hash_and_list_mutations_keep_the_ledger_exact caught the 48 B drift between the running ledger and a full recompute. All three now call the same boxed_payload_block the value itself uses.

every_boxed_payload_block_is_billed covers all six variants and fails on each of them without the fix.

  • storage: the HashWithTtl promotion/downgrade ledger balances. The ttls sidecar was allocated by promote_to_hash_with_ttl on the first HEXPIRE with no charge_memory, while remove_hot credits entry_overhead recomputed from the value as it stands at DEL time — which does include it. HSET h f v; HEXPIRE h 100 FIELDS 1 f; DEL h in a loop therefore walked used_memory down by 120 B every iteration (248 B from a HashListpack key), unbounded and reachable from any unprivileged connection. credit_memory saturates at 0, and once there --maxmemory can never bind again; recalculate_memory is a load-time healer only, so nothing repairs it at runtime. Same class as moon#788, moon#810 and moon#814.

Boxing ttls (above) contributes 48 B of that; the other 72 B is the sidecar entry itself, which estimate_memory has always billed and no writer has ever charged. Both are fixed together, in both directions: promote_to_hash_with_ttl now returns the signed delta it caused — O(1) and exact from a plain Hash (only the empty sidecar box is new), a before/after snapshot from a HashListpack, whose conversion changes the cost model outright and used to lose 320 B. Every site that drops a sidecar entry (hash_persist_field, hash_clear_field_ttls, hash_delete_field, hash_get_and_delete_field, the past-expiry short-circuit, and the active reap_expired_fields_one_hash sweep, which credited nothing at all) now credits the entry, plus the sidecar box on the downgrade back to plain Hash.

The active sweep needed the credit in both its outcomes. remove_hot credits entry_overhead recomputed from the value as it stands, and by the time the caller's db.remove() runs the maps are empty — so it credits the shell and never the fields the sweep dropped. A first cut skipped the credit on that path and stranded 384 B per two-field key; reap_key_deleted_does_not_double_credit is what caught it.

tests/hash_ttl_memory_accounting.rs holds the guard. Its oracle is recalculate_memory() — a full rescan by the same entry_overhead the running ledger claims to track incrementally — which is strictly stronger than checking deltas one at a time. All 9 tests fail on the parent commit; the two end-to-end cycle tests carry ballast keys deliberately, because on an empty database the ledger starts at 0, over-crediting saturates back to 0, and the repro passes against the very bug it exists to catch. - ci: the multi-shard consistency suite now runs before merge (#762). scripts/test-consistency.sh is the only harness that starts moon at --shards 1/4/12 and diffs behaviour across them, and no gate ever ran it: grep -rln "test-consistency" .github/ scripts/ci-local.sh returned exactly one hit, a manual checklist line in the PR template. Its 458 assertions — including FT.CREATE/SEARCH/AGGREGATE/DROPINDEX/INFO across shard counts — had never executed in CI. Wired into scripts/ci-local.sh, the local merge bar, and verified it can fail: mutating shard::scatter_aggregate to drop a partial makes the gate report AGG-03 cross-shard divergence with GATE_RC=1.

The gate lives in scripts/consistency-gate.sh — one implementation that both the host leg and the VM leg call, rather than the two hand-transcriptions of the same rc/count/waiver logic they started as, where fixing one and not the other failed silently. It checks two things an exit code cannot: that the suite reached its summary block, and that it ran >= 458 assertions. test-consistency.sh runs under set -euo pipefail, so an abort partway through leaves exit 0 and no FAIL: lines and is indistinguishable from a clean run — moon#634 is that exact defect, which had this script silently running about half its rows for months. Both checks run BEFORE the moon#536 waiver, so a truncated run can never be waived. --self-test covers six cases (clean / waived / different-single-failure / waived-plus-another / truncated / short) and runs in the hosted Lint job. - test: ten lib tests shared fixed temp paths, so two concurrent test runs corrupted each other (#822). std::env::temp_dir().join("moon_test_datafile") and nine siblings named their scratch path with a string literal, so every concurrent cargo test process on the host picked the same directory and one run's remove_dir_all teardown deleted it under another's feet. scripts/ci-local.sh runs the monoio and tokio VM suites concurrently by default, so this is the normal case. Measured at ae6cd003 on the macOS host, two processes looping one test: persistence::kv_page::tests::test_datafile_roundtrip 0/40 solo → 13/80 concurrent; tls::tests::test_reload_tls_config_swaps_config 0/20 solo → 24/50 concurrent. The TLS one is the dangerous case: it does not report a missing file, it reports TLS config: keys may not be consistent: KeyMismatch — one run reading another run's cert against its own key, a harness defect that reads as a TLS bug. tests/common/mod.rs::unique_test_dir had already solved this for the integration suites; lib tests cannot use tests/common, which is how the fix failed to reach them. Adds the same construction crate-side as crate::util::test_temp::unique_test_dir (pid + nanos + a process-local atomic counter — the counter is the part that cannot collide, because macOS SystemTime::now() has only microsecond resolution) and switches all ten call sites in src/tls.rs, src/persistence/kv_page.rs, src/admin/footprint.rs and src/config/conf_file.rs. Guarded by scripts/audit-test-tempdirs.sh, wired into the hosted Lint job and ci-local.sh; it ships a --self-test leg that plants a violation and asserts the scanner reports it, so a scanner that stopped matching fails instead of passing. Green after the fix: 0/80 and 0/50 under the same concurrency, and 3× two concurrent full --lib suites with only the pre-existing #856 failure.

  • ci: clippy now lints tests, benches and examples. Every clippy invocation in ci.yml was lib-only, so --all-targets code was never linted anywhere: moon#835's two errors (tests/busy_poll_idle.rs, src/io/fd_table.rs) survived for weeks, and moon#840 added a third (tests/restart_preserves_compact_encoding.rs, collapsible_if) the same day it merged. moon#849 fixed the two errors but added no gate, so nothing stopped the next one. This adds the --all-targets leg and fixes #840's lint. The leg runs on ubuntu deliberately: fd_table.rs is cfg(target_os = "linux") and is invisible to any macOS run, so a local lint pass cannot substitute for it. Verified it can fail: reverting the collapsed if gives RC=101, 2 errors.

  • ci: the memory steady-state gate never compared against its baseline, and its baseline was never comparable in the first place (moon#764). Three compounding defects, found by chasing the gate's own green:

  • --self-test exit 0'd immediately after proving its own injection check worked, one branch above the compare_snapshot call against tests/fixtures/memory-baseline.json -- the only place the committed baseline was ever read. ci.yml invokes exactly --self-test --skip-build, so that comparison had never executed in CI. Fixed: --self-test is now a phase, not a mode -- it falls through to the real comparison instead of returning early.
  • The committed baseline was captured on macOS aarch64/debug while the gate runs on ubuntu-latest (Linux), so even with (1) fixed the comparison would have measured the runner, not the code, on every PR. A new check_baseline_provenance refuses to compare (exit 2) when a baseline's platform doesn't match the runner's.
  • Regenerating a baseline on a real machine surfaced that os/arch-level provenance was not enough either: a GCE c3-standard-8 and GitHub's hosted ubuntu-latest are both Linux/x86_64 and still produced a false 243% allocator_overhead "regression" from machine differences alone (RSS +30.51%, allocator_overhead +243.22%, while dashtable -- unaffected by machine class -- held at -0.14% on the same run). check_baseline_provenance now also records and gates on cpu_count (hard, exit 2 on mismatch -- GitHub documents a fixed vCPU count per runner label) and records cpu_model/mem_total_kb/runner identity for warning and audit, without hard-failing on them (GitHub does not guarantee CPU-model stability within a runner label, so gating there risked trading a silent-wrong-machine-pass for a permanent false-red instead). Separately, real runs on that same GCE host showed hnsw and allocator_overhead are genuinely noisy run-to-run with zero code changes (~±16% / ~±19% measured over 10+ runs) -- the old flat ±5% tolerance would have failed the gate on those two kinds roughly every other real run for reasons having nothing to do with the PR under test. compare_snapshot now applies a wider, evidence-based floor to just those two kinds (±20% / ±25%); the other five kinds and RSS are unchanged. Every run now also uploads its captured snapshot as an artifact unconditionally (pass, fail, or baseline capture), and a workflow_dispatch input regenerates the baseline on the actual runner instead of an out-of-band machine.

  • ci: a memory drop is now a failure with its own name, rss judges the two directions at different thresholds, and the gate can no longer measure an empty server (moon#764). The first comparison that gate ever performed came back rss -14.02% / allocator_overhead -43.34% with dashtable flat at -0.06% and zero files changed under src/ -- so nothing had got leaner, and the cpu_model differs warning in the log was a red herring (a later capture on the baseline's own EPYC 9V74 measured -11.82% against it). The cause was the harness measuring itself: --memory-arenas-cap 2 (jemalloc 8 arenas -> 2) was added to start_server() one commit after the baseline was captured, so the gate was comparing two different measurement configurations. allocator_overhead was never a second opinion corroborating it -- it is rss - tracked_sum (src/command/server_admin.rs), a residual that moves with RSS by construction. Two FAIL lines, one measurement. Fixes:

  • compare_snapshot classifies every out-of-band delta as GREW or SHRANK. Both still fail -- a shrink is never silently accepted -- but a shrink-only failure prints === BASELINE NO LONGER DESCRIBES THIS BUILD === and names the two causes worth checking (a stale baseline, including one invalidated by a harness/server-flag change; or a workload that never ran, which is moon#764's vacuous gate wearing a different hat).
  • rss grows at +5% and shrinks at -10%; the growth threshold is not relaxed. Nine hosted-runner samples of an identical source tree, over four CPU SKUs the ubuntu-latest pool hands out, span 7.96% (119,943,168 .. 129,495,040) while dashtable over the same nine spans 0.11%. A symmetric ±5% band is 10 points wide and cannot hold an 8% spread: the first re-baseline left 0.63% of margin under the observed minimum and the next run measured -4.39%. Rather than widen the band -- every point added to it is a real regression the gate stops catching -- the noise slack goes only to the direction that can never be a regression. A shrink still has to clear -10%, which moon#764's own -11.82%..-14.02% step does not. The baseline is re-captured on a hosted runner with the current harness and chosen, from those nine samples, as the one that balances headroom (+3.70% above the observed maximum, +4.04% below the observed minimum); the full per-candidate table is in the fixture README.
  • check_workload_ran asserts absolute, baseline-independent floors on dashtable/hnsw/csr before anything is compared or written as a baseline. Every other check is relative, so all of them shared one blind spot: an empty capture and an empty measurement compare at 0% and stay green forever.
  • tests/memory_gate_compare_selftest.sh sources the gate (which now has a lib-only BASH_SOURCE guard) and drives the comparison against synthetic snapshots -- 16 checks, no server, no build, ~1s, run in CI before the build. Verified it can fail, by mutating the gate: accepting shrinks silently -> 2 failures; compare_snapshot hard-wired to return 0 (the original #764 bug) -> 5 failures; workload floors zeroed -> 1 failure; growth relaxed to the shrink floor -> 1 failure; shrink floor widened past the -14% step -> 2 failures. End-to-end on the real runner, against the real committed-baseline path: a baseline with dashtable cut 15% made the job exit 1 with FAIL (GREW): dashtable delta=17.64% (run 34191154173); the same commit with the injection removed passed with every kind in band (run 34191132296). The stale baseline this entry is about produced === BASELINE NO LONGER DESCRIBES THIS BUILD === / FAIL (SHRANK): rss delta=-13.42% on run 34190674369 -- the new classification, from a real run, not a mock.

  • ci: allocator_overhead's shrink floor is derived from RSS's, and the memory gate fails closed on a delta that does not compute (moon#764 follow-up to #854). Found by a real pre-merge dispatch, not by reasoning: run 34193633946 measured rss delta=-7.26% -- comfortably inside the shrink floor, a green whole-process measurement -- and went red on FAIL (SHRANK): allocator_overhead delta=-30.83%. allocator_overhead is rss - tracked_sum (src/command/server_admin.rs), and since the tracked kinds are stable to 0.11% run-to-run it absorbs essentially all of RSS's absolute jitter while being only ~23% of RSS's magnitude -- so a percentage band on it is ~4.3x tighter than the same percentage on RSS, in the only unit that matters. ao_shrink_threshold() now expresses RSS's own shrink allowance in this kind's units (max(kind_threshold, (baseline_rss * RSS_SHRINK_FLOOR / 100) / baseline_ao * 100), = -43.05% for the committed baseline): never gated tighter, in bytes, than the number it is derived from, never looser than its own measured floor. No detection power is given up -- untracked memory can only leak by growing, growth is untouched at +25%, and RSS at +5% (~6.4 MB) is a tighter absolute growth check than +25% of allocator_overhead (~7.4 MB), so RSS sees such a leak first. Replayed against all eleven real post-#854 snapshots: every one passes (worst margins 2.74 points on rss, 12.2 on allocator_overhead), while the pre---memory-arenas-cap baseline replayed as a measurement is still caught at rss +13.16%. Two fail-closed fixes alongside: a delta or verdict that does not compute (malformed snapshot, jq null, python traceback) reported delta=% and fell through to the OK branch -- it is now FAIL (UNEVALUATED), because "did not evaluate" must never render as "in tolerance" in a gate whose original bug was exactly that; and sourceing the script inherited the caller's positional parameters and died on Unknown option before defining a function, so the argument loop moves behind the same BASH_SOURCE guard as main(). The offline self-test grows to 21 checks, pinning the derivation itself and the unevaluable-snapshot case.

Documentation

  • BENCHMARK.md re-measured end to end and cut from 2137 lines to 285. The throughput and memory numbers now all come from one run on 2026-09-08 (the AOF, vector, graph and full-text results are carried forward with their own dates and are labelled as not re-measured): moon ae6cd003 (v0.8.9) against Redis 7.0.15 on GCE c3-standard-8 (x86_64) and t2a-standard-8 (aarch64), --shards 1, n=5 interleaved reps with Redis restarted and re-measured every rep as a live drift control, and a per-row noise floor below which a ratio is reported as a tie rather than a result. The previous report is archived verbatim at docs/internal/benchmark-history.md — it is kept because several of its sections exist to document how an earlier published number turned out to be wrong, and deleting them would delete the correction with the error. Raw CSVs in docs/internal/bench-data/2026-09-08/.

Findings: the GET/SET inline-path split is confirmed and remains the headline caveat (GET/SET 1.6x at p=64; every other family 0.45-0.87x at p>=8, both arches). The write-path regression the old §2.13 recorded between v0.6.0 and v0.8.7 is reversed — a v0.8.7 rebuild on the same hosts puts main ahead on every non-inlined family, twelve rows of twelve on both arches, by +5 to +23% (x86) and +5 to +18% (ARM) measured in moon's raw ops/s -- the ratio column is also published, but where the two disagree the raw one is the honest number, because per-cell Redis drift reaches 12.8% and lands on the largest ratio gains. The one loss that did not come back is SET at p=64. Memory: moon uses less per key at every size from 8 B to 1 KB on both arches (8-22% on x86, 8-16% on ARM), which does not reproduce the archived "11-51% worse at 32 B" (measured 0.83x here) or "empty-server RSS 1.7x Redis" (a tie here); both older runs used Redis 7.4.2 rather than 7.0.15, and both readings stay on the record.

  • The "27-35% less memory" claim is corrected to a measured 15-17%, on Linux, against a jemalloc Redis (#817, #821). The figure was re-measured on GCE c3-standard-8 (Linux 6.17, x86_64, 8 vCPU / 31 GB), moon d5f3501b at --shards 1, via scripts/bench-resources.sh with a fresh server per data point and redis-benchmark -r N for unique keys. Per-key = (loaded RSS - baseline RSS) / DBSIZE. The full twelve-point table and the method are in BENCHMARK.md §3, and the user-facing documents (README.md, docs/index.md, docs/benchmarks.md, docs/journey.md, docs/design-advantages.md, docs/comparison-valkey.md) state the measured number with its provenance and link there. The three example programs that republished the retired figure as a retrieval document at runtime (examples/ai-agent-tools/agent_tools.py, examples/rag-quickstart/ rag_quickstart.py, examples/graphrag/graphrag.py) now carry the measured figure, its scope, and the loss at 32 B. The retired number was also baked into the pixels of docs/images/diagrams/moon-memory-engine.png (published on docs/design-advantages.md directly above the paragraph that corrects it); that box is removed from the image and from its regeneration prompt (docs/images/diagrams/prompts/2-memory.txt), so a regenerate no longer reproduces it. The historical CHANGELOG entries that recorded the claim at the time are left as a dated record, alongside this retraction.

Result: 15-17% less memory per key at values >= 1 KB (9.5% at the smallest key count tested, 63K x 1 KB); a tie at 256 B (0.92-1.02x); and a loss of 11-51% at 32 B, which is published alongside the win rather than omitted.

The old number was wrong because of the oracle, not because of moon. The previously published 1M x 1 KB row was Redis 1,571 B/key against moon 1,153 B/key. Re-measured: Redis 1,380 B/key, moon 1,172 B/key. moon's own figure moved 1.6%; Redis's moved 12%. Both Redis binaries already installed on the benchmark host were libc-malloc builds, which inflate Redis RSS and would have biased the comparison in moon's favour, so Redis 7.4.2 was rebuilt from source against jemalloc 5.3.0 for this run. The inflation in the retired claim came from a Redis baseline measured on macOS and/or without jemalloc.

Scope, stated so it is not over-read: --shards 1, x86_64 only — nothing here may be restated as an ARM result — string values, one data point per (value size, key count) cell with no repetitions, and no sampling between 32 B and 256 B. Throughput and CPU columns from the same harness run are not published: that harness measures them incidentally and they are not a clean benchmark.

  • BENCHMARK.md §3.1's empty-server RSS is corrected, against moon's interest (#821). The published "Redis 7.0 MB / moon 1 shard 7.0 MB — identical" was an Apple M4 Pro development reference. Measured on all 12 points of the run above: Redis 7.5-7.7 MB, moon 12.6-12.9 MB — moon is ~1.7x worse. The cause is not known and is tracked in #821. The "moon (12 shards) 15.7 MB" row is kept but marked unverified / stale: this run did not measure it.

BENCHMARK.md §3.4's TTL-overhead claim is marked unverified for the same reason it could not be confirmed: that section of scripts/bench-resources.sh omits redis-benchmark -r, so every SETEX hits __rand_key__ and one key is loaded instead of 500,000. The structural description is retained as a reading of the source, not as a measurement.

  • Every published performance claim now names the host it was measured on, and the macOS-sourced ones are no longer presented as Linux results (#817). docs/benchmarks.md:9 already labelled its per-table figures an Apple M4 Pro development reference; docs/journey.md re-published the same values under a preamble asserting they came from GCloud Linux hosts, and README.md, docs/index.md, docs/design-advantages.md and docs/comparison-valkey.md quoted them with no host at all. CLAUDE.md and docs/PRODUCTION-CONTRACT.md both require every benchmark number to come from a Linux host.

Two claims are withdrawn rather than relabelled, because no defensible provenance exists for them:

  • "45x / 23x better CPU" — macOS-measured, sampled with ps -o %cpu= (a process-lifetime average, not steady-state load), and computed from a CPU numerator and an RPS denominator taken in different runs. The Linux comparison (§2.14) is a tie: 10.55 vs 11.33 us/op, inside Redis's own 11.9% spread. The underlying table stays in BENCHMARK.md §5.1 as a development record with no ratio derived from it.
  • "p50 latency 8-10x lower" — macOS-measured and taken with redis-benchmark, a closed-loop tool. No Linux latency comparison exists. Raw numbers retained in BENCHMARK.md §9.1.

Also corrected: docs/journey.md's retracted "1->8 shards is flat-to-slightly negative" now states the corrected Linux scaling of 1.42x / 2.14x / 3.79x (§2.14); README.md's "hash-field TTL (Valkey-parity)" now says feature parity and 4-10% behind Valkey, which is what docs/perf/2026-05-27-hash-ttl-3way-bench.md measured; the Moon-vs-Valkey throughput ratio is withdrawn as structurally asymmetric (different hardware, pipeline depth, payload and thread counts); and docs/references.md no longer claims open-loop methodology for figures produced by a closed-loop tool. docs/index.md gained the scope caveat it had none of. For the changes in this bullet no new measurements were taken and no number was invented, adjusted or extrapolated — every surviving figure already existed in the tree. The per-key memory and baseline-RSS corrections above are the exception: those are a new Linux measurement, and their host, build, oracle and method are stated in BENCHMARK.md §3.

Fixed

  • scripting/persistence: a Lua or Functions script write under appendfsync always is acknowledged only after the fsync that covers it (#831). EVAL/EVALSHA carry no WRITE flag by design — the bridge emits one effect record per successful inner redis.call, fire-and-forget into the shard's AOF writer — but the script arms never joined the batch-end barrier set that every ordinary write joins, so a batch holding only script writes issued zero fsyncs before its reply. The records were durable by accident of the writer's policy-driven fsync landing first, which is exactly what #763's waiter-gated fsync removes; #831 therefore blocked #763.

Reproduced on main@b04e8990 (macOS aarch64, monoio, --shards 1): with a second connection keeping the writer busy, an EVAL that SETs was acked :1 and the key was gone after kill -9 + restart in 9 of 16 attempts; the identical shape with a plain SET lost 0 of 16. Under MOON_TEST_AOF_FSYNC_FAIL=1 a plain SET answered AOF_FSYNC_ERR while EVAL, EVALSHA and FCALL answered their script's :1 — 32 of 32 at --shards 4, routed arms included.

The fix: every script arm reads the bridge's per-script write flag (take_script_had_write, consumed immediately after the VM returns and before any .await) and, when the script wrote and its reply is not an error, pushes the reply slot into local_leg_write_idxs so resolve_local_leg_barrier covers it with the batch's ONE fsync_barrier — the same group commit a SET gets, on both the monoio and the tokio sharded handler. A script routed to the shard owning its keys reports the flag in ExecReply::script_wrote, and the originator awaits fsync_barrier(owner) before serializing the reply. Read-only scripts and EVAL_RO/FCALL_RO are untouched. After the fix: 0 of 16 acked script writes lost across kill -9; every write-script arm answers the injected fsync failure. Pinned by tests/script_write_fsync_barrier_831.rs.

  • replication/persistence: a non-deterministic write now propagates its EFFECT, not itself (#825). SPOP and XADD key * were written to the AOF and streamed to replicas verbatim; replay re-rolled the RNG and re-read the clock, so the member the client was told it removed came back after a restart (and a different one vanished), and a stored stream ID could never find its entry — with no error anywhere. Enumerated from the state writers rather than the two reported names, the same class covered EXPIRE/PEXPIRE with a NX|XX|GT|LT flag (the frame-only expire rewrite bails on the fourth argument), HEXPIRE/HPEXPIRE/HGETEX EX|PX (relative hash-field deadlines, never rewritten) and RESTORE with a relative TTL. Orthogonally, the owner-shard SPSC arm, both MULTI/EXEC executors, the Lua effect plane and the embedded single-shard handler serialized the raw frame and bypassed the existing expire rewrite entirely, so every relative-TTL command was propagated with a restarting countdown on the cross-shard, transactional and scripted paths. A new reply-aware rewrite (replication::effect_rewrite, (frame, reply, now_ms)) evaluated on the executing shard turns each into its deterministic form — SREM key <popped>, XADD key <assigned-id>, PEXPIREAT/DEL, HPEXPIREAT … FIELDS <touched>, RESTORE … ABSTTL — or into nothing when the reply proves nothing was written; every propagation site now routes through aof::serialize_effect_for_log, which chains it with the expire rewrite. Covered by tests/nondeterministic_propagation_825.rs (restart, DEBUG DIGEST as the aggregate oracle, single- and four-shard). Not touched, by design: INCRBYFLOAT/HINCRBYFLOAT (moon's f64 replay is deterministic — measured stable across a restart) and the XCLAIM/XAUTOCLAIM/XREADGROUP PEL delivery-time family (needs XCLAIM … TIME/FORCE/JUSTID support first; filed separately).
  • storage: a read no longer flattens a compact encoding (#832). Database::get_promoted returns K::Shared<'_> -- a shared reference -- but obtaining one called K::upgrade unconditionally, and nothing in the tree ever downgrades. So a single read taken on the mutable dispatch path (inside MULTI/EXEC, inside a Lua script, or from try_inline_dispatch) permanently converted a listpack/intset key to its full form for the rest of its life -- and whether that happened depended on which of moon's three dispatch paths the command took, not on what the command did. OBJECT ENCODING became observation-dependent, and used_memory moved on a command the user believed was read-only.

Measured on the unmodified binary (--shards 1, used_memory ledger): 1000 eight-member integer sets went 333,055 -> 1,149,055 bytes (3.45x) after one SCARD each, intset -> hashtable; a three-element list went listpack -> linkedlist after one LLEN. That put a ceiling on every container memory figure, making them write-only-workload numbers.

All fourteen affected read handlers (SMEMBERS, SCARD, SISMEMBER, SMISMEMBER, SINTER, SUNION, SDIFF, SRANDMEMBER, SSCAN, SINTERCARD, LLEN, LRANGE, LINDEX, LPOS) now take their read through the &Database implementation that already backed dispatch_read, so the compiler makes the rewrite unrepresentable rather than merely discouraged -- and the two implementations of each command collapse into one, which is the divergence #610 came from. Five -WRONGTYPE/presence probes in the blocking machinery, which flattened a list purely to look at its type before parking a client, do the same. After the fix the same probe measures 1.00x.

The hash family was already correct (it moved to get_hash_ref_if_alive earlier) and is covered by the new gate too. tests/read_preserves_compact_encoding.rs asserts the encoding survives on all three dispatch paths; 19 of its 52 cases fail on the pre-fix binary. - server: a reader inside its own TXN no longer sees another connection's uncommitted write through the inline GET fast path (#807). Inside an open cross-store transaction the generic read leg runs the write-intent visibility filter (KvWriteIntents::is_key_visible) and answers (nil) for a key carrying a foreign transaction's uncommitted intent; try_inline_dispatch never consults the intent table, and can_inline_reads carried no in_cross_txn term, so the same reader got "modified" from the fast path and (nil) from generic dispatch for the same key. Measured on b04e8990 at --shards 1 with one reader speaking both framings back to back. The read gate now stands down inside a TXN, the term the write gate has carried since #660 — one Option::is_some() per batch, already loaded for the write gate. Scope: this closes the divergence between the dispatch paths. A reader that is NOT in a transaction sees the uncommitted value on both paths — the generic filter is gated on the reader's own transaction and non-transactional operations bypass the intent table by design — and that contract is unchanged here. tests/inline_read_txn_visibility_807.rs (monoio; red on b04e8990). - storage: ZADD reaches its listpack encoding -- small sorted sets stop being skiplists (moon#787). ZADD reported skiplist from its first member where Redis keeps a zset in a listpack up to zset-max-listpack-entries (128) / zset-max-listpack-value (64). SortedSetListpack was wired end to end -- the value codec, the RDB and AOF writers, DEBUG DIGEST, MEMORY USAGE, and the read-only SortedSetRef::Listpack arm all handled it -- but no accessor ever produced one, so the variant was unreachable at runtime and every zset paid the full B+tree-plus-HashMap cost. ZADD now routes through get_or_create_zset_listpack below both thresholds and promotes past either; SortedSetKind::upgrade gained the listpack arm that keeps every other zset command correct on a key ZADD created compact; and the restart-side re-derivation (compact_after_decode) gained its zset arm in the same change, flipping the tripwire the restart fix pinned -- so a listpack zset now survives a restart instead of reloading as a skiplist. The unit test test_object_encoding_sorted_set asserted the divergence and is corrected.

The listpack branch sits below the moon#814/#820 validation pre-pass, so an erroring ZADD still creates no key, writes no prefix and strands no charge; its first draft sat above it and re-created that regression on the new path, and ledger_consistency_788 now pins both halves for the listpack form. Scores are stored as the canonical text ZSCORE replies with (storage::zset_score::render_score, byte-identical to format_score_bytes and round-trip exact, both pinned by tests) into a stack buffer; the member lookup is a borrowed scan (iter_pair_refs), so the path allocates nothing per entry walked. The listpack -> B+tree swing is billed by SortedSetKind::upgrade from the arena's real capacity (moon#788/#810).

Scope -- what this does and does not buy (moon#832). ZADD is the only zset command that mutates a listpack in place. Everything on the mutable dispatch path -- ZREM, ZINCRBY, ZPOPMIN/ZPOPMAX, the store commands, every zset command inside MULTI/EXEC or a Lua script, and the reads that are not in dispatch_read at all (ZRANGEBYLEX, ZREVRANGEBYLEX, ZRANDMEMBER, ZINTERCARD) -- reaches the value through get_promoted, which upgrades unconditionally, and nothing ever downgrades. A zset is therefore flattened to a skiplist on its first such touch. The saving is a write-only-workload figure. Measured on Linux (moon-bench-x86, x86_64, load < 1.0, --shards 1, fresh server per row, 200 000 keys x ZADD z:i 1 alpha 2.5 beta 3 gamma, RSS delta / key, two reps): 4165 -> 295 B/key (-93%); Redis 7.4.2 measures 108 B/key on the same probe, the remaining gap being moon's per-key envelope, not the encoding. ZADD z XX 1 a on a missing key still leaves an empty zset behind -- the moon#830 get_or_create_* class, now with one more site. A replica also loses the encoding across a FULLRESYNC (redis_rdb::load_rdb rebuilds RDB_TYPE_ZSET_2 as a SortedSetBPTree) -- pre-existing and class-wide, tracked as moon#863.

tests/container_growth_memory_accounting.rs's test_sorted_set_arena_is_visible_to_used_memory (moon#788) had its premise removed here: it billed 500 ONE-member zsets and required >= 2048 B/key to prove the B+tree node arena reaches used_memory, and a one-member zset is now a listpack that owns no arena at all (measured 386 B/key -- a correct figure for the new encoding, and a red test). The invariant is not dropped: the fixture is moved past the listpack boundary with a single 65-byte member (one over zset-max-listpack-value), which is still a ONE-member zset, so the fixed arena cost still dominates the per-key figure and the 2048 B floor still discriminates. The test now asserts OBJECT ENCODING is skiplist before it measures, so a future boundary change fails loudly instead of measuring the wrong container. Proved it can still fail: dropping tree.memory_bytes() from zset_table_bytes takes it to 778 B/key, red.

  • persistence: a restart no longer flattens every compact encoding. RDB decode rebuilt each container in its full form, so a listpack hash, a listpack list, an intset and a set listpack all came back as hashtable/quicklist/hashtable on the first reload. Redis preserves all of them across DEBUG RELOAD. The memory cost is not incidental: the SADD-builds-a-listpack win (set 404.5 B/key) reverted to 978.2 B/key -- 1.96x -> 4.48x vs redis 7.4.2 -- the moment the server came back up.

The fix re-derives the compact encoding on the decode side against thresholds already in the tree (LISTPACK_MAX_ENTRIES, LISTPACK_MAX_ELEMENT_SIZE, INTSET_MAX_ENTRIES), with no wire-format change: every listpack variant already maps to the same ValueType tag as its full form (value_codec::value_type_of), so files stay readable in both directions -- an RDB written before this change reloads compact, and one written after loads on a build without it. The files are not byte-identical, because a listpack preserves insertion order where a HashMap does not; the type tag and the field/value encoding are unchanged.

It is deliberately not wired into the cold/spill path. ValueKind::classify_cold accepts only the canonical full forms, so a cold-decoded SetListpack would fall through its _ => Err(WrongType) arm and answer WRONGTYPE for a perfectly valid set -- the first attempt at this fix did exactly that and turned 11 cold-tier tests red. Compaction is therefore opt-in via decode_value_body_compacting, used only on the RDB path; HashWithTtl is also left alone, because a listpack carries no TTL sidecar. Widening classify_cold so the cold tier compacts too is a named follow-up.

Zsets were deliberately not compacted on reload by this fix (the exclusion is lifted by the ZADD-listpack entry above), and a unit test pinned the exclusion so lifting it was a decision: SortedSetKind::project_mut / project_ref accept only SortedSetBPTree and SortedSetKind::upgrade leaves SortedSetListpack alone, so a zset compacted here would answer WRONGTYPE to every zset command on the mutable dispatch path -- ZADD included -- after the very restart that compacted it. SortedSetListpack is unreachable from every load path today, so the arm would create that hazard rather than inherit it. It belongs with the write-side upgrade arm (moon#787, the ZADD-listpack change); until that lands the zset memory win is restart-transient.

INTSET_MAX_ENTRIES (512) now has one definition, exported from storage::db, read by both the SADD path and the decode-side re-derivation -- previously two private literals and a comment promising they agree.

Guarded by tests/restart_preserves_compact_encoding.rs, which writes one key of each compact type plus a 200-field hash as a negative control (without it, a bug that compacted everything would pass), BGREWRITEAOFs and waits for the AOF manifest to publish the new base (INFO persistence is not a completion signal: aof_base_size did not move across a rewrite), restarts a real server on the same --dir, and asserts the target encoding of all five. The test is not #[ignore]d: every --ignored invocation in the workflows names a specific --test target, so an ignored test here would have run nowhere. Shown red against the pre-fix binary (h: was listpack, after restart hashtable, l: ... linkedlist, si: ... hashtable) and green with the fix.

Changed

  • docs: --profile standalone is scoped to low connection counts (moon#772). The preset's p=1 win is a single-connection result; moon#772 measured it at c=200 on GCE t2a-standard-8: -80% vs stock --shards N at p=1 and -57% at p=16, because the preset's --shards 1 forfeits the other cores and its busy-poll spin steals cycles from an already-saturated shard rather than trading idle CPU for latency. Clarified in the --profile doc comment (src/config.rs), a second startup tracing::warn! (src/main.rs), README.md, docs/guides/tuning.md, docs/configuration.md, and docs/runbooks/upgrade-to-v0.6.0.md. No behavior change — flags set by the preset are unchanged.
  • metrics: the per-command observability counters no longer bounce one cache line across every shard core (moon#774). keyspace_hits, keyspace_misses, expired_keys, total_dispatch_cross_spsc and total_dispatch_cross_read_fast were plain global atomics — the first three 8 bytes apart in one line — RMW'd once per command by every shard thread. Measured shape: GET p=1 c200 at --shards 8 moved 332,779 ops/s with 87.54% cross-shard, i.e. ~291K contended RMWs/s onto that line, plus one keyspace RMW per lookup.

They now live in a 64-byte-aligned per-thread slot, reusing the scheme the total-commands counter has used since QW4; readers sum the slots and the sum is exact. All five share ONE line per thread on purpose: a cross-shard GET bumps a keyspace counter and a dispatch counter in the same command, and splitting them would trade one contended line for two private misses.

The increments stay ungated. These atomics back INFO fields that must be right with --admin-port 0, so gating them on METRICS_INITIALIZED would swap a performance defect for a correctness one. What changed is the contention, not the cost of the instruction. No INFO field changes value.

  • server: the monoio connection handler batches its dispatch-path counters, as the tokio one already did (moon#774). Batched recorders landed on handler_sharded and never on handler_monoio — the runtime that ships — so the runtime nobody deploys got the optimisation and the deployed one paid N global atomics per pipeline batch instead of one. Both handlers now accumulate per batch and flush once.

Nothing caught this: both spellings produce the same counter value, so every INFO assertion and Prometheus scrape stayed green. tests/dispatch_recorder_drift.rs now pins it as a source-level invariant, since the call site is the only observable.

A record_dispatch_cross_read_fast_batch counterpart was added for the L4 shared-guard read path, which had no batched variant.

Removed

  • shard: CoalescedReadBatch, the cross-connection read-coalescing type that never had a producer or a consumer (moon#773). Defined in PR #177 with the xshard-read-fastpath C1 types and deferred at C3, it sat in src/shard/dispatch.rs for three months with no ShardMessage arm, no spsc_handler arm, no config and no metric — its only references were its own definition and a type-existence test. In the type index and in code review it read as a shipped cross-shard optimisation. It was not one.

It is deleted rather than wired up because all three arguments for building it have since failed. Its stated premise — "a foreign lock-free read is storage-impossible in place, so SPSC is the only door" — stopped being true when L4 made Database Send + Sync and try_foreign_db_read began serving foreign reads on the calling thread with one CAS (#777, default auto since #785); coalescing would now be optimising the fallback. It also attacks the wrong term: it reduces msgs/cmd, whose fitted coefficient in the cost model is ~0, while every connection still parks on its own ResponseSlot — that is dead end #1, predicted 0.99x. And the variant that batches the wake instead was pre-flighted and retired in #778, because monoio's EventWaker already coalesces 87% of cross-thread wakes.

The reasoning, the load-bearing ordering invariant the type's doc comment asserted, and the pointer to where the remaining cross-shard work actually is (writes at 0.875 parks/cmd, the tokio handler, the read-side declines) are preserved in docs/internal/cross-shard-cost-model.md §9. No behaviour changes: nothing constructed, sent or matched on this type.

Fixed

  • bench: scripts/gcloud-xshard-absolute.sh could not toggle the cross-shard fast path, so its s4-c1-GET cell silently stopped measuring the thing it is named for (moon#416). start_moon passed a fixed server-argument list. Since #785 flipped --cross-shard-fast-path to auto by default, that cell on a populated keyspace is served in place by the L4 fast path — it was reporting the fast path while the harness, the CSV and the frozen XSHARD-READ-01 contract row all still called it the SPSC hop.

start_moon and cell now take an optional extra-server-args string, and --extra-args / MOON_EXTRA_ARGS plumb it from the command line, so the hop and the fast path can be measured as two cells of one A/B instead of one ambiguous number. The default cell set now runs s4-c1-GET-hop (--cross-shard-fast-path off, the actual hop) alongside s4-c1-GET (default auto), and a --self-test gate fails closed if the passthrough ever stops reaching the server process. - perf(server): the per-write eviction gate is keyed on STATE — a configured maxmemory or per-db quota — not on whether a spill sender is wired. eviction::write_gate_active() is now the one predicate behind all four write-path gates: the monoio batch flag (batch_eviction_active), the SPSC drain (evict_active), the Lua bridge (LuaEvictionCtx::gate) and the tokio per-write check (which had no gate at all). Each of the first three carried spill_sender.is_some() as a term, which is true on every default server (--disk-offload enable), so every non-inline write paid run_write_eviction_gate — a RuntimeConfig read-lock pair, an elastic_budget load, an EvictionRun build and evict_to_budget's early return — for nothing: a wired sender only changes where a VICTIM goes (spilled vs plain-dropped), and evict_to_budget selects no victim when maxmemory == 0. Under any configured limit nothing changes: the gate runs, builds the same EvictionRun for the same sender, and routes exactly as before. Also drops the per-batch runtime_config.read() the monoio flag took. Scope, measured: a default moon auto-caps maxmemory at ~80% of RAM (--maxmemory omitted != --maxmemory 0), so on a default server the gate is active through maxmemory_is_set() regardless of this change and its 2.8-3.5% + 0.7-1.3% of shard-0 cycles on -t hset (a configured but un-breached limit) remain — client-driven -t hset is 920K rps on both binaries. Under an explicit --maxmemory 0 the gate symbol disappears from the profile. Recovering the default-server cost needs the finer per-write "configured but un-breached" skip, which is a separate change. Guarded by gate_is_skipped_with_spill_sender_when_no_limit_is_configured on a test-only evict_to_budget entry probe; tests/oom_bypass_closure.rs still enforces the publish contract the predicate relies on. - perf(shard): the owner's inline GET takes the SHARED guard on its own database. try_inline_dispatch's GET ran under with_shard_db — since the L4 plane that is ShardDbSet::write, an exclusive guard for a lookup that mutates nothing (get_if_alive neither expires nor promotes nor touches LRU). While it was held, every foreign shard's try_foreign_db_read of that database declined and parked on SPSC, and the owner itself waited behind any foreign reader already inside. The lookup now runs under with_shard_db_read; the one mutating case (a key mid-spill, #459) is detected under the shared guard and retaken exclusive, so it still answers. Pinned by test_inline_get_does_not_wait_behind_a_shared_reader and test_inline_get_answers_a_mid_spill_key. - perf(shard): the per-batch clock refresh probes under a shared guard and takes the exclusive guard only when the clock moved. The monoio handler refreshed the database clock through with_shard_db once or twice per batch — per command at p=1 — for a value that changes once per millisecond. slice::refresh_db_clock replaces both sites. Pinned by refresh_db_clock_takes_no_exclusive_guard_when_the_clock_is_unchanged. - perf(storage): DashTable::split_segment repoints its directory in O(block), not O(directory). Every segment split scanned the whole directory (for i in 0..directory.len()) to find the slots routing to the split segment — O(2^depth) per split, O(N^2 / segment_capacity) over a fill. Extendible hashing guarantees those slots are one aligned block of 2^(depth - local_depth) entries, of which only the upper half moves, so the update is now one or two writes. Measured on moon-bench-x86: a 2M-key fill spent 65% of its wall clock in that scan (G1 §6.4); with the fix the fill goes 2,414/2,398 ms -> 953/954 ms (2.53x) and lands within 14% of a presized zero-split table, with identical split_count (58,900), segment_count (58,901) and len; split_segment falls from 7.4-7.7% to 0.6-0.7% of growth-phase shard-0 samples; a client-driven 2M fresh-key SET fill goes 664K -> 997K rps (+50%, spread < 0.1% over 3 fresh-server reps). Steady state is unchanged (the scan was 0.7-0.9% there); every bulk load, restart and fresh-key benchmark paid it. Under cfg(test) the old scan is kept as the oracle and every split is checked against it (split_directory_repoint_matches_full_scan_across_200k_fill); the aligned-block invariant is asserted from first principles alongside. examples/dashtable_growth.rs is the growth micro-bench that measured it. DashTable::split_count's doc no longer claims MEMORY DOCTOR reads it (nothing on the command path does). - storage: a heap string is one allocation, and nothing on the heap is read to drop or overwrite it (perf campaign G1 §5, lever 5). A string value longer than the 12-byte SSO cutoff used to be Box<HeapString(Box<[u8]>)>: a 16-byte wrapper allocation holding the pointer and length of the data allocation. Database::set's overwrite closure — the largest symbol in every steady-state -t set profile — dereferenced that wrapper to bill and then to free the old value: a second cache miss per overwrite, plus one malloc/free pair per write for the wrapper itself.

CompactValue now stores the string's thin pointer and its length in its own 16 bytes: the previously unread 4-byte prefix in payload[0..4] holds the low 32 bits of the length and the 28 LEN_MASK bits of len_and_tag the high bits (60 bits; the 256 MiB cliff a 28-bit length would have had does not exist). The kind of a heap value moved out of the pointer's low bits into len_and_tag's high nibble (0xE = heap string, 0xF = heap collection), so the string pointer is untagged and the layout assumes nothing about the alignment the allocator gives an align-1 [u8]. Drop rebuilds Box<[u8]> from (ptr, len); sdallocx needs nothing from the object.

used_memory bills a heap string as size_class(len) through mem_size (#810), 16 bytes less per key than before, and tests/compact_value_alloc_accounting.rs proves against a counting allocator that the value is exactly one request of len bytes, freed exactly once, at every jemalloc class boundary. The string path's unsafe surface went from nine blocks to three (heap_str, heap_str_mut, take_heap_string), each citing the numbered invariants on the type; every constructor establishes them in one function.

Added

  • test/bench: the shipped default is now measured, and its dispatch path is guarded (moon#833). Every general-purpose benchmark script started moon with --disk-offload disable; the default is enable. That one flag is what hid moon#812 for months — on a default server plain SET never took the inline fast path (can_inline_writes carried ctx.spill_sender.is_none(), a config predicate standing in for a state one), a ~53% SET deficit that no benchmark could see because none ran the default.

tests/default_config_dispatch_path_833.rs starts moon with no tuning flags (only --port, --admin-port as the instrument, and --dir to a fresh tempdir) and asserts from moon_dispatch_path_total that plain SET, plain GET, and a pipelined SET+GET batch are served local_inline at the default --shards 1. Proved to redden against the pre-moon#812 binary (29fc5fce): both SET tests fail with local_inline +0, local +200; GET stays green, as it should — can_inline_reads never carried the term.

scripts/bench-compare.sh now runs a third server, Moon default (moon --port N --dir <tempdir>, nothing else), and reports it in every table beside the tuned row with a default/tuned ratio, so the delta between the configuration we publish numbers for and the one we ship is visible by construction. --skip-default restores the old two-server run. scripts/bench-production.sh gains --default-config (its ten scenarios are too entangled with one port to host a third server per row): run it twice and diff.

Fixed

  • lint: cargo clippy --all-targets -- -D warnings failed on Linux on pristine main — clippy::collapsible_if at tests/busy_poll_idle.rs:83 and clippy::manual_contains at src/io/fd_table.rs:153, both in Linux-only files, so a macOS run was green and the hosted Check job (which runs clippy without --all-targets) never saw them. Two one-line fixes.
  • bench: scripts/bench-compare.sh aborted its first bench() call on macOS — rflag[@]: unbound variable: an empty array is "unbound" to set -u on bash < 4.4, and macOS ships 3.2 (the #634 class). The -r expansion is now ${rflag[@]+"${rflag[@]}"}.
  • server: a default-config server refused pipelined writes with -MOONERR AOF backpressure (moon#838, regression from moon#812 in v0.8.9). With no flags at all (--appendonly yes, everysec, --disk-offload enable) a routine redis-benchmark -t set -c 50 -P 16 aborted 5/5 on GCE; no benchmark script could see it because every one passes --appendonly no (moon#833).

Two write paths, one condition, two bounds: the inline SET fast path enqueued its AOF record with a synchronous 5 ms bounded block on the shard thread (AOF_SPSC_BACKPRESSURE_BOUND) and failed loud AFTER applying the write, while the generic leg awaits the same channel for --aof-fsync-timeout-ms (2 s) and blocks nothing. Until #812 no default-config SET could reach the 5 ms bound; #812 put every default SET on it, and any writer hiccup of ~10 ms at ~1M rps (a 50 ms idle-escalated poll step on the first burst after idle, a slow everysec fsync) filled the 10k channel and refused the burst.

The inline path now probes the writer (AofWriterPool::append_would_block) BEFORE consuming the command bytes and stands down to generic dispatch when the channel is full — the moon#660 pre-gate shape. It never applies a write it cannot queue; the leg that can wait takes it, and the bytes inline again the moment a slot frees. Not consulted under --appendonly no. The 5 ms bound itself is unchanged: raising it would trade a loud refusal for a shard thread parked for seconds (moon#769's SPSC-arm bound is a queue-depth × fsync-stall problem with its own design).

Guarded by tests/default_config_aof_backpressure_838.rs (shipped flags, hook-free first-burst-after-idle; a deterministic everysec stall via the new writer-side test hook MOON_TEST_AOF_FSYNC_STALL_MS; and a pin that a default server's refusal, when a stall outlasts --aof-fsync-timeout-ms, comes from the generic leg, is applied-first and loud) plus unit tests on the probe and on the inline path standing down with read_buf intact and the key untouched. All red on 6251429f. - storage: SADD reaches its listpack encoding -- small string sets stop being hashtables (moon#787). SetListpack existed as a storage encoding and SADD already had an intset path, but nothing ever created a string set listpack: a small set of non-integer members went straight to hashtable, where Redis keeps one in a listpack up to set-max-listpack-entries (128) / set-max-listpack-value (64). The guard does not live in OwnedKind::upgrade (every impl is unconditional) -- it lives one level up, in a per-type accessor the command layer calls instead of the owned accessor, and the set one was missing:

get_or_create_intset            EXISTS   <- SADD, integer members
get_or_create_hash_listpack     EXISTS   <- HSET
get_or_create_list_listpack     EXISTS   <- RPUSH
get_or_create_set_listpack      MISSING

A unit test had pinned the defect in its own name (test_object_encoding_set_hashtable, asserting "SADD with non-integer members should create hashtable"); it is corrected, with a new test keeping the past-threshold hashtable case covered. Promotion past either threshold is one-way, matching Redis; the listpack -> IndexSet swing is billed by SetKind::upgrade (the same conversion get_or_create runs), so the entries Vec and index table are charged from their real capacity (moon#788/#810), and a ledger_consistency_788 case walks create / duplicate / promote / delete against a full recompute. The membership scan uses Listpack::contains_element -- the borrowed comparison moon#801 introduced for exactly this lookup -- so the path allocates nothing per entry walked.

Scope -- what this does and does not buy (moon#832). Database::get_set is get_promoted, which calls SetKind::upgrade unconditionally, and nothing ever downgrades: any set command on the mutable dispatch path -- SCARD/SISMEMBER/SMEMBERS outside dispatch_read, every command inside MULTI/EXEC or a Lua script, SREM, SPOP, SMOVE, the store commands -- permanently flattens the listpack on its first touch. The saving is a write-only-workload figure; a read-mixed workload reverts key by key. An existing intset that receives a string member still promotes straight to hashtable (Redis 7.2+ converts it to a listpack) -- a named residual, not covered here. A replica loses the encoding across a FULLRESYNC: redis_rdb::load_rdb rebuilds RDB_TYPE_SET as RedisValue::Set, so the key comes back a hashtable until it is rewritten. That is pre-existing and class-wide -- intsets, hash listpacks and list listpacks are flattened by the same loader on today's main -- so it is filed as moon#863 rather than fixed here; this change only widens the set of keys it reaches. Measured on Linux (moon-bench-x86, x86_64, load < 1.0, --shards 1, fresh server per row, 200 000 keys x SADD s:i alpha beta gamma, RSS delta / key, two reps): 562 -> 264 B/key (-53%); Redis 7.4.2 measures 106 B/key on the same probe, the remaining gap being moon's per-key envelope, not the encoding.

[0.8.9] — 2026-09-04

Added

  • test: the hot/cold/WAL reconciliation invariant, as a property (moon#660 step 2). Disk offload is a two-source-of-truth durability path. The hazard is not a double-write conflict with the WAL — spilled segments are independently self-durable — it is RECONCILIATION: recovery runs Phase 3 (rebuild cold_index from the manifest) then Phase 4 (WAL replay on top, hot shadowing cold). Every bug found in that seam so far has been silent-data-loss class — DEL/FLUSH resurrection and expired-cold leak (#212), BITOP/COPY/DEL/UNLINK resurrection (#213), a spill completion resurrecting a DEL'd key (#459) — and every one was caught by soak or adversarial review, never by a proof that the invariant holds in general.

tests/cold_reconciliation_property_660.rs is that proof: a seeded generator drives writes, deletes and expiries under real memory pressure and asserts the server's answer for every key matches a model, both while running and again after SIGKILL + full Phase-3/Phase-4 recovery. Failures are named by shape (RESURRECTION, EXPIRED-COLD LEAK, LOST WRITE) and every seed is replayable via MOON_660_SEEDS.

It earned its keep immediately: it is what surfaced the COPY/BITOP single-shard durability bug fixed in this same release, as a deterministic 3-of-3 CI failure rather than a soak-only ghost.

The generator is hand-rolled rather than proptest on purpose — a durability default is not the place to also widen the supply chain, and the part that matters here is reproducibility, not shrinking.

Changed

  • config: --disk-offload now rejects values other than enable and disable. Only the exact string enable ever turned the tier on, so a typo (--disk-offload enabled) silently meant "without the tier". Failing at parse time is the difference between a startup error and a cluster quietly holding its whole keyspace in RAM.

The default is unchanged (enable). moon#660 proposes making the tier opt-in; that change is held back until it can fail loudly rather than silently — an operator upgrading with existing offload state on disk currently gets no warning, no error and no INFO field, only a smaller keyspace. Tracked separately.

  • test: tests/vector_db_isolation.rs pins --disk-offload explicitly rather than inheriting the default, and adds ft_index_survives_restart_without_disk_offload. FT index definitions persist via the offload dir when the tier is on; with it off they need --appendonly yes or --save. Measured: the index survives with either backstop, and is lost only when the operator configured neither.

Fixed

  • server: a blocking pop propagated NOTHING outside MULTI — acked writes undone by restart, replicas permanently diverged (moon#827). BLPOP, BRPOP, BLMOVE, BRPOPLPUSH, BZPOPMIN, BZPOPMAX, BLMPOP and BZMPOP mutate the keyspace and ack the client, but the blocking path is an INTERCEPT: it short-circuits the dispatch exit where every other write meets the AOF and the replication stream. Nothing fed them. Measured across all eight commands on both the immediately-satisfiable and the parked-then-woken path — sixteen cases, sixteen losses; the incremental AOF held the sixteen setup records and not one pop. A queue consumer on BLPOP — the overwhelmingly common use — redelivered every message it had already consumed after a master restart or on failover.

The propagated record is the SYNTHESISED non-blocking sibling (BLPOP -> LPOP, BZPOPMIN -> ZPOPMIN, BLMOVE -> LMOVE, …), never the command itself: a replica applying a literal BLPOP would park its apply loop and an AOF replaying one would stall recovery. It is derived from the REPLY, not the arguments — a blocking pop takes many keys and only the reply says which one served — which also makes the immediate and parked paths one case rather than two, since both converge on the same reply frame. For BLMPOP/BZMPOP the record carries the count ACTUALLY popped, never the requested one: a COUNT 10 that popped 3 must not replay as 10 against a replica that has since received more elements.

moon already knew this rule and already implemented it — but only for the queued-in-MULTI path (blocking_txn.rs). The A/B that localised it: same server, same command, only MULTI differs — BLPOP standalone 0 logged nothing, MULTI / BLPOP intxn 0 / EXEC logged LPOP intxn. Wired at both runtime handlers, beside the moon#644 tracking invalidation that fixed the same structural gap on the same path. Non-writes (timeout, WRONGTYPE, a miss) still reach neither plane, and the moon#539 phantom-key guard is unaffected — both verified against the AOF bytes.

  • stream: every rejected XADD ID left a phantom stream in the keyspace that was charged and never persisted (#823). db.get_or_create_stream(key) ran BEFORE the ID was parsed, so it inserted the entry, charged entry_overhead and burned a birth version; the ID errors then returned, and propagation is gated on the reply not being an error, so nothing reached the AOF or a replica. Measured: all five of bogus, 0-0, 1-1-1, abc-1 and -5 created the key, and 3,000 rejected XADDs cost 1.46 MB that nothing credits back — an unbounded memory-growth primitive available to any client, through a command that only ever answers an error. Unlike the Lua case above it needs no odd frame: plain redis-cli reaches it. Real Redis parses the ID first and creates nothing.

Fixed by resolving the ID before the key can be created, peeking last_id from the existing stream (0-0 when absent). StreamData::validate_explicit_id now delegates to a free validate_explicit_id_against(last_id, id) that the pre-check also calls, so the two cannot drift apart. * still resolves after creation — it reads the shard clock and cannot fail.

  • scripting: a non-string Lua argument reached command argv, and nine commands wrote part of a command before refusing it — silent data loss across restart, and permanent replica divergence (#823). redis.call/redis.pcall converted a Lua nil, boolean or table argument into Frame::Null/Frame::Integer and handed it to dispatch. No wire client can put those shapes in an argv, so command parsers do not expect them: extract_bytes returns None, and HSET, HMSET, LPUSH, RPUSH, LPUSHX, RPUSHX, ZREM, MSET and MSETNX each discovered that from INSIDE their mutation loop and returned an arity error with part of the command already written.

That alone would be the memory-ledger drift of #814. It was worse, because propagation is gated on the reply not being an error: the partial write was applied on the master and never appended to the AOF and never sent to a replica. Measured, single shard, --appendonly yes:

EVAL "return redis.pcall('LPUSH','mylist','a','b',true)" 0
  -> ERR wrong number of arguments for 'lpush' command
LLEN mylist = 2          <- the error was a lie; two elements are resident
[restart]
LLEN mylist = 0          <- gone. An unrelated SET in the same session survived.

EVAL is how this was found, but it is not the only trigger. validate_frame (src/protocol/parse.rs) accepts +, -, (, :, , and # as top-level array elements, so a RAW WIRE client reaches the same code with no Lua at all: *5\r\n$4\r\nMSET\r\n$3\r\nwa1\r\n$1\r\n1\r\n$3\r\nwb1\r\n:7\r\n is delivered to mset, where real Redis answers ERR Protocol error: expected '$', got ':' and closes the connection. That protocol-layer divergence is filed separately and is NOT fixed here — which is precisely why the per-command validation matters independently of the Lua boundary. Verified: after this change the wire probe above leaves nothing behind.

Real Redis 8.6.1 refuses the Lua call outright — ERR Lua redis lib command arguments must be strings or integers — writing nothing; moon now returns that same error from the same boundary.

Fixed in two independent layers, because they protect different things. lua_arg_to_frame refuses every non-string, non-number argument, which is Redis parity and closes the whole class at the one place that should never have let the shape through. The nine commands additionally validate their whole argv before the mutation window opens, so the keyspace is safe even if some future caller reintroduces the shape. Mutation-tested both ways: with only the boundary reverted, zero partial writes occur and only the error message regresses; with both reverted, all nine commands write and the restart check reports nine divergences.

The old fall-through was reasoned about — the comment argued a nil/boolean argument is un-nameable in a KEY position and the ACL walker denies it. That is true, and it is why this survived. It says nothing about the VALUE position, which is where the writes happened.

Found by the adversarial review of #820, which observed that #814 was one instance of a class.

  • storage: ZADD and GEOADD stranded their memory charge on any argument error, driving used_memory monotonically to zero (#814). Both commands parsed and mutated in the same loop and returned on the first bad argument from inside the table_before … db.charge_memory() window. Members already inserted stayed in the keyspace with their charge never applied.

It compounded through DELETE, which is what made it unbounded rather than a one-off under-count: db.remove credits an entry_overhead recomputed from the CURRENT value, so a delete afterwards credited back memory that was never charged. Measured, 300 unrelated keys held constant, rounds of 50 x (partial ZADD + DEL):

floor with 300 held keys = 6,487,257
  after churn round 1: used_memory=5,602,457   drift=  -884,800
  after churn round 2: used_memory=4,717,657   drift=-1,769,600
  after churn round 3: used_memory=3,832,857   drift=-2,654,400

-884,800 B per round = exactly 50 x 17,696, monotone and unrecoverable, and driveable by any unprivileged client. used_memory reaches 0 with an arbitrarily large keyspace resident, at which point --maxmemory never fires. A single 100-pair ZADD with a bad tail billed 3,846 B — the empty container cost alone — against 21,541 B of real content.

Fixed by validating every pair/triple before touching the keyspace, which is also Redis parity twice over: real Redis's ZADD is all-or-nothing, and it does not create the key when the command errors. moon previously did both. Two passes rather than a parsed Vec, because src/command/ is a no-allocation path and parsing an f64 twice is far cheaper than the B+tree insert it guards.

Found by a post-hoc adversarial review of #810, which had been admin-merged with no human reader. ledger_consistency_788 drained every container but never exercised an error-mid-command path, which is why it was green; it now has four cases that do, including the delete-churn compounding.

  • shard: COPY and BITOP were silently lost across restart at --shards 1 (data loss). coordinate_copy and coordinate_bitop each opened with a num_shards == 1 fast path that returned run_local(..) directly — "zero coordinator overhead on the 1-shard hot path". run_local mutates the keyspace and returns; it takes no aof_pool and no repl_state, so it appended nothing to the AOF/WAL and issued no replication LSN. The write was acked, read back correctly, and was gone after a restart.

This violated a contract documented 300 lines above it in the same file, on persist_local_leg: "the coordinator's in-process local legs (run_local, ..) MUST call this or their writes are lost on restart while the remote legs survive."

Measured at --shards 1 --appendonly yes --appendfsync everysec, isolated per-runtime binaries, same host and script:

runtime COPY destination lost after SIGKILL + recovery
monoio 0 / 6
tokio 6 / 6

On tokio, after SET k / COPY k k2 / SET k3 the AOF is 55 bytes holding only the two sets, and recovery logs replayed 2 AOF commands for three writes. BITOP behaves identically. The monoio connection handler appends these commands on its own path, which is why the same probe is 0-of-6 there and why only the tokio CI leg ever saw it.

Fixed with run_local_persist: dispatch locally, then honour the same persist_local_leg contract the multi-shard same-owner branch already honours via run_on_owner_persist. COPY persists only on :1 (:0 is a refusal that writes nothing); BITOP persists on any non-error, since it always writes DEST (SET, or DEL when the combine is empty).

  • test: the local-leg durability suite could not see the bug it was named for. tests/coordinator_local_leg_durability.rs already contained copy_dst_local_leg_persists_across_restart and had been green throughout, because const SHARDS: u32 = 4 scoped every case to the one shard count where the coordinator routes through run_on_owner_persist and the bug cannot occur. That constant is now SHARD_MATRIX = [1, 4] and all seven cases run at both counts (14 tests). With the fix reverted, exactly copy_dst_local_leg_..._s1 and bitop_dest_local_leg_..._s1 fail while all twelve others pass, including both _s4 controls.

  • test: the PERF-08 single-probe timing net was measuring a page-fault artifact, not the optimisation (moon#789). tests/perf_v0112_insert_or_update_single_probe.rs timed the legacy get_mut+insert control and the fused insert_or_update back to back in one process, always control first, and kept the best of five ratios. Its comment claimed runner noise "can only inflate the ratio toward 1.0, never deflate it". That was false. The first timed loop in the process takes ~5,000 minor page faults that no later loop takes (measured: minflt = 5061 for loop #1, 0 afterwards, on both arches), so the control was handicapped by ~27% on rep 0 — and best-of-K deterministically selected that rep. Swapping the loop order moved the artifact onto the test loop instead (rep 0 ratio 0.779 → 1.190), proving it positional.

Consequences, both measured on GCE (t2a-standard-8 Neoverse-N1 and c3-standard-8 Xeon 8481C), 1M keys, warmed, isolated, order-alternating:

  • The aarch64 pass was the artifact. Warm the process up and the old mixed workload gives a median of 0.975 on aarch64 — above its own 0.95 threshold. The regression net had been green on aarch64 for the wrong reason since April 2026, so a real regression could have landed unnoticed.
  • The wall-clock win is arch-specific. On the miss path — the only path PERF-08 changes, compared at identical work — the fused probe is 0.893× (≈11% faster) on aarch64 and 1.166× (≈17% slower) on x86_64, despite issuing strictly fewer SIMD probes. This is codegen, not extra work; tracked in moon#789. Database::set still issues one probe everywhere, so nothing is functionally wrong.

A first fix (warm up, alternate order, take the median of whole-loop timings) removed the bias but not the variance, and false-failed on a 12-core macOS host under parallel-build load. Measuring the A/A noise floor — the same legacy loop timed against itself, where the true ratio is exactly 1.000 — showed why a median could not save it:

estimator host condition A/A min..max spread
whole-loop t2a aarch64 idle 0.973..1.046 0.073
whole-loop t2a aarch64 16 spinners / 8 vCPU 0.880..1.179 0.299
whole-loop c3 x86_64 idle 0.899..1.045 0.146
whole-loop c3 x86_64 16 spinners / 8 vCPU 0.559..1.795 1.236
chunked t2a aarch64 16 spinners / 8 vCPU 0.950..1.057 0.107
chunked c3 x86_64 16 spinners / 8 vCPU 0.867..1.093 0.227

Under load a whole-loop estimator returns ratios from 0.56 to 1.80 on identical code — so the previously committed x86_64 ceiling of 1.30 was not safe either; it merely happened to pass on an idle host. The fix is a better estimator, not more reps: the fill is split into 64 slices and the two sides alternate slice by slice, so a scheduler steal lands inside both sides and cancels in the paired difference. That is 3x tighter on aarch64 and 5x tighter on x86_64 under load, and its A/A ratio centres on 1.00 (0.997 / 0.998), so it is unbiased rather than merely quiet.

Validated by running the real test 8 times per arch, idle and under 2x CPU oversubscription: 16/16 green, aarch64 medians 0.873-0.894 (threshold 0.95), x86_64 medians 1.110-1.156 (ceiling 1.30). The medians barely move between idle and loaded, which is the paired estimator doing its job.

PERF-08's claim moved off the wall clock entirely. Two new deterministic tests count control-byte group scans and assert the fused path scans strictly fewer on a miss than get_mut + insert: one at the Segment::insert_or_update_at level and one at the DashTable::insert_or_update level that Database::set actually calls. The second exists because the first would not notice DashTable::insert_or_update being reimplemented as get_mut + insert on top of a healthy segment helper — the very regression PERF-08 prevents. Both are exact, arch-independent, load-independent and run in microseconds; forcing the PERF-09 fallback scan makes them report "fused scanned 6, legacy scanned 5" and fail on both arches.

To make that assertion possible the #[cfg(test)] SIMD-probe counter was extended from insert_or_update_at to find and find_free_slot_in_group via a thread-local, with a #[cfg(not(test))] #[inline(always)] empty no-op so it compiles to nothing in production — it sits in the hottest probe loops in the codebase. The old probe_count <= 6 bound could be satisfied by re-implementing insert_or_update_at in terms of find, because find was not counted; it is tightened to the true <= 4 and joined by the two cross-pattern comparisons.

With the claim held deterministically, both arms of the timing test are now blow-up ceilings (aarch64 1.02, x86_64 1.30) rather than one floor and one ceiling. The earlier asymmetry — asserting a win on aarch64, the arch that runs ci-local.sh's VM suites and a hosted macOS leg that cannot be validated from Linux — put a wall-clock floor exactly where a loaded runner does the most damage. The ceilings still catch what the counters cannot: same scan count, more time per scan (a lost #[inline], an added copy). The win itself is measured, documented and printed, but no longer asserted by a clock. The PERF-08 speed claim in Database::set's docs is corrected to say the probe-count reduction is universal and the wall-clock win is not. - storage: used_memory under-reported every container type, up to 15.2x (#788). The ledger behind the global --maxmemory gate and the per-db quotas was built from per-element constants that did not correspond to any real allocation. Measured on Linux (GCE c3-standard-8, 8 shards, 50k keys, RSS growth over what the ledger charged): a one-member sorted set read 15.23x low, a 64-member set 4.66x, hashes and lists 1.2–2.1x.

What that cost the gate, stated precisely rather than as the worst case: the eviction budget is already divided by maxmemory_footprint_correction, a measured RSS-over-accounted ratio republished once a second — so an under-report inside that correction's range was absorbed. The correction is clamped to 8.0. A 15.23x under-report is not, so a sorted-set workload could hold roughly 1.9x the configured --maxmemory before the gate fired, and every under-report additionally overshoots by its own factor within the correction's 1-second staleness window. Per-db quotas and the used_memory an operator reads in INFO have no correction at all and were wrong by the full factor.

Root causes, all four of them accounting that never matched an allocator:

  • zset_member_cost billed member.len() + 120 for a SortedSetBPTree whose arena slot is 784 bytes and whose minimum allocation is four of them — an empty B+tree already owns 3584 bytes. tree_mem = tree.len() * 80 charged per ENTRY for a cost that is per NODE.
  • set_member_cost = member.len() + 24 modelled neither of the IndexSet's two tables (a 40-byte-per-slot entries Vec and a power-of-two hashbrown index).
  • hash_field_cost_len and list_elem_cost billed element buffers at len rather than at the jemalloc size class actually handed out, and understated the slot (24 where the tuple is 32, 32 where Bytes is 40).
  • The Box<RedisValue> behind every collection value — 128 bytes on every hash, list, set, sorted set and stream key — was never billed at all.

Fixed by sizing every charge from the allocator's own shape, in a new storage::mem_size module (jemalloc size classes, hashbrown's power-of-two 7/8-load table, Vec doubling). Where a container exposes an O(1) capacity() the table is charged EXACTLY, snapshotted before/after each mutation — the same pattern the listpack and intset paths already used. Where no snapshot exists at every site the table share is folded into the per-element constant as the growth-cycle average, rounded up; a single linear constant cannot track a doubling allocator, so that term oscillates between 1.14x and 2.29x of the slot and is deliberately biased to over-report. Over-reporting wastes headroom; under-reporting is what makes the gate fail open. Nothing added to the write path allocates, locks, or walks the heap.

Two mutation sites charged nothing at all and are now covered: GEOADD (a 200-member geo key moved the ledger by 131 bytes against 26,931 bytes real) and the vector-search session recorder.

Three further defects surfaced while testing the fix, each stranding bytes that nothing gives back:

  • A hashbrown capacity() SHRINKS as entries are erased (56 → 21 over a 50-member drain), so skipping the table adjust on the "container went empty" branch left 1504 bytes charged per create/drain cycle.
  • SPOP and SMOVE skipped the member credit on that same branch, relying on a db.remove recompute that runs after the member is already gone.
  • A compact→full encoding upgrade (OwnedKind::upgrade) charged nothing for the size change, so LPOP credited full-encoding element costs against a listpack-sized charge — 2536 bytes over-credited on a 60-element list, an under-report of live memory.

Post-fix on the same Linux host (RSS growth / bytes charged; Redis 8.x on the same box for scale):

type elems before after redis
list 1/4/16/64 2.11 / 1.96 / 1.44 / 1.34 1.26 / 1.20 / 1.32 / 1.15 1.01–1.02
hash 1/4/16/64 1.71 / 1.59 / 1.39 / 1.22 1.34 / 1.11 / 1.14 / 1.12 1.01–1.02
set 1/4/16/64 3.51 / 3.79 / 4.35 / 4.66 1.16 / 1.50 / 1.54 / 1.56 1.03
zset 1/4/16/64 15.23 / 6.87 / 3.03 / 2.80 1.00 / 1.09 / 1.22 / 1.14 1.01–1.02

Sets remain the one type outside Redis's ~1.0–1.25x band; the residual is not yet attributed and is tracked separately. Plain strings gain only the size class of their heap buffer (+2 B/key on this workload); their RSS column could not be resolved on a contended host — interleaved base/fix pairs drifted further apart than the two binaries differ — so no string claim is made here.

MEMORY USAGE and the src/admin/ memory treemap use a different estimator (estimate_serialized_length) and are untouched, so their pre-existing divergence from used_memory widens rather than narrows.

  • storage: numeric strings with leading zeros or a leading + are no longer rewritten (data loss). SADD s 000000012345 followed by SMEMBERS returned 12345 — the caller's bytes were gone, not merely reformatted. The same applied to HSET/HGET and RPUSH/LRANGE, so it affected intsets, hash listpacks and list listpacks alike. Root cause: try_encode_as_integer (listpack) and try_parse_i64 (SADD) used a bare parse::<i64>(), which accepts 000000012345, +5 and -0, then stored the parsed value. Redis stores an integer encoding only when the decimal rendering reproduces the input exactly.

Two further sites were found by sweeping for the same predicate rather than fixing only the two the report named:

  • SetRef::contains parsed the query the same way, so SISMEMBER s 000000012345 returned 1 against a stored 12345 — a false positive that survived fixing the insert path.
  • OBJECT ENCODING reported int for "000000012345", where Redis reports embstr.

All five now route through storage::numeric::canonical_i64, which accepts a value only if itoa renders it back byte-identically. Verified 13/13 against a live redis-server 8.6.1, including guards that canonical values still take the compact encodings (OBJECT ENCODING on a pure-integer set is still intset). New listpack_roundtrip fuzz target, registered in both matrices in fuzz.yml; it reproduces the bug against the pre-fix code in ~1,500 execs ("00" → "0") and runs 600k execs clean against the fix.

This was also a merge blocker for the SADD listpack work: a mixed set such as SADD z 000000012345 abcdefgh is preserved on main only because it becomes a hashtable, and routing it into a listpack would have converted a latent encoding bug into a new data-loss path.

Performance

  • server: inline writes are no longer disabled by the default --disk-offload enable. can_inline_writes carried the term ctx.spill_sender.is_none(). --disk-offload defaults to enable, which spawns a per-shard SpillThread and hands every connection a live sender — so that term was false out of the box and the inline SET fast path never ran in the default configuration. (GET was unaffected: it is gated by can_inline_reads, which never carried the term.)

The term was a config predicate standing in for a state one. What a live sender changes is eviction ROUTING, and only that: with one, run_write_eviction_gate builds EvictionRun::async_spill, whose victims are handed to the SpillThread under --appendonly yes; the inline path can only build EvictionRun::plain, whose victims are DELETED. Inlining a write while eviction fires would silently substitute a drop for a spill.

Nothing else diverges. string::set's args.len() == 2 fast path and the inline path build the same Entry, queue the same set keyspace notification, and call the same Database::set — which is where the cold-tier obligations live (spill_inflight_forget retires an in-flight spill payload; the Updated arm drops a stale cold_index shadow). Neither path consults or promotes the cold tier on a write.

So the gate now enforces the actual invariant — the inline write path may run only when eviction provably will not fire — instead of a proxy for it. The lock-free inline_write_can_skip_eviction pre-gate the path already ran is hoisted above the point where the command bytes leave read_buf, so under pressure with a live sender the path returns "not handled" with the buffer byte-for-byte intact and generic dispatch executes the write with the spill-aware evictor. Bailing after the split would have consumed a command nothing then ran — a silently lost write — so the ordering, not a comment, is what enforces it. Cost on the hot path is one bool parameter and one predictable branch; no lock, no allocation.

Deliberately not used as the condition: disk_offload_spill_inert() (a startup-config predicate that is true for exactly the --appendonly no benchmark shape and false for the durable production one — and unsound across a restart that changes appendonly on a dir holding cold data), and maxmemory_is_set() (moon's auto-maxmemory default is 75% of RAM, so it is true out of the box and discriminates nothing).

Measured on the GCE x86_64 Linux host (8 vCPU), --shards 1 --appendonly no --protected-mode no --disk-free-min-pct 0, disk-offload left at its default, redis-benchmark -n 400000 -c 50 -P 16 -r 100000 SET k:__rand_int__ v after a 100k warm-up, A and B legs interleaved, 5 reps:

rep before after ratio
1 770,713 1,384,083 1.80x
2 760,456 1,369,863 1.80x
3 754,717 1,360,544 1.80x
4 759,013 1,369,863 1.81x
5 766,284 1,365,188 1.78x

Mean 762,237 -> 1,369,908 ops/s, 1.80x, no overlap between the two sets. Under --appendonly yes (5 further interleaved reps): 620,777 -> 1,266,722, 2.04x. The mechanism is confirmed, not inferred: moon_dispatch_path_total{path="local_inline"} reads 0 on every before-leg and exactly 500,000 (100k warm-up + 400k measured) on every after-leg. HSET has no inline path and is unaffected.

A first measurement run on the same host read 351,518 -> 614,799 (1.75x) with a 15-minute load average of 1.89 — the host was not idle, and a follow-up probe on the quiet host read 2.5x the absolute throughput for the same command. Interleaving preserved the RATIO across both runs; only the table above, taken with the load recorded per leg (0.38-0.67) and zero leaked server processes, is quoted as the absolute number.

Known consequence, fail-loud not silent. A faster write path can outrun the AOF writer. On --appendonly yes at --shards 1, a sustained pipelined write burst now reaches AofWriterPool's channel bound often enough to surface -MOONERR AOF backpressure: write applied in memory but not queued for persistence — the existing PR #211 behaviour, which answers an error rather than a lying +OK. Observed on roughly 1 test run in 3 locally, and once as an aborted redis-benchmark leg on the Linux host (0 of 16 legs on an idle host). It is not new code and not data loss, but it is newly REACHABLE, and an operator running --appendonly yes at this throughput may see it. The AOF writer's capacity, not this gate, is the thing to raise.

New suite tests/inline_write_spill_gate_660.rs, fifteen tests over --shards 1 and --shards 4, resting on the fact that evicted_keys (key left the keyspace) and spilled_keys (key moved to disk, still readable) are never both incremented for one victim. Widening the gate WITHOUT the bail-out turns it red at 1,093 keys deleted where they should have been spilled, spilled_keys flat at 0; restoring the old gate turns it red with no connection inlining at all, and takes every gate-term test down at its own vacuity control.

Review follow-up, all of it driven by findings rather than by design. A test-integrity pass showed the two most recent fixes — refreshing the shard database's cached clock before an inline write, and counting inline writes in total_commands_processed — had shipped with NO guard: deleting both left the whole suite green. GROUP 7 covers them (OBJECT IDLETIME after a 5 s idle; the command counter against 200 inline SETs). G1 gained an inline-path CONTROL, having previously passed unchanged with inline writes disabled outright, and spill_sender_active: true — the operand that makes the bail-out fire — gained its first unit test; every existing one passed false. The unit tests that drive try_inline_dispatch are now serialised, because the CLIENT PAUSE guard mutates process-global state that the other twenty-one read.

Two claims were WITHDRAWN rather than defended. G2's slip bound now applies at --shards 1 only: the elastic budget lets a lone hot shard borrow its idle siblings' headroom, so at --shards 4 it spends most of a window legitimately under budget and inlines most of it — measured on the Linux gate at 7,719 of 8,000 writes, with at most 15 plain drops against 156 spills. That is the pre-gate working. G2's safety assertion, the one that guards against silent data loss, still runs at both shard counts. And G3 is documented as what it is: remove_cold_only is unreachable from try_inline_dispatch, so G3 reddens identically on merge-base and is a cold-plane guard riding this fixture, not evidence for the inline change.

Two harness defects surfaced by the same gate are fixed here: spawn_moon reserved its admin port OUTSIDE the retry loop (a held admin port made moon exit at start-up while spawn_listening — which polls the child once, moon#811 — handed back the corpse), and the crash-restart leg had no readiness check at all, only Client::connect, which proves a listener accepts and, under SO_REUSEPORT, not even that the peer is the process just spawned. Both now wait for a real +PONG.

A term had to be ADDED, and two more are covered for the first time, because this change is what makes them reachable. can_inline_writes is a conjunction, and ctx.spill_sender.is_none() was false in the shipped default — so the whole conjunction was false, the inline write path never ran, and every other term in it was dead code unless an operator passed --disk-offload disable.

The added term is !conn.in_cross_txn(), and unlike the rest this defect is introduced by this change, not merely exposed by it. Inside an open TXN the generic write leg captures an undo record (txn.kv_undo.record_update) and a write intent (s.kv_write_intents.record_write) before dispatching; try_inline_dispatch does neither. Without the undo record TXN ABORT restores nothing, and without the write intent the MVCC snapshot-visibility filter cannot hide the uncommitted value from a foreign transaction. Measured on the release binary at --shards 1, stock config, one connection — SET k original; TXN BEGIN; SET k modified; TXN ABORT; GET k answers "original" with the term and "modified" without it, the in-TXN SET having been inlined (local_inline +1). An acked TXN ABORT that rolls back nothing. Found by security review of this branch before it was opened, and guarded by the two g6_* tests.

A second obligation had to be added for the same reason, and this one was caught by the existing suite rather than by review. segment_stall::stall_refusal — the only producer of -MOONERR memfull, the MA12 disk-free refusal and the moon#718 segment-stall refusal — has exactly two call sites, both in generic dispatch. try_inline_dispatch has none, so an inlined SET answered +OK for a write the server had already committed to refusing under memory or disk pressure. Merge-base 7678156f passes tests/mem_watchdog.rs and tests/compaction_escape_hatch_718.rs; this branch failed all three of their cases until the inline path grew a is_any_write_stall_active() bail-out next to the eviction one. It BAILS to generic dispatch rather than answering the error itself: stall_refusal exempts the commands that are a stall's own remedy, and re-deriving that here is the drift the shared helper exists to prevent.

Adversarial review then found two more, and a performance review two beyond that. Same class every time: an obligation the generic leg carries that the inline path does not.

  • CLIENT PAUSE was bypassed. check_pause is consulted once per frame in the generic loop, which sits BELOW the inline block — and that block continues when it consumed the buffer. Measured on one binary under CLIENT PAUSE 3000 WRITE: inline SET returned in 0.027 s, generic SET (MONITOR attached, forcing the slow path) in 2.999 s, HSET in 2.002 s. The pause worked; the inline leg escaped it. This silently voids the one guarantee the command exists to provide before a failover or backup.
  • The -LOADING gate was bypassed. Across a restart with a 40k-document FT index, 6/6 probes while loading:1 answered +OK here and -LOADING on main. SET touches no index, so this is contract violation rather than corruption — but a client keying "server ready" on -LOADING proceeds against a half-recovered server.
  • Inlined commands were invisible to total_commands_processed — 200 plain SETs moved it by 0. Not merely an INFO inaccuracy: this_thread_commands is the adaptive idle park's (#373) activity signal, so a shard serving nothing but inlined commands read zero commands/tick and could be classified IDLE under full load.
  • The shard's cached clock was never refreshed on the inline path. The generic leg calls refresh_now_from_cache once per batch; the inline path called set_last_access(db.now()) against whatever the stored clock last held. Measured: idle 10 s, then SET stale:k v immediately followed by OBJECT IDLETIME stale:k answered 18. Under allkeys-lru every inline-written key carried a frozen stamp, degrading victim selection exactly under memory pressure. Fixed at the same per-batch cadence the generic leg uses.

Scope of the win, measured rather than assumed. Interleaved A/B against 5092e4db in the shipped default, with moon_dispatch_path_total proving each leg took the path it claims (noise floor 0.934 A-vs-A):

config ratio
--shards 1, P=16 1.84–1.97x
--shards 1, P=16, --appendonly yes 1.79x
--shards 1, P=1 0.97 — no win
--shards 4, P=16 0.996 — no win

The multi-shard result is structural, not tuning: try_inline_dispatch returns 0 on a remote key and the loop BREAKS, so one remote key in a batch ends inlining for the rest of it. At --shards 4, P=16 only 2.0% of writes inline (10,400 of 520,000), matching the (1/N)/(1-1/N)/P prediction of 2.08%. This change helps --shards 1 pipelined workloads and is neutral elsewhere — it is not a general throughput win, and the tuning guide's --shards 4 recommendation for 8+ connections lands exactly where the benefit is absent. (macOS A/B; every published figure must come from Linux.)

The two previously-dead terms had no test at all, and both fail dangerously:

  • !fanout_hint_active() — deleting it makes a master with an attached replica ack 50 plain SETs with +OK and deliver none of them. That is not a hypothetical: the task #34 comment above the gate records the same bug being found and fixed once already, in the only configuration that could then reach it.
  • !is_replica — deleting it makes a client SET against a read-only replica answer +OK and land. try_inline_dispatch has no read-only guard of its own; the -READONLY error is produced exclusively by generic dispatch, so this term is the only thing enforcing it.

The suite also now carries a crate-level #![cfg(feature = "runtime-monoio")]. record_dispatch_local_inline has exactly one production call site, in the monoio handler, so under a tokio build the local_inline counter is permanently 0 and every control block in the file would have failed. Gating the suite rather than each assertion is deliberate: a test that cannot observe the mechanism it asserts on is not a weaker guard, it is a false one.

  • storage: the listpack scan no longer allocates once per element walked. Every listpack lookup decoded each entry it passed into an owned ListpackEntry — copying string payloads with to_vec(), and calling encode_backlen() purely to read .len(), which heap-allocated a Vec on every entry decoded in nine separate decode arms — then compared it and dropped it. One HSET missing on a 128-field listpack ran on the order of several hundred malloc/free pairs to answer one lookup, inside src/command/ where hot-path allocation is forbidden.

A borrowing twin of the decoder (ListpackRef, iter_refs, iter_pair_refs, find_pair_index, contains_element, and a backlen_size() that computes the width by arithmetic) now serves the lookup paths: HSET, HMSET, HGET, ZSCORE on a listpack zset, and both hash-TTL field probes. Only the entry that actually matches is materialized. Integer entries still match only their canonical decimal spelling, rendered through itoa into a stack buffer.

Measured on GCE x86_64 (8 vCPU, --shards 1, redis-benchmark -n 150000 -c 50 -P 16), interleaved A/B against the parent commit:

HSET, fields before after speedup
8 545,455 641,026 1.18x
32 311,203 517,241 1.66x
64 225,225 414,365 1.84x
128 (threshold) 131,234 297,619 2.27x
192 (promoted to hashtable) 697,674 691,244 0.99x

HGET against a preloaded 121-field listpack improves 3.99x (141,376 -> 563,910) — the read path materialized both field and value for every pair walked. The 192-field row is the negative control: once the hash promotes and no listpack scan runs, the change does nothing. SET/GET/INCR/LPUSH/ SADD are unchanged at 0.99-1.02x. - storage: SPOP and SRANDMEMBER become O(1) in set size. SRANDMEMBER on a 100,000-member set ran at 506 ops/s against Redis's 129,032 — 255x slower, and perfectly linear in set size. Both dispatch paths were O(n): the mutable one deep-cloned the entire set (s.clone()) and then collected every member into a Vec to choose one; the read-only one skipped the clone but called SetRef::members(), which clones every member anyway.

Redis is flat at any size because dictGetRandomKey samples a random bucket. std::collections::HashSet exposes no way to address the i-th element, so the representation had to change: RedisValue::Set is now SetValue = IndexSet<Bytes>, and a new SetRef::nth(idx) addresses a member on all four set representations without materializing any of them.

Measured on GCE x86_64 (--shards 1, redis-benchmark -n 20000 -c 20), interleaved A/B against the parent commit:

set size before after speedup vs redis 7.4.2
100 125,000 125,000 1.0x 0.97x
1,000 39,526 125,786 3.2x 0.97x
10,000 5,057 125,000 24.7x 0.97x
100,000 505 126,582 250.9x 0.97x

Throughput is now flat in set size, as Redis's is. SPOP improves 27.1x at 10,000 members (1.05x Redis) and 272.4x at 100,000 (0.83x). Ordinary set operations are unchanged: SADD 1.01x, SISMEMBER 0.99x, SREM 1.02x, SCARD 0.99x.

Cost, disclosed: IndexSet keeps an entry vector plus an index table where HashSet keeps one table, so set memory rises +8.6% on large sets (95.7 -> 103.9 B/member, 20 x 50,000) and +12% on many small ones (766.1 -> 858.5 B/key, 200,000 x 5). Once the set-listpack encoding of #787 lands, sets up to 128 members will not use this representation at all, confining the cost to large sets — which is exactly where the 250x matters.

Removal moves to swap_remove (O(1), reorders) rather than shift_remove (O(n), order-preserving). Redis set iteration order is unspecified, and all 5,136 lib tests pass, so nothing depended on it. - storage: the block a container key allocates is priced by the variant it holds, not by the widest variant in the enum. CompactValue::from_redis_value stores a collection as a Box<RedisValue>, so the block charged to a container key was size_of::<RedisValue>() — the width of the enum's widest variant — no matter which variant the key actually holds. One variant set that width:

variant payload
HashListpack / ListListpack / SetListpack 24
String(Bytes) / List / SetIntset 32
Hash 48
Set (IndexSet) 72
SortedSet { members, scores } 72
HashWithTtl { fields, ttls, min } 104
SortedSetBPTree { tree, members } 128 <- sets the enum

A hash small enough to live in a listpack — the common case — was billed the 128 bytes that a large B+tree zset needs. The five fat payloads are now boxed, taking the enum from 128 B to 40 B and its jemalloc class from 128 to 48.

What that is worth, per container key, is not one number. Boxing a payload does not delete it: it moves it into a second block. Summing the blocks a key really holds:

key holds before after delta
HashListpack / ListListpack / SetListpack / SortedSetListpack 128 48 -80
SetIntset 128 48 -80
List (VecDeque, stays inline) 128 48 -80
Stream (already boxed; its 112 B block was never billed) 240 160 -80
Hash 128 48 + 48 = 96 -32
Set (IndexSet, 72 -> class 80) 128 48 + 80 = 128 0
SortedSet (legacy) 128 48 + 48 + 32 = 128 0
HashWithTtl 128 48 + 48 + 48 = 144 +16
SortedSetBPTree 128 48 + 80 + 48 = 176 +48

So the win is real and 80 B on the compact encodings, plain lists and streams — where the overwhelming majority of small keys live — and the two regressions land on the two heaviest variants, where 16 B and 48 B sit against a table that is already kilobytes. String(Bytes) is unaffected: a string value is never stored behind a Box<RedisValue> at all.

String(Bytes), the three listpack variants and SetIntset deliberately stay inline: strings are the hot path and the one dimension moon already wins on, and the compact encodings are where small collections live. HashWithTtl::min_expiry_ms stays inline too, so the "has any field expired?" fast path still reads a plain u64 with no pointer chase. test_hot_variants_stay_inline pins that at compile time by binding each payload to its exact unboxed type, and a const assertion plus test_redis_value_fits_48_byte_size_class pin the 48-byte ceiling.

The boxes are on the fields, not on newtype wrappers around each struct variant, so every existing match arm binds the same names and Box<T> derefs to T at each use; only construction sites changed. The cost is one extra malloc when a collection is promoted out of its listpack encoding — a cold, once-per-key event.

No on-disk or wire format changes: RedisValue is in-memory only.

  • storage: an empty DashTable no longer reserves 16 segments to hold one. SegmentSlab::new seeded its first slab at 16 segments, so DashTable::new allocated 16 x size_of::<Segment<CompactKey, CompactEntry>>() = 55,296 B to store 3,456 B. moon builds --databases (16) of these per shard at boot, all empty, so every shard reserved 884,736 B for segments that never exist on an idle server — measured with a counting allocator at 84% of all per-shard heap reservation, four times the entire SPSC mesh. The first slab is now sized to demand: 1 segment for new(), exactly dir_size for with_capacity() (previously that path also rounded up to the fixed slab and fragmented its segments across ~log2 slabs). The existing doubling growth is unchanged, so a table that fills sees the same amortised curve, and slabs are still never reallocated — segment pointers stay stable.

Measured, structural (allocation sizes; host-independent), for --appendonly no --disk-offload disable with default features:

per-shard term before after
one empty Database 55,432 B 3,592 B
16 Databases (per shard) 893,976 B 64,536 B
ChannelMesh::new (unchanged) 135,984 B 135,984 B
model total per extra shard 1,030,432 B 200,992 B

The RSS effect is not claimed here. Reserved bytes are an upper bound on resident — untouched pages of a fresh mapping never become resident — and moon's idle-RSS numbers must come from Linux, where jemalloc's background_thread exists and retention differs. tests/shard_idle_alloc_attribution.rs is the regression gate and prints the full attribution.

[0.8.8] — 2026-09-01

Changed

  • Every stored string value above the 12-byte SSO cutoff is 16 bytes/key smaller. The heap-string wrapper Box<HeapString> held a Vec<u8> — pointer, capacity and length, 24 bytes. jemalloc's 64-bit small size classes are 8, 16, 32, 48, 64, 80, ... (verified with nallocx): there is no 24-byte class, so every one of those wrappers was billed in the 32-byte class and 8 bytes of it was pure slack. HeapString now holds a Box<[u8]> — pointer and length, 16 bytes — landing exactly in the 16-byte class. The saving is flat, per key, on every value longer than 12 bytes, and it drops the capacity field so a value built from an over-allocated buffer can no longer strand its excess for the lifetime of the key. Nothing mutated the buffer's length in place — as_bytes_mut had no callers at all and now hands out &mut [u8], which also removes a latent hazard, since CompactValue::len_and_tag caches the length beside the pointer and an in-place resize would have desynchronised the two. estimate_memory (and through it used_memory, the --maxmemory gate and per-db quotas) was charging a hardcoded 32 and now reports the wrapper's real size.

Calibrated against the Linux per-key measurement of 2026-09-01 (moon --shards 1 90.9 B/key, Redis 7.0.15 103.2 B/key at 1M keys), a 16-byte key with a 64-byte value goes from 186.9 to 170.9 bytes/key against Redis's 168.4 — from 11% behind to 1.5% behind. The band it actually flips is 13..128 bytes at the sizes that sit on a jemalloc class boundary (13-16, 32, 45-48, 80, and everything from 128 up), because Redis's embstr (OBJ_ENCODING_EMBSTR_SIZE_LIMIT 44) packs robj and sds into ONE rounded allocation where moon pays two. Values just above a boundary (20-24, 40-44) stay 14-15% behind; closing those needs the wrapper removed entirely, which needs unsafe and has not been done. Analysis in tmp/WAVE3_M_MEMORY.md. Measured on Linux only: jemalloc on macOS lacks background_thread and retains freed pages.

Documentation

  • BENCHMARK.md §2.14 — moon --shards 8 vs Redis io-threads 8 on all three dimensions, and two retractions. Throughput is a win at pipeline depth on both architectures (aarch64 1.26x / 2.74x, x86_64 1.32x / 2.91x at p=8 / p=64, 8/8 families on x86) and a tie at p=1 — the rig is bimodal for Redis as well as for moon, so p=1 must be mode-matched and never averaged. CPU per op is a tie (10.55 vs 11.33 us; moon is 6.9% lower but that sits inside Redis's own 11.9% run-to-run spread, and the durable difference is stability: moon's CV is 2.0%). Memory is not won: 0.94x at 8-byte values, 1.16x worse at 64-byte values, 0.97x at 256-byte, 1.26x worse on idle RSS.

Two earlier claims are retracted in-tree, with their raw data kept. (1) "moon gains nothing from eight shards" (s8/s1 = 0.97x) was a harness artifact: redis-benchmark -t lpush|sadd|hset|zadd drives ONE literal key and -r randomises the element, not the key, so eight of twelve families were asked to parallelise a single key. Re-run with explicit __rand_int__ keys and a DBSIZE >= 50000 guard proven to fire, real scaling is 1.42x / 2.14x / 3.79x. (2) The "0.90x per-key memory win" measured redis-benchmark's default 3-byte value: both legs sat below their own arithmetic floor, which is impossible. Re-measured across 8/64/256-byte values under a key + value + 24 floor check — verified to reject both historical numbers before being trusted — the win is a band, not a trend, and it exists only below the 12-byte CompactValue inline cutoff.

Performance

  • Ordinary keyspace commands skip the connection handler's name-dependent intercept gates. Every command walked 26 gates before dispatch; only six carried a cmd_len pre-guard and twelve are async fns whose futures were built and polled purely to return false. CommandFlags::NO_INTERCEPT (free bit 14 of the existing u16) is read from the same COMMAND_META entry the arity check already consults, and replaces 21 of those gate calls with one bit test. The gates and their order are unchanged — the ordering comments in handler_monoio/mod.rs record real bugs (ACL above privileged intercepts, workspace rewrite above key-readers, MULTI queue below ACL), and reordering to save a branch would re-open them. The four state gates (ACL, cluster routing, readonly, disk-full) apply to every command regardless of name and are deliberately not guarded. The flag's sense is inverted on purpose: unmarked means "take the slow path", so adding a command or an intercept can cost throughput but never correctness. tests/intercept_flag_drift.rs guards the other direction and was verified to FAIL when WAIT is marked. Measured on Linux/ARM (GCE t2a, two-box, --shards 1, 6 interleaved reps, arms alternating every rep, noise floors 1.3-5.2%): INCR +11.5% at p8 and +16.8% at p64, HSET +6.5/+12.0/+11.9% at p1/p8/p64, LPUSH +8.7/+10.9% at p8/p64, SADD +7.6/+11.0% at p8/p64 — geometric mean +3.9% over all 18 cells, and the base/ni distributions do not overlap at p64. GET and SET are unchanged (they are bimodal in both arms; the inline byte path is what serves them). This is larger than the 3-5% originally predicted, but it does not close the 0.43x deficit recorded in BENCHMARK.md §2.12.
  • A flat multibulk is now parsed in one pass instead of two. parse() ran validate_frame over the request bytes to compute the frame's total length — computing every argument offset on the way and throwing all of them away — and then ran parse_frame_zerocopy over the same bytes to re-derive exactly those offsets. For a top-level *N of $-bulks, which is the shape of essentially every client command, a new scan_flat_multibulk records each argument's span into a stack SmallVec while it validates, and the Frame is built straight from the spans: one memchr walk and one strict_atoi per token instead of two, and no recursive per-element call. Everything else — RESP3 containers, nested arrays, null bulks ($-1), null arrays (*-1), inline commands, and reply parsing — is untouched and still takes the two-pass path; the fast path declines on anything it does not handle exactly.

Correctness is asserted differentially, not argued: parse_reference_two_pass (the old pipeline, compiled only under cfg(test)/feature = "fuzzing") is compared against parse() over a corpus of ~50 hand-picked cases and every truncation of each, under four ParseConfigs including degenerate limits, and across argc 0–20 × payload lengths 0–300. Five deliberate mutations of the scanner were confirmed to fail those tests, so they are not vacuous. A new resp_parse_fused fuzz target runs the same differential and is registered in both matrices in .github/workflows/fuzz.yml.

The reserve for the scanned spans is capped at buf.len() / 6, not at the count off the wire. Six bytes is the shortest an element can be, so the cap can never under-allocate a scan that succeeds — and without it *1048576\r\n, ten bytes against the default 1Mi max_array_length, would have reserved 8 MiB before the scan discovered the frame was incomplete. The two-pass path never had that amplification because it reaches FrameVec::with_capacity only after validate_frame has proved the whole frame is present.

  • benches/resp_parsing.rs gains the argc > 4 pair. FrameVec is Box<SmallVec<[Frame; 4]>>, so FrameVec::with_capacity(count) heap-spills past four elements and a *5 command pays two allocations where a *3 pays one. BENCHMARK.md attributes the SET k v (2.08x) vs SET k v EX 100 (0.87x) step to the inline byte path, but a second, independent step sits at exactly that boundary and nothing has separated them. parse_set_ex_5arg (*5) pairs with the existing parse_set_single (*3), and parse_hset_4arg / parse_hset_6arg are the clean control — same command, same work per argument, only argc differs, and the inline path never touches HSET. No numbers: these must be run on a Linux host.

  • INCR/DECR/INCRBY/DECRBY no longer allocate a String per operation. The new value was stored as Entry::new_string(Bytes::from(new_val.to_string())) — a heap allocation on the command hot path, which CLAUDE.md forbids outright — and CompactValue then copied the digits out of it and freed it again immediately: a counter of up to twelve digits inlines into the 12-byte SSO payload, so the allocation was never even the storage. Now itoa::Buffer formats into a stack buffer and the new Entry::new_string_from_slice{,_with_expiry} / CompactValue::from_slice constructors take it by reference. Unit-tested against i64::to_string at every length from 0 to 32 bytes and at both i64 extremes, on both the plain and the TTL-preserving arm — the arms differ, and the 12/13-byte SSO boundary sits inside the range an INCR can reach.

  • Database::set borrows its key instead of demanding an owned Bytes. Every use of key inside set was already by reference — spill_inflight_forget, entry_overhead, hash_expiry_index_note_value, CompactKey::from (which copies the bytes either way), ColdIndex::remove, and both expiry-index writers all take &[u8], and the Bytes was never moved anywhere. The owned signature forced a key.clone() at each write command: one shared_v_clone in and one shared_v_drop out, per command, producing nothing. set, set_string and set_string_with_expiry now take &[u8], which removes 32 Some(k) => k.clone() key extractions across the string, hash, list, set and sorted-set write families — the ten families Wave 0 measured as losing. It also deletes a real allocation (not just a refcount) from RESTORE, COPY, RENAME, MOVE, the cold-tier promote, WAL v3 replay and replication apply, all of which were building a throwaway Bytes::copy_from_slice(key) purely to satisfy the signature. Six call sites genuinely need ownership and keep their clone; the compiler identified them.

  • The monoio local write path no longer clones the whole reply to read one bit. After every local dispatch the handler built response_frame by cloning the DispatchResult's Frame, then used it for exactly one thing — matches!(response_frame, Frame::Error(_)) — and dropped it. response_frame had those two occurrences and no others. For an Array reply that clone is a fresh FrameVec box plus one Bytes refcount bump per element plus the matching drops, paid per command on the write path that the inline byte path never touches. Replaced by DispatchResult::is_error(), a borrow-only matches! over both variants. Semantics are unchanged by construction; the new method is unit-tested over Response/Quit × error/non-error.

Changed

  • The cross-shard read fast path ships enabled (--cross-shard-fast-path auto). At --shards 8 on a populated keyspace it now serves 100% of foreign reads on the calling thread and takes parks/cmd for GET at p=1 from 0.87336 to 0.00023 — same binary, one flag apart, measured from INFO stats (total_dispatch_cross_read_fast / total_dispatch_cross_spsc / total_remote_awaits_parked). docs/internal/cross-shard-cost-model.md prices a park at ~24.9 core-us and at 85% of p=1 cost, so this is the largest single lever on the cross-shard read path. Reads only — a cross-shard write still parks, unchanged at 0.875/cmd.

It shipped off because the only evidence was moon#768's -8.61% CPU/op and a doubled s8 p16 variance. Both readings came from a half-populated keyspace. The path declines a key that is not resident — dispatch_read cannot consult the cold tier, the moon#610 class — and a declined read falls back to the SPSC hop, so the measured "in-place rate" was tracking the benchmark's key hit rate, and a wandering hit rate is precisely the variance that held the default down. Reproduced against DBSIZE: 63,114 keys resident gives 62.9% in place, 98,169 gives 98.2%, 100,000 gives 100.0%. This also retires the standing question of why #768 measured 50.5% in place where the model predicted 87.5%.

auto declines where the path cannot fire — --shards 1 (every key is local) and the tokio leg (handler_sharded has no fast-path site, moon#776) — so the switch never reads as enabled on a leg that would ignore it. on forces it; off is the rollback and is pinned by l4_cross_shard_read_fastpath::the_fast_path_stays_dark_when_the_flag_is_off.

Documentation

  • First Moon-vs-Redis benchmark since v0.6.0, and it corrects the headline framing. BENCHMARK.md §2.12 records v0.8.7 measured on both GCE arches (7 interleaved reps, Redis 7.0.15 re-measured every rep, 0 failures). Moon wins GET (2.40× x86 / 2.29× ARM) and SET (1.78× / 2.02×) at P=64 and is at parity at p=1 — but every non-inlined command family (INCR/LPUSH/SPOP/HSET) runs 0.40–0.67× Redis at p≥8. The boundary is the two-command inline byte path, not the engine: SET k v runs 2.08× Redis while SET k v EX 100 — same work, one disqualifying option — runs 0.87×. §1, the exec-summary table and the README benchmark section now carry that scope. The p=1 --io-busy-poll-us 40 claim (1.65–1.66× x86) is annotated as requiring dedicated cores; on shared-tenant instances it measures 1.06–1.08× x86 and within-noise on ARM, because the contention governor self-gates.
  • §2.13: a 9–21% write-path regression since v0.6.0, in three steps. A 9-point rebuild sweep across the 326 commits in v0.6.0..d63ffcd8 shows flat throughput for a month then three separate drops; git bisect would have named the first and missed two thirds of the loss. GET is the only read in the grid and the only family that lost nothing. perf attributes +5.5 points of per-command cost to the dispatch/intercept region. Four candidate mechanisms were tested and refuted rather than published — the intercept chain shrank (28→26 gates), the cmd_len == 4 guard theory reverses under an INCR-vs-INCRBY probe, the with_shard_db fallback is never taken, and ShardSlice did not grow. Attribution is bounded by fat LTO absorbing inlined closures into the with_shard symbol.

  • The moon-dev recreate recipe now works end to end. docs/internal/orbstack-linux-parity.md was missing four packages that are each load-bearing (git, python3-redis, libicu-dev, cargo-nextest) and omitted the GitHub Actions runner entirely — it dies with the VM, and without step 5 the hosted monoio and client-compat legs queue forever with no error. Also records that the machine has now vanished five times, that ci-local refuses with exit 4 when it has, and that orb create ubuntu tracks the current release (24.04 when the doc was written, 26.04.1 now).

  • CONTRIBUTING.md fuzz and CI sections re-checked against the workflows. The target list said 7 (there are 20); "add the target to the matrix" was singular where there are two and a target missing from one silently never runs; the CI table claimed Fuzz (PR) blocks merge (it is opt-in behind the ci-fuzz label), that the nightly budget is 6h (5h — 6h hits the hosted job ceiling and loses the corpus), and that CodeQL blocks merge (its PR runs were removed in 2026-07). Adds the rule that cost two fuzz targets: never hand-copy a ShardSliceInit literal, build slices with test_support::make_init.

Fixed

  • *-\r\n would have parsed as an empty array on the new fast path. Found by resp_parse_fused within 90 seconds of its first run, before the change shipped. strict_atoi reads a lone - (and -0) as zero, while parse()'s is_null_multibulk gate keys on the raw byte buf[1] == b'-' rather than on the parsed count — so the two-pass path silently consumes *-\r\n and reports no frame, where a count-based check would have produced *0. The scanner now declines on the byte. No released version is affected; recorded because the shape (is_null_multibulk testing a byte, not a number) is a trap for any future fast path over the same bytes.

  • Fuzz crash reproducers are archived instead of discarded. fuzz.yml uploaded only fuzz/corpus/, so the minimal reproducer libFuzzer writes to fuzz/artifacts/<target>/crash-<sha> was destroyed with the runner on every finding — the recent term_fst_sidecar OOM had to be diagnosed from the single ASan line left in the log, and its reproducer is gone for good. Both the nightly and the ci-fuzz PR job now upload it on failure. Deliberately a separate artifact rather than a second entry in the corpus step's path:: upload-artifact@v4 re-roots a multi-path archive at the common ancestor, which would have made the next run's download land the corpus in fuzz/corpus/corpus/ and silently stop seeding while still reporting success.

Changed

  • ci-local.sh runs the two VM suites concurrently. They were sequential because parallel builds of two feature sets were expected to contend on memory and on the shared-volume virtiofs — a rule set before the 6-CPU cap and the moon#735 debug-info cuts and never re-measured since. Measured ABBA on moon-dev with warm target dirs, 2 reps per arm: sequential 469s/453s (mean 461s) vs concurrent 256s/256s (mean 256s), -44.5%, with zero flaky tests and zero retries in all four arms. The mechanism shows in the per-suite times — each suite slows only ~9% when sharing the VM (monoio 229 -> 250s, tokio 228 -> 249s) because neither saturates the 6 vCPUs; both spend most of their wall clock waiting on spawned servers, and overlapping those waits is the other half. CI_LOCAL_VM_SEQUENTIAL=1 restores the old order without an edit.
  • ci-local.sh: the macOS host tokio leg runs under nextest. It was the one leg still on a plain cargo test, which finishes each test binary before starting the next. Measured same-tree with target-tokio warm (the build was a 1.46s no-op, so this is execution time only): 264 test binaries, 561.5s sequential -> 148.4s under nextest (-73.6%), 5354 passed with no retries. Doctests moved to their own step, because nextest does not run them and the switch would otherwise have dropped that coverage silently -- ci.yml already splits them the same way.

Fixed

  • ci-local.sh refuses to start when the moon-dev VM is unreachable (new exit 4). The disk pre-flight deliberately steps aside when it cannot measure the VM -- an unreadable df must not ground a healthy run -- but the same branch also swallowed a VM that was gone. Measured 2026-08-31, the 4th time moon-dev vanished: the pre-flight printed ok (0s) for a machine that did not exist, all four VM legs failed at 0s, the run then spent 1184s on the macOS suite, and the verdict read "re-run failing suites in isolation" -- pointing at the tests rather than the missing machine. The two cases are now separated, and the refusal names the recreate recipe and --native as the meanwhile gate.
  • deserialize_term_fst_sidecar no longer pre-allocates from an unchecked term_count. The per-field term count is a full u32 read straight off disk and went to Vec::with_capacity unvalidated, so a corrupt or truncated sidecar asked the allocator for ~101 GB before reading a single term. read_bytes already refused to read past the end of the input; what it could not do is stop the reservation. The count is now bounded by the bytes actually remaining (6 per term, the tightest legal encoding) and fails closed with InvalidData, matching the decoder's documented contract. Found by the nightly term_fst_sidecar fuzz target, which has reported AddressSanitizer: out of memory on every run since at least 2026-08-26.
  • mq_registry_blob fuzz target compiles again. It hand-copied a 20-field ShardSliceInit literal, and that literal rotted twice without any gate noticing: ShardStoreMemory gained lua_vm and then pagecache, and separately databases became Arc<ShardDbSet> under the L4 read plane. The target has failed to build since at least 2026-07-17, running zero executions on every nightly since; fuzz/Cargo.lock was stale enough to still list stop-words, dropped from moon by the #690 stoplist fix. Nothing in CI compiles fuzz/, so none of it was visible.

Rather than patch the literal a third time, the target now shares moon's own shard::slice::test_support::make_init fixture, exposed through a new additive fuzzing feature that only fuzz/ enables. A new field on any of those structs now breaks the build in-tree, where a gate can see it. ShardStoreMemory also derives Default, and the two remaining literals construct through it.

Removed

  • .github/workflows/claude.yml. It triggered on every issue comment, review, review comment and issue open, then skipped unless the body contained @claude. Across the last 60 triggering events it ran 0 times and skipped 60. Recoverable from git history.

Added

  • CI type-checks the fuzz/ crate on every PR. Nothing else in CI compiled it, which is how two fuzz targets rotted unnoticed. Added to the Lint job, where it costs no PR wall clock: Lint ran 51s against Check's 585s in the same run, so it has ~9 minutes of slack before it could become the critical path.
  • D3 W1: try_foreign_db_write, the exclusive twin of try_foreign_db_read. Mutates a foreign shard's database on the calling thread under a non-blocking exclusive guard, or returns None so the caller falls through to SPSC. No production callers yet by design — a KV write also owes a WAL append, an AOF append, a backlog append, an offset advance keyed by the owner's shard id, and an ordered replica fan-out, and those must be enqueued inside the guard. Design: docs/internal/d3-concurrent-keyspace.md.
  • ShardDbSet::try_write_foreign. try_write's "try" is only its index bounds check — it calls write() and BLOCKS on a held database, which is exactly what a foreign thread must never do. The new method is a single non-blocking attempt, matching try_read.
  • --cross-shard-fast-path: serve a foreign shard's read without an SPSC hop (L4, S4).

A read whose key hashes to another shard previously always crossed that shard's SPSC channel. It can now be served on the receiving thread under a shared read guard on the owner's database — one CAS, no park.

docs/production-guide.md has documented this flag, an auto default, and three Prometheus metrics for some time. None of it existed: the dispatch site was disabled (handler_monoio/mod.rs, "ShardSlice is thread-local; foreign-shard data can only be read via SPSC hop"), the metrics had zero references in src/, and the server rejected the flag outright — an operator following that tuning advice got error: unexpected argument '--cross-shard-fast-path' found instead of a running server. The L4 plane removes the stated blocker, so this implements the flag under the documented name and corrects the section to match what ships.

Default off. Acceptance, two-box GCE ARM, same binary flag off-vs-on, ABBA-ordered, n=10 reps/cell: at --shards 8 pipeline 1, −8.61% server CPU/op (95% CI −11.55 … −5.67) and +12.03% throughput, cheaper in 10 of 10 reps (sign test p=0.002), with 50.5% of reads served in place. The --shards 1 negative control — where every read is already local — fires the path 0.0% of the time and shows −0.05% (CI −0.82 … +0.72), as it must. At pipeline 16 the effect is not measurable at this n and the enabled leg's variance doubles rather than shifting, which is why the default stays off.

The path also declines, falling back to SPSC, whenever:

  • the connection has in-flight remote work on that shard (read-your-own-writes, moon#507 / moon#512);
  • the command spans shards — placement is read from multikey_placement, the same function the router uses, so the two cannot drift (moon#592);
  • the command is multi-key (conservative for v1: an all-on-one-shard MGET can still have some keys cold, and the residency probe only covers the primary key);
  • the key is not resident, since the read-only path does not consult the cold tier (the moon#610 class);
  • the owner holds the write lock — the guard is attempted, never waited on.

Observability: moon_dispatch_path_total{path="cross_read_fast"} and INFO stats → total_dispatch_cross_read_fast. Measured at --shards 4, 200 scattered GETs moved the counter 0 → 147 (73.5%, the expected 3-of-4 foreign fraction) while total_dispatch_cross_spsc stayed flat.

Performance

  • Owner reads take a shared guard instead of an exclusive one (L4, S3).

with_shard_db_read was added by the L4 plane skeleton and then had zero callers — every owner read still went through with_shard_db, taking the db's write lock to serve a GET. This wires the four owner read sites (handler_monoio and handler_sharded, each a single-key read and the local part of a spanning multi-key read) onto the shared guard.

dispatch_read was already written for this access mode: its hot-key sketch uses a relaxed fetch_add and a try-lock that drops the sample under contention, documented as being so that "concurrent cross-shard fast-path reads never block here". The conversion is enforced by the type system — the closure goes from &mut Database to &Database, so anything needing mutation fails to compile.

No behaviour changes: an exclusive holder is replaced by a shared one on a thread that is the sole writer, so nothing that was previously serialised becomes concurrent yet. The win is unlocked by S4, which lets a foreign shard serve a read of this db instead of hopping to the owner.

This does not extend to the SPSC batch loop. That loop uses the full dispatch, which needs &mut Database; routing it through dispatch_read to justify a shared guard would reintroduce the moon#610 cold-tier read-bug class.

  • A spanning multi-key READ no longer cuts the pipeline batch (moon#513, A2a).

Since moon#512 fixed the moon#507 write-loss inversion, a multi-key command mid-pipeline forced the batch tail to defer one dispatch boundary. moon#721 (A1) took the case where one shard owns every key. This takes the genuinely SPANNING case for the two per-key decomposable READS, MGET and EXISTS.

Measured at --shards 4, 32 interleavings of SET,SET,MGET per shard pair, deferrals read from INFO stats (macOS, placement-only — no throughput number is claimed here):

                    before   after
  pair (0,1)         32/32     0/32
  pair (0,2)         32/32     0/32
  pair (0,3)         32/32     0/32
  pair (1,2)         32/32     0/32
  pair (1,3)         32/32     0/32
  pair (2,3)         32/32     0/32

The command is split per owner shard; each part is routed exactly like an ordinary single-shard command at the command's own position in the batch — appended to remote_groups[owner], or executed inline against the local slice — and the parts are folded back into one client reply at the drain. That removes the coordinator's second, immediate SPSC push, which was the moon#507 inversion, rather than merely declining to wait for it. Ordering is per-key: a key has exactly one owner, every operation on it in the batch lands in that owner's vector in loop order, and one thread executes that vector in order.

must_wait_for_pending_remote and the routing side now consult ONE function, multikey_placement, so the guard cannot say "safe" about a command routing still runs inline. Everything uncertain resolves to Coordinator and keeps waiting: workspace connections, an untrustworthy key mask, anything inside a cross-shard TXN, MSETNX/BITOP/ COPY (not decomposable at all), and MSET/DEL/UNLINK — decomposable, but splitting a write adds per-part AOF, replication, group-commit and tracking-invalidation work and is held back to A2b so a durability bug and a throughput change cannot bisect to one commit.

A part that fails for any reason — backpressure give-up, reply timeout, the shard's own error, or a slot nothing ever wrote — fails the whole reply rather than assembling a partial answer. Partial failure across parts is possible; coordinate_multi_key has the identical exposure today for these same commands, and they are reads, so nothing is applied either way.

INFO stats gains total_pipeline_multikey_fanout (moon_pipeline_multikey_fanout_total), one per fanned-out command. It exists because the deferral counter alone cannot separate a real fix from a dangerous one: a change that merely stopped waiting while the command still ran inline would drive deferrals to zero too, and that is moon#507 reopened. Mutation-tested — disabling the fan-out routing while leaving the guard relaxed passes the deferral assertion and is caught only by this counter (and, one test later, by an MGET reading stale values from its own batch).

Changed

  • Shard databases moved behind a per-(shard, db) lock plane (L4 step 2, moon#513). ShardSlice.databases changed from Box<[Database]> to Arc<ShardDbSet>, a registry of CachePadded<RwLock<Database>> built on the main thread before any shard thread spawns.

This step is behaviour-preserving and has no flag: every site that had exclusive access before has exclusive access after. It exists so a later step can let a foreign shard serve a READ of a key this shard owns without the cross-shard hop — and the park that costs. Measured on GCE t2a-standard-8 (aarch64, 8 vCPU), fitting per-command CPU against pipeline depth gives cost = 0.413 - 0.046*msgs/cmd + 2.488*parks/cmd (CPU%/kops): 2.49 per park, ~zero per message, with the park term at 85% of the p=1 cost.

Why per-(shard, db) rather than one lock per shard: a write to db 0 must not exclude a read of db 3, and no lock is taken on any per-key path — one per command, not per key. s8 x 16 dbs costs 8 KB of padding and nothing per key. Why locks at all rather than a seqlock or epoch/RCU: values are heap-owning (Bytes, HashMap, Listpack) and get_mut mutates them in place, so both alternatives need copy-on-write on the write hot path.

No unsafe is introduced. Database is Send + Sync, and the static L4_REGISTRY now pins that as a compile-time contract: adding a non-Sync field to Database breaks the build rather than silently making the plane unsound. The slice's !Send marker is untouched, and vector, text, graph and registry state never cross a thread.

Guards are handed out through FnOnce closures, so a guard cannot escape and cannot cross an .await — "never hold a lock across .await" is enforced by the type system here, not by review. A thread-local mask restores the loud failure the RefCell used to give: re-acquiring a database inside its own guard would DEADLOCK on a real RwLock, so it panics instead.

144 call sites across 20 files were converted. Four multi-database helpers (flush_every_database, rdb::save_to_bytes, redis_rdb::load_rdb, snapshot::shard_snapshot_load) became generic over Borrow/BorrowMut<Database>, which left every existing caller unchanged.

  • CI gained a runtime-tokio + text-index lint leg. handler_sharded is cfg(runtime-tokio), so the default leg never compiles it; the tokio leg drops text-index, so it never compiles the parts behind that cfg. Their intersection was covered by no gate, local or hosted. Measured: a deliberate type error at handler_sharded/ft.rs:498 left BOTH standing legs reporting zero errors. A real defect had been sitting in that blind spot.

  • New docs/internal/cross-shard-cost-model.md records the per-park cost model, the profile breakdown (83% of moon's user time is not Redis work), seven measured dead ends, and five retracted claims — so none of them are re-derived.

Fixed

  • Cross-shard fast path: a thread outside the registered database plane could serve a foreign read off the registry. try_foreign_db_read resolved shard_dbs(shard) from the process-wide L4 registry without first checking that the calling thread is part of that registry. init_shard publishes MY_DB_SET only when the registry holds the very same set the slice does, so a thread holding a slice built beside the registry would read the registry's Database while its own writes landed in a different one — stale data returned under a +OK, silently. The gate is now cached_db_set()? first; unregistered threads fall through to SPSC, which is what they did before S4 landed. Covered by foreign_read_refuses_from_a_thread_outside_the_registered_plane, which fails with Some(1) against the previous code.
  • ResponseSlot unit tests raced the process-global park counters. test_future_resolves_after_fill and test_concurrent_fill_from_another_thread poll a ResponseSlotFuture, moving REMOTE_AWAITS/REMOTE_AWAITS_PARKED, while the two counter assertions ran in parallel in the same binary. Both now take the same PARK_COUNTERS lock the counter tests hold.
  • The inline write path enforced a different maxmemory than every other dispatch path (moon#475).

maxmemory is scaled by the OS-footprint correction so it bounds what the process actually costs the OS rather than the allocator's accounting figure — a 2.3x gap on the instance that motivated PR #478, most of it swapped. evict_to_budget applied that divide. The lock-free pre-gate on the inline write path, can_skip_eviction, open-coded the budget and omitted it:

let budget = if elastic_budget > 0 { elastic_budget.min(mm) } else { per_shard };
estimated_memory <= budget          // no `/ footprint_correction()`

So one server enforced two effective caps at once. The pre-gate is reachable only from server::conn::try_inline_dispatch, so a plain SET was held to maxmemory while HSET and EVAL — which reach command::dispatch / dispatch_read and call evict_to_budget unconditionally — were held to maxmemory / ratio. At the measured 2.3x that is a 2.3x difference in the limit an operator configured, with no error and no log line.

Scope of the divergence, stated precisely: under an evicting policy the 100ms periodic tick still runs the corrected slow path, so the overshoot is a lag bounded by that tick, not unbounded growth. Under noeviction it is not covered at all, because the tick discards the result (let _ = evict_to_budget(...), src/shard/timers.rs:207) — an inline write that should have been refused with -OOM is silently accepted.

Both paths now call one effective_budget(elastic, mm, per_shard, ratio). They do not apply matching formulas, they execute the same function, so the pre-gate cannot drift from the limiter again. The pre-gate stays lock-free: the correction is read with the same cached Relaxed atomic load evict_to_budget already used (footprint_correction()), adding one load and no sampling — PR #510 moved that sampling off the write path after it measured -57% on SET at c=8 P=16, and this does not reintroduce it.

The gate had ~20 assertions and every one of them pinned the neutral correction 1.0, which is why this shipped. test_inline_pre_gate_agrees_with_slow_path_under_footprint_correction sweeps ratios across the clamped [1.0, 8.0] range and asserts the pre-gate's decision equals the slow path's for every combination of elastic budget and estimated memory. - Test harnesses reported spurious failures for any long-output command (set -o pipefail + grep -q).

grep -q exits on the FIRST match, closing the pipe. Under set -o pipefail that SIGPIPEs the writer, so the pipeline reports failure even though grep itself succeeded:

set -euo pipefail
if echo "$long_output" | grep -qvE "^(\(error\)|ERR )"; then PASS; else FAIL; fi
#    ^ SIGPIPE -> pipefail -> non-zero pipeline -> spurious FAIL

Any assertion whose output exceeded a pipe buffer failed regardless of what moon returned, and both match polarities were affected — a positive-match assert whose pattern hits early closes the pipe just as readily. The failure could only manufacture false failures, never false passes, so nothing was being hidden; but it silently understated the pass rate on every long-output command.

Confirmed against moon#683's remaining failures: SSCAN and COMMAND were both reported as errors while returning correct results (SSCAN -> the cursor plus all three members; COMMAND -> 3,948 lines across 273 commands). Both were preceded in the log by echo: write error: Broken pipe.

Fixed with a qgrep helper that drains stdin before matching, so the writer always completes and only grep's own verdict is reported. Applied to 156 call sites across test-commands.sh (128), test-consistency.sh (17), test-vector-clients.sh (6), gcloud-benchmark.sh (3) and audit-unwrap.sh (2) — every harness that runs under pipefail. Remaining piped grep -q uses match single-line output (PING | grep -q PONG) and cannot trip the bug.

scripts/test-commands.sh --skip-bench, same binary, before -> after: 517 total / 514 passed / 3 failed -> 517 total / 516 passed / 1 failed. The total is unchanged, so no row was skipped; the one remaining failure is ROLE, tracked separately as moon#536.

This is the same family as the set -u defect that had test-consistency.sh silently running about half its rows.

  • TXN ABORT left torn state after a multi-key write (moon#500).

The undo capture in both connection handlers recorded a single key per command:

} else if let Some(key) = shared::extract_primary_key(cmd, cmd_args) {
    match db.get(key.as_ref()).cloned() {
        None => txn.kv_undo.record_insert(key.clone()),
        Some(entry) => txn.kv_undo.record_update(key.clone(), entry),
    }
}

extract_primary_key returns exactly one key. MSET k1 v1 k2 v2 k3 v3 mutates three and logged an undo record for one, so TXN ABORT restored the first and left the rest at their new values — an acked abort landing a keyspace that is neither the pre-TXN nor the post-TXN image, produced by the one operation whose purpose is to prevent partial state. No crash, no disk pressure and no concurrency required.

Measured across the matrix before the fix — keys restored out of 4:

                    MSET      multi-key DEL
  monoio shards=1    1/4          4/4
  monoio shards=4    0/4          0/4
  tokio  shards=1    0/4          0/4      <- even DEL, which has its own arm
  tokio  shards=4    0/4          0/4

Three causes hid in that table. The capture recorded one key (the 1/4). At --shards > 1 the command was intercepted by coordinate_multi_key before the capture ran, which also bypassed the cross-shard TXN guard that exists precisely because there is no cross-shard undo log (the 0/4). And handler_sharded's multi-key branch, unlike its monoio twin, has no num_shards <= 1 early return, so on tokio the interception happened at one shard too and swallowed DEL as well.

The capture now walks written_keys — the same key spec the ACL and cache-invalidation paths use, filtered to KeyRole::Write so reads stay out of kv_write_intents, the cross-shard conflict surface — and falls back to the old single-key capture for an argv the walker cannot enumerate. A multi-key write inside a TXN no longer enters coordinate_multi_key: a single-owner key set falls through to ordinary routing (captured when the owner is local, refused by the existing cross-shard guard when it is not), and only a genuinely spread key set is refused, with the same error and the same #499 poisoning a cross-shard single-key write already got.

All four matrix cells now restore 4/4 for both MSET and DEL. Read-only MGET inside a TXN is unaffected, and MSET outside a TXN is untouched.

Behaviour change: a multi-key write whose keys span shards is now refused inside a TXN (ERR TXN does not support cross-shard writes -- use hash tags {tag} to co-locate keys) instead of silently succeeding with no rollback. Co-locate the keys with a hash tag, or issue the write outside a transaction.

  • Test fault injection leaked between concurrently-running unit tests (moon#750).

ManifestIo::persist read three process-global #[cfg(test)] statics — an injected error flag, an injected delay, and a persist counter. Rust runs a crate's unit tests concurrently in one process, so those statics were a channel between unrelated tests: while the injecting test held the error flag true, any other test's commit() failed with injected persist failure (test). That is the panic moon#750 recorded, in test_overflow_compaction_bounds_growth — a test that has nothing to do with fault injection and never touches the flag.

TEST_SYNC_KNOB_LOCK did not prevent it and could not: it serialised the tests that set the knobs against each other, while every test that merely commits a manifest is an unguarded reader. A lock only helps when every participant takes it.

The knobs now live on the ShardManifest that armed them (TestKnobs, held as shared handles so they survive enable_deferred_sync moving the io to the sync thread), so one test's injection cannot reach another manifest — and a test added later cannot re-open the race by forgetting a lock. The lock survives, renamed TEST_AGENT_REGISTRY_LOCK, for the state that genuinely must stay global: the manifest-sync agent registry and the flush_all_agents() barrier that walks it. flush_all_agents_is_noop_when_idle now takes it too — it was asserting on global registry state without holding it, the same unguarded-reader shape.

Also fixed a vacuous assertion this uncovered: the coalescing test bounded the persist count only from above (< 20), so a counter that always read 0 would have passed it.

  • A server that failed to bind its port exited 0 (moon#751).

main() logged a fatal listener failure and then fell through to the ordinary shutdown sequence and Ok(()):

if let Err(e) = server::listener::run_sharded(...).await {
    tracing::error!("Listener error: {}", e);   // logged, not propagated
}
...
info!("Server shut down");
Ok(())                                          // <- exit 0, always

So a server that never accepted a single connection produced exactly the same observable as a clean SIGTERM stop: shard-shutdown lines, Server shut down, exit 0. Under systemd Restart=on-failure it is never restarted; an orchestrator records a successful run; a test readiness probe sees a graceful shutdown. Reproducible in one command — moon --bind 203.0.113.7 (an unroutable TEST-NET-3 address) logged Listener error: Can't assign requested address and exited 0.

A fatal listener failure is now recorded and returned after the shutdown sequence has run, so the exit status is non-zero while the data path still flushes and joins exactly as before. Only the reported status changes: this path was already terminating the process. A clean SIGTERM shutdown still exits 0, which tests/listener_bind_failure_exit_code.rs pins alongside the failure case — "make bind failures non-zero" must not become "make everything non-zero".

This is what moon#751 was really looking at. That issue observed a test server shutting down gracefully during startup and reasoned that exit 0 proved the process had been asked to stop, which sent the investigation hunting a phantom SIGTERM sender. The inference did not hold. The two cases are now distinguishable by exit code, and the readiness failure message in tests/sigterm_shutdown.rs says which log line to look for in each case. - FT.CACHESEARCH and FT.SEARCH ... RANGE were inverted on COSINE / IP indexes (moon#748).

Three places tested distance >= threshold for the unit-sphere metrics, on the belief that those return a similarity where higher is closer. They do not. Cosine and InnerProduct are normalized at encode time and every dense scoring path runs through l2_f32, so the score is ‖a−b‖² = 2 − 2·cos — a monotonically increasing function of cosine distance, ranked ascending exactly like L2. An exact self-match scores ~0.006 (not 0.0: vectors are stored quantized) and an orthogonal one ~2.0.

The effects were exact inversions: - FT.CACHESEARCH reported a cache MISS for a query identical to a cached entry and a HIT for an unrelated one. At the moondb default threshold=0.95 a COSINE semantic cache was wrong for essentially every query. - Among entries that were inside the threshold, the best-candidate comparison picked the FARTHEST one, so a hit could return the wrong cached answer with cache_hit: true. That failure is silent — no flag reveals it. - FT.SEARCH ... RANGE <t> discarded every near match and returned only the vectors farthest from the query.

L2 was always correct and is unchanged.

apply_range_filter turned out to be the harder half. It is applied to three different score conventions while taking its direction from the index's dense metric in all of them, which is meaningless for two: sparse results are raw dot products (higher is better) and RRF fusion results are -score (negated on purpose so SearchResult::Ord sorts best-first, i.e. lower is better). It now takes an explicit ScoreOrder naming the convention at each call site, so each one is individually checkable instead of inheriting an unrelated field's metric.

The DistanceMetric doc comments asserted the same wrong thing and are corrected; that is where the belief propagated from.

CI could not catch any of this: the only tests called the predicate with hand-written numbers chosen to match the wrong assumption, so they passed while the product was broken. tests/ft_cachesearch_metric_748.rs closes the gap end-to-end against a live server — hit direction, miss direction, which entry a hit returns, dense RANGE across all three metrics, and hybrid RANGE to pin the negated-RRF direction.

Two things found on the way that are NOT fixed here, because they are separate problems: FT.CACHESEARCH is local-only in every dispatch path (at --shards > 1 it cannot see a cache entry that lives on another shard), and SparseStore::insert has no production caller at all — HSET never populates a sparse store, so the sparse-only search path always returns nothing.

Note for existing users: threshold is a distance ceiling, so the moondb default of 0.95 is very loose under the corrected semantics — on COSINE it admits entries roughly orthogonal to the query. The default is unchanged (it is a public SDK API), but the docstring now says so; pass an explicit value, typically 0.05-0.2. - GRAPH.ADDNODE/GRAPH.ADDEDGE no longer corrupt integer-like string properties that overflow i64 (moon#724).

parse_property_value coerced any bulk string that parsed as i64 or f64 into a number. A value too long for i64 fell through to f64, which keeps 15–17 significant digits and zeroes the rest — so a 32-digit key stored through GRAPH.ADDNODE came back as 12345678901234567000000000000000, and the node could no longer be found by the key it was created with. Irreversible, and silent.

An integer-syntax string that does not fit i64 is an identifier, not a number, so it now stays a string. Verified end-to-end: the reported 32-digit key round-trips exactly and MATCH (n:Doc {_key: '<key>'}) finds it, where before it returned nothing.

Only the integer case is withheld. moon#724 suggested refusing whenever the round-trip is inexact (format!("{f}") != s), but that also demotes "3.0" and "1e5" to strings and breaks MATCH (n {x: 3.0}), so genuine float syntax still coerces exactly as before. "42" is still Int(42), "3.14" still Float(3.14); i64::MAX still coerces and i64::MAX + 1 no longer does.

The Cypher path was never affected — it is fed by the grammar, which already knows the type. - FT.INFO reported Unknown Index name for an index that exists, whenever any single shard had lost its registration (moon#745).

merge_ft_info_responses propagated the first Frame::Error it saw from any shard. A shard that holds none of an index's documents can lose its registration without changing the search result at all, so FT._LIST (answered from the connection's own shard) listed the index, FT.SEARCH returned every document correctly, and FT.INFO said the index did not exist — on the same connection, to the same instance. The mirror case was equally wrong: an error from the local shard was returned even when remotes had the index, making the answer depend on which shard happened to accept the connection.

Unknown Index name from a strict subset of shards is now treated as a coverage gap rather than an error: the merge answers from the first shard that has the index and reports the gap as shards_missing_index, emitted only when non-zero so a healthy cluster's reply shape is unchanged (and so single-shard FT.INFO, which never reaches the merge, does not differ). When every shard reports it the index really is absent and the error is returned. Any other error stays fail-loud, including alongside a missing shard.

  • INFO persistence reported a hardcoded loading:0 while shards were still rebuilding their indexes (moon#744).

shard::loading::any_shard_loading() exists for exactly this — its own doc says the process-wide counter "exists only so INFO can answer 'is anything still loading?' from any thread" — but nothing ever read it. An operator (or client) polling loading: during a restart was reading a constant. Redis sets loading:1 while the dataset is being read from disk, so this is parity as well as diagnosability. Verified end-to-end: a 4-shard server with 30k vector documents now reports loading:1 through index recovery, then loading:0.

tests/ft_search_star_vector_only_695.rs now gates its post-restart assertion on that flag instead of on DBSIZE alone. The old gate proved the keyspace recovered, not the index — moon#743's permanently-broken instances all reported a full DBSIZE while the index was missing documents forever, which is why a durability bug reached us disguised as an intermittent test failure. That assertion also gained the as_err() guard its pre-restart counterpart already had, so a refused query reads as a refused query instead of dying inside Resp::total().

  • FT.CREATE acked +OK before every shard had persisted the index definition, so a restart permanently lost part of the index (moon#743).

VectorStore::save_index_meta_sidecar is a no-op while persist_dir is None, and persist_dir was assigned only inside recover_indexes_task — which moon#476 made spawned rather than inline, so the shard starts serving before it runs. A shard that handled FT.CREATE in that window created the index in memory, wrote no vector-indexes.meta, and still acked +OK.

The -LOADING gate could not cover it: shard::loading::is_loading() is a thread-local read on the connection's shard, while FT.CREATE fans out to every shard via broadcast_vector_command. The connection's own shard being ready says nothing about shard 3.

On the next start those shards found no index metadata, so recovery skipped the HASH rescan entirely and their documents never re-entered the index — permanently, and with no log line saying so. Measured symptom: an 8-document vector-only index restarting as FT.SEARCH vidx "*" → 3 documents, forever, on 5 of 6 concurrently started servers.

persist_dir is now assigned synchronously, before the recovery task is spawned, so no shard can serve FT.CREATE without one; the path is resolved once by a shared vector_persist_dir_for helper rather than computed in two places. Separately, a sidecar that fails to read now logs a warning instead of being swallowed by a catch-all _ => None — an empty sidecar is ordinary first-boot state, but a read error costs that shard every index it had and must not be silent.

Regression test: tests/ft_create_durable_at_ack_743.rs asserts the ack is durable on every shard. It repeats the spawn because the defect is a startup race — pre-fix detection was 3/12 spawns at --shards 4, 5/10 at 8, 7/10 at 16, so the test uses 8 shards × 6 rounds (~1.6% escape). Post-fix it is deterministic: 18/18 spawns, 144 shard-draws, zero misses.

Changed

  • Bump the async-runtime group: tokio 1.52.3 -> 1.53.1, tokio-stream 0.1.18 -> 0.1.19, tokio-util 0.7.18 -> 0.7.19 (moon#692).

Lockfile only — no Cargo.toml requirement moved, and cargo update resolved these as the "latest Rust 1.94 compatible versions", so the MSRV pin still binds the resolution.

The one behavioral change in tokio 1.53.0 does not reach moon: mpsc::{Receiver, UnboundedReceiver} now drop their waker on drop even while senders live (tokio#8095), and moon uses tokio::sync::{broadcast, watch} plus std::sync::mpsc, never tokio::sync::mpsc. The queued-reserve[_many] wake fix (tokio#8260) is likewise mpsc-only.

Taking 1.53.1 rather than 1.53.0 is load-bearing: 1.53.1 restores MSRV by removing OnceLock::wait from the Windows signal handler (tokio#8300), and moon pins MSRV 1.94 and ships Windows.

Transitively pulls socket2 0.6.5 alongside the existing 0.5.10 (distinct dependents) and windows-sys 0.61.2 in winsplit.

Fixed

  • Test harness: two #[test] threads could be handed the same --dir, and moon's instance flock correctly refused the second (#741). Suites built their data dir as {prefix}-{pid}-{nanos}. Both tests in a binary are threads of one process, so pid does not discriminate them, and SystemTime::now() resolves to microseconds on macOS — every nanos value ends in 000. Two tests entering the spawn helper in the same microsecond got the identical path; the second server exited 1 before binding. Four things then hid the cause: spawn_listening's panic asserted "not a port race" without admitting a directory race; the retried closure captured the colliding dir so all three attempts failed identically; the suite spawned with Stdio::null() so moon's Error: line was discarded even though the panic tells you to read it; and the winning test's Drop deleted the shared dir. tests/common now provides unique_test_dir() (per-process AtomicU64 counter — cannot collide whatever the clock does) and server_stderr() (append-mode log in the test's own --dir), the panic text names both failure modes, and the three suites with no per-test discriminator (lua_vm_memory_published, dir_deleted_degraded, txn_kv_wiring) were converted. Measured on one macOS host under 6-way load, 120 runs per binary: 14 failures before, 0 after.

wc4 waiter-cannibalisation test no longer reports a false failure when the CI box is loaded.

wc4_a_list_push_must_not_destroy_a_zset_waiter_on_the_same_key sent BZPOPMIN, slept 150ms hoping the registration reached the owning shard, then raced an RPUSH against it. Under a loaded full-suite run (5309 tests, 6-CPU VM) that guess is wrong: the push lands first, the key becomes a list, and the pop then answers -WRONGTYPE immediately — moon#556's correct behavior — which the test read as a cannibalised waiter. It failed 3-of-3 that way while passing 11-of-11 in isolation.

The test now asks the server instead of guessing, waiting on blocked_clients from INFO clients — the same handshake list_move_cross_shard.rs already uses, whose helper doc says outright why a sleep cannot do that job. record_client_blocked() fires when the waiter takes a vacant wait_keys slot in the owning shard's registry, so the gauge moves on registration, which is exactly the event being waited for, and it works cross-shard because the counter is one process-wide atomic.

The gauge can only be read as "one more than before", so every blocker connection is now held for the whole test rather than dropped per iteration — a dropped blocker unblocks asynchronously, and under load those decrements land after the baseline is sampled. The first version of this fix did drop them, and an interleaved under-load A/B caught it failing 2-of-4 on the sequence before=3 -> 3 stale reaps -> 0 -> new waiter -> 1, where 1 > 3 is never true. With the blockers held: 6-of-6 at load average 10.5, and 14.79s instead of 17.5s since the handshake beats twelve 150ms sleeps.

The guard was verified to still catch what it exists for: reverting the moon#535 fix in try_wake_list_waiter (pop_front_of_family(List) back to the blind pop_front) makes the test fail on 12 of 12 keys with waiters woken *-1. That is a different signature from the load artifact above (-WRONGTYPE, 2 of 12), which is what distinguishes the two.

Changed

  • Dependency debug info is dropped from dev/test builds, halving the target dir again and cutting the cold build 60% (moon#735).

moon#655 took a CI target dir from 39.1 to 14.9 GiB. What remained was still overwhelmingly debug info: a median 53.8 MB test binary was .debug_line 18.0 + .debug_str 13.8 + .debug_info 12.3 + .debug_ranges 4.7 + .debug_aranges 2.5 MB against 2.9 MB of .text — 96% debug info, 5% code. The cause is structural: 272 test executables each statically link the whole dependency graph, amplifying a 0.61 GiB shared rlib set into 10.43 GiB.

A six-leg same-tree matrix on moon-dev measured each lever separately (cold dir per leg, identical cargo nextest run --profile ci --no-run):

leg total build
baseline 13.88 GiB 362s
CARGO_INCREMENTAL=0 11.67 GiB 277s
+ [profile.dev.package."*"] debug = 0 8.48 GiB 239s
split-debuginfo = "unpacked" alone 9.13 GiB 193s
both 6.81 GiB 145s
debug = 0 entirely (floor, not a candidate) 2.43 GiB 114s

All three ship. The real cold ci-local legs that followed landed at 7.01 GiB (monoio) and 6.81 GiB (tokio) against 6.81 predicted, with cold suites at 453s and 446s versus 588s and 572s before. Cumulatively with moon#655: 41,993 MB → 7,313 MB (−82.6%) and 449s → 145s (−67.7%).

debug/incremental was 2.18 GiB of cache that a one-shot CI leg never reads back, and building it cost 85s per leg. Local development keeps its cache — only CI legs opt out.

What it costs, verified rather than assumed. A real moon test binary was built under the new configuration and panicked: moon's own frames keep symbol and at file:line, std frames keep theirs (std ships its own debug info, which package."*" does not touch), and third-party crate frames keep symbols but lose file:line — the intended trade. Panic messages are unaffected via #[track_caller]. That probe binary went 54 MB → 6.1 MB. Separately confirmed that split-debuginfo = "unpacked" does not put line tables out of reach: backtraces still resolve file:line with every .dwo file deleted, because line-tables-only keeps .debug_line in the binary.

split-debuginfo is set per-leg, not as a profile key. In [profile.dev] it would also apply to Windows, where MSVC uses PDB and support differs, and to macOS, where dev already defaults to unpacked — risking a platform for no gain. It is a CARGO_PROFILE_DEV_SPLIT_DEBUGINFO env var on the Linux legs only: ci-local.sh's two VM suites and the four Linux jobs in ci.yml (check, check-monoio, check-console, client-compat), set at job level because a workflow-level env: merges into every job and none can unset it. Windows, macOS, lint, msrv and the memory gate are untouched.

ci-local.sh's disk pre-flight constants are re-measured from the real leg sizes (cold-leg need 18G → 9G, full-dir size 3G → 1.5G); the warm-headroom warn line stays 15G, because it describes free space a running suite wants underneath it, which this change does not touch. scripts/test-ci-local-preflight.sh gains rows at the moved boundaries — including a 2G target dir, the size that separates a partially-warmed cache from a stub only under the new threshold — and all six constants were mutated back to wrong values to confirm the 31-row harness reports them rather than passing regardless.

  • Dev/test builds emit line tables instead of full DWARF, cutting the CI target dir 62% and the cold build 38% (moon#655).

tests/ holds 253 integration files. Each is its own crate statically linked against the whole moon lib and every dependency, so the dependency graph's debug info is duplicated across 258 test executables — the reason a single ci-local.sh leg's target dir had grown to 39 GiB, two legs filled moon-dev's 159G root, and a full root does not fail loudly: the leg exits rc=1 printing nothing and OrbStack itself wedges (moon#658, moon#661).

[profile.dev] debug = "line-tables-only" (inherited by profile.test) keeps exactly the information a RUST_BACKTRACE frame resolves against and drops the rest. Measured on moon-dev as a same-tree A/B — identical checkout, identical cargo nextest run --profile ci --no-run, cold target dir on both legs, the debug setting the only difference:

debug = 2 line-tables-only delta
target dir 39.1 GiB 14.9 GiB −62.0%
258 test executables 34.1 GiB 11.0 GiB −67.8%
cold build 449 s 279 s −37.9%

Nothing in the tree consumes full DWARF: no debugger, no core-dump inspection, no backtrace crate. Panic messages keep file:line via #[track_caller] regardless, and a panicking two-frame program built both ways on Linux resolves every backtrace frame to the same symbol, file, line and column — only addresses and symbol hashes differ.

scripts/ci-local.sh's disk pre-flight hardcoded the pre-change sizes, so a healthy cold leg would have been refused for wanting 36G it no longer needs. The two constants that describe what a build writes are re-measured here (cold-leg need 36G → 18G, full-dir size 5G → 3G); the warm-headroom warn line stays at 15G, because it describes free space a running suite wants underneath it, which this change did not touch. scripts/test-ci-local-preflight.sh gains rows at the moved boundaries — including a 4G target dir, the size that separates a partially-warmed cache from a stub only under the new threshold. Every constant was then mutated back to a wrong value to confirm the harness reports it rather than passing regardless.

The effect is ELF-only: on macOS DWARF lives outside the executable, so a Mach-O A/B produces byte-identical binaries and no saving. - The nightly fuzz soak no longer runs through the working day (moon#732).

fuzz.yml was scheduled '0 2 * * *'. This project's timezone is UTC+7, so that is 09:00 local — the nightly matrix is 18 jobs with fail-fast: false and no max-parallel, each budgeted 5 hours, so it occupied 18 hosted runners from 09:00 until ~14:50 local every day.

Measured 2026-08-26 with the run 2.5h in and all 18 jobs live: two CI runs sat queued for 35+ minutes, while Check (Windows) — which draws from a different runner pool — completed normally. Every PR raised in the morning paid that queue, and it read as "CI is slow" rather than as a scheduling collision.

Now '0 16 * * *' — 23:00 local, finishing ~04:50 local. Deliberately a reschedule rather than a max-parallel cap: capping to 9 would double the wall clock and push the tail back into working hours. crash-matrix.yml has the same 10:17-local schedule but runs entirely on the self-hosted VM, so it costs no hosted capacity; its cron is also load-bearing (recall-canaries matches the literal string), so it is left alone.

fuzz.yml also gains workflow_dispatch, with an optional duration_seconds. It had only schedule and a ci-fuzz-labelled PR, so a nightly that needed to be cancelled — exactly what the 2026-08-26 queue collision forced — could not be relaunched, only gh run rerun'd. The input is clamped to the existing 18000s budget rather than merely defaulted to it: the job's 350-minute ceiling has to outlast fuzzing plus the corpus archive, and a manual trigger must not be able to reintroduce the overrun that ceiling exists to prevent.

  • The pre-merge hosted matrix now runs only what no local gate can produce (moon#732).

A merge cost ~59 minutes of gating: ci-local ~26m, then a push, then the PR gate 8.6m, then a wait for it to settle, then a dispatch matrix of 24.5m. Nothing overlapped.

Measured per-job on PR #731, seven of the nine dispatch legs were re-running what ci-local had just finished — and two of them, check-monoio and client-compat, are runs-on: [self-hosted, moon-dev], so the matrix queued on the very VM that had produced the same results minutes earlier.

workflow_dispatch now runs Windows (835s, the only leg with no local equivalent and the floor of any hosted matrix), MSRV, and the memory steady-state gate. macOS, monoio, tokio-Check, client-compat, console and Lint stay on main-push as a post-merge net.

That is a real shift of responsibility, so two things move with it:

  • scripts/ci-local.sh with no arguments is now --full — it runs client-compat and the macOS suite, because after this change nothing else does before merge. client-compat is not a formality; it is what caught the v0.8.6 inline-GET ACL bypass. --fast is the old default, kept for iterating.
  • Every mode's verdict now states what it did NOT run. Previously --quick — six lint gates and not one test — printed the same RESULT: PASS line as a complete run.

tests/ci_covers_monoio.rs gains a test that follows the coverage rather than the job: it asserts ci-local still runs a monoio suite, so the shipped runtime cannot fall out of both gates at once. Mutation-tested — renaming the step makes it fail.

  • A multi-key command whose keys all live on one shard now joins the slotted batch instead of cutting it (moon#513).

moon#512 made a pipelined command that does not route by its own single key wait for the batch's already-deferred remote commands, because such a command executed INLINE while earlier writes in the same batch were still undispatched — silent write loss, not a stale read. moon#513's mask refinement then stopped the cut when the command's keys avoid every pending shard. The CO-LOCATED case was deliberately left alone: an MGET reading the very shard the pending writes are going to genuinely could not be waved through while it still ran inline.

The fix is not a further relaxation of the guard but a change to where the command runs. A multi-key command whose key mask has exactly one bit has a single owner shard, so the multi-key coordinator now stands aside and ordinary routing appends it to remote_groups[owner] like a single-key command. extract_primary_key hashes its first key, which by definition names that same owner, so the two decisions cannot disagree. Ordering then comes from the slotted batch itself — the property that has always made SET k + GET k correct — so the wait is discharged rather than skipped.

This is the shape CLAUDE.md tells users to adopt ({tag} co-location), so it is the shape most likely to sit in a hot pipeline. Measured on Linux (moon-dev, aarch64, --shards 4, interleaved control/fix legs, scripts/bench-single-owner-multikey.py), a read-modify-write loop over rotating tags:

shape control fix delta
round-trip (one batch pass per group) 41,699 ops/s, 4,499 cuts/leg 69,014 ops/s, 0 cuts +65.5% (noise 8.4%)
one huge write 1,470,418 ops/s, 8 cuts/leg 1,516,643 ops/s, 0 cuts +3.1% — inside the 11.9% noise floor

Reproduced at +64.7% on a second run. The 4,499 cuts per 6,000 groups is 75%, exactly the 1 - 1/shards remote fraction at four shards, which is the harness saying it measured placement rather than noise.

Two honest caveats. The bulk shape shows no effect because one huge write means few batch passes and therefore few cuts — that is a fact about the shape, not about the fix. And a connection that uses a FIXED tag self-heals on Linux without this change at all: the connection-affinity tracker migrates it onto the owner shard after ~16 samples, every key goes local, and the cuts stop. So the change earns its keep on connections migration cannot help — ones spanning many tags (a cache layer keyed by user), or ineligible for migration (inside MULTI, subscribed, tracking, replica). Connection migration is Linux-only, so on macOS the fixed-tag case cuts on every interleaving.

A key set that genuinely spans shards is unchanged: it is still consumed inline by the coordinator and still waits when it meets a pending shard.

The workspace exemption is now an explicit parameter rather than a u64::MAX sentinel passed through pending. A workspace connection rewrites every key to a {<32-hex>}: hash TAG after the guard runs, so raw names read at the guard hash to shards the command will never touch — and a single-owner mask read off them would be the most confident wrong answer available. Making it a parameter means the compiler names every call site when a new mask-based relaxation is added, instead of the relaxation silently escaping the sentinel.

Fixed

  • A flush issued by a ROUTED script or transaction now reaches every shard (moon#705).

EVAL "return redis.call('FLUSHALL')" 1 <key> at --shards 4 answered +OK while leaving 7 to 10 of 12 keys alive. Which answer a caller got depended on where a key the script never touches happened to hash — an implementation detail no client can see or control.

moon#685 gave a script-issued flush both dimensions the Lua bridge cannot reach on its own, but the broadcast half stopped at the four routed entry points in shard/spsc_handler.rs: the owner shard cannot fan out from inside its own message loop. It no longer tries to. The flush now rides home on ExecReply::script_flush and route_script_elsewhere completes it in async context — the same division of labour the MULTI path has used since c10k E2.

Two further defects were found and fixed while getting there:

  • coordinate_flush_broadcast overloaded ONE parameter as both the SPSC sender and the leg to skip. They coincide for a flush typed on a connection and differ for one performed on an owner shard, so the first cut sent from a shard the caller was not running on and left the caller's own slice full. They are now separate parameters.
  • broadcast_txn_flushes had no sender parameter at all, so a routed MULTI carrying a FLUSHDB/FLUSHALL had the same defect independently of scripts. Measured by reverting only that fix: EXEC answers [ok, ok] with 3 keys surviving.

The SPSC mesh has no self-loop (ChannelMesh::target_index debug-asserts against it and in release silently picks a neighbour), so the one leg that now lands on the caller's own shard goes through the thread-local self queue the event loop already drains.

Routed scripts also never fired the client-side-cache invalidation a flush owes tracking clients; completing the flush through finish_script_flush fixes that in the same change.

A fifth Execute reply consumer lives on the admin console plane (admin/console_gateway.rs) and compiles only under the console feature, which neither the default nor the tokio build covers. It now completes the flush over its own producers rather than dropping it — the one plane that runs scripts trusted was otherwise the one place the partial flush would have survived this fix.

  • FT.SEARCH on a missing index no longer reads as an empty listing in a build without text-index (moon#728).

A build compiled without the text-index feature — what CI's tokio leg compiles — answered a text-shaped FT.SEARCH on an unknown index two different ways depending on a deployment knob the caller cannot see: ERR invalid KNN query syntax at --shards 1, and a SUCCESSFUL empty array at --shards > 1. The second is the damaging one, because it is indistinguishable from an index that exists and holds nothing.

Both answers came from the same root: with no text engine, the handler's single-shard text fast path is compiled out entirely, so the query fell through to the KNN parser, while the multi-shard branch still routed to scatter_text_search, whose #[cfg(not(text-index))] arm skips the index-existence check its twin performs. The DFS scatter then ran against permanently empty per-shard text stores and merge_text_results summed the replies into [0].

The multi-shard branch no longer routes to the text scatter when there is no text engine, so both shard counts now converge on the KNN parser, and both of its entry points consult one shared predicate (text_engine_absent_refusal) that names the real reason: ERR text-index feature not enabled. Shard-count independence is now structural rather than two error strings that happen to match. The text-index build is untouched, and the moon#695 vector-only * enumeration — which deliberately needs no text engine — still answers in both feature sets at both shard counts.

  • FT.SEARCH <index> "*" now enumerates a VECTOR-only index (moon#695).

moon#693 made a bare * the match-all query and answered it from the inverted index, because that is where the document registry lives. Every index with a TEXT, TAG or NUMERIC field gained match-all — including mixed VECTOR+TEXT schemas. An index built from VECTOR fields alone has no inverted index at all, so * fell through to the text engine and answered ERR no such index for an index FT._LIST happily listed. Before #693 the same call said ERR invalid KNN query syntax. Both were wrong, and neither told the user their index had no way to be enumerated.

The vector engine has its own registry — VectorIndex::key_hash_to_key — and it is live: filled on index, pruned in the same function body as the segment tombstone, so the two cannot disagree. It is also the exact map KNN already resolves hits through, which bounds this honestly: a document match-all misses is one KNN would report as a synthetic vec:<id> rather than by name.

The routing, not the data, was the work. is_text_query("*") is true, so * reaches the text path at all four FT.SEARCH sites, and a bare * is also the leading token of a HYBRID or SPARSE query whose retriever clauses live in separate args. The gate therefore rejects HYBRID/SPARSE itself rather than relying on where it is called from, and declines any index the text store knows — so mixed schemas keep the path they already use. Multi-shard fans the same enumeration out over the generic per-shard command channel and reuses the text merge, which already sums per-shard totals; each shard is capped at offset+count rather than handed the caller's own LIMIT, so paging cannot re-serve the same documents.

Verified at --shards 1 and --shards 4, which is not ceremony: the first cut wired the multi-shard branch and missed the single-shard one, and returning nothing at one shard count would have been worse than the honest error it replaced. Also verified across a restart, where * agrees with KNN exactly. - A compaction backlog no longer refuses the commands that cure it (moon#718).

Reported as a permanent write outage with no in-band recovery. Background merges keep failing, immutable segments accumulate to --max-unflushed-immutable-segments, and every foreground write is refused with MOONERR busy: compaction backlog. Reads keep working, and so does the server — but both documented escapes are themselves foreground writes. FT.COMPACT drains the backlog; FT.CONFIG SET <idx> MERGE_RECALL_TOLERANCE 0 relaxes the recall gate a repeatedly-failing merge is stuck behind, and is the command the rejection log line recommends by name. The registry flags both W, so the backlog refused its own remedy and the only exit was a restart, which replays into the same state.

src/shard/segment_stall.rs had claimed this exemption already existed — "FT.COMPACT / GRAPH.COMPACT commands bypass this guard so compaction can always proceed to drain the backlog". It did not, and GRAPH.COMPACT is not a command at all. The same false claim appeared in config.rs and in the sharded dispatch path; all three are corrected.

The decision now lives in one place, segment_stall::stall_refusal, which both dispatch paths call — they previously carried byte-identical copies of the message ladder, so a fix applied to one would silently have missed the other.

The exemption is deliberately narrow: it applies to the segment backlog only. dirmissing, diskfull and memfull are answered first, so a compaction that would write new segment files onto a full disk is still refused. A remedy for one stall is not a licence to ignore the others. VACUUM needed no change — it is flagged A, not W, so no stall ever saw it.

Ordinary writes are unaffected: SET, HSET, FT.CREATE and FT.SEARCH are still refused by the backlog, which is what makes this a fix rather than a hole in the backpressure guard.

Note for anyone hitting the reported recall 0.0000: that specific trigger was fixed after 0.8.5 by moon#588 (the gate scored an ID-set overlap, which is undefined when distances tie). This change is still needed, because any repeatedly-failing merge reaches the same trap.

  • FT._LIST enumerates both index stores, so a TEXT-only index is no longer invisible (moon#709).

An index whose schema carries no VECTOR field lives only in the TextStore. FT._LIST enumerated the vector store alone, so such an index never appeared — even though FT.INFO and FT.SEARCH both work on it. FT._LIST is how tools and the Moon Console discover indexes, so a TEXT-only index could not be listed, inspected in a UI, or picked up by anything that enumerates before acting. It also silently broke any harness that used FT._LIST to verify index creation: the one that found this reported "built 0 indexes" after 50 successful FT.CREATEs.

The two stores are now unioned by sort-then-dedup rather than a membership scan — it collapses the both-stores duplicate (an index carrying TEXT and VECTOR fields is registered in both and must appear once) in O(n log n) instead of O(n^2), and it gives the result a stable order that neither store's hashing provides on its own.

The text_store parameter was added rather than reached for locally, so the compiler enumerated all seven call sites instead of leaving one behind. The sweep the issue asked for came back clean: FT.INFO and FLUSHALL/FLUSHDB index-clearing already consult the pair, and FT._LIST was the only enumerator reading one store.

Proven against a pre-fix binary rather than assumed: the e2e suite fails on main's binary with left: ["both", "vec"], right: ["both", "txt", "vec"], and both of its premise guards (FT.INFO and FT.SEARCH working on the TEXT-only index) pass on that same binary — which is the issue's point. The suite also restarts on the same dir, so the TEXT-only index has to come back from its sidecar and re-register.

  • Cypher writes narrow through the property index instead of scanning every node of the label (moon#719).

The write executor reached IndexScan through the same match arm as NodeScan and threw the planner's narrowing away — it visited every node carrying the label, then let the residual Filter cut the result down. Output was therefore always correct, which is exactly why this survived: a MATCH (n:L {k: '…'}) SET … and its read-only twin return identical rows, and only the write one costs O(nodes already present). Building a graph one node at a time — the normal shape of knowledge-graph ingest — was quadratic. The reporter measured a 50-statement batch growing 9.07x between a 1,000-node and a 10,000-node graph while the byte-identical read grew 1.35x, and GRAPH.PROFILE could not show it because it refuses write statements outright.

The IndexScan arm now calls the read path's own index_scan_keys — shared, not a second copy that can drift — with a context mirroring the visibility arguments the label scan passed one arm up (u64::MAX, txn 0, no valid-time), so the candidate set is unchanged and only its size moves.

Because both plans return the same rows, correctness cannot catch a regression here. ExecResult gained a nodes_scanned counter — the observable difference, and useful to any profiling caller that wants to see a plan silently fall back — and the test asserts on it rather than on a timer: one matching node among 5,000, scanned 5,000 before and fewer than 64 after. The test also guards its own premise, failing if the planner ever stops emitting IndexScan for the shape.

The issue's second suspect — the property-index update on SET being O(N) — was checked and ruled out: MutablePropertyIndex::remove is O(bucket), and a unique key's bucket holds one entry.

End-to-end, scripts/bench-cypher-write-index-719.py drives the issue's own discriminator against two separately-built binaries (aarch64 macOS dev host; these are growth ratios, not throughput numbers, and the legs are interleaved so a drifting machine cannot fake a win):

50-statement batch 1,000 nodes 10,000 nodes growth
write, before 5.52 / 7.49 ms 84.67 / 105.35 ms 14.69x
write, after 0.75 / 0.92 ms 0.99 / 1.22 ms 1.33x
read (in-run control, both) ~0.8 ms ~0.8 ms ~1.0x

The write now tracks the read. The 10,000-node batch is 86x faster, and the harness refuses to print any of it if the two binaries share a sha256, if the port is answered by something that is not moon, if the two binaries disagree on matched rows, or if the control binary fails to exhibit the defect it is supposed to demonstrate.

  • 26 integration suites now own their server through a Drop guard, so a failing assert cannot orphan it.

Every one of these suites killed its server on the last line of the test body. That line is only reached when the test passes: a failing assert! unwinds straight past it, the Child is dropped without kill() (Rust's Child::drop deliberately does not reap), and the moon process is reparented to init. It then runs forever — this is not a tidy-up nicety. Five such orphans were measured on the dev host burning ~170% CPU each (834% combined) for 7.5 hours, one thread showing 13:42 system time against 0:09 user.

common::ServerGuard owns the child and SIGKILLs it in Drop, so the reap happens on the unwind path and on the success path alike; kill_now() is idempotent, and take() hands ownership back for the suites that deliberately outlive their server. The 14 suites that go through common::spawn_listening use the new spawn_listening_guarded; the other 12 wrap their own Command::spawn() chains. Seven now-dead file-local sigkill copies are deleted.

tests/server_guard_contract.rs proves the guard rather than assuming it: a deliberately panicking test must leave no live pid, an explicit kill_now() followed by Drop must not double-kill, and take() must transfer ownership. The panic test is #[cfg(unix)] and its ps probe expects — on Windows the suite really runs, and a probe that could not spawn would have reported "not alive" and passed vacuously. Mutation-checked: mem::forget(guard) fails it with "server pid N survived a panicking test".

  • mq_integration no longer races its own server on port or on startup.

The fixture reserved a port by binding 127.0.0.1:0 and immediately dropping the listener, then waited for the server with a fixed sleep(200ms). Both are bets. The dropped probe hands back a port the OS is free to reissue — and 17 tests in this file each start their own server, under cargo test concurrently in one process and under nextest in 17 concurrent processes. The fixed sleep bets that shard startup fits in 200 ms on a loaded machine.

Observed 2026-08-25 on a 12-core macOS host immediately after a full suite: 12 of 17 tests failed with Connection refused (os error 61) at connect(), three of them surviving all three nextest retries; all 17 passed on the same commit once the host was quiet. The trigger was not reproducible on demand (synthetic CPU load and a concurrent cold build both failed to provoke it), so the mechanism below is what the old code made possible, not a mechanism caught in the act.

The fixture now uses common::reserve_port() — intra-process dedupe plus a cross-process DirLock claim, the helper the rest of the suite already uses — and polls connect() to a 30 s deadline instead of sleeping. Removing the wait entirely fails the suite, so the poll is load-bearing, not decoration. - scripts/ci-local.sh --native runs the long legs on the macOS host, with the gaps printed.

The merge bar's two full suites and the client-compat harness live in the moon-dev VM, so when the VM is unavailable the long legs simply cannot run — and a testing phase that wants to stay on one machine had no supported way to do it. --native runs both suites (monoio on kqueue, tokio) plus the compat harness against a host redis-server.

It is deliberately not sold as equivalent. macOS does not have io_uring, so the shipped Linux path is untested by a native run, as are cfg(target_os = "linux") code, Windows, and the MSRV pin. The mode ends by naming those instead of printing a bare RESULT: PASS, and a missing redis-server oracle refuses with exit 2 rather than skipping — a differential harness with no oracle proves nothing, and a green skip would be a lie.

Covered by four new cases in scripts/test-ci-local-preflight.sh; both guards were mutation- checked (deleting the oracle refusal, and hoisting the native summary above the failure check so a failing run would exit 0 — each makes its test fail). - A failing loading_state_476 test no longer orphans its server (#476 follow-up).

The kill lived on the last line of the test body, so a failing assert! unwound straight past it and left the server reparented to init. Measured after one afternoon of deliberately failing mutation runs: five orphans, ~170% CPU each (834% combined) for 7.5 hours, answering nothing on their ports and spending the time in syscalls (13:42 system vs 0:09 user). The cost is not only the cores — every wall-clock-sensitive test that ran afterwards did so on a machine under invisible load.

The child is now owned by a ServerGuard that kills it on drop, so an unwinding panic reaps it. A regression test panics deliberately and asserts on the PID rather than the port: a dead server frees its port either way, so connect-refused proves nothing about whether anything is still running. Replacing the guard with mem::forget fails that test.

  • A restarting server tells clients it is loading instead of hanging on them (#476).

The per-shard listener was bound and its accept task spawned before index recovery, but recovery was fully synchronous — no .await — so on a single-threaded shard runtime neither the accept task nor the event loop was ever scheduled until it returned. The kernel completed handshakes from the backlog by itself, so a client connected successfully and then waited out the whole recovery in silence. Measured on moon-dev over 83,828 keys:

connect() succeeded at     148.2 ms
PING sent at               148.3 ms
first byte back at        1376.6 ms   -> 1228 ms on an ACCEPTED socket

Not refused, no error — the client holds what looks like a healthy connection and cannot tell loading from wedged, so a health check hangs for as long as the store takes. The reported production case was ~94 minutes.

Recovery now runs as a task on the shard's local executor while the event loop serves normally. During it, commands that would read a half-built index get Redis's -LOADING moon is loading the dataset in memory; PING/INFO/CLIENT/CONFIG/AUTH/HELLO and the pub-sub family still answer, so clients can authenticate, diagnose, and fail over at once. The reconcile loop yields between 1024-key chunks — with_shard takes a synchronous closure, so chunking is what makes an .await possible at all — and the loading state is held by a guard whose Drop clears it, so a panicking recovery task cannot leave a shard refusing every command for the life of the process.

Both halves are proven necessary by mutation: disabling the gate fails the acceptance test (and the command is then served off a half-built index at 166 ms), and running recovery inline instead of spawned restores the original 715 ms hang. Green under tokio and under monoio/io_uring, the shipped runtime.

  • Recovery reports its progress instead of going silent for the whole reconcile (#546).

A production restart spent ~94 minutes inside the keyspace reconcile loop and logged nothing between restoring 2311 text index(es) from sidecar and the final count. An operator watching that log cannot tell a slow repair from a wedged process, and the two call for opposite actions. Recovery now emits a line every 10 seconds with the key it is on, the total for that db, the running rate, and elapsed time.

The trigger is wall-clock, not "every N keys": the recoveries that need a line are exactly the ones whose keys are individually slow, and a key-count rule stays silent through precisely those. A recovery that finishes inside one interval stays silent — verified end-to-end both ways on a 20,000-key store (19 lines with the interval forced to zero; 0 lines at the shipped 10s).

While measuring this, the other two hypotheses in #546 were re-tested and do not reproduce on current main. Restart-to-PONG is now linear in FT index count — 50→800 TEXT indexes holds flat at ~0.5 ms/index (0.03s → 0.41s for 800 indexes over 32,000 docs) — and stripping every warm segment's mvcc.mpf, the state that left 897 segments unregistered, costs nothing measurable: A/B on one built store, 4–32 indexes over 8,000–84,000 keys, at both DIM 8 and DIM 384, stripped/control ratios spanning 0.78–1.35 with the stripped leg faster as often as slower (n=1 per configuration; consistent with run-to-run noise, and nowhere near the reported behaviour). The 94-minute cost was dominated by TextIndex::remove_field, which #613 made O(terms the document carried) rather than O(vocabulary). What remains of #546 is the observability this adds, plus the partially-GC'd segment directories themselves, which are still left on disk and re-warned about every restart.

Measured on moon-dev (aarch64, 6 vCPU) with --shards 1; harness verifies its own instrument (index and key counts read back after each restart, and the stripped leg asserted to actually log one leaving unregistered per stripped segment) so a run that silently indexed nothing cannot report a fast restart as a win. - A warm segment directory whose mvcc.mpf is gone is now retired instead of re-warned about forever (#546).

register_warm_segments decides which index owns a recovered warm segment by reading exactly one file — mvcc.mpf, via peek_key_hashes. When that file is missing, no evidence exists, no owner can be chosen, and the directory was left on disk with a warning. Nothing ever removed it, so every subsequent restart re-read it and re-warned. Reproduced by stripping mvcc.mpf from 6 of 12 warm segments and restarting three times: 6 warnings each time, 6/6 directories still present, DBSIZE unchanged at 2400 — the warning was permanent and the disk cost was permanent with it. On the store that prompted the report this was 897 directories.

Such a directory is what a crash mid-creation, or a GC pass that emptied a directory without removing it, leaves behind. It can never be attached to an index; the keys it might hold are recovered by the keyspace rescan regardless. It is now removed, counted, and reported as retired as orphans with no mvcc.mpf in the startup summary.

Retirement is gated strictly on ErrorKind::NotFound. Every other IO error keeps the old warn-and-leave behavior: a transient EIO or a permissions problem can sit over perfectly good vectors, and deleting on those would turn a recoverable blip into data loss. A present-but- unparseable mvcc.mpf — truncated or corrupt — does not error at all; it reads as zero ids and is likewise kept, because an unreadable id sidecar is not evidence that the codes beside it are worthless. Both of those cases are pinned by tests that fail if the rule is widened.

  • A pipelined multi-key command no longer cuts the batch over shards it never touches (#513).

must_wait_for_pending_remote's multi-key arm answered "wait" on the command NAME, without asking where its keys were. But remote_groups only ever holds FOREIGN shards, so the #507 hazard — reading state a pending command is about to write, or writing state it then overwrites — requires the two to meet on the same shard. An MGET reading shards the batch has no pending work for was being cut for nothing.

The cut is not free: it ends the batch pass, and the phase-2b drain then dispatches one PipelineBatchSlotted per target shard and awaits each reply slot in turn.

Measured on moon-dev (aarch64, 6 vCPU), --shards 4, 32 interleavings of SET,SET,MGET, six fresh server starts per side interleaved:

shape before after
MGET reads shards the writes never touch 41,600 ops/s, 64 deferrals 86,500 ops/s, 0
MGET reads the shards being written 34,300 ops/s, 64 deferrals unchanged

Fresh starts per measurement because SO_REUSEPORT decides which shard the connection lands on, and that changes the shape's cost as much as the code does — one start per side compares placements as much as binaries. The deferral counts are placement-independent and were 64/64 before and 0/0 after in every round.

Only the multi-key arm is refined; the other two still always wait, because neither can be bounded by a key mask. An inline-intercepted command (EVAL, SWAPDB, …) executes against the local slice whatever keys it declares, and a keyless command (FLUSHALL, KEYS, SCAN) touches every shard. A key layout the shared walker cannot enumerate — SORT ... BY w_*, a key position holding a non-string, more shards than the mask has bits — also still waits: wrongly waiting costs a batch boundary, wrongly proceeding corrupts data.

A co-located {tag} multi-key command still defers whenever its one shard is the shard with pending work, and must: the coordinator executes a multi-key command inline rather than slotting it, so skipping the wait there would re-open #507. Routing a single-owner multi-key command into the slotted batch is the remaining scope of #513.

Workspace connections keep the old always-wait behaviour. workspace_rewrite_args rebinds the argv below this guard, and the guard cannot move down there — the connection-level intercepts it exists to hold back run in between — so the keys visible to it are the raw ones. The workspace prefix is a hash tag, so every key in a workspace routes to one shard however the raw names scatter, and a mask read off raw names called commands disjoint from the very shard their writes were pending on (5 of 12 connections lost an MGET's own batch writes before this was caught). Treating every shard as pending makes the predicate answer exactly as it did before the mask existed.

Added

  • INFO stats reports what the pipeline ordering guarantee costs (total_pipeline_remote_defer, and the moon_pipeline_remote_defer_total Prometheus counter) — groundwork for #513.

#512 made a pipelined command that cannot route by its own single key wait for the batch's pending cross-shard commands, which is what stopped the silent write loss of #507. It is also expensive, and nothing said so: the only symptom was throughput that looked bad for no visible reason.

Two things bound the cost, and the counter is what made both checkable:

  • A deferral needs an undispatched cross-shard command already in the batch — the guard is !remote_groups.is_empty() && must_wait_for_pending_remote(..). A shard-spanning MGET on its own never defers: 64 spread MGETs with no preceding writes measure 0. A preceding foreign read counts too, since the E2 read fast path is disabled and foreign reads are slotted alongside writes.
  • At most one deferral per batch pass. The cut re-parses the tail with remote_groups cleared, so the command at the head of the next pass runs inline whatever its shape.

Together those explain why 64 interleavings produce fewer than 64 deferrals, and fewer still at --shards 2 (48) than at --shards 4 (55) — with two shards, more of the preceding SETs land locally and never reach remote_groups at all. The counts are shape- and placement-specific, not constants: the same shape re-measured on a different key set gave 59.

Measured on moon-dev (aarch64, 6 vCPU), one connection, 9 reps alternating leg order, median:

pipeline shape shards=1 shards=2 shards=4
MGET after every 2 SETs 1,296,360 ops/s (0 deferrals) 49,203 (48) 38,856 (55)
128 SETs then one MGET 1,156,693 (0) 659,436 (1) 539,996 (1)
SET,SET,GET — routes by its own key 1,700,287 (0) 1,122,573 (0) 993,784 (0)

The deferral counts are the server's own, not inferred: reading the code suggested 64 for the first shape and the counter says 48, which is exactly why it exists. The --shards 1 column is the control — remote_groups is always empty there, so the guard structurally cannot fire.

Changed

  • Sharded pub/sub channels are no longer workspace-scoped (#703, fallout of #668). The hand-rolled workspace key walker had PUBLISH/SUBSCRIBE/PSUBSCRIBE in its no-key list, so their channels were never scoped — while SPUBLISH/SSUBSCRIBE/SUNSUBSCRIBE were not in that list, so its single-key default prefixed args[0], their channel. The result was that sharded pub/sub was workspace-isolated and ordinary pub/sub was not, purely as an artefact of which names somebody had remembered to add to a list. Nothing tested either half.

A channel is not a key, the shared walker names none, and all six are global now. Whether a workspace should be a pub/sub namespace is a product decision, tracked in #703 along with the two other things a workspace does not partition (MQ queue names, and FUNCTION libraries — which the old default broke outright inside a workspace by prefixing the SUBCOMMAND, and which work now).

Fixed

  • A flush issued from Lua reached neither every database nor every shard (#685). redis.call('FLUSHALL') cleared only the script's own database, and both redis.call('FLUSHALL') and redis.call('FLUSHDB') cleared only the shard the script happened to run on. Measured on main:
--shards 1, five databases seeded with one key each, from db0
  plain FLUSHALL:                          db0=0 db1=0 db3=0 db7=0 db15=0
  EVAL "return redis.call('FLUSHALL')" 0:  db0=0 db1=1 db3=1 db7=1 db15=1

--shards 4, db3 seeded with 40 keys
  EVAL "return redis.call('FLUSHDB')"  0:  40 -> 29   (one shard of four)
  EVAL "return redis.call('FLUSHALL')" 0:  29 -> 29   (nothing further)

Both halves are structural, not a missing call: the bridge reaches the keyspace through ONE &mut Database on ONE shard, so it can express neither "every database" (that needs the slice its borrow came out of — the moon#677 shape) nor "every shard" (that needs to .await a broadcast, inside a synchronous redis.call). The bridge now records what the script asked for and the caller finishes it one frame up, where both are in scope again, through a single run_and_complete helper the twelve script entry points (EVAL, EVALSHA, FCALL, FCALL_RO × three handlers) all go through — FCALL/FCALL_RO carried the same bug and are not mentioned in the issue.

Reaching every database is only safe together with the rest of what a flush means, so the same completion drops the vector/text index CONTENTS (R3 — FT.CREATE definitions survive, matching restart semantics) and tombstones the durable MQ streams (task #46). Without those, a script's flush would have emptied sixteen databases on every shard while leaving every flushed hash searchable as a ghost — a wider version of exactly the inconsistency R3 exists to prevent.

The live server also disagreed with its own log. The AOF and replication planes were already correct — the record says FLUSHALL and replay completes it across the set — so a restart silently deleted four databases that had been readable a moment earlier, and a primary and its replica diverged on any script that flushed. A script's flush now also pushes the RESP3 flush invalidation to tracking clients, as the typed command has — unconditionally, because the flush is real however the script ends: a script that flushed and then returned an error left every tracking client caching keys that no longer existed.

Still uncovered, deliberately: a script ROUTED to another shard for its declared keys gets the database half and not the broadcast — handle_shard_message_shared is synchronous and holds neither the SPSC producers nor the notifiers a broadcast needs, and fanning out from inside a shard's own message loop is the wait cycle the MULTI path documents. Measured at --shards 4: 6 to 10 of 12 keys survive, depending only on where the script's declared key happens to hash. Tracked as #705.

  • Three dispatch paths reached the keyspace without workspace key prefixing (#702, #668). A workspace-bound connection that wrapped its commands in MULTI/EXEC addressed the raw, unprefixed keyspace — so it could read and overwrite any other workspace's keys, and the global keyspace with them. Measured on main: tenant B ran MULTI; GET {<tenant-a-hex>}:secret; SET {<tenant-a-hex>}:secret OVERWRITTEN-BY-B; EXEC and both succeeded; tenant A's own GET secret then answered OVERWRITTEN-BY-B. A Lua script's whole KEYS[] vector escaped the same way.

The rewrite sat ~450 lines below the ACL gate behind a comment claiming it was "the ONLY code path where workspace prefixing occurs". Three paths reached the keyspace above it and prefixed nothing: the MULTI queue gate (the queue stores FRAMES and EXEC replays those, not the shadowed cmd_args), the EVAL/EVALSHA intercepts, and cluster routing (which computed a slot from a key name that does not exist). A simpler symptom of the first, with no second tenant involved: MULTI; SET k v; EXEC inside a workspace wrote the GLOBAL k, and the connection's own later GET k — which is prefixed — answered nil. The transaction wrote somewhere its own writer could not read.

The rewrite is now a labelled gate of its own, directly below the ACL gate and directly above the MULTI queue gate, in both handlers. Above it, ACL keeps matching the UNPREFIXED key the user typed (test_workspace_acl_grant pins that). Below it, every path — queue gate, script intercepts, cluster routing, blocking, ordinary dispatch — sees the prefixed argv. The one thing the hoist cannot fix on its own is the queue itself, which stores frames rather than argvs: those are now rebuilt from the rewritten cmd_args. The blocking branch of the same gate was broken identically and is fixed by the hoist alone, since queued_blocking_frame already rebuilds from cmd_args (moon#524) — measured pre-fix, a queued BLPOP popped off the GLOBAL list.

An argv whose key positions are indeterminate poisons the transaction at queue time like any other queue-time rejection, so EXEC cannot apply the valid half against the global keyspace. All eleven cases in tests/workspace_isolation_668.rs were confirmed to FAIL against a build of the previous main.

  • A workspace-bound connection can no longer reach the global keyspace (#668). src/workspace/mod.rs rewrote a command's key positions to prefix them with the workspace hash tag using a SECOND, hand-rolled key walker — four hardcoded name lists plus a handful of special cases, with a single-key args[0] default underneath. It had drifted from acl::keyspec::command_key_positions, the shared walker moon#582 consolidated the ACL layer and cache-invalidation onto.

A registry-wide sweep — not a re-reading of the file — put the drift at 50 commands, in two flavours. Positions the list walker MISSED were left unprefixed, i.e. read or written in the global keyspace from inside a workspace: BITOP (every position), SORT ... STORE, GEORADIUS*/... STORE|STOREDIST, the SOURCE of ZRANGESTORE / GEOSEARCHSTORE / LCS, everything after the first key of BLPOP / BRPOP / BLMOVE / BZPOPMIN / TOUCH / EXISTS, all of EVAL / EVALSHA's KEYS[], and the key behind the subcommand of OBJECT / XINFO / MEMORY. Since README sells workspaces as enforced multi-tenant isolation (GA in v0.6.0), each of those is a tenant break, not a cosmetic gap: SORT k STORE dst in workspace A wrote a global dst that workspace B could read and overwrite.

Positions it INVENTED were the other half: the single-key default prefixed args[0] on commands whose args[0] is not a key, so the command simply did not work inside a workspace — the numkeys count of ZUNION / ZINTER / ZDIFF / ZINTERCARD / SINTERCARD / LMPOP / ZMPOP / BLMPOP / BZMPOP, BITOP's operation name, XGROUP's subcommand, FCALL's function name, and SCAN's cursor.

The lists are gone; the file now reads command_key_positions and prefixes exactly the positions it reports. Three things a workspace partitions that ACL key patterns do not are applied as explicit policy on top, each documented where it is applied: FT.* / GRAPH.* index and graph names; the globs of KEYS and SCAN ... MATCH (a glob is not a key, but an unscoped one enumerates every other tenant — SCAN gets a workspace-scoped MATCH injected when the client supplied none); and SORT ... BY|GET <pattern>, whose run-time-computed lookups land inside the workspace once the PATTERN is prefixed, since it is expanded by plain * substitution. BY nosort and GET # are sentinels and are left alone.

An argv whose key positions cannot be determined is now REFUSED (ERR workspace: refusing '<CMD>': its key positions could not be determined from this argv) rather than run against the global keyspace. In practice that is reachable only through a malformed argv the command itself would have rejected a moment later; an unregistered command name still passes through, so dispatch answers unknown command.

Fenced by tests/workspace_key_walker_668.rs (a registry-wide sweep, so the next command added cannot re-open the gap by being forgotten, plus a curated corpus for every numkeys-counted layout a synthetic argv cannot reach) and tests/workspace_isolation_668.rs (a real server; every case uses a third, unbound connection as the oracle, so an escaped key is observed where it escaped TO). All eleven of the latter were confirmed to FAIL against a build of the previous main, each reproducing its own defect — BITOP answering ERR BITOP requires AND, OR, XOR, or NOT because its OPERATION had been prefixed, ZINTERCARD answering ERR numkeys can't be non-positive value, TOUCH answering 1 for a key that exists only in the global keyspace. - scripts/test-commands.sh can no longer run a whole suite against a server that never started. Both startup guards were structurally unreachable: rcli/mcli end in || true — right for the ~500 assertion rows, fatal for the health checks built on them, since mcli PING >/dev/null 2>&1 || exit 1 can never take its failure branch.

Measured against the pre-fix script with a binary that never listens: 86 rows ran and 10 of them PASSED, including FT.SEARCH does not error and FT.SEARCH stop-words-only returns no documents — an empty reply contains neither "err" nor "doc:", so those rows are satisfied by nothing at all. The same condition now aborts with 0 rows attempted and moon failed to start (no PONG on its port after 10s).

The probes match the actual PONG rather than an exit code, so a foreign process holding the port cannot fake liveness either, and they poll instead of a fixed sleep 1 — a cold first exec of a freshly built binary can take longer than that, which was the other half of the same bug. RUST_BINARY is now ${MOON_BIN:-./target/release/moon}, which makes the guard reproducibly testable (MOON_BIN=/usr/bin/false) and lets the suite be pointed at a known build instead of whatever happens to sit in target/release. - FT query errors are classifiable by a client (#691). FT.SEARCH / FT.AGGREGATE query errors reached the wire as bare snake_case tokens with no prefix — -numeric_filter_invalid, -syntax_error, -unknown_field.

A RESP error's first word is its code, and every client branches on it: redis-py surfaces it through ResponseError, redis-rs matches ErrorKind off the leading token, and all of them have rules for ERR, WRONGTYPE, NOSCRIPT, MOVED, LOADING. None has a rule for numeric_filter_invalid, so the entire message was swallowed as the code and nothing downstream could tell a client-side query mistake from a server fault.

They now read ERR <token>: <what went wrong> — e.g. ERR numeric_filter_invalid: numeric filter bounds must be numbers with min <= max, and the unknown-field error names the field the user actually typed, which matters in a query with several @field: clauses. The five tokens are a frozen contract, so they are kept verbatim inside the message rather than renamed out from under whatever still greps for them; only the prefix and the detail are new.

The echoed field name is user input going into a frame that is written to the wire raw, so every control byte in it is substituted. is_term_byte already excludes ASCII whitespace, so a CR/LF could not reach it today — but a reply that can desync a connection is not an edge worth resting on a rule enforced somewhere else. - FT.SEARCH <index> "*" — the match-all query (#693). * is how RediSearch says "every document in this index", and it is what every "show me what is in here" example uses. moon refused it on every index, and the refusal named KNN (ERR invalid KNN query syntax) even for an index with no vector field — so a user trying to enumerate a TEXT index was told the vector syntax was wrong.

* now routes to the inverted index, which is where the document registry lives, and answers from that registry rather than from a posting list. That matters for enumeration: a document whose text analyzes to nothing is still in the index, and * lists it where no term query can reach it.

Deliberately narrow. alp* is still a prefix query — only the token that is exactly * is match-all — and @field:* keeps its existing syntax error rather than silently widening to every document, since RediSearch has no field-scoped match-all either. *=>[KNN 10 @vec $q] is untouched; it is still routed by the [KNN marker.

Verified end-to-end at --shards 4 as well as --shards 1, so the scatter path is covered and not just the local one.

Still open, tracked as #695: an index built from VECTOR fields alone has no inverted index, so * there answers ERR no such index. Enumerating it means routing the vector engine's live key map through both the local handler path and scatter_text_search — and a fix that covered only the local path would pass at --shards 1 and silently return nothing at --shards 4. FT.AGGREGATE idx "*" was already correct (it short-circuits to the registry before the parser) and is not touched. - FT.SEARCH can find ordinary English words again (#690). Two independent defects produced one symptom — an empty result set, with no error and nothing for the user to act on.

The analyzer filtered against stop_words::get(LANGUAGE::English), which resolves to the stopwords-iso list: 1,298 entries, against RediSearch's 33. hello, world, test, name, order, open and index were among the ~1,265 extra words discarded at index time, so a document whose only word was hello indexed zero terms and was unreachable by its own content. moon now ships RediSearch's 33-word list, spelled out in text::analyzer::DEFAULT_STOP_WORDS rather than pulled from a crate, so what moon discards is greppable and cannot change under a dependency bump. The stop-words dependency is dropped.

Separately, a query token that analyzed to nothing evaluated to ∅ and was then intersected into its conjunction, so a stop word anywhere in a query zeroed the entire result set: on a corpus containing no stop words at all, alpha matched 2 documents and alpha the matched 0. RediSearch removes stop words from the query; moon now does too. The discriminator is "analyzed to nothing", not "matched nothing" — alpha zzz still correctly matches nothing, and TAG/NUMERIC/fuzzy/prefix leaves are never dropped.

Each half is load-bearing and was proven so: with the list swapped but the query path untouched, alpha the still returned 0; with the query path fixed but the list untouched, hello was still unindexable.

Existing indexes are not rewritten. An index built before this change is missing the ~1,265 terms the old list dropped; queries for them return 0 until the index is rebuilt (FT.DROPINDEX + FT.CREATE, or a restart that re-indexes).

Not addressed here: FT.CREATE ... STOPWORDS <count> <word>... is still unimplemented, so the list is not yet per-index configurable.

Added

  • DEBUG DIGEST (#636). A SHA1 fingerprint of the whole dataset, so two servers can be compared in one round trip instead of key by key. The crash-matrix and replication suites want it for moon-vs-moon (pre-crash vs recovered, master vs replica); the client-compat harness now uses it for moon-vs-redis.

The digest is byte-compatible with redis-server 8.6.1, which is the entire point: a moon-private digest could only ever compare moon to moon, and an almost-compatible one is worse than none — it reports two identical datasets as different and sends someone hunting corruption that is not there. Every constant was verified against a live redis rather than read off the source and assumed, which caught two details that are not guessable from the names: mixDigest is SHA1(digest ^ SHA1(data)), not SHA1(digest ‖ data) (the natural reading yields a codec that is perfectly self-consistent and wrong against every real redis, so a round-trip test cannot catch it), and the per-key fold hashes the 20-byte digest before xoring it rather than xoring the digests directly. Scores use shortest-round-trip formatting, not %.17g — 1.5/2.25/-0.125 agree under both and prove nothing; 3.3 is the only discriminating case and has its own test.

Verified live against redis 8.6.1 across all six value types, key TTLs, database placement, list order, and the empty dataset — identical digests at both --shards 1 and --shards 4. Sharding works because within a database every key contributes by XOR, which is commutative: each shard accumulates independently and the coordinator merges, folding each database index exactly once per server. The digest also does not depend on which storage plane a key lives in — proven with 32 keys spilled to the cold tier, same digest before and after (a digest that walked only the hot table would report a server as differing from its own replica purely because the two had evicted different keys).

Because the command answers from an intercept — dispatch has no context that spans every database and shard — it sits below each handler's ACL gate rather than above it. An intercept placed earlier answers before NOPERM is ever considered, and a whole- dataset fingerprint handed to a user who was denied DEBUG is the worst instance of that class: the command works, the digest is correct, and only a restricted user would ever notice. tests/admin_intercept_after_acl_shape.rs pins the ordering in all three connection handlers so the next privileged intercept cannot land above the gate.

The parity rows found a real divergence on their first run, which is the case for building an aggregate oracle rather than more per-command assertions: moon's FLUSHALL clears only the selected database (#677), so an "empty" moon still held four of the five databases the row had written. Around 450 per-command parity assertions had never seen it, because they all live in db0 where the wrong behaviour and the right one are identical. Both harnesses clear database by database until #677 lands, with a comment saying so.

Validating those rows also turned up three set -e defects in scripts/test-commands.sh itself, fixed here only because they decide whether the script can report anything at all: four redis-cli -p "$PORT" sites named a variable the script never defines (the #634 class again), two unguarded pkill calls that exit 1 when nothing matched, and hard-coded ports — so on a machine where another checkout already holds 6399, the "expected" side of every row silently came from a second moon. The remaining reasons that script cannot finish a run on macOS, and the vector rows that have never passed, are filed as #679 rather than fixed here.

Known difference, stated rather than papered over: redis serves DEBUG DIGEST inside MULTI/EXEC and Lua; moon does not. Those contexts execute inline and cannot reach the coordinator that spans every database and shard, so moon returns a clear error naming the limitation instead of a digest computed from the one database that path can see. DEBUG DIGEST-VALUE is deliberately not implemented yet — it needs per-key cross-shard routing that DEBUG's registry entry does not express, and one that silently answered for the wrong shard would be the same almost-compatible hazard.

  • DUMP and RESTORE (#636). moon can now serialize a value to a redis-compatible payload and load one back, which is the primitive behind key migration and per-key backup. Both were deregistered on 2026-08-15 because they were advertised in COMMAND while dispatching nowhere; they return here the only way that table allows — with the arms that serve them.

The payload is redis's own framing, <type><value><rdb version u16><crc64 u64>, verified byte-for-byte against redis-server 8.6.1: moon's existing crc64_jones reproduces redis's checksum exactly, and the checksum covers the version bytes as well as the body — a detail a round-trip test cannot catch, because a codec that gets it wrong still round-trips perfectly against itself.

Interoperability is one-directional, and the PR says so rather than implying more. A moon payload restores into real redis: measured, all five types, values identical after the hop (redis re-encodes the plain RDB types into its own listpack forms on ingest). The reverse does not work for collections — redis 8 emits listpack and quicklist encodings that moon has no decoder for. Those are refused with a distinct error naming the encoding, not with the checksum error, because the payload is not corrupt and sending an operator to hunt for corruption that is not there is its own bug. A payload stamped with a newer RDB version is refused before its type byte is ever read, which is the mechanism that keeps a redis 13 payload from reaching a decoder that cannot read it.

Semantics were read off redis 8.6.1 rather than inferred, including three that are easy to guess wrong: RESTORE validates its TTL before its payload; an ABSTTL in the past is +OK with the key already gone, not an error; and IDLETIME/FREQ are range-checked and then discarded, so FREQ 300 is refused even though a valid value changes nothing.

New fuzz target dump_payload (registered in both matrices in fuzz.yml) covers the decoder, which any client permitted to run RESTORE feeds attacker-chosen bytes. 1.92M executions, no crashes, flat memory — the guards that matter are the length-driven allocation bounds, since a valid CRC over a body claiming four billion elements is trivially constructible.

  • EVAL_RO and EVALSHA_RO (#636). The read-only script forms clients use to send scripts to a replica, or to prove to themselves that a script cannot write. Previously -ERR unknown command.

Both share EVAL/EVALSHA's entire path — parsing, cross-shard routing, the server-wide body fan-out, the caller's ACL — and differ in one bit handed to the executor: a write attempted from the script body is refused at the first redis.call, before it lands, not reported after the fact. The refusal machinery already existed for FCALL_RO; this wires the EVAL family into it. The read-only bit travels with a routed script, so EVAL_RO stays read-only when its keys live on another shard — the place it would most easily have been dropped.

Verified against redis-server 8.6.1 at --shards 1 and --shards 4: reads answered, writes refused, the value provably unchanged after each refusal, and plain EVAL still writing as the control.

  • HSTRLEN and MODULE (#636). Two commands clients reach for that moon answered with -ERR unknown command.

HSTRLEN key field returns the length of the field's value — not the field name, the one thing an implementation can plausibly get backwards — and 0 for a missing field, a missing key, or an empty value alike, exactly as redis does. It is wired into both the mutable and the read-only dispatch tracks.

MODULE matters because clients feature-detect with MODULE LIST on connect, and read -ERR unknown command as a broken server. moon has no module loader and is not growing one, so MODULE LIST answers the truth — an empty array — while LOAD/LOADEX/UNLOAD are refused in redis's own words (the text a stock redis gives when enable-module-command is unset). The container distinguishes redis's three refusals: a bare MODULE is a container arity error, MODULE LIST extra is a subcommand arity error named module|list, and anything else is ERR unknown subcommand '<as sent>'. Try MODULE HELP.

Both were diffed command-by-command against redis-server 8.6.1 — 17 probes byte-identical, including ACL enforcement, COMMAND GETKEYS, RESP3, and every arity and wrong-type edge. MODULE LIST is asserted moon-only rather than by parity: redis 8.x ships the vectorset module built in, so its list is legitimately non-empty. Two gaps found while measuring are filed separately, not papered over: unknown container subcommands are queued inside MULTI instead of aborting the transaction (#670, systemic across all six containers tested). - INFO MoonStore now reports the size and shape of the KV cold tier (#656). The whole # MoonStore section was one boolean, disk_offload_enabled, while the cold tier was the single largest consumer in the data directory. On the instance that motivated this — 1.43M keys, used_memory 1.49 GB, disk-offload on — the 15 GB data dir held 7.6 GB of heap-*.mpf, and nothing reported how big the cold tier was or how much of it was dead.

Two existing fields look like they answer this and do not. reclamation_cold_segments and its siblings sit under -- Vector segment tiers -- and count vector segments, so on a KV-only instance they read 0 while the KV cold tier holds gigabytes — worse than absent, because a zero reads as an answer. spilled_keys is a monotonic count of keys ever spilled: a rate, never a level.

# MoonStore
disk_offload_enabled:1
cold_keys:...                 live keys resident only on disk
cold_disk_bytes:...           on-disk bytes of live KvLeaf files
cold_files:...                heap files the manifest lists as live
cold_files_referenced:...     files holding at least one live key
cold_files_dead:...           on disk, referenced by nothing
cold_files_pending_unlink:... awaiting unlink by the sweep
cold_index_bytes:...          RAM the index costs (NOT disk)

This exposes existing state rather than adding bookkeeping: ColdIndex already maintained every input, and the manifest already recorded each file's byte_size. The values are published by the cold orphan sweep, i.e. once per --cold-orphan-sweep-interval-secs (60s default) rather than on the 100ms memory tick — summing the manifest's file entries is O(files), which on a G2-sized cold tier (100k+ heap files) is milliseconds of shard-thread time per tick for a number nobody reads at that resolution. All fields read 0 until the first sweep, and stay 0 when disk-offload is off.

cold_files - cold_files_referenced is the dead-space answer at the granularity reclaim actually operates on. Per-key dead space inside a partially-live file is not derivable: ColdLocation records where an entry is, not how many bytes it occupies.

The fields split by what their source can answer. cold_keys, cold_files_referenced, cold_files_pending_unlink and cold_index_bytes come from ColdIndex, which is per-shard state needing no manifest, and publish whenever a sweep has run. cold_disk_bytes, cold_files and cold_files_dead come from the manifest and are omitted if any shard's manifest failed to open — a reachable state, since ShardManifest::open failing leaves the shard running with disk-offload on, and spill still populates ColdIndex because the insert sits outside the manifest branch. Reporting those as 0 would place cold_disk_bytes:0 beside a non-zero cold_keys and present it as an answer. That shard also logs a warning, because cold keys with no manifest will not survive a restart — rebuild_from_manifest is the only thing that re-indexes them.

Fixed

  • FUNCTION dispatches inside MULTI instead of reporting an unknown command (#697). Every FUNCTION subcommand — valid or not — was queued by MULTI and then answered at EXEC with ERR unknown command 'FUNCTION', with args beginning with:, while the same commands worked outside a transaction. FUNCTION was the only container with this behaviour.

It broke the queue gate's safety argument, which is queueable iff dispatchable: COMMAND_META has FUNCTION, so the gate queued it, but shared::is_txn_connection_intercept did not list it, so the executor left no placeholder for the connection-owning caller to fill and the command fell through to the keyspace dispatch(), which has no FUNCTION arm. The fallback then claimed the command did not exist rather than failing loudly — the #639 class, with a lie on the end of it.

The registry could not simply be rebuilt where it was needed. Its RefCell is shared with the shard thread's SPSC drain loop, which applies inbound fan-outs, so it is per-shard-thread and that sharing is what makes a loaded library visible server-wide. It is threaded through try_handle_multi_exec → fill_txn_intercept_slots → run_txn_connection_intercept in both runtimes, and the EXEC-side intercept runs the same function_fanout_op + function_registry_fanout the live path does — a FUNCTION LOAD that dispatched but skipped the broadcast would apply to one shard and answer with the library name anyway. One implementation in shared.rs serves both runtimes rather than a copy each.

FUNCTION consequently joins #670's queue-time gate: FUNCTION BOGUS is now refused before it is stored and EXEC answers -EXECABORT. Measured against redis-server 8.6.1, that is what Redis does — #697's own table claimed Redis answers *1 + the error instead, and taking it at face value would have left FUNCTION ungated "to match Redis" and diverged in the opposite direction.

  • Every container command answers HELP, in Redis's shape (#698). Redis gives all 13 container commands a HELP subcommand; Moon served it on five and refused it on eight, each with a different error. Two of those (PUBSUB, XINFO) reported an arity problem, which reads to a client as "the subcommand exists, you called it wrong", and three (CONFIG, COMMAND, FUNCTION) refused the request with a message telling the client to run the exact command it had just refused. HELP is how redis-cli and several driver test suites discover a container's surface, and it is the fallback a user reaches for after a typo — precisely the moment Moon's own error text pointed them at it.

The five that "already worked" were only correct under a type-blind reading. Measured against redis-server 8.6.1 by RESP type, every Redis help reply is an array of simple strings opening with <CONTAINER> <subcommand> [<arg> [value] [opt] ...]. Subcommands are: and closing with HELP / Print this help.; OBJECT, MEMORY and SLOWLOG emitted bulk strings with no header line. All 13 now route through one constructor that emits the header and footer itself, so a container with a divergent shape is unrepresentable rather than merely untested.

The help body advertises what Moon dispatches, not what Redis does: copying Redis's text verbatim would have advertised four CLIENT KILL filters Moon silently ignores (parse_kill_args supports ID/ADDR/USER and the legacy addr:port only), plus COMMAND GETKEYSANDFLAGS and the recognised-but-unimplemented FUNCTION DUMP/RESTORE/STATS. Tests walk SUBCOMMAND_META in both directions so the two tables cannot drift. Also removes a dead second MEMORY HELP in key_extra.rs that had no callers and a different body, surviving only because pub silences dead_code.

Because the HELP name now appears in SUBCOMMAND_META, this also removes #670's one known consequence: <CONTAINER> HELP inside MULTI queues and runs instead of aborting the transaction. FUNCTION is the exception, blocked on #697, and is fenced by a test that fails once #697 is fixed rather than silently skipping forever.

  • Unknown container subcommands are refused at MULTI queue time, and every container refuses them with Redis's wording (#670). Redis validates a container's subcommand before storing the command in a transaction: CONFIG BOGUS is refused on the MULTI connection and the block is poisoned, so EXEC answers -EXECABORT and nothing runs. Moon replied +QUEUED and only noticed at EXEC, so the transaction ran — a client that treats +QUEUED as "this command is valid", which is what Redis guarantees, sent the rest of a transaction Redis would have refused wholesale and then applied the partial result.

The issue reported six containers. Sweeping all fourteen against a live redis-server 8.6.1 found the gate missed fourteen, and that ten of them also spelled the rejection differently: COMMAND said Unknown with a capital U (a one-character difference is still a different string to a client matching on it), SLOWLOG and XGROUP never named the offending subcommand at all (XGROUP reported a literal 'UNKNOWN'), and OBJECT and XINFO reported an arity error — which reads to a client as "the subcommand exists, you called it wrong".

All of it now goes through one err_unknown_subcommand helper producing Redis's single shape, ERR unknown subcommand '<as sent>'. Try <CONTAINER> HELP., with the echoed name control-byte-substituted (a subcommand arrives as a bulk string, so an un-substituted CRLF would end the error frame early and let the client read the remainder as a second, attacker-chosen reply).

The queue gate and each container's own dispatch guard now read one predicate, is_known_subcommand. That shared reading is the safety argument: the gate's contract is queueable iff dispatchable, and gating on the raw SUBCOMMAND_META table instead would have broken it — that table is a publication contract for COMMAND DOCS/INFO and deliberately omits FUNCTION DUMP, which dispatch accepts and answers.

Two containers are deliberately ungated, each fenced by a test rather than a comment: CLUSTER, because with cluster support disabled every subcommand including a bogus one is answered "cluster support disabled" so dispatch never reports an unknown subcommand; and FUNCTION, because of #697. Known consequence: <CONTAINER> HELP now aborts a transaction on Moon where Redis would run it — that is Moon's missing HELP (#698) surfacing earlier, and is the same treatment Moon already gave unknown top-level commands. - scripts/test-commands.sh: 26 failures down to 4, and the 4 name their issue (#683). The suite's failures were mostly the suite. Eight distinct defects, each verified against a live server before touching the row:

  • 9 TXN rows probed a transaction across separate connections. mcli spawns a fresh redis-cli per command, so MULTI, the queued commands and EXEC each landed on a different one; MULTI state is connection-scoped, so every row asserted against a transaction the server had already discarded. Added msession (one connection, one reply per line) and mreply N.
  • 10 vector rows queried with truncated blobs. The float32 vectors were built as "$(printf '\x00\x00\x80\x3f...')" and command substitution drops NUL bytes — a 16-byte vector arrived as 2. Replaced with NUL-free patterns at file scope behind a length guard that aborts the run if the shell ever eats one again.
  • MIXED %%machne%% deep asserted that AND behaves like OR. Measured against its own corpus: machine→fz:1, deep→fz:2, so the conjunction is empty by construction. Now uses a satisfiable one (%%machne%% learning→fz:1) and keeps the impossible one as the counter-assertion that proves AND is really AND.
  • TTL after EXPIREAT compared a live countdown across two redis-cli invocations ~100ms apart and failed whenever a second ticked over between them. Sampling both servers back-to-back 16 times (8 in each order) gave diff=0 every time, so there is no rounding difference to catch; new assert_match_countdown allows ±1.
  • FT.INFO after FT.DROPINDEX grepped for "err" or "not found". moon answers -Unknown Index name — a real error frame carrying RediSearch's exact text.
  • TAG-05 demanded the rejection moon used to emit for multi-tag OR; the feature landed and the row went stale. Now asserts the union equals the sum of the two single-tag queries.
  • NUMERIC-05 grepped for "min > max"; moon refuses the inverted range as -numeric_filter_invalid (#691). Now asserts what the row is for: refused, not executed.
  • The stop-words row unconditionally failed the 0 its own comment called acceptable.

What is left is 4 rows against real defects, each printing its issue number in the failure line: #536 (ROLE offset), #693 (no match-all query), and two on #690 (the 1,298-word stoplist). Triaging them is what surfaced #690, #691 and #693.

  • Client-compat freshness guard blamed files the binary never compiled (#687). The guard walks src/ and refuses under --strict when any .rs is newer than the binary — the right instinct (moon#461: a harness that tests the wrong binary reports a confident, false green) applied to the wrong set of files. src/ is not what went into the binary: handler_single.rs is #[cfg(feature = "runtime-tokio")], and the compat job builds the default monoio runtime, so rustc never opens it. A checkout stamped it 13 minutes newer than a binary cargo reported Finished ... in 0.57s for, and the gate refused to run — twice, on two separate dispatches, on a binary that was correct.

The guard now reads rustc's own dep-info (<binary>.d) — the record of every file it actually opened — and compares mtimes only against those. On the real CI inputs that is 438 of 478 src/ files; the other 40 are the tokio handlers, the console admin tree, the x86 SIMD kernels, the GPU module, and #[cfg(test)] files, none of which are in the binary under test. Verified both directions against the exact runner paths that failed: the CI binary is now accepted, and the same binary backdated three weeks is still refused, naming a compiled file (src/shard/spsc_handler.rs).

Missing or unparsable dep-info falls back to the old full-tree walk rather than to an empty set — an empty comparison set would read as "nothing can be stale", which is the failure the guard exists to prevent. A newly added source is still caught: rustc cannot see a file until some existing file declares mod for it, and that file is in the dep-info.

  • scripts/test-commands.sh runs to completion and prints its summary (#679). The suite CLAUDE.md points at for every new command had never reported totals: it died partway through, silently, and the operator saw a truncated log rather than a result. It now finishes — 504 rows, 478 passing — which is also how the 26 long-standing failures in it became visible for the first time (#683 tracks them; none are regressions from this change).

Every abort was the same bash shape: a command whose non-zero status set -euo pipefail turns into a silent exit. grep exits 1 when it matches nothing, lsof exits 1 when a port is free, pkill exits 1 when nothing matched, and a shell function whose last statement is a false if returns 1 too. Because the failure prints nothing at all, each instance hid the next — which is why this took several rounds and why two of the fixes are for aborts introduced by the earlier fixes:

  • 26 command substitutions ending in grep (the reported NUMERIC-07 site was one of a class, so the whole class is guarded);
  • 6 raw redis-cli pipelines and all four client wrappers, so a dead server produces failing rows and a summary instead of a truncated log;
  • cargo build ... 2>/dev/null, which discarded the reason a build failed and left the log reading Building moon... and nothing else;
  • the cleanup trap, which returned the status of its last kill rather than the script's — reporting a clean run as a failure.

Three defects behind wrong results rather than aborts:

  • grep -Pzo "(?s)A.*B" (13 call sites) is GNU-only. On a macOS host grep is ugrep, which rejects -P and exits 2 — and since the rows compared output rather than status, that 2 was reported as moon's answer. Replaced with a portable spans() helper.
  • No --dir, so moon used the CWD. The suite wrote appendonlydir/ and moon.lock into the repo root and reloaded the previous run's FT indexes, so a second run failed with ERR Index already exists. Each run now gets a fresh temp dir, removed on exit.
  • No port pre-flight. A leftover server from an unrelated run answers, and every row silently compares against it. This is not hypothetical — it produced a full run of MOONERR diskfull failures traced to another session's moon holding the port. The suite now refuses to start on an occupied port and names the holder.

The FT.CREATE ... VECTOR FLAT row expected OK, but moon has only ever implemented HNSW (ERR expected HNSW algorithm, in ft_create.rs since #27), so it had failed from the day it was written and took four dependent rows down with it. It now creates an HNSW index, and the FLAT gap is asserted explicitly instead of hiding inside a row that expected success. Added a regression row for #681 that checks the server is still alive after a truncated FT.CREATE — proven in both directions: it fails against a pre-#682 binary and passes after. - FLUSHALL clears every database, not just the selected one (#677). It behaved exactly like FLUSHDB: an operator who ran it believed the instance was empty while INFO keyspace still listed the other fifteen databases, each with its keys intact. Reproduced at --shards 1 and --shards 4 — not a routing problem.

dispatch hands every command a single &mut Database, so server_admin::flushall could not express "every database" no matter what it did, and the cross-shard broadcast faithfully replayed the same one-database flush on every other shard. The keyspace half now runs where the whole set is in scope, next to the index hook that already distinguished the two commands — which is what made the bug visible in the first place: a FLUSHALL dropped the search indexes for db1..db15 while keeping their keys.

Fixed at every site that can execute a flush, because a missing arm in one of them is invisible to CI (only the monoio handler ships, and only the tokio one runs on PRs): the three connection handlers, the sharded and single-threaded MULTI/EXEC paths, and all six SPSC arms — including MultiExecute, which is the arm the flush broadcast lands on and therefore the one that empties the other shards.

Two legs beyond the live path, both of which would have silently undone the fix:

  • AOF replay. The record is logged with the writer's selected db, so replaying it through dispatch restored every database the flush had emptied. Now intercepted before dispatch, the same way SWAPDB already was and for the same reason. A restart test covers it, and it fails against the un-intercepted replay.
  • Replication. apply.rs handled FLUSHDB and FLUSHALL with one arm; a replica applying a streamed FLUSHALL cleared one database and kept the rest, diverging from its master in fifteen databases until somebody SELECTed one.

Database::clear already drops the cold-tier index and the in-flight spill record along with the hot table, so a flushed key cannot return through a read-through or through a spill that lands after the flush.

FLUSHDB keeps its single-database scope, and there is a counter-test for that on both the local and the replicated path — "clear every database" is a one-character mistake away from making FLUSHDB destructive.

Known gap, tracked as #685: redis.call('FLUSHALL') inside a Lua script still clears only the script's database. The script bridge holds one &mut Database and already runs inside with_shard, so reaching the full set from there is a re-entrancy problem rather than a missing call.

  • A truncated FT.CREATE no longer aborts the server (#681). FT.CREATE idx ON HASH PREFIX 1 d: SCHEMA v VECTOR HNSW — the argument list cut off right after the algorithm keyword — indexed one past the end of argv and panicked. The panic ran on a shard thread, and moon deliberately escalates a shard panic to a whole-process abort rather than serve on with a dead shard, so one short line from any client took the entire server down: every database, every other connection. No auth and no large payload required.

The parameter loop below the fault already guarded both ends (*pos + 1 < param_end && *pos + 1 < args.len()), so the value read for every keyword was safe; the parameter count read was the single unguarded one. That was measured, not assumed — six truncation shapes were probed against freshly spawned, listener-PID-checked servers, and only this one killed the process. It reports ERR invalid param count, the same error an unparseable count already produced, so the two ways of failing to supply a count are indistinguishable to a client. There is no redis oracle for the string: the redis-server checked against has no query engine, so FT.CREATE is unknown command there.

FT.CREATE argument parsing had no fuzz target for its whole life, which is how a one-line remote crash survived in it; fuzz/fuzz_targets/ft_create_args.rs now drives the real entry point with arbitrary argv and is listed in both matrices in fuzz.yml (an unlisted target never runs — #576).

  • Lua script errors reached the client as an unparseable RESP frame (#672). mlua's Display carries a multi-line Lua traceback, and a RESP simple error may not contain CR or LF anywhere — so every runtime error (redis.call('INCR', k) on a string, error('boom'), a nil index, an unknown command) produced a frame redis-cli rejected with Bad simple string value. The client never saw what went wrong. The same frame quoted a moon source path (src/scripting/mod.rs:496:1) back at the client.

Three parts, each measured against redis-server 8.6.1:

  1. The frame is now a single line — the traceback's remaining frames say nothing a client can act on, which is what the NOPERM arm had always assumed.
  2. The Lua chunk is named (@user_script, @user_function — the names redis uses), so no moon source path can appear in an error at all.
  3. A redis error code raised by redis.call now leads the reply, as it does on redis. WRONGTYPE, OOM, NOSCRIPT and friends were buried behind ERR Error running script: runtime error:, so a client testing for WRONGTYPE saw a plain ERR and could not tell a type clash from a bug. moon already special-cased NOPERM and BUSY for this reason; the rule is now general, keyed on redis's own error-code shape rather than an allowlist that someone must remember to extend.

FUNCTION LOAD carried an independent copy of the same defect and is fixed with the same helper. - EVALSHA <sha> with no numkeys answered NOSCRIPT instead of an arity error (#636). redis rejects on arity (-3) before it ever looks the sha up. A client told NOSCRIPT re-runs SCRIPT LOAD and retries the same malformed call — forever. - ci-local.sh's disk pre-flight passed while the volume backing the VM was full (#661). The guard added in #659 sampled free space inside moon-dev and nowhere else. On 2026-08-22 it printed local-monoio target 41G free 18G OK — the VM was telling the truth — while the macOS volume holding the VM's auto-expanding disk image sat at 4.1G of 460G. Nothing could be written, so both suites exited rc=1 with no error text at all, which reads exactly like a test failure. A guard that goes green on the one signature it exists to catch is worse than no guard.

The pre-flight now judges both sides and refuses on either. Three details that each cost a measurement:

  • df -PB1 is a GNU spelling. macOS df rejects -B and prints usage, so a host probe written that way reads empty forever — and under the existing "cannot measure, step aside" policy that is a check which never runs. The probe uses df -Pk. When the host volume genuinely cannot be read the pre-flight now says host NOT CHECKED out loud rather than passing quietly. The test for this is deliberately platform-aware: the arm that pins the spelling points at /, which every host mounts, and the macOS-only volume is asserted only where it exists. A first draft asserted the default path unconditionally — true on the machine that runs ci-local, false on the Linux runner that gates the PR.
  • Not df /. / is the sealed system snapshot. It shares the APFS container's free space so Available matches, but Capacity does not (24% vs 93% on the same machine) — anything keyed on the percentage is wrong there.
  • A warm leg is cheap. Measured: one warm tokio leg (5143 tests) grew the image 90.2 → 90.5 GB. An earlier draft of this guard charged 12G per warm leg — inferred from a 124G → 147G growth figure that was really a cold run — and refused to start on a host with 33G free, an hour after ci-local had passed on that same host. Only cold legs, which materialize a whole ~34G target dir, are charged the full amount.

Also: on any failing leg the summary now prints host and VM free space before the re-run advice, and the reclaim instructions no longer promise what fstrim cannot deliver — it reports every unallocated block it trims, which is not what the host gets back (82.8 GiB reported, under 1G returned).

  • MQ POP no longer strands messages it claimed but never returned (#652). MQ POP <key> COUNT <n> asks read_group_new for n + max_delivery_count entries so that dead letters do not consume the caller's budget, then returns at most n. The surplus was left in the group PEL with last_delivered_id advanced past it — and MQ reads only > (new) entries and exposes no XCLAIM/XAUTOCLAIM path, so those entries were unreachable forever. No error, no DLQ entry, no metric.

With the default MAXDELIVERY 3, a COUNT 1 polling loop — the ordinary shape for a queue worker — destroyed up to 3 messages per call against any backlog deeper than 1. The reporter measured 28–62% of items never processed, daily, for over a month. A message pushed later is delivered normally, which is what made it read as intermittent loss rather than a dead queue.

handle_pop now walks the claim in id order and releases everything from the first entry it can neither deliver nor dead-letter: removed from the PEL and the consumer's pending set, with the group cursor rewound to the last entry actually kept. The MqPop WAL record carries only the kept ids and the post-rewind cursor, so replay cannot resurrect the stranding — proven by a kill -9 round trip, since the in-memory state and the WAL record have to agree and only a restart shows whether they do.

Two tests that had been loosened around the bug are tightened: the COUNT test now asserts conservation against the pushed set instead of len() >= 2 (its comment had documented the over-claim as expected behaviour), and a new test drains an 8-deep backlog with COUNT 1 and requires every id back, in order.

Not fixed here, and now tracked as #663: max_delivery_count is meaningless in MQ — nothing ever redelivers, so the value tested against it is always 1, and the shipped >= comparison makes MAXDELIVERY 1 dead-letter every first delivery so such a queue never delivers anything. Correcting the comparison alone would make the DLQ branch unreachable for every ceiling; it needs a redelivery path first. The predicate is extracted as should_dead_letter with that trap documented at the one place it must be fixed. - FT.SEARCH KNN prefilter: an inverted numeric range no longer aborts the server, and an unreadable filter is an error instead of an unfiltered search (#664, #648).

FT.SEARCH has three numeric-range grammars. The full query grammar and the FieldFilter grammar both read (-prefixed exclusive bounds and reject an inverted range. The KNN-prefilter grammar did neither — its numeric branch was a bare parse::<f64>().

#664 (process abort). @field:[300 100] reached BTreeMap::range(min..=max), which panics by contract when start > end, and a shard-thread panic aborts the whole process. Any client able to issue FT.SEARCH could kill the server — every connection and every other shard — with one command. Measured on main: the reply never arrives, the moon process is gone. This violated the parser-defensiveness rule that malformed client input must never crash the server.

#648 (silent widening). On any bound it could not read, the parser returned None for the entire filter expression, and the caller read None as "no filter was asked for" and ran an unfiltered KNN. @vt:[100 (300] against three documents returned 3 rows instead of 2 — no error, no warning, no metric. A filter that has silently stopped filtering is indistinguishable from one that legitimately matched everything, so the caller cannot detect it. On a scoped or multi-tenant store that is a confidentiality bug, not a recall bug.

Three changes:

  • FilterExpr::NumRange carries min_excl / max_excl, matching QueryNode::Numeric and FieldFilter::NumericRange, so all three grammars now agree on what [100 (300] means. [v v] collapses to NumEq only when both bounds are inclusive — [(50 50] is the empty set, not "equals 50".
  • The prefilter parser rejects an inverted range with the same rule the full grammar uses, and the evaluator returns an empty bitmap rather than calling BTreeMap::range with an impossible range. The parser is the contract and the evaluator is the backstop; this bug existed because there was only ever one of the two, and FilterExpr is reachable from more than one parser.
  • parse_filter_clause / parse_inline_filter return a tri-state (Absent / Parsed / Invalid) instead of an Option that conflated "no filter" with "unreadable filter". Invalid becomes ERR invalid FILTER expression. It also short-circuits rather than falling through from an unreadable explicit FILTER to the inline prefix.

⚠ Behaviour change: a query whose prefilter cannot be parsed now fails instead of returning results. Queries that appeared to work were returning a wider result set than they asked for. - FT.SEARCH ... HYBRID ... FILTER NUMERIC with an inverted range no longer aborts the server (#669). The same defect as #664, one command away, and still live after #664's fix: that fix hardened the vector payload index, while HYBRID's FILTER NUMERIC <field> <min> <max> filters through the text index's numeric BTree — a different evaluator with the same BTreeMap::range call and no guard. Its parser checked that each bound was finite and never that they were ordered.

Measured at --shards 1 on a build with only this fix reverted: one query produced Abort trap: 6, thread 'shard-0' panicked ... range start is greater than range end in BTreeMap, and Connection refused for every client thereafter.

Fixed in two layers, because guarding the caller is what left this reachable in the first place: TextIndex::search_numeric_range is now total — an impossible range is an empty result, compared in OrderedFloat space so the infinities and a stray NaN are covered — and the FILTER NUMERIC parser refuses min > max by name (ERR FILTER NUMERIC min is greater than max), since silence would tell the caller their filter matched nothing.

The inverted-range rule also lost its min.is_finite() && max.is_finite() && conjunct in the two grammars that carried it. The conjunct excluded nothing real — [-inf +inf] and every half-open form already satisfy min <= max — while letting [+inf 5] through, so the three FT.SEARCH numeric grammars disagreed on exactly the case #648 exists to unify.

Changed

  • scripts/ci-local.sh now pre-flights the VM's free disk before it starts (#658). A full VM root does not announce itself: a run reported VM tokio suite FAILED (rc=1) with no error text under it, and the next orb run answered sconrpc ready event fired but socket was not connectible — OrbStack had wedged on a 97%-full root. The visible symptom was three TRY 1 FAIL / TRY 2 PASS retries that belonged to a different leg, the one that had passed 5885/5885. After reclaiming space the tokio leg ran clean in isolation (5110 passed, 254 skipped).

The check is now the first gate, ahead of cargo fmt, so a doomed run stops in about a second instead of forty minutes in. Thresholds are measured, not guessed (moon-dev: a 124G root, and each of the two leg target dirs 34G warm): refuse below the 8G floor where OrbStack itself wedges; require 36G for a leg whose target dir is absent or a stub (<5G); warn below 15G of headroom for a warm leg. A cold leg is charged against the second leg's budget, since it consumes its build before that leg starts. On refusal the biggest consumers and the exact reclaim commands are printed.

Two deliberate non-behaviours: failure to measure never blocks — a pre-flight that cannot read the numbers steps aside rather than grounding a healthy run — and --quick skips it, since that mode runs no VM legs. The decision is a pure function, so scripts/test-ci-local-preflight.sh exercises every band (including the 3G case that produced the original wedge) without filling a disk, and it runs in the Lint job so the thresholds cannot rot.

[0.8.7] — 2026-08-22

Fixed

  • GEORADIUS/GEORADIUSBYMEMBER implement STORE/STOREDIST (#645). Both answered a generic ERR syntax error for a clause COMMAND DOCS advertises and COMMAND GETKEYS already resolved the destination key for. The _RO twins refused it by name — with a bespoke message that read as though the writable forms accepted it — so the surface was self-inconsistent three ways. STORE writes the 52-bit geohash score, STOREDIST writes the distance in the unit the query used, and an empty result deletes the destination and answers :0, matching redis-server 8.6.1 value-for-value.

Two adjacent gaps in the same clause family went with it:

  • GEOSEARCHSTORE accepted WITHCOORD/WITHDIST/WITHHASH and silently dropped them. It now answers redis's ERR GEOSEARCHSTORE is not compatible with WITHDIST, WITHHASH and WITHCOORD options, and supports the bare STOREDIST flag it had never parsed.
  • GEOSEARCHSTORE read only the match list and threw away the reply frame, so a parse error answered :0 and deleted the destination key. It now reports the error and leaves the key alone.

The read-only twins still refuse the clause — that check is load-bearing now that the writable forms really do write — but in redis's own words (ERR syntax error) rather than moon's.

Because these are two-key writes whose destination is not the routing key, both commands joined the moon#592 cross-shard guard in the same change: at --shards > 1 a destination owned by another shard is refused with CROSSSLOT before anything is read or written, never misplaced. Measured with the guard removed: 24 of 24 constructed cross-shard placements acked :2 and left the destination empty.

One deliberate limit: STOREDIST is the only geo reply that exposes a raw f64 rather than a %.4f rendering, and moon's haversine can land ~1 ULP from redis's (190.44242984775795 vs 190.44242984775784 for Palermo; Catania matches exactly). The operation order already matches geohashGetDistance; the residue is below every other geo reply's precision.

  • Eight command families executed at queue time inside MULTI instead of at EXEC (#639). CONFIG, CLIENT, ACL, CLUSTER, SCRIPT, WAIT, PUBSUB and AUTH/HELLO are connection-level intercepts: they are answered above dispatch() because they need the connection's own state. The MULTI queue gate sat below them and consulted an exemption list, so a queued CONFIG SET maxmemory 123mb ran immediately, answered +OK where the client expected +QUEUED, and stayed in effect through DISCARD. EXEC then returned an array with no slot for it — a client that indexes the EXEC reply by queue position read every subsequent result off by one.
MULTI                          redis 8.6.1      moon before      moon after
CONFIG SET maxmemory 123mb     +QUEUED          +OK  (applied)   +QUEUED
SET k v                        +QUEUED          +QUEUED          +QUEUED
DISCARD                        +OK              +OK              +OK
CONFIG GET maxmemory           unchanged        123mb (leaked)   unchanged
GET k                          $-1              $-1              $-1

The gate now exempts only MULTI/EXEC/DISCARD/WATCH/UNWATCH, and AUTH/HELLO skip their own above-the-gate intercept while a transaction is open. The executor leaves one placeholder per queued intercept, and the connection — which is the only thing that can run them — fills those slots after the body commits. Running them after rather than before is what keeps an aborted EXEC (a broken WATCH) side-effect free: it returns *-1 and no intercept runs at all. The residual is a documented ordering divergence — an intercept's side effect lands after the keyspace body rather than interleaved with it — which no reply the client sees can distinguish.

handler_single (the single-shard tokio / embedded library path, not the shipped server) keeps the old queue-time behaviour for the six families it intercepts above its own gate; embedded.rs already steers transactional embedders to the sharded handler. SUBSCRIBE inside MULTI remains the inverse divergence: Moon refuses it at queue time, Redis queues it.

A queued HELLO records its protocol switch at the index the EXEC reply occupies in the outer batch, not at index 0. The intercept post-pass writes into a local one-element buffer, so the index it would otherwise derive is the START of the batch — and encode_response_batch would then re-encode replies the client had already been promised in RESP2. COMMAND DOCS, XINFO GROUPS and XPENDING all flip shape when that happens; HGETALL, GEOPOS, ZSCORE and CONFIG GET do not, which is why me12g uses the first group.

Two client-compat waivers are retired and become live parity assertions (multi_queues_config_get, empty_config_get_in_multi), and shapes_acl_getuser, auth_without_password_configured, verbatim_info, verbatim_client_list, acl_users_lists_registered_usernames and acl_help_is_not_a_missing_subcommand regain the multi context they had been excluded from for this reason.

  • Every blocking pop invalidated nothing, so client-side caches went permanently stale (#644). try_handle_blocking serves the keyspace-modifying half of BLPOP/BRPOP/BLMPOP/ BZPOPMIN/BZPOPMAX/BZMPOP/BLMOVE/BRPOPLPUSH and pushes its own reply — and it was the one write path with no invalidate_after_write call. That call is hand-copied at twelve other sites; the blocking path was a thirteenth nobody added it to, which is the same shape as #623 (eight hand-copied wake hooks) one layer over. A RESP3 client with CLIENT TRACKING ON that cached a list and had BLPOP drain it kept serving the old value forever. Measured against redis-server 8.6.1 at --shards 4, hash-tagged keys, three non-blocking controls proving the instrument:
                   before            after / redis 8.6.1
LPOP     [control] PUSH              PUSH
ZPOPMIN  [control] PUSH              PUSH
LMPOP    [control] PUSH              PUSH
BLPOP              (nothing)         PUSH
BRPOP              (nothing)         PUSH
BLMPOP             (nothing)         PUSH
BZPOPMIN           (nothing)         PUSH
BZMPOP             (nothing)         PUSH
BLMOVE src         (nothing)         PUSH
BLMOVE dst         (nothing)         PUSH
BRPOPLPUSH src     (nothing)         PUSH

The keys come from the reply, not the arguments, because a blocking command's arguments name candidates and only one is served. Redis invalidates the served key alone and nothing at all on timeout — both measured, both pinned: BLPOP k1 k2 with k1 empty leaves a client's cache of k1 intact, and a timed-out BLPOP pushes nothing. Reusing written_keys here would have re-introduced #584 (invalidating keys a command only read) on a new path.

This also corrects PR #583, which listed BLMPOP and BZMPOP as covered: the key extractor was taught about them and has passing unit tests, but the execution path never called it — which is why a unit test could not see the gap and an end-to-end probe could.

Changed

  • The consistency suite's null-type probes use hash-tagged keys (#637). BLMOVE %K %K-d and BRPOPLPUSH %K %K-d generated two keys that hash to different shards, so at --shards >= 2 moon correctly refused the move cross-shard (#570/#591) and the probe compared a routing refusal against redis's *-1 — a null-type assertion that was really measuring routing, failing for a reason it was never written to test. nulltype:{N} co-locates both keys at any shard count. Nine new tracking rows cover the blocking family, including the two directions a fix must NOT break: an unserved candidate key, and a timed-out pop.

Performance

  • Retiring a vector document from the payload index was quadratic in the field's cardinality (#614). PayloadIndex::remove and remove_field walked every distinct value bitmap of a field, because nothing recorded which bitmaps the document was actually in. That is harmless on color or status, and quadratic on sku, price, or any coordinate — fields whose distinct value count grows with the corpus — so a bulk delete or re-index got slower as the index got bigger. A per-document forward index (internal_id -> field -> values written) now makes the retire path proportional to what the document actually wrote, and emptied bitmaps are dropped rather than left behind to be walked forever (the moon#613 growth bug, in the sibling index).

Measured with the in-tree harness (bench_payload_removal_cost_vs_cardinality, #[ignore]d — run it with --nocapture), timing the removal of every document one at a time. Per-document cost used to double as the corpus doubled; it is now flat:

docs tag µs/doc before after numeric µs/doc before after
250 0.98 0.28 0.86 0.28
500 1.71 0.23 1.59 0.28
1000 3.28 0.30 3.06 0.35
2000 6.54 0.24 5.56 0.27
4000 13.50 0.26 11.74 0.28

Total time to retire 4000 documents: 53.98ms -> 1.03ms (tag), 46.94ms -> 1.13ms (numeric). The trade-off is memory: one entry per (document, field) written, holding the values. Tag values are Bytes sharing the buffer already stored as the bitmap map's key; numerics are 8 bytes each. - Every client disconnect swept the entire tracking table, holding the process-wide tracking mutex (#649). TrackingTable::untrack_all had no record of which keys a client had tracked, so it walked all of them — even for a client that had tracked nothing. The tracking table is deliberately one global mutex-guarded instance (a per-shard table drops cross-shard invalidations), so the sweep did not merely cost the departing client: every shard's invalidation path blocked behind it, and the stall grew with the table, capped at one million keys. untrack_all runs from eight disconnect / CLIENT TRACKING OFF sites, so any client closing its socket reached it.

A reverse index (client_id -> keys it tracks) makes teardown proportional to what the client actually tracked. Measured with bench_untrack_all_cost_vs_table_size (#[ignore]d), timing 100 disconnects of clients that tracked nothing, against a table of n keys:

tracked keys µs/disconnect before after
1,000 0.97 0.07
2,000 2.72 0.03
4,000 6.30 0.03
8,000 11.91 0.03
16,000 31.77 0.03

Linear before, flat after — so the win keeps growing toward the 1M cap, where the old path extrapolates to roughly 2ms of global-lock hold per hangup. Same remedy as #613 (text index) and #614 (payload index); the reverse index costs one entry per (client, key) currently tracked, which is the same bound the forward map already carries. - Every disconnect swept the whole pub/sub registry three times, under its exclusive lock (#651). unsubscribe_all / sunsubscribe_all had no record of which channels a subscriber had joined, so teardown walked every channel in the registry — including for a connection subscribed to nothing — while holding pubsub_registry.write(), which every shard's PUBLISH fan-out blocks behind. Two things made it worse than the tracking-table analogue (#649): the channel map has no cap, so there was no ceiling on the stall; and bare UNSUBSCRIBE reaches the same sweep, making it a command-rate path and not only a connection-churn one.

A reverse index (subscriber id -> channels joined) for both the exact and sharded namespaces. Measured with bench_unsubscribe_all_cost_vs_channels (#[ignore]d), 100 disconnects of connections subscribed to nothing against a registry of n channels:

channels µs/disconnect before after
1,000 1.50 0.007
2,000 3.04 0.005
4,000 6.01 0.005
8,000 12.89 0.006
16,000 25.40 0.006

patterns still sweeps: it is a Vec bounded by distinct patterns, not by subscribers.

[0.8.6] — 2026-08-22

Security

  • A blocking command with an oversized timeout killed the whole server. BLPOP k 1e300 built its deadline with Duration::from_secs_f64, which panics on a value it cannot represent, and moon turns a panic on a shard thread into a process-wide abort — so one unauthenticated line took the server down, every shard and every other client with it:
$ redis-cli -p 7833 BLPOP nokey 1e300
Error: Server closed the connection
cannot convert float seconds to Duration: value is either too big or NaN
FATAL: thread 'shard-0' panicked; aborting the whole process rather than serving with a
dead shard or a dead cluster control plane

Found while adding XREAD BLOCK (#595), which routes into the same deadline code — the defect itself predates it on the BLPOP float-seconds path. Both parsers now port the tail of Redis's getTimeoutFromObjectOrReply exactly, including the ORDER of its checks, which is observable: it converts to milliseconds first and only then judges the sign, so a negative timeout is -ERR timeout is negative and not the "not a float" text moon used to send. Measured against redis-server 8.6.1 with one fresh connection per case (a case that BLOCKS leaves its reply in the socket and silently desynchronises every later read on a shared one — the first run of this table was fiction):

                                  before              after / redis 8.6.1
BLPOP nokey 1e300                 process abort       -ERR timeout is out of range
BLPOP nokey 1e16                  process abort       -ERR timeout is out of range
BLPOP nokey 1e15                  parks               parks
BLPOP nokey -0.5                  not a float …       -ERR timeout is negative
XREAD BLOCK 9223372036854775807   parked forever      -ERR timeout is out of range
XREAD BLOCK 9223372036854775      parks               parks

Deadline construction itself is now total as well (deadline_after returns None instead of panicking), so nothing that slips past a parser can reach the abort. Both sides of the boundary are asserted, in-process and end-to-end: the unit tests pin every error text, and bsr12_oversized_timeout_errors_and_the_server_survives PINGs a live server after each case, because a panicking build passes every in-process assertion right up to the moment it aborts.

  • h2 DoS advisory RUSTSEC-2026-0258 and event-listener unsoundness RUSTSEC-2026-0221 were both live on main. h2 0.4.14 accepted and queued empty DATA frames without limit — a remote memory-exhaustion vector in a networked server. event-listener 5.4.1 allowed !Send tags to cross thread boundaries via StackSlot, which is uncomfortably adjacent to this codebase's core concurrency assumption: monoio tasks are !Send and moon's cross-shard reply path is built entirely on the premise that they never cross threads. Bumped to h2 0.4.18 and event-listener 5.4.2 (lockfile only — no source, API, or MSRV change; cargo check and cargo audit both clean, 401 crates scanned).

These went unnoticed because cargo audit / cargo deny do NOT run on ordinary pull requests — they run weekly, or when a PR happens to touch Cargo.toml. moon#598 touched Cargo.toml to register a bench, which is the only reason the advisories surfaced at all.

  • redis.call() inside EVAL/FCALL bypassed ACL key patterns entirely (#569). The dispatcher gates EVAL script numkeys k1 ... on the keys an invocation DECLARES — but a script need not declare anything, and redis.call/redis.pcall went straight to Database::execute_command with no permission check at all. numkeys 0 therefore made the outer key check a no-op and everything the script touched was unchecked: a ~app:* +@all user read and wrote any key in the keyspace with EVAL "return redis.call('GET','secret:x')" 0, and could run commands their -del/-flushall rules denied. Every inner command now runs under the calling user's identity, resolved once per script (acl::ScriptAcl) and re-checked per redis.call against the same check_command_permission + check_key_permission the dispatcher uses, with the same fail-closed acl::keyspec walker — so movable-key layouts (LMPOP, ZMPOP, SINTERCARD, ZUNIONSTORE, SORT ... STORE, GEORADIUS ... STORE, COPY, SMOVE, BITOP, XREAD STREAMS) are covered and an argv the walker cannot enumerate — including SORT k BY w_*, whose weight keys are computed at runtime — is DENIED. The check sits at the single choke point every script-issued command passes through, before the SELECT/read-only/replica gates, before the eviction gate and COW pre-image capture and before MONITOR is fed, so a denied command produces no side effect and no laundering shape helps: literal keys, .. concatenation, string.format, ARGV, table lookups, loops, closures, redis.pcall and nested Lua pcall all refuse identically, on EVAL, EVALSHA, FCALL and FCALL_RO. The caller's identity also rides the SPSC message to the shard a script routes to (#508), so the cross-shard copy is not a hole; the fail-closed default (ScriptAcl::deny) means a future script runner that forgets to supply an identity refuses every command instead of inheriting ~*. Denials answer -NOPERM ACL failure in script: ... on both surfaces. Unrestricted users pay one enum test plus one Acquire load of the ACL version counter per redis.call — no lock, no allocation, no key extraction: +1.9% per call in an interleaved same-machine A/B against the pre-fix binary (20 000 redis.call('GET') per sample, 9 reps, min-of-reps, empty-loop cost subtracted; dev profile, so the absolute +86 ns is inflated several-fold and only the ratio transfers). A key-restricted user pays the full check — one uncontended ACL read-lock plus the same key extraction and glob match the dispatcher already runs for every non-script command — measured at +33% per call on the same harness. That version check is what makes the cached verdict safe: an ACL SETUSER landing mid-script is honoured by the very next redis.call rather than a script outliving its own revocation.
  • Fuzz the shared key-position walker, and stop a bare .gitignore entry from hiding new fuzzers (#576). acl::keyspec::command_key_positions parses attacker-controlled argv on behalf of three consumers — ACL key-pattern enforcement, client-side cache invalidation, and command introspection — so one bounds bug is a remote panic in three places at once. PR #571's review had already found exactly that: a numkeys usize overflow that wrapped first + nk in release builds and sliced &args[1..0], reachable by any key-restricted authenticated user. The new acl_keyspec target asserts the properties those callers rely on — every reported position indexes args; At is never empty; and, the security-relevant one, Unknown and AtPlusComputed must reach ACL as Indeterminate, so a ~pattern user can never be granted a key whose name is computed at runtime (SORT k BY w_*). Proven non-vacuous by reverting the checked_add guard: the target reproduces the #571 crash from the seed corpus alone, minimizing to the original attack string LMPOP 18446744073709551615 a LEFT. Clean over 3.27M executions after restoring it. Two coverage gaps closed alongside: .gitignore matched a bare fuzz, which ignores only NEW files (already-tracked ones are unaffected), so the 17 existing targets stayed visible while the 18th would have committed clean locally and failed CI as "no such fuzz target"; and term_fst_sidecar had been present in the tree but listed in neither CI matrix, so it had never actually run.

  • ACL ~pattern restrictions were silently unenforced for most multi-key commands (#566). AclTable::check_key_permission read a command's keys from extract_command_keys, a hand-maintained match on the command name whose fallthrough returned an EMPTY key list — and an empty key list is not "checked less precisely", it makes the permission loop a no-op, so every ~pattern was ignored outright for any command the list forgot. Measured against a live server, a user restricted to ~app:* reached arbitrary keys through 21 distinct commands, in both --shards 1 and --shards 4: COPY, ZRANGESTORE (both positions), SMOVE's DESTINATION (it was listed, but as a single-key command), LMPOP/ZMPOP/BLMPOP/BZMPOP, SINTERCARD, ZDIFF/ZINTER/ZUNION/ZINTERCARD, SORT ... STORE, SORT ... BY <pattern>, GEORADIUS ... STORE, EVAL's declared keys and MEMORY USAGE. Key extraction is now derived from the command registry's key specs (COMMAND_META first_key/last_key/step), so a command that declares its keys is enforced automatically; hand-written arms remain only for layouts a fixed spec cannot express (numkeys vectors, positional STORE clauses, the STREAMS token, subcommand-shaped key positions). Extraction fails closed: a command that names keys but whose argv cannot be enumerated — or a command missing from the registry entirely — is DENIED with the standard NOPERM error and logged once per command name, so the next command that ships without a key spec fails safe instead of falling open. SORT's BY/GET patterns read key names computed at runtime and are therefore refused for key-restricted users (BY nosort / GET # are unaffected). Commands that genuinely name no key (PING, CONFIG, SUBSCRIBE, KEYS, the FT.*/GRAPH.* families, ...) are unaffected, and a registry sweep test now fails if a NEW command declares no keys without being reviewed. Unrestricted and ~* users still short-circuit before any extraction; the new path borrows key slices into a SmallVec and no longer heap-allocates per command.

  • A remote client could hang a shard thread and exhaust memory with a 4-byte inline command (#487). split_args_quoted — the path any inline command containing a quote takes — skipped only ' ' and '\t' between arguments, while its token loop terminated on the wider ' ' \t \n \r \x0b \x0c. A byte in the difference (\n, \r, vertical tab, form feed) sitting at a token boundary therefore made no progress at all: the token loop returned without advancing, an empty argument was appended, and the outer loop restarted at the same offset — forever, with the argument vector growing until the process died. A lone \r reaches the line because the CRLF scanner only terminates on the \r\n pair. Parsing happens before dispatch, so this was reachable pre-authentication: \r" followed by CRLF was enough. Both loops now read one is_inline_space definition, so they cannot drift apart again — the same six bytes C's isspace() accepts, which is what Redis's own sdssplitargs skips with. Found by the resp_parse, resp_parse_differential and inline_parse fuzz targets; guarded by an exhaustive test over every short line built from the bytes the splitter branches on (271k cases, ~0.03s), which without the fix allocates until the OS kills the process.

  • ACL bypass on the inline GET fast path (monoio runtime). An authenticated but restricted user could read any key with plain GET: try_inline_dispatch answered the *2 $3 GET shape straight from the shard map, running neither the ACL command check nor the ACL key-pattern check. A -@all user read arbitrary keys by name, and a ~app:* user read outside its pattern; no ACL LOG entry was produced. Writes were already gated on can_inline_writes (which folds in conn.acl_skip_allowed()) — reads were gated on nothing. Not single-shard-only: at --shards 4 every key hashing to the connection's own shard leaked (measured 24/160 and 44/160 in the new suite); --shards 1 — the config recommended for non-pipelined workloads — leaked 160/160. Reads are now gated on the same acl_skip_allowed() latch, so a restricted connection falls through to generic dispatch where both ACL checks run. The tokio handlers were never affected (they gate correctly at handler_single.rs / handler_sharded/mod.rs), which is why no CI job caught this: every CI test job builds tokio. New suite: tests/acl_inline_read_enforcement.rs (deny-all and key-pattern users, at --shards 1 and --shards 4).

Fixed

  • COMMAND LIST published no container subcommands, and omitted four commands Moon answers (#635). Redis enumerates every container subcommand as its own container|sub entry — 137 of its 411 COMMAND LIST names — which is why its LIST is larger than its COUNT. Moon emitted zero names containing |, so the two returned the identical number, and that equality was the tell:
                 COMMAND COUNT   COMMAND LIST   of which container|sub
  redis 8.6.1          274            411               137
  moon (before)        263            263                 0
  moon (after)         267            350                83

A client library builds its local command table from COMMAND LIST. Against the old Moon it refused to send PUBSUB CHANNELS (works, was unlisted — a pub/sub client that feature-detects concluded the server had no pub/sub introspection at all), and it never learned that CONFIG, CLIENT, XINFO and ten other containers take subcommands with their own arity. ASKING, READONLY and READWRITE were unlisted for the same reason. COMMAND INFO container|sub now resolves too, and a container's own spec carries its subcommands' specs the way redis's does.

COMMAND DOCS follows redis's split, measured on 8.6.1: an explicitly named container|sub resolves to a doc map, while the bare COMMAND DOCS keys only the top-level names (redis's answer has 274 keys and not one contains a |, because it nests subcommand docs under each container rather than flattening them the way LIST does).

COMMAND COUNT deliberately does not grow to match LIST — redis counts only top-level commands, and the two agreeing is precisely the symptom this fixes.

The 83 names are measured, not copied from redis. Each was probed against a live Moon with a per-container BOGUS control, because a container's refusal of an unknown subcommand is not one string but three shapes: MEMORY says "not supported", XGROUP says "not recognized", and OBJECT / XINFO / HOTKEYS answer an ARITY error, which a naive probe reads as "it exists". The first sweep believed all 31 of redis's CLUSTER subcommands were present — with cluster support disabled every one answers "This instance has cluster support disabled", hiding the real count of 18.

A third shape cost a real entry, caught in review: bare COMMAND GETKEYS answers ERR Unknown subcommand **or wrong number of arguments** for 'GETKEYS' — one string carrying both conditions — so the sweep read the first half and recorded an implemented subcommand as absent. Every redis-minus-Moon name was then re-probed WITH real arguments against a same-shape bogus control, which confirmed the other 54 absences and recovered this one.

Names Moon does not implement (ACL DRYRUN, CLIENT UNBLOCK, COMMAND GETKEYSANDFLAGS, SCRIPT KILL, every |HELP, the LATENCY and MODULE families, FUNCTION DUMP/RESTORE/ STATS, 13 CLUSTER subcommands) are deliberately absent, and command_list_subcommands::cl3 SENDS every published name and fails if one stops dispatching. HOTKEYS publishes none: redis models it as a container, Moon takes HOTKEYS [COUNT n] with no subcommands at all.

  • AUTH <password> answered +OK on a server with no password configured (#640). That is the standard probe for "does this server require authentication?", so a deployment that MEANT to set requirepass and did not looked correctly secured from the client side. The single-argument form is now refused when the default user carries nopass, in Redis 8.6.1's exact wording:
                            before          after / redis 8.6.1
AUTH wrongpass              +OK             -ERR AUTH <password> called without any password
                                             configured for the default user. Are you sure your
                                             configuration is correct?
AUTH default anything       +OK             +OK          (default is nopass -- redis agrees)
AUTH <user> <wrong>         WRONGPASS       WRONGPASS    (unchanged)

Only the one-argument form changed. With --requirepass set, moon already matched redis byte-for-byte on all three cases and still does — acl_auth_surface::au4 is the control that keeps the fix from becoming a regression on the path that was already correct.

  • ACL USERS and ACL HELP were unimplemented (#640) — and moon's own unknown-subcommand error said Try ACL HELP., pointing the reader at the one subcommand that did not exist. ACL USERS now lists the registered usernames (the same sorted table walk ACL LIST already does, rendering only the name). ACL HELP answers a help array that advertises only what moon dispatches: Redis also lists DRYRUN, which moon does not implement, and listing it would reproduce the same defect one level up. acl_auth_surface::au6 walks every name in the help text and asserts it dispatches, so the two cannot drift.

  • ACL GETUSER, XINFO GROUPS, COMMAND DOCS and COMMAND INFO sent flat Arrays where RESP3 Redis sends Maps and Sets (#631). Unlike #462 no shared conversion could fix these: each reply is built by its own handler and the structures are nested rather than flat-to-map — ACL GETUSER is a map whose values have per-key types, XINFO GROUPS an array OF maps, COMMAND DOCS a map OF maps. The handlers now emit the right Frame and the serializer downgrades it for RESP2, so one construction serves both protocols. Types swept on the wire against redis-server 8.6.1:

                     redis 8.6.1                              moon before
ACL GETUSER default  %{flags:~[…], passwords:*[], commands:$,  *[$,$, $,*[…], $,*[], $,*[$],
                       keys:$, channels:$, selectors:*[]}         $,*[$], $,$]
XINFO GROUPS st      *[%{name:$, consumers::, pending::,       *[*[$,$,$,:,$,:,$,$]]
                        last-delivered-id:$, entries-read:_,
                        lag::}]
COMMAND DOCS GET     %{get:%{summary:$, since:$, …}}           *[$, %{summary:$, since:$, arity::}]
COMMAND INFO GET     *[*[$,:,~[…],:,:,:,~[…],~[],~[…],~[]]]    *[*[$,:,*[…],:,:,:,*[…],*[],*[],*[]]]

Three of these carried field-set defects that a Map makes worse than the Array did, because a client reads a map by name rather than by position — so the fields were corrected in the same change, on both protocols:

ACL GETUSER    dropped `username`   Redis has not sent it since 7.0
               added `selectors`    was missing entirely; now the empty list Moon can honestly
                                    report, so `.get("selectors")` means "none" not "unknown"
               keys, channels       ONE space-joined string each, not a list of patterns
XINFO GROUPS   added `entries-read` Null — Moon does not track it, and Redis sends Null itself
                                    whenever it cannot determine the value
               added `lag`          computed by counting entries past the group's cursor
COMMAND DOCS   dropped `arity`      Redis puts arity in COMMAND INFO and never in DOCS

COMMAND INFO's Set-vs-Array halves were item (3) of the identity_command_info_known_and_unknown client-compat waiver; that item is now retired and pinned live instead. What remains waived there is content, not typing: Moon's registry carries thin acl_categories and no key specs. COMMAND DOCS likewise still omits group, complexity and arguments, which Moon's registry does not carry — omitting a field it has no value for is the gap; inventing one is not.

  • RANDOMKEY answered Null while DBSIZE reported keys, and never sampled more than one shard (#629). It was missing from the cross-shard coordinator, so it read the shard the connection happened to sit on. With one key in the db that is deterministic — on --shards 4, three connections in four never see it:
HELLO 3 · FLUSHALL · SET k v
DBSIZE     -> :1
EXISTS k   -> :1
KEYS *     -> *1 ["k"]
RANDOMKEY  -> _        <- 20 of 20 draws, one connection

KEYS, SCAN and DBSIZE all scatter-gather, so they disagreed with RANDOMKEY at the same instant and the whole thing read as data loss. Fresh redis-cli invocations HID it: each opens its own connection and SO_REUSEPORT spreads those over the shards, so only a client that keeps one connection — every real client — saw the run of empty replies.

RANDOMKEY now fans out, asking each shard for both its key count and one candidate of its own, and draws a shard with probability proportional to its count. Weighting by count rather than picking a shard uniformly is the difference between sampling the KEYSPACE and sampling the SHARDS, and {hash tag} co-location makes unequal shards the normal case. Null is still the answer when — and only when — no shard holds a key.

  • RANDOMKEY returned the same key for every call inside the same millisecond (#629). The draw indexed with current_time_ms() % total, not a random number, so a client polling it in a loop — the normal way to sample a keyspace — got roughly one distinct name per millisecond of wall time however fast it asked. Measured on --shards 1, where no coordinator is involved, so this is the clock defect alone:
64 keys, 300 draws on one connection      distinct keys
before                     (17 ms)        10 of 64
after                      (18 ms)        62 of 64

It now draws from the thread RNG, like HRANDFIELD, SRANDMEMBER and ZRANDMEMBER already did.

  • scripts/test-consistency.sh died 40% in and reported success. TOTAL was incremented by the null-type probes but never initialised, so set -u aborted the run at the first TOTAL=$((TOTAL + 1)) — and the EXIT trap's own last command then set the exit status, so the script exited 0 while never reaching the #592 two-key sweep, script routing, SWAPDB, the FT.* block, RESET or any restart loop. Every run since the moon#594 probes landed was reporting on roughly the first half of the suite. Both halves are fixed, and the run now completes: 390 checks at --shards 4, 388 at --shards 1.

Un-hiding the tail surfaced 8 pre-existing failures at --shards 4 and 2 at --shards 1, all present on the merge-base binary too (same-script A/B). They are tracked separately: the ROLE/master_repl_offset divergence under #536, COMMAND LIST omitting container subcommands, and the cross-shard BLMOVE and movablekeys-tracking cases under the #592 family. None is introduced here.

  • INFO, CLIENT LIST, LOLWUT and MEMORY DOCTOR sent a BulkString where RESP3 clients are told to expect a VerbatimString (#462). Swept every reply type against redis-server 8.6.1 on the wire, at --shards 1 and --shards 4, RESP2 and RESP3:
                       redis 8.6.1        moon (before)      moon (after)
  INFO [section...]    = (verbatim)       $ (bulk)           =
  CLIENT LIST [TYPE]   =                  $                  =
  LOLWUT [VERSION n]   =                  $                  =
  MEMORY DOCTOR        =                  $                  =
  MEMORY MALLOC-STATS  $                  (unimplemented)    (unimplemented)
  CLIENT INFO          =                  =                  =

RESP2 is untouched — all four stay BulkStrings there, which is what every existing driver reads. MEMORY MALLOC-STATS is in the table because it looks like exactly the same kind of human-readable blob and is NOT verbatim on 8.6.1; the policy is transcribed from the oracle, never inferred from what a reply obviously ought to be.

  • Command intercepts can no longer forget the RESP2→RESP3 conversion, because they can no longer perform it (#462). Two of the four above were not a missing table entry — they were the structural hole the issue names. A try_handle_* intercept answers the command itself and short-circuits the dispatch exit where the reply-shape policy runs, so each intercept had to remember to apply the policy by hand, and nothing made forgetting visible. CONFIG GET reached the wire as a flat Array where Redis sends a Map from the day RESP3 landed; CLIENT INFO was patched by hand and CLIENT LIST, one if below it in the same function, was not.

Intercepts now receive an InterceptReplies sink (src/server/conn/intercept.rs) instead of a &mut Vec<Frame>, and its push applies the policy unconditionally. The hand-written conversions at CONFIG and CLIENT INFO are deleted — they are the sink's job now, which the mutation test below proves. 41 intercepts across both shipped handlers were converted by changing the parameter type and letting the compiler name every call site. Three paths keep a plain vector, each with the reason written at its signature: the AUTH gate (runs before the command name is even extracted), the blocking pops (encode straight into the write buffer after the batch has flushed), and the cross-shard batch (reply arrives later; shape rides in RemoteMeta as a one-byte tag).

Cost on RESP2 is nil — the conversion returns before classifying when proto < 3. Proven load-bearing by mutation: with InterceptReplies::push stripped of its conversion, INFO, CLIENT LIST, CLIENT INFO and CONFIG GET all revert to the wrong wire type while LOLWUT, MEMORY DOCTOR and HGETALL — which reach the dispatch exit — stay correct. - HSCAN ... NOVALUES was accepted and silently ignored, so clients that asked for field names got values instead (#630). The option is how a client says "send me the fields, not the pairs", and it is what redis-py's hscan(no_values=True) and hscan_iter(no_values=True) pass. Moon returned the field/value interleave anyway:

HSET h f1 v1 f2 v2
  redis 8.6.1   HSCAN h 0 NOVALUES  ->  "0"  ["f1", "f2"]
  moon (before) HSCAN h 0 NOVALUES  ->  "0"  ["f1", "v1", "f2", "v2"]

Nothing errored, so the caller read v1 and v2 as field names and any HGET/HDEL built from them addressed fields that do not exist. This is the failure mode that does not look like a bug from either side of the connection.

  • The whole SCAN family accepted unknown options instead of refusing them (#630). That is what kept the NOVALUES drop invisible: it was never "unimplemented and rejected", it was silently agreed to — and every option added later would have inherited the same silence.
                            moon (before)   redis 8.6.1 / moon (after)
  HSCAN h 0 BOGUSTOKEN      full reply      ERR syntax error
  SSCAN s 0 BOGUSTOKEN      full reply      ERR syntax error
  ZSCAN z 0 BOGUSTOKEN      full reply      ERR syntax error
  SCAN 0 BOGUSTOKEN         full reply      ERR syntax error
  SCAN 0 COUNT 0            full reply      ERR syntax error
  SCAN 0 COUNT abc          full reply      ERR value is not an integer or out of range
  HSCAN h 0 MATCH           full reply      ERR syntax error
  SSCAN s 0 NOVALUES        full reply      ERR NOVALUES option can only be used in HSCAN
  SCAN 0 TYPE nosuchtype    empty result    empty result   <- still accepted, matches nothing

SCAN, HSCAN, SSCAN and ZSCAN had eight hand-copied option parsers between them — a _readonly twin of each on top of the four — and they had already drifted: ZSCAN alone refused a non-numeric COUNT, and it answered wrong number of arguments where Redis answers syntax error for a dangling MATCH. All eight are replaced by one parser in src/command/scan_options.rs, which knows per command which options are legal (TYPE is SCAN-only, NOVALUES is HSCAN-only) and reproduces Redis's error text exactly, including the detail that COUNT abc and COUNT 0 give different messages because Redis parses the integer before it judges the range.

  • A fired MQ trigger reached only subscribers that happened to land on the queue's home shard (#474). fire_pending_mq_triggers delivered its callback notification with a single publish_shared into the shard's OWN pub/sub registry. Triggers do fire on the queue's home shard — a hash tag routes every workspace key there — but that is a fact about the KEY, not about the SUBSCRIBER: a consumer that subscribes to mq:trigger:<queue> lands on whichever shard accepted its connection, and moon keeps one registry per shard. So a consumer was woken only when it had connected to the right shard: always at --shards 1, roughly a 1-in-N chance at --shards N. Silent by construction — publish_shared returns a subscriber count of 0 and the path only debug!s it, so a durable-queue consumer simply never woke, with no error and no metric. The trigger now goes out through the same fan-out PUBLISH uses (local registry + the remote-subscriber map + one batched NotifyPublish per target shard), extracted as notify_fanout::publish_fanout so keyspace notifications and MQ triggers share one copy of the rule rather than growing a second. Measured at --shards 4 with eight queues and one subscriber connection: 4 of 8 delivered before, 8 of 8 after (0 of 8 before under the tokio runtime).

  • XREADGROUP across shards claimed entries it never delivered (#605). moon routes a command to ONE shard and that shard runs the whole command against its own slice, so the other streams of a multi-stream read are invisible — read as "does not exist". For XREAD that was an under-read. For XREADGROUP it was data loss. Measured at --shards 4, 12 stream pairs:

XREADGROUP GROUP g cc COUNT 10 STREAMS a b > >
  -> -ERR The XREADGROUP subcommand requires the key to exist.
  ... and a's entry is now PENDING for consumer cc

read_group_new claims the local stream's entries into the PEL and THEN the command fails on the stream this shard cannot see. The client is handed an error and never receives the entry, but the entry is marked delivered-and-unacked: XREADGROUP > never returns it again and only XPENDING/XAUTOCLAIM recovers it. 10 of 12 pairs stranded an entry that way; plain XREAD answered from a subset in 10 of 12.

Behaviour change. A multi-stream XREAD/XREADGROUP whose streams do not all hash to one shard is now refused before routing, with the same CROSSSLOT refusal two-key writes already get (#592) — the refusal happens before execution, which is what stops the claim. Non-cluster Redis serves such a read; moon cannot, and answering from a subset is the worse of the two failures. Co-locate the streams with a {hash} tag, or run --shards 1, and they are served in full as before. Single-stream reads are unaffected: one key cannot disagree with itself. - Test port reservations are now exclusive across processes (#489). reserve_port deduped within one test binary, but cargo runs test binaries concurrently and the probe listener must be dropped before the server can bind it — a held plain TcpListener is exactly what stops a SO_REUSEPORT bind — so two binaries could hand out the same port. moon's client listeners use SO_REUSEPORT, so the second server then binds successfully: both processes stay alive and the kernel splits or hijacks the connections, with no bind error, no log line and no panic. Every reservation now also takes an exclusive flock under $TMPDIR/moon-test-ports/<port>/ via persistence::dir_lock::acquire, which the kernel releases when the holder dies — including SIGKILL, which the crash-matrix suites do deliberately, so a killed binary frees its ports with no cleanup step to forget. Cluster reservations claim the bus sibling too. Unix only; dir_lock is a documented no-op on Windows, where the in-process set still applies.

  • A blocked XREAD could be answered with another client's entries, or never answered at all (#620). wait_id is (shard_id << 48) | counter, minted by the registry of the shard the connection lives on, while the waiter is queued on the shard that owns the key — so two readers parked on one key from connections on different shards arrive with ids from different high-order spaces, and the queue is not in id order. try_wake_stream_waiter assumed it was, in two places: it handed take_waits an unsorted slice to binary-search (a search that misses ids which are present, leaving those readers parked until their BLOCK budget ran out), and it paired the survivors with their decisions by position, sending the entries computed for one reader's cursor to a different reader's socket. In a debug build the mispairing tripped a debug_assert_eq! instead, and moon turns a shard-thread panic into a process-wide abort — so every connected client was dropped:
thread '<unnamed>' panicked at tests/common/mod.rs:527:26:
server closed after 0 bytes while 1 replies were expected: ""

The ids are now sorted before the lookup and each entry is paired with its decision by wait_id, never by position. Only reachable at --shards >= 2; the list and zset wakers consume one waiter at a time and were never affected. - Eleven more test binaries drew ports from an unverified :0 probe (#489). Each carried its own copy of the same unique_port() — bind 127.0.0.1:0, read the kernel's port, drop the listener, hand it out — with no record of what it had already handed out and no floor on the value. Two draws in one binary can return the same number, and moon's SO_REUSEPORT listeners bind an already-taken port without an error, so the collision produces no bind failure and no log: the kernel simply splits (Linux) or redirects (macOS) one test's connections into another test's server. All eleven now call common::reserve_port, whose process-wide dedup set makes an in-process repeat impossible and which floors draws at 20000.

crash_matrix_per_shard_bgrewriteaof.rs carried a comment asserting that "two independent kernel allocations never return the same port" — they can, which is the whole reason the shared helper keeps a dedup set. Corrected in place.

Files: crash_recovery_wal_recycle_legacy, crash_matrix_per_shard_bgrewriteaof, cold_shadow_overwrite_resurrection, crash_recovery_orphan_sweep_readiness, vector_idle_unload, crash_recovery_vector_durability, crash_recovery_graph_durability, crash_recovery_disk_offload_no_aof, aof_multidb_kill9, crash_recovery_cold_del_resurrection, wal_group_commit. - A write inside MULTI/EXEC never woke a client blocked on the key it wrote (#606). Measured against redis-server 8.6.1 — one client on BLPOP k 5, another running MULTI ; LPUSH k v ; EXEC:

redis 8.6.1 :  woken at 0.514s   <- the push woke it
moon        :  5.013s            <- its own timeout, not the push

The EXEC executor runs body commands through its own path and reached none of the live write path's wake hooks, so the consumer was never wrong-answered — only late by its entire timeout, which turns a queue into a poller and reads as a flake. All three families were affected (BLPOP/BLMOVE, BZPOPMIN, XREAD BLOCK), 8 of 8 trials each.

The executor now records each producer's key and hands it to the caller, which raises the wake after the body — the same deferral Redis uses, so a waiter sees the whole transaction applied rather than a half-built body. All three EXEC callers are wired, including the forwarded (cross-shard) arm in spsc_handler, which is where a transaction lands most often at --shards >= 2 because TxnLocality routes a body to the shard that owns its keys.

  • Crash-matrix tests could silently run two servers on one port (#489). tests/crash_matrix_per_shard_aof.rs took a port from a local unique_port() — bind :0, read the port, drop the listener — and then offset it by +1/+2/+3 for its other three tests. Two tests could therefore land on the SAME port, and moon's per-shard SO_REUSEPORT listeners bind an already-taken port without an error, so one test's redis-cli traffic silently entered another test's server. The symptom was a "200 missing" total-loss false alarm in a crash-recovery assertion — a failure that reads as a durability bug.

The file's SERVER_TEST_LOCK could never close this: it serializes tests within one binary, and cargo runs test binaries in parallel. Converted to common::spawn_listening, which pairs reserve_port (process-wide dedup, floored at 20000) with a wait-until-it-actually-accepts loop that respawns on a fresh port when a child loses the bind race. The lock is kept, for resource contention rather than correctness, and its doc comment now says so.

Restarts inside a test keep the original port deliberately — that is the point of a crash-recovery test, and by then the port's previous owner is dead. - Two cluster nodes could silently share one client port (#505). cluster_client_bootstrap.rs and cluster_formation.rs each carried their own reserve_cluster_ports, which probed candidates by binding them, pushed the listeners into a local held vec — and then returned, dropping every listener before any moon process bound the port. The scan was also unbounded ((start..40000).step_by(3) inside a nominally 300-port window), so a test whose window was busy walked into its neighbour's window by construction.

moon's client listeners use SO_REUSEPORT: a second node on an already-taken port binds successfully and the kernel splits incoming connections between the two. A client that connected to what it believed was node A intermittently reached node B, which reset it — no bind error, no log, no panic. cluster_client_bootstrap flaked at 25–33% per run, with a different test failing each time.

Replaced both copies with common::reserve_cluster_port + common::spawn_listening_cluster. The reservation records the client port AND its +10000 bus sibling in the same process-wide dedup set reserve_port already used, draws client ports from [20000, 22700) so the bus space [30000, 32700) neither overlaps it nor reaches Linux's ephemeral floor (32768), and panics when its own 100-candidate window is exhausted instead of wandering. The spawn helper adds a post-accept liveness window: moon's cluster bus binds plainly and main.rs exits(1) on EADDRINUSE by design, so the loser of a cross-process client port collision dies on its own and is respawned on a fresh pair.

The fix the issue proposed — hold the probe listeners until the child has bound — is not available: a plain listener held by the test is exactly what prevents a SO_REUSEPORT socket from binding that port. In-process collisions are closed by construction; the cross-process residue is narrowed, not eliminated, and the helper says so.

  • Eighteen read commands answered as though a TIERED key did not exist (#610). STRLEN, TYPE, TTL, PTTL, EXPIRETIME, PEXPIRETIME, OBJECT ENCODING, DEBUG OBJECT, MEMORY USAGE, GETRANGE, SUBSTR, GETBIT, BITCOUNT, BITPOS, BITFIELD_RO, LCS, SORT_RO and MGET all read through Database::get_if_alive, which probes the HOT plane and stops. PFCOUNT had the same bug through a loader of its own (load_hll_readonly), reporting cardinality 0 for a spilled HyperLogLog that EXISTS answers 1 for.

A source-grep shape pin (tests/cold_read_through_shape.rs) now DENIES get_if_alive( inside src/command/, with a reasoned allowlist for the three genuinely hot-only readers (KEYS, SCAN, RANDOMKEY are plane-partitioned and enumerate the cold plane separately; GET does its own cold read after the hot miss). A second test fails if an allowlist entry outlives the call site it excuses. Both were mutation-verified: planting a hot-only read in a non-allowlisted handler fails naming the file and line, and a stale allowlist entry fails naming the file. A key that eviction spilled to the cold tier is absent from data, so every one of them returned the missing-key answer for a key that is present and readable. Measured on one instance, one shard, --maxmemory 4mb --maxmemory-policy volatile-ttl --disk-offload enable, 6,000 × 1KB values (spilled_keys 2,599) against a HOT control key with an identical value and TTL: STRLEN 1000 vs 0, TYPE string vs none, TTL 3598 vs -2, BITCOUNT 4000 vs 0, MGET the value vs nil — while EXISTS answered 1 and GET served the value in full on the same key.

0 and -2 are the worst possible shape: both are exactly what a MISSING key returns, so nothing distinguishes "tiered" from "absent". EXISTS said 1 while TTL said -2 in the same breath, and MGET contradicted GET on the same key. Which keys are affected is decided by the eviction sweep, so no caller can predict or avoid it.

get_readonly was the ONLY one of twenty-six call sites with a cold fallback, hand-written locally; Database::get (the &mut path) promotes via promote_cold_if_present, so the same command was right or wrong depending on which dispatch path served it. The fix therefore belongs to the accessor, not the call sites: a new get_if_alive_any_plane returns an EntryView from whichever plane holds the key, and Derefs to Entry so each handler is a one-word rename rather than a rewrite. Per-call-site fallback is precisely how this drifted — GET got fixed, the other twenty-five did not, and nothing failed.

The first cut of the fix traded one wrong answer for another: taking the value through get_cold_value alone drops the deadline, so TTL answered -1 ("no expiry") for a key written with EX. The in-flight spill plane (#459) is now consulted first because it is the only accessor that returns a whole Entry, TTL included. That regression was caught by measurement against a hot control — it is invisible to review, and the first A/B probe missed it too, because probing GET PROMOTES the key and every later row then measures a hot key. Re-run with one distinct tiered key per command, all nine wrong answers match the hot control and the two that were already correct are unchanged.

A first pass converted only the commands that first probe happened to cover, which left the TTL fix in place while its own siblings EXPIRETIME/PEXPIRETIME still answered -2, plus BITFIELD_RO 0, LCS 0 and MEMORY USAGE an error. Sweeping the accessor's remaining call sites rather than the command list closed those. KEYS and SCAN were checked and are NOT affected — both list all 6,000 keys — so they are deliberately left alone rather than converted on suspicion.

Two of the remaining differences turned out to be the PROBE, not the server: comparing absolute EXPIRETIME across two keys written at different moments, and comparing MEMORY USAGE between keys whose NAMES differ in length. A hot-vs-hot control settles both — two resident keys differ by exactly the 1-byte name-length delta, and judged against its own TTL a tiered key's EXPIRETIME - now equals its TTL exactly, matching the hot control. - The client-compat harness could test a binary that was not the code under test (#461). MOON_BIN defaulted to the first of target/release/moon / target-fast/release-fast/moon that existed — which is target/release/moon, the one this repo explicitly quarantines, while everyday work builds release-fast. The harness then reported the resulting divergences as ordinary FAILs with no hint that the binary was stale. Measured cost (2026-08-10, on a tree where the fix was already green): PASS=94 FAIL=32, including a divergence from a function deleted hours earlier; the same tree with MOON_BIN pinned gave PASS=128 FAIL=0. For a harness whose entire value is being the trustworthy oracle-backed verifier, that is the one failure mode it must not have.

Three changes: the default now resolves to the newest of the two build layouts rather than the first found; the resolved path, build time and size are always printed before the spawn (so a later crash still leaves the provenance on screen); and a binary older than the newest .rs under src/ warns loudly, or under --strict refuses with ERR_STALE_BINARY and exit 2. That mirrors the existing ERR_NO_ORACLE rule — a differential harness with no oracle proves nothing, and one testing the wrong binary is the same lie with extra steps. Non-strict warns rather than blocks, because an intentionally older binary is a legitimate run (a bisect, or a deliberately pinned MOON_BIN).

A source tree that cannot be scanned reports NOT stale rather than stale, so "unknown" never masquerades as a verdict.

  • XREAD BLOCK / XREADGROUP BLOCK were a silent no-op: never blocked, never woken (#595). Both readers parsed BLOCK, stepped over its value with an // ignored comment, and answered the null array immediately. That is worse than an unimplemented command, because the reply is indistinguishable from a legitimate "nothing new": every consumer loop in the wild — redis-py, go-redis, lettuce — degrades from a parked connection into a hot spin, re-issuing at full speed and generating unbounded request volume where Redis would have slept. Measured raw-socket against redis-server 8.6.1, both runtimes, --shards 1 and --shards 4:
XADD bt 1-1 f v ; XREAD BLOCK 2000 STREAMS bt $
  redis 8.6.1 :  *-1  after 2.044s     moon (before) :  *-1  after 0.000s
parked XREAD BLOCK 5000 STREAMS blk $ , XADD blk 2-1 g w at t=600ms
  redis 8.6.1 :  the entry, at 0.605s  moon (before) :  *-1  at 0.000s

crate::blocking::WaitFamily::Stream and a stream waker already existed; what was missing was everything that would have put a waiter in front of them. Four defects, each independently sufficient to keep the command broken:

  1. Nothing ever registered. is_blocking_command answers from the command NAME, which is enough for the eight blocking pops but not for XREAD — the same name is a plain read without BLOCK. Added is_blocking_command_args, used at every site that decides whether a command parks. The two MULTI queue gates deliberately keep the name-only predicate: inside a transaction a stream read is queued as its ordinary self and executed by the dispatch table at EXEC, which already ignores BLOCK (measured: MULTI; XREAD BLOCK 3000 …; EXEC → *1 *-1).
  2. The local write path could not wake anything. The wakeup gate was open-coded at ten dispatch sites; eight said is_list_producer || ZADD || XADD while the two connection handlers said only is_list_producer || ZADD. A reader blocked on a key its own shard owned was therefore unwakeable, while the same XADD arriving over SPSC woke it — a routing-dependent hang that reads as a flake. All ten now share wakeup::is_producer.
  3. The waker was built for a destructive pop. It stopped at the first served waiter and ran remove_wait + send(None) on every waiter it popped, servable or not. Both halves are wrong for streams: an XADD wakes every parked reader (measured — two clients on XREAD BLOCK 5000 STREAMS k $ both receive the entry), and a reader this XADD cannot serve must stay parked. The send(None) also fired on the re-check that runs immediately after a remote registration, so a cross-shard XREAD BLOCK would have answered null the instant it registered. It now decides each waiter while it is still queued and removes only the ones it can answer.
  4. $ had nowhere correct to be resolved. The parsing connection may not own the stream, so $ now travels as StreamSince::Latest and is bound by the shard that owns the key, with no .await between reading last_id and registering — which is the lost-wakeup argument, not a stylistic preference. Redis binds it the same way, and it is observable: block on $, DEL the stream, re-XADD at a LOWER id, and the waiter still times out.

Errors are answered at once rather than parked: -WRONGTYPE for a stream read on a string key, and XREADGROUP's missing-key / missing-group errors — the latter checked on the shard that owns the key, because otherwise XREADGROUP GROUP nope c BLOCK 800 STREAMS k > answered -NOGROUP when k hashed to the client's own shard and parked for the full budget when it did not.

The gate is a whitelist, and it covers ONE stream. Everything it does not positively recognise keeps reaching the dispatch table and answers exactly as it did before this change, so each clause is a narrowing and never a regression:

  • One stream. Moon cannot decide a multi-key command on one shard, so a multi-stream read would be deferred to one BlockRegister per key, on a different thread each — and every owner shard's post-registration re-check MUTATES: read_group_new moves entries into the consumer's PEL. The client's coordinator keeps the first reply and drops the rest, stranding the sibling's entries as delivered-and-unacked to a consumer that never received them (XREADGROUP > never returns them again; only XPENDING / XAUTOCLAIM does). Measured on --shards 4, 10 key pairs each holding one undelivered entry: 5 lost an entry that way, against 10/10 correct on redis 8.6.1 and on --shards 1. Gating on placement instead would make one command behave two ways depending on a hash, so the gate reads the argument the client wrote. Multi-stream BLOCK therefore still answers immediately, as it did before — no worse, and never lossy. Lifting it needs multi-key stream reads decided by a single owner, the shape #602 gave two-key writes; filed as a follow-up.
  • Wakeable ids only. $, or an explicit id, for XREAD; >, and only >, for XREADGROUP (a history read answers immediately even with BLOCK — measured). + is excluded because StreamId::parse maps it to StreamId::MAX, so a waiter parked on it could never be served by any XADD and would sit there until its client disconnected — a permanent park where moon used to answer instantly. > on a plain XREAD is excluded because registering it would bind the waiter at 0-0 and replay the whole stream on the first write.
  • Valid arguments only. A malformed COUNT must be an error, not a 300 ms wait ending in the null array. BLOCK is validated with Redis's own texts (-ERR timeout is not an integer or out of range, -ERR timeout is negative) and with Redis's own string2ll accept-set, which takes neither a leading + nor leading zeros — moon parked the full budget on BLOCK +300 and BLOCK 0300, both of which Redis rejects outright. BLOCK 0 still blocks forever.

The waker walks each queue once. Deciding waiter-by-waiter through an id-then-lookup loop was quadratic in the number of clients tailing one stream — the canonical fan-out workload — and XADD p50 against parked, never-servable readers crossed over Redis at about 2000 tailers:

waiters                0     500    1000    2000    3000
moon (ids-then-lookup) 48.9  314.7  1005.2  3511.9  7601.3 us
redis 8.6.1           209.8  878.1   717.9  3032.9  5259.0 us
moon (one pass)        46.0  167.8   335.9   685.9   530.2 us

At 3000 tailers a single XADD had been holding the shard thread for 7.6 ms p50, which on a thread-per-core runtime stalls every other client on that shard.

Non-vacuity proven by mutation: each component was reverted in turn and the corresponding tests confirmed to fail (predicate off → 7 of 11; XADD dropped from the producer gate → 4; premature-null restored → 4; $ binding disabled → 3; the whitelist widened → the two gate tests). handler_single (the embedded server) is unchanged — it implements no blocking commands at all, BLPOP included.

Not fixed here, and pre-existing: a multi-key stream read whose keys span shards is answered by the routing key's shard alone, which cannot see the others. XREADGROUP claims the local stream's entries and then under-reports or errors, stranding them in the PEL — measured identically on a binary built from the merge-base and on this branch (7 of 24 both ways; redis 8.6.1: 0 of 24). It needs the same single-owner design as the multi-stream BLOCK case and is filed with it.

  • XREAD emitted an unserved stream with an empty entry list; Redis omits it (#594). Measured against redis-server 8.6.1 after XADD pelA 1-1 f v and XADD pelB 1-1 f v:
XREAD COUNT 10 STREAMS pelA pelB 0 99999
  redis 8.6.1 : *1 …pelA…
  moon        : *2 …pelA… *2 $4 pelB *0        <- the extra

A client iterating the reply saw a stream it had to special-case as present-but-empty. Under RESP3 it is sharper still: since #593 the container is a Map keyed by stream name, where membership is the natural "was this stream served?" test, so an entry with an empty list asserts the stream is quiet — a claim the serving shard is not entitled to make. Fixed in both xread and xread_readonly, so the reply does not depend on which dispatch path served it. The all-unserved miss is unchanged (still the null array, moon#482), and this is not the XREADGROUP history rule: XREADGROUP … 0 on an empty PEL genuinely does answer {name: []} in Redis (verified), so the omission belongs to plain XREAD's "did anything arrive after this id" question alone. The parity suites never called XREAD with a mix of served and unserved streams, which is why it survived them.

One consequence worth naming: a stream on ANOTHER shard is invisible to the routed shard, which reads it as "does not exist", so a cross-shard multi-stream XREAD now omits it instead of reporting [name, []]. Both are wrong against Redis, which serves the entries — but the empty list was an active lie (it asserts the stream is quiet), where the omission merely says nothing. The underlying routing gap is pre-existing and filed with the multi-key work above.

  • Scripting state never left the shard that created it, so EVALSHA and FCALL failed depending on where the key landed (#515, #514). Moon keeps the EVAL script cache and the Functions library registry per shard. SCRIPT LOAD fanned out; nothing else did — so the state a command needs was, like moon#592's second key, simply not where the routing key sent the command. To an application author that is indistinguishable from corruption. Measured at --shards 4 on fresh connections, before the fix:

  • EVAL did not publish its body (#515). Redis caches an EVAL'd script server-wide, which makes "EVAL once, then EVALSHA by sha" a supported and very common idiom — it is what redis-py's Script wrapper does. One bare EVAL followed by EVALSHA of that sha on twelve other keys: ok=2, NOSCRIPT=10. The same body via SCRIPT LOAD first: 12/12.

  • FCALL was three defects stacked, each hiding the next (#514). One FUNCTION LOAD then eight single-key FCALLs: 5/8 CROSSSLOT, 3/8 ERR Function not found, 0/8 succeeded, and FUNCTION LIST on a fresh connection returned *0.

    1. FCALL still used validate_keys_same_shard, which demands every key hash to the connection's shard — the exact defect #508 fixed for EVAL. A single key cannot cross a slot; it just lives elsewhere.
    2. The FunctionRegistry was built per connection, so a loaded library was invisible to the very next connection, and FUNCTION LOAD had no fan-out either. This half reproduced at --shards 1 too.
    3. Callbacks were invoked with no arguments, so the documented function(keys, args) signature received nil and any idiomatic library died with attempt to index a nil value (local 'keys'). Only visible once (1) and (2) stopped erroring first.

The fix. EVAL publishes its body to every other shard once per distinct script, and FUNCTION LOAD/DELETE/FLUSH replay on every other shard (DELETE/FLUSH for the same reason as LOAD in reverse: a delete that reaches one shard leaves the library callable elsewhere, which is a worse lie than not deleting it). The registry is now per shard — a thread-local shared by every connection on that shard and by the shard's SPSC drain loop. FCALL/FCALL_RO route to the shard owning their key through the same helper EVAL uses, so there is one routing policy rather than two that can drift. call_function passes (keys, args); the KEYS/ARGV globals are still set, so EVAL-style bodies keep working. A genuinely cross-shard key set is still refused CROSSSLOT before anything is touched — a function runs against one shard's database and cannot reach another's, so deleting that check instead of routing around it would have converted #514 straight into #592.

The replay is acknowledged, and a partial replay is reported, not swallowed. Pushing to a ring only enqueues; the target applies it in its own drain loop. Returning +OK at push time makes FUNCTION LOAD a lie for as long as that queue takes to drain — a client that loads and immediately calls can have its FCALL reach a shard that has not installed the library yet. So each replay carries an ack that reports whether the shard applied the op (delivery is not application: a shard can receive a library and still reject it), and the whole ack loop runs under a single 2s budget rather than a fresh 30s per shard. When some shard did not apply it, a FUNCTION mutation answers MOONERR partialfanout FUNCTION <verb> applied on N of M shards — its ops are idempotent under retry (DELETE/FLUSH outright, LOAD when re-sent with REPLACE), so the truth plus a retry beats +OK over a registry that answers differently per shard. A partial EVAL publish does not fail the EVAL: the script still runs, only a later EVALSHA routed to the shard that missed it is affected, and NOSCRIPT is exactly what client libraries already handle by re-sending the full EVAL — which republishes, because the cache tracks "published" separately from "cached". Every failure is also loud (warn! + moon_xshard_fanout_drop_total{kind}).

A repair leg was designed first and rejected under adversarial review: a later routed call that hit NOSCRIPT / ERR Function not found would re-push the state from a shard that had it. Both errors are also what a shard says when it has ALREADY applied a FUNCTION DELETE the origin has not seen yet, so the repair could not tell "you never got it" from "you already dropped it" and would resurrect deleted libraries — spreading the divergence, since a diverged shard keeps re-seeding whichever shard owns the next key. The same mechanism un-did SCRIPT FLUSH, which is local-only. Divergence is now prevented at publish time and reported, never patched up after.

Known limitations, stated rather than implied. Registry mutations have no total order: two FUNCTION mutations issued concurrently from connections on different shards can apply in different orders on different shards, and nothing reconciles them — issue them one at a time until an ordering authority exists. SCRIPT FLUSH still does not fan out (pre-existing, genuinely unchanged here). And a client that interpolates values into script bodies now replicates every distinct body to all N shards instead of one, so that anti-pattern costs N× the cache it used to.

Covered in handler_monoio (the shipped runtime), handler_sharded (tokio) and the shard-side SPSC drain; handler_single is excluded deliberately — it only runs at --shards 1, where there is no other shard to publish to. Each half is proven load-bearing by mutation: deleting the EVAL publish, reverting the registry to per-connection, deleting the FCALL routing, restoring the zero-argument callback, swallowing a partial fan-out, and claiming the publish regardless of the outcome each fail a distinct subset of tests/script_function_fanout.rs; restored, it passes 13/13. Two vacuous tests were found and fixed that way — the first draft took the sha from SCRIPT LOAD on the server under test, the one command that already fanned out (it now uses a hard-coded SHA1 moon cannot influence), and the republish assertion counted log lines where one round already emits one per peer. - volatile-ttl eviction could spin the shard thread forever on an index-only victim (#600). find_victim_volatile_ttl returned the head of the deadline index verbatim, with no cross-check against the keyspace — the only sampler that does, because it is the only one that reads a maintained index instead of iterating occupied DashTable slots. Every caller's sole way to make progress is db.remove, so a (deadline, key) pair whose hot entry is gone is not a victim, it is a non-terminating loop: evict_one_with_spill returned true while removing nothing, before == after, current_total never moved, and the next iteration peeked the very same head pair. The failure mode is silent and total — no OOM reply, no log line, no metric: the shard thread spins at 100%, the instance stays over maxmemory, and every client on that shard simply stops being served. Reproduced with a unit test that runs evict_to_budget on a worker thread behind a 10 s watchdog: before the fix the watchdog fires (the call never returns); after it, the same call completes in 0.1 s having evicted every live victim.

Two layered guarantees, because "the sampler is now correct" and "the loop cannot hang" are different claims and the second must not depend on the first:

  • The sampler verifies and self-heals. find_victim_volatile_ttl now walks the index head and reaps any pair that is provably stale, using byte-for-byte the predicate the active expiry sweep already uses (src/server/expiration.rs): the entry is gone, or it carries a different deadline than the pair claims. Anything else is left untouched — in particular a valid pair that merely is not due yet, which is why no clock is read here at all. Reaping, not skipping, is what Database::drop_expiry_index_pair exists for and what its doc comment already warned about: without it the next peek walks the same head again. Termination is structural — each non-returning iteration removes exactly the pair it peeked, so the walk is bounded by the index length — and a further per-tick cap of 64 keeps a large leaked backlog off the 100 ms sweep; the reaping is permanent, so successive ticks converge, and exhausting the cap logs a warn! naming the writer-coverage bug rather than failing quietly.

  • The loop terminates regardless of the sampler. evict_to_budget now judges progress on an OBSERVATION rather than on the sink's claim: (estimated_memory, pending_spill_bytes, hot key count) sampled between victims. Sixteen consecutive victims that move none of the three answers OOM with a warn! instead of looping. A bounded, visible, retryable error is strictly better than a spun shard. The probe is a triple and not a byte count on purpose — async spill (moon#466) moves a value from the hot plane to the pending plane and leaves estimated_memory flat, so a bytes-only test would call a working reclaim stalled and convert it into a rejection; the counter also resets on the first real progress, so slow reclaim can never accumulate into a false OOM. Verified by mutation: with the sampler fix removed the loop stops spinning and answers OOM in 0.14 s, and with both removed it hangs.

    The backstop carries its OWN tests rather than only that combined mutation. Disabling it while leaving the sampler correct left the entire eviction suite green — no sampler can reach the backstop today, which is exactly why a defence built for the NEXT index-only victim would have rotted unnoticed before that victim arrived. The stall decision is therefore a pure function over two probes and a counter, and three tests pin the three properties the constant claims: it fires on the 16th consecutive no-op and not the 15th, each of the three fields counts as progress on its own (a bytes-only test would reject a working async spill), and progress RESETS the count, so intermittent-but-real reclaim never accumulates into a rejection. Both mutations — backstop never fires, and progress judged on bytes alone — are caught.

moon#599's accounting distinction is preserved exactly, and neither change touches it: reaping a stale pair is a repair, not an eviction — the key it names is already absent from the keyspace, so DBSIZE cannot move and no counter is recorded. Confirmed end to end on a live server in both destinations: with --disk-offload disable 2,415 victims were DROPPED (evicted_keys 2,415, spilled_keys 0, DBSIZE 3,585), and with disk offload on the same 2,415 were TIERED (spilled_keys 2,415, evicted_keys 0, DBSIZE unchanged at 6,000 and the tiered keys still readable) — evicted_keys + DBSIZE == keys written in both. noeviction is unaffected: its gate runs before any victim is selected, so it still answers OOM immediately and, deliberately, does not even reap the stale pair — it promises not to mutate the keyspace.

No producer of a stale pair is known today: self.data.remove appears exactly once in the codebase (inside Database::remove_hot, which un-indexes the entry's own pair), and debug_expiry_index_consistent is an oracle over every TTL writer. This is the defence being present on a path where the equivalent path already had one, because the cost of being wrong is an unreachable shard rather than a wrong number. - The two-key write family acknowledged success while destroying the data at --shards > 1 (#592). Moon routes a command to ONE shard — the owner of the key extract_primary_key picks — and that shard then executes the whole command against its own slice. first_key names the routing key, not every key the command writes, so for this family the other key was read from and written to the routing key's table, under the right name but on the wrong shard, where every normally-routed access is blind to it. RENAME alpha omega answered +OK with alpha destroyed and omega never created — the value simply gone, with no error, no log line and no metric a client could detect. Measured at --shards 4 against 12 pairs constructed to straddle a shard boundary: 12 of 12 lost the data, for all twelve commands — RENAME, RENAMENX, SMOVE, SINTERSTORE, SUNIONSTORE, SDIFFSTORE, ZRANGESTORE, ZUNIONSTORE, ZINTERSTORE, PFMERGE, GEOSEARCHSTORE and SORT ... STORE. At --shards 1 the loss rate is 0. A thirteenth shape had the same defect through Lua: EVAL "redis.call('RENAME', KEYS[1], ARGV[1])" 1 src dst — routing never sees a key that arrives via ARGV, so the script ran on the source's shard and returned +OK with the value gone. Such a command is now refused before anything is read, written or deleted, from the two key names alone: CROSSSLOT Keys in request don't hash to the same shard; co-locate every key of the command with a {hash} tag. Refusing is the only answer that cannot lose data — there is no window in which the value exists in neither place, no undo to get wrong under a race or a timeout, and no partial state to recover after a crash; a shard hop would trade the loss for a non-atomic RENAME plus two independent AOF records with no shared commit point. It is also what Moon already answers for every other operation it cannot perform atomically across shards (a MULTI/EXEC body spanning shards, a script spanning shards, MSETNX). Guarded on all three dispatch paths that can reach the keyspace at shards > 1 — handler_monoio (the shipped runtime), handler_sharded (tokio), and redis.call from Lua — each proven independently load-bearing by mutation; handler_single has no multi-shard mode. MULTI ... EXEC was already safe (analyze_txn_locality walks every key of every queued command) and is now covered by a regression test. Commands NOT affected and deliberately untouched: MGET/MSET/MSETNX/DEL/UNLINK/EXISTS/BITOP/COPY, which shard::coordinator already dispatches a leg per owning shard; and LMOVE/RPOPLPUSH/BLMOVE/ BRPOPLPUSH, the same defect owned by #570.

⚠ Breaking change. At --shards > 1, these commands now return an error where they previously returned success. Co-locate the keys with a {hash} tag — RENAME {user:42}:staging {user:42}:live — and the command works exactly as it does at --shards 1. Every affected call was silently destroying data before, so no working deployment loses behaviour; a --shards 1 deployment is unaffected entirely. - Lua numbers passed to redis.call produced an argv no command could parse (#569). Argument conversion reused the RETURN-value converter, so redis.call('SET', 'k', 5) handed the dispatcher a Frame::Integer — something no wire client can send — and the command answered wrong number of arguments. redis.call('SET', 'app:k'..i, i) in a loop, and redis.call('LMPOP', 1, 'k', 'LEFT'), both failed on that. Upstream Redis stringifies Lua numbers for exactly this reason; Moon now does too (float truncation unchanged: 3.7 still becomes 3). This is also why the #569 ACL gate had to be paired with it — acl::keyspec reads keys and numkeys out of the argv as bytes, so an integer frame in either slot was un-enumerable and a LEGITIMATE in-pattern LMPOP would have failed closed. - The hash-field TTL sweep was unbudgeted, expiry ticks ran with nothing due, and eviction spilled keys about to expire (#543, #552, #553) — three CPU/IO wastes on the 100ms shard-loop timer, the latency-critical path. - #543 — sweep 2 is now deadline-indexed and bounded. The hash-field reap collected EVERY HashWithTtl key in the database with a full table scan and reaped all of them, with no time budget, no sampling and no batch cap, every tick. It now pops due (min_expiry_ms, key) pairs off a hash_expiry_index — the sibling of #541's whole-key index — and is bounded three ways: the same 1ms wall-clock budget as sweep 1, 256 hashes visited per tick, and 256 fields drained per visit. A partly-reaped hash is re-armed at its still-due minimum and resumed. The index holds a LOWER BOUND on each hash's minimum, which makes it self-healing: only hash_set_field_ttl can lower a minimum (so only it must index), while a raised TTL or a deleted hash leaves a stale-EARLY pair that the sweep pops, finds nothing due for, and re-arms or drops. Deferring a reap is invisible to clients — the read path already filters ttls against the shard clock, so only the memory reclaim is deferred. The reap now also takes its clock from the caller: sweeping against a clock ahead of Database::cached_now_ms (which the hash read path filters with) could physically drop a field reads still considered live. - #552 — a tick with nothing due costs two head-peeks. The maybe_has_expiring_keys latch only ever saved a database with ZERO TTL'd keys; a TTL-heavy one entered the cycle every 100ms just to re-derive "nothing due". Both indexes are deadline-ordered, so one O(log n) head-peek each answers that definitively, and the gate asks each sweep against the clock IT reaps with. - #553 — eviction stops spilling keys that are about to expire. A victim under SPILL_TTL_FLOOR_MS (1s) from expiry cost a datafile write, an fsync, a manifest commit and a cold-index entry for a value reclaimed moments later — and cold TTLs are lazy-only, so nothing sweeps it afterwards. volatile-ttl made this pathological by construction: it selects the NEAREST deadline in the database, i.e. the victim least worth persisting. Such victims are now plain-dropped instead (same RAM reclaimed, zero IO) on both the sync-batch and async-spill paths, reported through the same on_plain_drop sink a plain eviction uses so the dual-plane DEL still reaches replicas and the AOF, and counted as progress so an all-drops batch is not mistaken for "cannot evict". Eviction already DELETES victims under every non-noeviction policy in redis, so this gives up at most one second of extra visibility on a key that is about to become invisible anyway; noeviction never reaches the path. New metric moon_eviction_expiring_drop_total counts the spills avoided. - Found and fixed alongside: a HashWithTtl arriving as a WHOLE value (RESTORE, replication apply, cold promotion) never raised maybe_has_expiring_keys, so unless some unrelated key in the same database happened to hold a TTL, its expired fields were never actively reaped at all. - New benches/expiry_sweep.rs measures one tick in all three shapes; the bounded-sweep and spill-guard tests were each verified to FAIL against the pre-fix behaviour. - Recovery re-walked and re-warned about dead warm-segment manifest entries on every boot, forever (#546). A warm vector segment is retired by deleting its directory, but the manifest entry was left Active on the stated grounds that recovery "already tolerates" a manifest entry whose directory is missing. It does tolerate it — by logging manifest references warm segment N but directory missing and moving on, every single boot, because recovery opened the manifest read-only and never wrote back. Nothing ever retired the entry, so the dead entries could only accumulate: a live instance logged 13,214 of these warnings against 55 real segment directories, and separately 5,096 failed to read mvcc.mpf for ownership attribution for the same vanished segments. Recovery is the only component positioned to observe the discrepancy, so it now heals it: entries whose directory is gone are tombstoned via the existing remove_file path and committed in a single manifest generation (one dual-root swap for the whole batch, not one per entry — a store with thousands of stale entries would otherwise pay thousands of swaps on the boot path). The existing two-axis retention in gc_tombstones still governs when entries are physically pruned. Tombstoning is safe in exactly the way deleting the directory already was: the data is gone, so no reader can be served from the entry either way. A commit failure is non-fatal and logged — recovery proceeds identically, since the entries are skipped regardless. Live segments are untouched, which the new test asserts explicitly in both directions (dead entries must be retired AND the live entry must stay Active), re-opening the manifest from its path so an in-memory-only retirement cannot pass. - GraphUnion segment merges were rejected forever on any index carrying duplicate vectors, so segments accumulated without bound (#546). verify_merge_recall gates every merge on a recall measurement that scored an ID-set overlap: it took the brute-force top-k ids and intersected them with the ids HNSW returned. That measure is undefined when distances TIE. Real corpora tie constantly — repeated text chunks, boilerplate, and hash fields whose vector was never written all produce exact-duplicate embeddings — and when a point has 99 duplicates the ground truth takes an ARBITRARY 10 of the 100 tied candidates while the graph takes a DIFFERENT arbitrary 10. Both answers are exactly correct; the overlap scores them 0.0, the merge aborts, and it aborts again on every retry because the input never changes. The live store recorded 5,596 merge attempts with zero successes, 1,367 of them at recall exactly 0.0000, with the retry backoff saturated at 480s — every affected index grew its segment count monotonically, which is upstream of both query cost and startup recovery time (recovery rebuilds per segment). Recall is now scored by distance equivalence: the k-th ground-truth distance is the acceptance threshold, so any neighbour at least that close counts, whichever of the tied ids it happens to be — which is what recall@k means when distances tie. The gate is not weakened: measured against a deliberately corrupted merge (layer-0 codes decoupled from their graph nodes) on an all-distinct corpus, the old id-overlap metric scores 0.9667 and the new distance metric scores 0.9673 — a 0.0006 difference — so the two are equally sensitive to real breakage and differ only on ties. The new test merges 8 segments drawn from 8 distinct vectors (100x tie sets), is red on the old metric at recall 0.0000, and asserts the merged segment is FUNCTIONALLY correct (every one of the k hits is an exact duplicate at ~0 distance) so a merely-laxer gate cannot pass it. - COMMAND GETKEYS answered "the command has no key arguments" for every movablekeys command (#537). getkeys() decided from meta.first_key <= 0, which is true of LMPOP, ZMPOP, SINTERCARD, ZINTERCARD, ZDIFF/ZINTER/ZUNION, BLMPOP/BZMPOP, EVAL/EVALSHA/ FCALL/FCALL_RO, XREAD/XREADGROUP, ZUNIONSTORE/ZINTERSTORE and the SORT/GEORADIUS STORE family. In redis's table — which COMMAND_META mirrors — that value means "the keys are not at a FIXED argument position", not "there are no keys", so Moon replied with an error whose text was itself false. COMMAND GETKEYS is precisely how a cluster-aware client routes a command it cannot parse itself, so such a client either refused to route these commands or guessed. getkeys() now delegates to the shared walker in acl::keyspec — the same one ACL ~pattern enforcement and client-side cache invalidation consume, so the three answers cannot drift apart again (the drift #537 called "the deeper problem"). All four of redis's error strings are reproduced, in redis's order: the no-keys check runs BEFORE arity, which is observable (COMMAND GETKEYS SELECT reports no-keys even though its arity is also wrong), and an argv whose keys cannot be enumerated now says Invalid arguments specified for command rather than lying about the command having none. EVAL/EVALSHA/FCALL/FCALL_RO carry redis's no-mandatory-keys flag, so EVAL <script> 0 — and an unparsable count — reply with an empty array instead. Verified on the wire against redis-server 8.6.1: 87/87 probes byte-identical on both runtimes (monoio and tokio) at 1 and 4 shards. Two deliberate divergences remain, both on argv that cannot execute anyway: a repeated STORE reports every destination where redis keeps only the last (a superset, so ACL still checks both), and a container command missing its key (OBJECT ENCODING) reports the no-keys error where redis reports its subcommand's arity error. - CLIENT TRACKING over-invalidated: the read-only SOURCES of *STORE commands were pushed as changed (#584). The invalidation hook used every key a write command NAMED; redis pushes only what it MODIFIES (signalModifiedKey). A client caching a was told a had changed by ZUNIONSTORE d 2 a b, SINTERSTORE d a b, BITOP AND d a b, PFMERGE d a b, COPY a d, ZRANGESTORE d a ..., GEOSEARCHSTORE d a ... or SORT a STORE d; worse, a plain SORT src (a write-flagged command that writes nothing without STORE) invalidated src unconditionally. Each spurious push costs the client a dropped cache entry and a refetch. The shared walker now reports a per-position KeyRole — the RO/OW distinction redis carries in its key specs and first_key/last_key/step cannot express — and invalidate_after_write uses only the writable positions. ACL is unchanged by construction: it ignores the role, because a ~pattern gates reads and writes alike. Roles are deliberately over-inclusive where the argv genuinely cannot say: LMPOP 2 a b LEFT pops from whichever key is non-empty at execution time and EVAL may write any key it was handed, so those stay writable — over-invalidating costs a refetch, under-invalidating leaves a cache stale forever (#582). Verified against redis-server 8.6.1 with destination CONTROLS on every case (a fix that simply stopped pushing would fail all of them): 25/27 identical on both runtimes at 1 and 4 shards, the two residuals being the known LMPOP limit above and GEORADIUS ... STORE, which Moon rejects as a syntax error today (a pre-existing gap, unrelated to this change). - evicted_keys counted keys that had NOT left the keyspace, so it contradicted DBSIZE (#585). On a live instance evicted_keys climbed by 456,018 while DBSIZE never moved. The two numbers were counting different things: with --disk-offload enable (the default) the durable batch spiller tiers a victim — it frees the hot copy, registers the key in the cold index, and the key stays readable and stays in DBSIZE (#355) — yet it recorded that as an eviction. The one metric an operator instinctively checks therefore stayed flat while the counter that is supposed to explain it ran away, which reads exactly like "DBSIZE never decrements". evicted_keys now counts only keys REMOVED FROM THE KEYSPACE, so it and DBSIZE move together; tiered keys are reported in a new INFO stats field, spilled_keys (Prometheus moon_spilled_keys_total), which is the counter to watch for memory-pressure activity on a disk-offload instance. Two further inconsistencies closed on the way: the async spill path recorded nothing at all (two spill paths, two different answers), and the plain-drop path incremented evicted_keys even when its sampled victim turned out not to be there to remove. Invariant now locked by test — DBSIZE falls by exactly the evicted_keys delta, and does not move for a spilled_keys delta. - CONFIG SET maxmemory 4gb was rejected (#586). Memory-unit suffixes are the spelling redis documents and every runbook, Helm chart and tuning guide uses; an operator applying a cap under memory pressure got ERR Invalid argument and had to convert to bytes by hand — which is exactly when they are least able to. maxmemory, db-maxmemory's byte half, the --maxmemory CLI flag and therefore moon.conf now share one parser reproducing redis's memtoull() exactly, including its two-scale rule (1k = 1000 but 1kb = 1024, same for m/mb and g/gb), its rejections (1.5gb, 4 gb, -1, +4gb, unknown units) and — verified against a live redis-server 8.6.1 rather than read off its source — its refusal of surrounding whitespace: redis rejects " 4gb", "4gb " and even "1024 ", so Moon does too. This also fixes startup: moon.conf synthesises CLI arguments, so the maxmemory 1gb line in Moon's own conf-file documentation used to fail the parse (the conf-file parser trims each value itself, so padded conf lines keep working). One deliberate deviation, in the safe direction: redis does not check strtoull for ERANGE, so it saturates on a bare integer and wraps with a suffix — CONFIG SET maxmemory 17179869184gb returns OK and reads back 0, silently disarming the memory limit. Moon rejects overflow instead, surfacing the typo. - used_memory did not count spilled-but-not-yet-written values (#466). A victim queued for async spill has its whole entry cost credited back to the ledger while a full copy of the value is still pinned in RAM — by the queued SpillRequest and, since #459, by the in-flight plane until the completion is applied. used_memory now carries that term (pending_spill_bytes, maintained O(1) at mark and at every retire), paired with the rule that makes it safe: once the pending bytes cover the overshoot, evict_to_budget treats the tick as done rather than selecting more victims to chase memory that cannot drop yet. Without the pairing the honest figure would have turned the eviction loop into a runaway that drains the whole database; that is why the visibility fix was deferred out of #459, and both halves now ship together under test. - A list MOVE to a destination on another shard acked the element and then discarded it (#570). BLMOVE/BRPOPLPUSH — and, unreported until now, their non-blocking twins LMOVE and RPOPLPUSH — are routed by their PRIMARY key, the source, and were then executed whole on that one shard: the pop came off the real source list and the push went into the SOURCE owner's slice under the DESTINATION's name. Every normally-routed read of the destination goes to the destination's OWNER and is blind to that entry, so the client received the moved element as its reply while the element left the keyspace. Measured against the pre-fix binary at --shards 4: BLMOVE lost 10 of 12 key placements, RPOPLPUSH 11 of 12, LMOVE 6 of 6 — the survivors are the placements where source and destination happened to be co-located. Silent, acked data loss on a two-key write, the failure mode a durable queue exists to prevent. Moon is shared-nothing across shards: a shard can pop only from keys it owns and push only to keys it owns, and there is no cross-shard commit to split a move across. Such a move is now refused with CROSSSLOT Keys in request don't hash to the same shard; co-locate source and destination with a {hash} tag, decided from the two key NAMES before anything is popped — so there is no window in which the element exists in neither list, no undo to get wrong under a race or a timeout, and no partial state to recover after a crash. This is the answer Moon already gives for every other operation it cannot perform atomically across shards (a MULTI/EXEC body spanning shards, a script spanning shards, MSETNX), and cross-shard COPY already degrades to an error naming {hash} tags for the types it cannot carry. Splitting the move across a shard hop was rejected as the alternative: it would trade the loss for a non-atomic LMOVE — an intermediate state Redis never exposes — plus two independent AOF records with no shared commit point. For the blocking pair the refusal is delivered immediately, before a waiter is registered, rather than after the source is served: shard ownership is a static property of the key names, so a client that blocked to its timeout only to learn Moon cannot route its move would have waited for nothing. Three independent guards cover the three execution sites (pre-registration scan, the wake path in blocking::wakeup, and pre-routing for the non-blocking pair); each was proven load-bearing by disabling it alone and watching exactly the expected assertions go red. Narrowness is enforced by the same suite: {hash}-tagged pairs still move at --shards 4, every move still works at --shards 1 for arbitrary key names, and the rotate form (LMOVE k k) is never refused. Users running --shards > 1 who move between untagged lists must co-locate the pair with a {hash} tag — previously those moves reported success and destroyed the element. - XREAD/XREADGROUP answered a RESP2 Array to RESP3 clients instead of a Map (#577). Redis 7+ types a non-empty stream read as a RESP3 Map keyed by stream name (%N <name> => <entries>); Moon built Frame::Array unconditionally, in every protocol, so a client that dispatches on the reply CONTAINER was handed the wrong type. Now classified as Resp3Shape::StreamMap in src/protocol/resp3.rs, which means the conversion happens at the single choke point every dispatch path already shares — including the cross-shard path, where the 1-byte shape tag rides along on the batch. Only the container re-types: entry ids and their field/value pairs stay BulkStrings, the empty-PEL history stays %1{name: *0}, and the miss stays a null array (_ on RESP3, *-1 on RESP2) in both protocols. RESP2 output is unchanged. The client-compat probe parity_xreadgroup_history_resp3_reply_is_map is un-waived in scripts/client-compat/manifest.yaml and now compares exact bytes. - GEOSEARCH/GEORADIUS ... WITHCOORD coordinates were BulkStrings and rounded to 4 decimals (#568). Two independent defects in one reply. Redis emits the coordinate pair through addReplyHumanLongDouble, so on RESP3 it is a pair of Doubles — now classified as Resp3Shape::GeoCoords for GEOSEARCH, GEORADIUS(_RO) and GEORADIUSBYMEMBER(_RO) whenever WITHCOORD is present, converting the trailing coordinate element (Redis emits the extras in a fixed dist/hash/coord order, so coords are always last). Separately, the coordinates were formatted {:.4} — about 11 m of resolution — where Redis sends the full shortest round-tripping decimal; they now use the same fmt_geo_coord that GEOPOS already used. That half is protocol-independent, so RESP2 WITHCOORD bytes change too — from $7\r\n13.3614 to $18\r\n13.361389338970184, matching redis-server 8.6.1 exactly. WITHDIST is deliberately untouched: Redis builds it with addReplyDoubleDistance, a %.4f BulkString in both protocols. - CLIENT NO-EVICT and CLIENT NO-TOUCH accepted a missing argument and answered +OK (#580). Redis registers both subcommands with arity 3 (exact), so CLIENT NO-EVICT with no ON/OFF is -ERR wrong number of arguments for 'client|no-evict' command — and so is a fourth argument, while a present-but-unrecognised value is -ERR syntax error. Moon answered +OK to every one of those, telling a client the setting had been applied when nothing was ever parsed. All three dispatch paths now share one parser (command::client::no_evict_or_no_touch); this also fixes a third divergence found while fixing it — the single-shard tokio path (handler_single) did not implement either subcommand at all and answered -ERR unknown subcommand. - CLIENT TRACKING never invalidated for movablekeys commands, so client-side caches went permanently stale (#582). Commands whose keys are not at a fixed argument position carry first_key: 0 in COMMAND_META, mirroring redis's own table — that means "the keys are not at a FIXED position", not "there are no keys". The tracking hook's key extractor read it as the latter and returned an empty key list, which disabled BOTH halves of the protocol: a movablekeys read (SINTERCARD, ZINTERCARD, ZDIFF/ZINTER/ZUNION, XREAD) never registered the client, so it cached a value it would never be told about; and a movablekeys write (LMPOP, ZMPOP, BLMPOP, BZMPOP, XREADGROUP) never pushed an invalidate. SORT src ... STORE dst was a third shape: first_key = 1 names the SOURCE, so Moon invalidated a key it had not written and missed the one it had — likewise GEORADIUS/GEORADIUSBYMEMBER ... STORE/STOREDIST. Because client-side caching is a correctness contract (the client may serve its cached copy until told otherwise), a missed invalidation is unbounded stale data with no signal to the client, rather than a slow path. Measured against redis-server 8.0.5, which invalidates in all of these cases; now verified on the wire at --shards 1 and --shards 4, each case run beside a fixed-position control so a silent harness cannot pass for a fix. Key extraction here is now shared with acl::keyspec, which already understood every one of these layouts (numkeys vectors, the STREAMS token, positional STORE clauses, subcommand-shaped positions). The shared walker reports key POSITIONS and each consumer applies its own policy, because the consumers legitimately disagree: SORT ... BY <pattern> reads runtime-computed key names, which ACL must refuse outright while invalidation must still act on the keys that ARE named. ACL behaviour is unchanged — the fail-closed contract, its error text and every existing key-permission test are byte-identical.

  • Inline commands terminated by bare LF were never dispatched, and an interior LF silently merged two commands into one (#381). Moon's inline parser searched for \r\n and, on finding a lone \r, kept scanning; Redis's processInlineBuffer searches for \n and then strips at most one preceding \r. Two consequences, both measured against redis-server 8.0.5 over raw sockets: a bare-LF stream (shell loops, awk output, anything piped into redis-cli --pipe) never terminated a line, so the command was never dispatched and the client simply hung — and, worse, SET k v1\nSET k v2\r\n skipped the interior \n to reach the trailing \r\n and produced ONE command with the second command's arguments appended, a silent write loss rather than a rejection. The terminator is now \n with an optional preceding \r. The token-separator sets were corrected in the same pass, in both directions: \r now ends an unquoted token (Redis's sdssplitargs breaks on {space, \n, \r, \t}, so RPUSH k a\rb is two elements there and was one here), while \x0b/\x0c no longer do — they are isspace() but are NOT Redis token separators, so a line that merely contained a quote used to split a\x0bb while the same line without one did not. Redis reads two different sets here (isspace() to skip between tokens and to validate the byte after a closing quote; an explicit 5-byte list to end a token) and Moon now does too; the skip set remains a strict superset of the terminator set, which is the invariant that keeps the #487 livelock fixed. \0 is deliberately unchanged — redis-server returns nothing at all for an inline line containing a NUL, so there is no oracle to match. The inline fuzz target now drives the pipelined drain loop with a forward-progress assertion and replays every corpus entry a second time with CRLF rewritten to bare LF.
  • A blank inline line stalled the command queued behind it (#578). parse_inline correctly consumes an empty line and returns Ok(None), but Ok(None) also means "need more bytes", and every read loop acts on that second meaning — so \r\n\r\nPING\r\n arriving in ONE read left the PING unparsed until unrelated later traffic happened to wake the loop, where redis-server answers immediately. This needed no bare LF and predates #381. Fixed in protocol::parse, the single funnel the codec and all three connection handlers share, so the fix cannot be CI-invisible in whichever handler was missed; the re-dispatch goes back through the RESP/inline decision, because what follows a blank line is very often a RESP array.
  • Arity errors named the command in upper case, and spelled container subcommands with a space (#491). Redis formats this message with the command table's ->fullname, which is stored lower case, so echo, ECHO and EcHo all produce 'echo' — Moon normalised UP instead, and any client that string-matches error text saw a mismatch. Measured against redis-server 8.0.5, the container form differs too: 'client|setname', 'acl|getuser', 'memory|usage', 'xgroup|create', 'object|encoding' — a pipe, not a space. Underscores are not separators and stay put (sort_ro, bitfield_ro). Normalisation now happens inside err_wrong_args, which 703 call sites already share, plus 33 static literals and two HGETDEL/HPERSIST sites rerouted through it. Scoped to commands with a real redis-server oracle. Moon-native families (WS, MQ, GRAPH.*, TEMPORAL.*, FT.*, TXN) keep their current names — there is nothing to be compatible with, and the redis-server used as the oracle has no search module to answer for FT.*. The sibling rule is deliberately NOT swept: unknown command '<x>' echoes back what the client sent, case intact, because there is no registered name to normalise to. A guard test pins that, since a blanket lower-casing would have broken it.
  • MEMORY USAGE key SAMPLES 0 was rejected by a dead second parser (#519). key_extra::memory_usage had no callers at all — the live path is server_admin.rs — but it was stricter than Redis (SAMPLES 0 means "sample every nested value" and is valid) and sat next to copy/sort, which ARE wired up. A plausible-looking duplicate is how the next person wiring a new dispatch path silently narrows accepted syntax, so it is deleted rather than corrected. The live behaviour stays pinned by memory_usage_samples_zero_is_accepted.
  • A blocking pop's immediate path consulted the LOCAL shard slice for every key (#557), including keys owned by another shard. The pre-registration scan runs against the connection's own ShardSlice, where a remote key does not live, so at --shards N the fast path missed on (N-1)/N of keys and every one of them fell through to registration. Post-#560 that was harmless to the keyspace, but the scan can only ever produce a WRONG answer for a key it does not own — a stale local look-alike would be popped from the wrong shard, and since #556 it could even invent a -WRONGTYPE for a key whose real owner holds the right type. The scan now skips keys this shard does not own (key_to_shard, the same routing the slotted dispatch uses) and lets the owning shard answer them at BlockRegister time, which is how a remote blocking pop has always been served — a populated remote key still comes back in milliseconds, with no second client pushing.
  • A blocking pop on a key of the WRONG TYPE blocked instead of erroring (#556). SET s v then BLPOP s 0 sat there forever where Redis answers -WRONGTYPE immediately: the pop helpers reach the store through get_mut_if_present(..).ok()??, which collapses Err(WRONGTYPE) into None — indistinguishable from "empty" — so the connection registered a waiter on a key that can never serve it and the client eventually read a null as "queue empty". All eight blocking pops (BLPOP/BRPOP/BZPOPMIN/BZPOPMAX/BLMOVE/BRPOPLPUSH/BLMPOP/BZMPOP) now run a read-only type gate over their keys, in argument order, BEFORE the pop attempt and before any registration — the same gate the in-MULTI path has had since #524, so the live and transactional replies agree by construction. The value itself is untouched (the #560 guarantee holds). Keys the connection's own shard does not own are answered by their OWNING shard at BlockRegister time, so the fix covers --shards >= 2 and not just co-located keys; a multi-key waiter's remote keys keep the old behaviour on purpose (an error raised there would race a sibling key's real wake-up). BLMOVE/BRPOPLPUSH additionally check the DESTINATION's type once the move is about to happen, matching lmoveGenericCommand and moon's own non-blocking LMOVE: previously the element was popped from the source and then silently swallowed by list_push_*'s if let Ok(list), so the client received a value that had left the keyspace entirely — on the immediate path and on the wake path alike.
  • RESP3: keyed sorted-set pops sent their score as a BulkString (#559). BZPOPMIN, BZPOPMAX, ZMPOP and BZMPOP answered $3\r\n1.5 where redis-server 8.x answers the RESP3 Double ,1.5 — every one of these replies is built by Redis's genericZpopCommand, which emits the score through addReplyDouble, exactly like the ZPOPMIN Moon already got right. Two causes, both fixed at the single seam rather than per command: (1) the shape classifier (protocol::resp3::resp3_shape_of) lumped BZPOPMIN/BZPOPMAX in with ZPOPMIN, whose rule is "a second argument is a COUNT" — but a blocking pop's second argument is the TIMEOUT, so every call classified as pair-wrapped, and since the reply is [key, member, score] (odd) the pair-wrapper passed it through untouched; ZMPOP/BZMPOP had no rule at all. They now have their own shapes, KeyedScoredFlat and KeyedScoredPairs. (2) The monoio handler's blocking branch is an intercept: it short-circuits the dispatch exit where every other reply meets the RESP3 policy, and it never applied that policy itself (#462's class), so on the shipped runtime the whole blocking family answered RESP2 shapes to RESP3 clients while the tokio handler, which does convert there, was right. It now routes through the same choke point. The blocking-in-MULTI executor converts too, so standalone, pipelined, in-MULTI and cross-shard all answer the one shape. RESP2 is byte-identical — the conversion is gated on the negotiated protocol version, and the Double's text is the same shortest-repr formatting the bulk carried, asserted directly. Also pre-registered, inert until the commands land: ZRANK/ZREVRANK WITHSCORE (Double score) and ZADD ... INCR (Double reply), both classified positionally so a member literally named WITHSCORE or INCR is not mistaken for the modifier. Verified unchanged against the oracle: GEODIST and ZSCAN scores stay BulkStrings in RESP3.
  • ZREVRANGE key start stop WITHSCORES was rejected inside MULTI. Found by the #559 sweep: the command table gave ZREVRANGE arity 4 where Redis gives -4, and the MULTI queue gate is the only consumer of that number — so the optional WITHSCORES turned a legal command into a wrong-arity error at QUEUE time, aborting the whole transaction. Standalone it always worked, which is why no suite saw it.
  • Ordinary LOCAL writes skipped snapshot copy-on-write, double-applying on recovery (#558). spsc_handler::cow_intercept — the only pre-image capture ordinary commands had — is reachable exclusively from the routed/queued arms that run on the shard event loop's own stack, where &mut Option<SnapshotState> is in scope. Every LOCAL write reaches the database from a connection task instead: the monoio inline SET fast path frames the write straight from the read buffer, and the monoio/tokio local dispatch arms, both MULTI/EXEC executors and the coordinator scatter arms all call command::dispatch directly. None of them could capture anything. At --shards 1 that is every write; at --shards N it is the same-shard fraction. Consequence: an INCR issued while a BGSAVE was in flight, on a key whose segment had not been serialized yet, was written into the snapshot at its POST-increment value while the WAL still held the INCR — recovery loaded the snapshot and replayed the INCR on top of it, so a key that was 11 came back as 13. Silent, and worse the longer the snapshot runs. Capture is now wired at the choke point every non-routed write funnels through (command::dispatch) plus the inline SET path, reusing the per-shard thread-local queue #517 added for Lua writes (drained into the live SnapshotState by the persistence tick, before it advances another segment). Costs one thread-local bool load per command when no snapshot is in flight; the is_write lookup, key extraction and entry clone are all behind that gate. Double capture on the routed arms is harmless — SnapshotState::capture_cow is first-wins deduped.
  • Blocking pops queued inside MULTI answer the wrong reply SHAPE (#524). BLPOP/BRPOP/ BZPOPMIN/BZPOPMAX were rewritten at queue time into LPOP/RPOP/ZPOPMIN/ZPOPMAX, whose replies drop the key: MULTI; BLPOP q 0; EXEC answered ["v1"] where Redis 8.6.1 answers [["q","v1"]], and a miss came back as a null bulk ($-1) instead of a null array (*-1 / _ on RESP3). The key is not decoration — BLPOP takes many keys precisely so the caller can tell which one served, and a client unpacking the documented 2-element reply either errors or silently reads the value as the key. The rewrite was also wrong beyond the shape: BLPOP q1 q2 0 became LPOP q1 q2, where q2 is LPOP's optional COUNT — Moon popped up to two elements from q1 and never looked at q2. Those four commands are now queued unrewritten and executed at EXEC in immediate-only mode (Redis's CLIENT_DENY_BLOCKING path) through the same helper the live blocking fast path uses, so the in-MULTI reply is shape-identical to the standalone one by construction. Both transaction executors and both runtimes. The AOF/replication record is the synthesised single-key sibling scoped to the key that actually served — never the blocking command itself, which a replica would block its apply loop on. BLMOVE/BLMPOP/BZMPOP/ BRPOPLPUSH keep the rewrite: their siblings are shape- and arity-correct. Three side effects fall out: a queued blocking pop now appears in the MONITOR feed as the command the client actually sent; a miss no longer conjures the key it failed to find (the pop helpers are get_or_create, so an absent key was inserted as an empty collection and left behind, making EXISTS answer 1); and a multi-key blocking pop whose keys straddle shards is now refused CROSSSLOT like every other cross-shard MULTI body, because the queued frame finally declares the full key set — a loud refusal replacing a silent mis-execution against the first key alone.
  • RESP3: the inline GET fast path replied $-1 for a miss (#522), the RESP2 null bulk, where Redis replies _. The path frames its own reply bytes and so never reached the protocol-version-aware serializer. At --shards >= 2 this made the null spelling depend on the KEY: keys owned by the connection's own shard took the inline path and got $-1, keys that routed elsewhere came back through the serializer and got _ — on the same connection, so a client could not even cache "this server speaks RESP3". The inline path now takes the negotiated protocol version and selects between two static byte strings (io::static_responses::null_bulk, which also gives the previously caller-less NULL_BULK constant a home). +OK and -WRONGTYPE, the only other bytes this path frames, are spelled identically in both protocols.
  • TXN.COMMIT no longer answers +OK for a transaction that was applied in part (#499). A TXN body whose ops are rejected by a TXN guard — a cross-shard write at --shards > 1, MOVE, COPY ... DB, SWAPDB, a cross-shard Cypher write — used to commit the accepted subset and reply +OK. The per-op errors did go back to the client, but a driver inspects the COMMIT reply, not the replies of the body commands (exactly as it inspects EXEC and not the QUEUEDs), so a routing mistake became silent partial application. A rejection now poisons the transaction the way a queue-time error poisons MULTI: TXN.COMMIT rolls the whole transaction back through the TXN.ABORT path and answers EXECABORT TXN.COMMIT discarded because of previous errors: N operation(s) rejected inside the transaction (first: <CMD>) -- rolled back and NOT committed, leaving the connection out of the transaction. The wording stops at what the code can guarantee — the commit did not happen — because rollback is the same best-effort TXN.ABORT path whose undo capture still has gaps (#500). Nothing changes for a transaction whose every op was accepted. Both handlers (monoio and sharded/tokio) carry the check, and both are covered end-to-end.
  • INFO memory now reports used_memory_lua, and the Lua memory gauges stop reading ~48 bytes (#506). The reported defect was "monoio's shard event loop holds a different Lua VM than the one executing scripts". Instrumenting both sites disproved that: on monoio the tick's lua_rc and handle_eval's VM are the same Rc pointer, reporting the same 24165 → 25403 → 4219542 bytes the tokio leg reported. The real defect was the SOURCE of the published number. ShardStoreMemory::lua carried ScriptCache::resident_bytes() — the cached script text, which for EVAL "return 1" 0 is exactly 48 bytes (a 40-char SHA1 key plus an 8-byte body), the very figure the issue read as monoio's "wrong VM". Nothing anywhere sampled the interpreter heap, so moon_memory_bytes{kind="lua_scripts"}, moon_used_memory_bytes, INFO used_memory and MEMORY DOCTOR all under-reported Lua by three orders of magnitude, and a script that anchors megabytes of tables in _G was invisible to every one of them. The VM is now sampled into its own ShardStoreMemory::lua_vm atomic and summed into all four, with used_memory_lua emitted under Redis's field name (its client-compat waiver is deleted). The sample deliberately lives in persistence_tick::run_eviction_tick — the one tick body both event loops share — because the shard periodic tick has TWO arms in event_loop.rs (a tokio select! arm and a monoio counter arm) and a publish written into either one alone is invisible on the other runtime; that is the trap the original investigation fell into, and a guard test now fails if the sample migrates back into a runtime-specific arm.

  • ACL key patterns are now enforced for RPOPLPUSH and BRPOPLPUSH (#520). The ACL layer extracts a command's key arguments from a name table, and a command missing from that table falls through to an empty key list — which makes the permission loop a no-op, so every ~pattern is ignored rather than merely approximated. LMOVE/BLMOVE were listed; their RPOPLPUSH siblings, with the identical two-key layout, were not. A ~cache:* user could therefore move an element out of any key in the keyspace into any other. Found while wiring #520; BRPOPLPUSH was already reachable, so this closes a live hole as well as pre-empting one in the new command.

  • XREADGROUP in history mode answers the stream, not a null (#526). XREADGROUP ... STREAMS s 0 asks for the consumer's PENDING entries; Redis serves the stream before it knows whether the PEL slice has anything in it, so an empty PEL replies *1\r\n*2\r\n$1\r\ns\r\n*0\r\n — [["s", []]]. Moon replied $-1, so a client that iterates the returned stream list got a decode error (or a None branch) where Redis gives it zero iterations. This is a reply-SHAPE divergence, not a null-type one. The two request modes are now served on Redis's own conditions: history (an explicit ID) always contributes its stream, > contributes only when there ARE new entries — a > stream with nothing new is dropped from the reply rather than rendered as an empty entry list, which also fixes the mixed STREAMS a b > 0 case. The reply is the null array only when nothing at all was served, so the >-with-no-new-entries answer stays *-1 (#482).
  • LPOP/RPOP validate the optional count BEFORE the key lookup, with Redis's error text (#527). Both commands looked the key up first and returned the miss (*-1) without ever reaching the count parser, so LPOP nokey abc answered "no such list" and the identical command answered -ERR ... the moment somebody created the key — the same malformed request getting opposite replies depending on unrelated keyspace state. Redis parses argv[2] before lookupKeyWrite, so the argument error never depends on the key. The message now matches too: a non-integer and a negative count both answer ERR value is out of range, must be positive (Moon said ERR value is not an integer or out of range, which clients that classify retries by error string read as a different failure). A well-formed count on an absent key still answers the null array (#482), and validating first does not create the key. LPOS's RANK/COUNT/MAXLEN were already ordered and worded correctly and are untouched.
  • Blocking pops no longer create the key they miss on (#523, #539). BLPOP, BRPOP, BLMOVE (source), BRPOPLPUSH (source), BZPOPMIN, BZPOPMAX and BZMPOP reached their value through get_or_create_list / get_or_create_sorted_set, so every miss materialised an empty list/zset before discovering there was nothing to pop — a phantom key that EXISTS, TYPE, DBSIZE, KEYS and SCAN all reported, that the ordinary "block on a key another client is about to create" idiom then hit with WRONGTYPE on the producer's RPUSH, and that nothing ever removed (a worker polling rotating ids leaked one empty collection per timed-out poll). An empty list/zset is not a representable Redis value, so those phantoms were also leaking into RDB, AOF rewrite and replication. The four pop helpers now take a new non-creating accessor (Database::get_mut_if_present::<K>) — get_or_create minus the fabrication step, so expiry drop, cold-tier promotion, compact-encoding upgrade, WRONGTYPE classification and the become-empty ⇒ remove-the-key behaviour are all unchanged, but a missing key leaves the keyspace and used_memory byte-identical. BLMPOP was already correct (it length-checks through the read-only get_list) and is untouched. At --shards >= 2 the phantom appeared only for client-local keys (~1/N of them), which is why it read as a flake.
  • Lua script writes no longer skip the AOF plane under runtime-tokio, nor snapshot copy-on-write on either runtime (#517). Two independent holes, both pre-existing: (1) LuaEvictionCtx::emit_effect was #[cfg(feature = "runtime-monoio")]-gated in full, with a let _ = (db_index, cmd_and_args) arm that DISCARDED the effect off monoio — a runtime-tokio build ran EVAL "redis.call('SET',KEYS[1],'v')" 1 k, mutated the keyspace, answered OK, and wrote nothing to the AOF, so the write was gone after a restart while every ordinary write around it survived. The gate now lives per plane inside reason_del::record_bytes_conn: the AOF leg (a channel send to the writer pool, exactly what the tokio connection handler already does for ordinary writes) runs on every runtime; the replication leg stays monoio-only because it pushes through shard::self_msg and master-side PSYNC does not exist under tokio at all. (2) All three handle_eval call sites (local monoio, local tokio, routed ShardMessage::Execute) run inside a bare with_shard(...), and spsc_handler::cow_intercept is only reachable from the shard event loop's own stack — so a script writing while a BGSAVE was in flight left no pre-image, and the snapshot captured the POST-write value that WAL replay then applies a second time (INCR double-counts). Capture now happens in the scripting bridge at the redis.call level, which is also the only level where a real key exists — EVAL <script> <numkeys> k's own command[1] is the script body, so wrapping the call sites in cow_intercept would have stashed the script text under a bogus key. Pre-images go to a per-shard thread-local queue that the persistence tick folds into the live SnapshotState before it advances another segment (persistence::snapshot_cow); the whole path costs one thread-local bool load when no snapshot is in flight. SnapshotState::capture_cow is now first-capture-wins, which also fixes a pre-existing hole on the ordinary-command path: a key written twice in one epoch appended two overflow records and load took the LATER one — a value that already contained the first write.
  • EXPIRE/PEXPIRE/EXPIREAT/PEXPIREAT now accept the Redis 7.0 NX | XX | GT | LT conditions (#544) — previously any option token was rejected with a wrong-arity error, which breaks typed clients that call them directly (redis-py expire(k, t, nx=True)). Semantics match Redis exactly, oracle-diffed against a live redis-server: NX sets only when the key has no expiry; XX only when it has one; GT/LT only move the expiry later/earlier, with a no-TTL key counting as infinite (so GT never sets it and LT always does); XX composes with GT/LT; NX composes with nothing (ERR NX and XX, GT or LT options at the same time are not compatible), and GT+LT is its own error — Redis's exact strings. The gate runs after the invalid-expire-time range check and before the past-time delete, so EXPIRE k -1 GT on a TTL'd key refuses (:0, key intact) while EXPIRE k -1 LT deletes — Redis's evaluation order. COMMAND INFO arity for the four commands corrected from 3 to -3. In the same file: TTL now rounds to the nearest second like Redis ((ms+500)/1000) — the old floor answered 99 for a fresh EXPIRE k 100, a divergence the oracle probe surfaced and a pipelined /dev/tcp probe (spawn-latency-free) confirmed fixed byte-for-byte.
  • Lazy expiry now emits — reads no longer delete expired keys silently (#542). The removal Database::get/get_mut/exists performed when a read touched an expired key fired neither the expired keyspace notification nor the dual-plane (AOF + replication) DEL — only the active sweep did both. Two consequences: subscribers to __keyevent@<db>__:expired never heard about any key that was read after expiring (the common case for cache/session patterns, where the read is what discovers the expiry), and once the master lazily removed a key its sweep could never emit the authoritative DEL, so an attached replica kept the entry resident forever (logically hidden, but RAM/DBSIZE/used_memory diverging unboundedly). The lazy paths now HIDE the key (answer absent, leave it in place) and record it in a bounded per-db queue (PENDING_EXPIRED_CAP = 8192, adjacent-dedup); the shard's next active-expiry tick drains the queue, re-verifies expiry (a key overwritten in between is a new incarnation and is skipped), and deletes through the same emitting sink as sweep victims — identical notification + DEL treatment. This also fixes the replica-side divergence for free: a replica's reads now hide without deleting (Redis semantics — the replica waits for the master's DEL), its queue sits bounded at ≤ cap since replicas never drain, and promotion resumes draining with re-verification. Past the cap the lazy path still hides but stops recording; the probabilistic sweep remains the backstop. Write-path expired-entry replacement (get_or_create_*) is intentionally unchanged — the streamed write itself supersedes the key.
  • Commands whose key is not their first argument were routed by hashing a literal (#533, #534). extract_primary_key special-cases the commands whose key is not args[0] and falls through to args[0] for everything else. LMPOP, ZMPOP and SINTERCARD take numkeys first, and XREADGROUP takes the literal token GROUP, so each hashed a constant: every invocation landed on one fixed shard regardless of the key, and a key any other shard owned read as absent.

The failure is quiet, which is why it survived. A mis-routed LMPOP/ZMPOP answers *-1 and a mis-routed SINTERCARD answers :0 — byte-identical to the honest answer for a key that really is empty. XREADGROUP at least errors. BLMPOP/BZMPOP were fixed in the same arm for their other consumers (cluster slot resolution asks the same function, and hashing the timeout yields a wrong MOVED target); their blocking path extracts keys separately and was correct. Also corrected: the XREAD arm's len == 5 guard could never match XREADGROUP despite the comment above it naming both.

Two rules, both learned by having the previous instrument miss this class entirely, govern the regression tests in tests/shard_routing_parity.rs and the route_probe rows in scripts/test-consistency.sh: probe populated keys, because on an absent key a mis-routed command and a correct one return the same bytes; and probe many keys, because a constant route is still right for ~1/N of them, so a single-key test at --shards 4 passes a quarter of the time and reads as a flake. Against the pre-fix binary the four parity rows fail with 9, 11, 10 and 7 of 12 keys wrong; the fence rows (LPOP, ZDIFF, XREAD) pass on both. - A blocking waker destroyed waiters it could not serve (#535). try_wake_list_waiter, try_wake_zset_waiter and try_wake_stream_waiter each popped whatever waiter sat at the front of a key's queue. A command outside the waker's family fell through a _ => (None, None) arm into the cleanup that runs for every waiter it pops — remove_wait(wait_id) plus reply_tx.send(None) — so the registration was dropped across all its keys and the client was answered a null it never earned.

Two independent paths reached it, which is why the fix is in the wakers rather than at one call site. ShardMessage::BlockRegister registers a remote waiter and then, gated on the key existing, runs all three wakers in sequence, list first — so a BZPOPMIN on a populated zset owned by another shard was eaten by the list waker during its own registration and answered *-1 in microseconds instead of returning the member sitting right there. Separately, an ordinary RPUSH on a key name that also carried a zset waiter destroyed it the same way.

Measured at --shards 4 before the fix: 9, 8 and 10 of 12 populated zsets answered *-1 through BZPOPMIN, BZPOPMAX and BZMPOP — the ~3-of-4 signature of a key the client's own shard does not own — each in 130µs–1.5ms despite a 3-second timeout, because the null is sent during registration and the command never blocks.

BlockedCommand::family() now maps every variant to the one waker entitled to serve it, and pop_front_of_family takes only that family's waiters, leaving the rest queued in order. The mapping is exhaustive with no _ arm, so a new blocking command will not compile until it declares its family instead of silently becoming edible. The wakers' loop condition moved from has_waiters to the pop itself — a queue holding only foreign waiters is not empty, and the old condition would have spun forever.

  • RESP2 replies whose missing value is an array sent the null string instead (#482). RESP2 has two nulls — $-1 says "the string you asked for is missing", *-1 says "the array you asked for is missing" — and Frame had a single Null variant that always serialised $-1. A statically-typed client decodes the two differently, so this was a decode error client-side, not a cosmetic difference. Frame::NullArray now carries the second one: *-1 under RESP2, _ under RESP3 (which has only one null), and composable inside an array.

Seventeen sites were wrong, each measured against a live redis-server 8.6.1 rather than read from the docs: the eight blocking commands on timeout (BLPOP, BRPOP, BLMOVE, BRPOPLPUSH, BLMPOP, BZPOPMIN, BZPOPMAX, BZMPOP), LMPOP/ZMPOP finding nothing, XREAD and XREADGROUP with no data, GEOPOS of an absent member (nested inside the outer array), LPOP/RPOP with a count on an absent key, EXEC aborted by a broken WATCH, and the parser, which collapsed an inbound *-1 so a reply relayed from a peer went back out as $-1.

Two of those are worth calling out. EXEC-abort was the highest-impact one and was not in the original report: every client library decodes EXEC as an array, so the abort path — the one optimistic-locking code is written to handle — was the one that mis-decoded. And LPOP key <count> was never a Frame::Null at all; it returned an empty array (*0), so an audit that only rewrote null sites would have missed it. That cuts both ways: roughly thirty sites already answered correctly (nineteen $-1, twelve *0) and had to stay put — including GEOHASH, which answers $-1 for the same absent member on which GEOPOS answers *-1, in the same source file. tests/resp2_null_array.rs pins both directions, and scripts/test-consistency.sh now diffs the reply TYPE against a live Redis over a raw socket, because redis-cli renders both nulls as (nil) and so cannot see this class of defect at all.

One case in the original report is deliberately not fixed here. BLPOP inside MULTI still answers a null bulk, because that path rewrites the queued command to LPOP, whose own null correctly is $-1. Measuring the hit path showed the rewrite is wrong in shape rather than in null type — Redis answers [["q","v1"]], Moon answers ["v1"], dropping the key that BLPOP exists to report — so patching the null alone would have turned its test green over a command that still answers wrongly. Filed as #524; the test stays ignored with its reason moved there.

XREADGROUP was added to the list in review rather than by the original sweep: XREAD lives in stream_read.rs, which the audit walked, while XREADGROUP lives in stream_write.rs, which it did not. Splitting a command family across a read and a write file is exactly what makes a file-scoped audit miss a case.

  • MEMORY USAGE <key> reported existing keys as absent at --shards >= 2. It routed by hashing the literal subcommand "USAGE" rather than the key (#511), so every invocation went to one fixed shard and read that shard's slice for a key it does not own. Measured at --shards 4: 22 of 24 existing keys reported absent. The rate is 1 - 1/shards — the signature of a constant route, not a race — and at --shards 1 it was always correct, which is why it survived.

extract_primary_key treats args[0] as the key for anything not in its keyless list, and MEMORY was in neither that list nor the set of commands with an explicit subcommand arm (OBJECT, XGROUP, XINFO, BITOP, ZDIFF/ZINTER/ZUNION/ZINTERCARD, XREAD — each added for exactly this reason). MEMORY USAGE now routes by args[1], matched on the subcommand rather than blindly: USAGE is the only MEMORY subcommand that takes a key, so a bare args[1] would work today and silently hash a future subcommand's first option as a key tomorrow. DOCTOR, STATS, PURGE, MALLOC-STATS and HELP are keyless and now execute locally instead of routing by "DOCTOR"/"STATS"; they aggregate across shards internally, so the answer is unchanged.

Neither shell suite covered it — test-commands.sh tested only MEMORY DOCTOR, and test-consistency.sh explicitly noted MEMORY as out of scope — so a 75%-failure-rate bug was invisible to both. test-consistency.sh now sizes 20 keys at 1/4/12 shards (a single key passes by luck at 12 shards), and test-commands.sh asserts the single-shard shape via a new assert_moon_matches helper, since assert_moon_contains greps -F and cannot express "some positive integer".

  • A script whose keys all lived on one shard was refused instead of routed (CROSSSLOT). Reported as "EVALSHA of a single-key script fails with CROSSSLOT" (#508), with the guess that a 1-element key list was being mis-folded. It was not. validate_keys_same_shard required every key to hash to the shard the CONNECTION happened to occupy — so CROSSSLOT was standing in for "I cannot run this here", and nothing ever asked where it could run. One key cannot cross slots; it just lives somewhere else. Measured at --shards 4: 7 of 8 single-key EVALs rejected, only the key that happened to land on the connection's own shard running. numkeys=0 always worked, which is why the defect read as intermittent rather than total.

This broke redis.lock.Lock.release() — a single-key EVALSHA, and one of the most-used constructs in redis-py. Through the Script.__call__ wrapper the caller saw a NoScriptError followed by a cross-slot error, neither of which names the cause.

A script now ROUTES to the shard owning its keys (scripting::route_script_keys → coordinator::coordinate_script → a shard-side EVAL/EVALSHA arm on the SPSC Execute path), because a script executes against one shard's database and the correct place to run it is where that data is. Keys that GENUINELY span shards are still refused — there is no single target for them — and that refusal is now doubly held: the routing decision rejects it before the hop, and validate_keys_same_shard remains as the shard-side backstop, so the two must agree before a script can read another shard's (empty) view of a key. Shards build their Lua VM lazily on first use and share the one slot conn_accept fills, so a shard reached only by routing still gets exactly one VM.

One thing this changes rather than fixes: EVAL caches its script only on the shard that ran it and, unlike SCRIPT LOAD, never fans out — so a bare EVAL followed by a direct EVALSHA on a key owned by another shard now answers NOSCRIPT where it previously answered CROSSSLOT. Measured after the fix at --shards 4: one bare EVAL, then EVALSHA of that sha across 12 other keys — 4 ok, 8 NOSCRIPT. Neither ever worked, and NOSCRIPT is the better failure because every client library retries it — redis-py's Script.__call__ re-issues EVAL and self-heals, whereas CROSSSLOT was unrecoverable. Anything built on register_script/SCRIPT LOAD, including redis.lock.Lock, is unaffected (12/12). Filed as #515.

Applied to both handlers through one shared helper rather than two implementations that can drift. FCALL has the same defect plus a second one — FUNCTION LOAD never fans out to other shards, so routing alone would trade CROSSSLOT for ERR Function not found at the same rate. Filed as #514 rather than half-fixed here.

A routed script that gets no reply no longer claims it did not run. recv_reply_bounded returns the same error for a closed channel and for a 30s reply timeout, and the first cut of this path reported both as "cross-shard reply channel closed during script execution". Those have opposite retry semantics: a timeout means the target may still be executing, or may have already applied its writes, so a client told the script never ran will re-send a non-idempotent script that did. The two are now distinguished (ReplyFailure::{Closed, TimedOut}), both say execution status is unknown, and the timeout records moon_xshard_reply_timeout_total{kind="script"} so a wedged owner shard is visible in metrics like every other cross-shard reply path.

Not fixed here, and pre-existing rather than introduced: script writes skip cow_intercept on every path, and under runtime-tokio emit_effect is compiled out entirely, so script writes emit no AOF or replication record. Routed scripts persist exactly as well as local ones do — which is the bug. Filed as #517, since a fix needs replication parity, kill-9 durability and a VM A/B bench of its own.

  • A pipeline did not execute in order at --shards >= 2, and writes were silently lost. Reported as "MGET in the same pipeline as its SETs returns nulls" (#507). The cause is wider than the symptom: the sharded pipeline handlers DEFER a single-key command whose key lives on another shard into remote_groups, dispatching the whole group as one PipelineBatchSlotted at the end of the batch — while multi-key and keyless commands execute INLINE, mid-loop. An inline command therefore ran against a shard whose earlier writes in the same batch had not been sent yet.

Measured at --shards 2, 20 trials each, before the fix:

in one pipeline batch wrong consequence
SET a, MSET a 7/20 the MSET's value is lost — the earlier SET lands on top of it
SET, FLUSHALL 10/20 the key survives a flush that returned +OK
SET,SET, DEL 6/20 the keys survive a DEL that returned success
SET,SET, MGET 10/20 the reported symptom
SET, DBSIZE 12/20 also KEYS, RANDOMKEY, EXISTS, UNLINK, COPY, BITOP, INFO keyspace
SET, TOUCH k 0/20 single-key: always correct, and the shape of the fix

So this was not only the stale read it was filed as — same-key write ordering inverted, which is silent data loss. The rate rises with shard count (a key is remote with probability 1 - 1/shards).

The fix defers such a command and the unconsumed batch tail to the next iteration, reusing the mechanism #438 already built for early-flush commands: phase 2 resolves the pending remote replies first, and the tail re-parses with remote_groups empty, so it cannot loop. Applied to both sharded handlers (handler_monoio and handler_sharded); handler_single has no deferral and was never affected.

The predicate is keyed on ROUTABILITY, not on a list of command names: a command routed by its own single key needs no wait, because a key maps to exactly one shard — if that shard is local the key cannot be pending, and if it is remote the command is appended behind the pending ones and the slotted batch preserves order. Everything else waits. A name list written for an MGET bug would not have contained INFO keyspace or RANDOMKEY, both of which were wrong. The one case routability cannot see — commands intercepted inline BEFORE routing, which still have a key-shaped first argument (EVAL, SWAPDB, …) — is named explicitly and carries a test that fails if an entry is dropped.

Deferring is the conservative direction: a command sent down this path unnecessarily is merely executed at the start of the next batch, which is always correct.

This costs throughput, and the cost is per interleaving rather than per pipeline — each deferral is one extra shard dispatch/await boundary (~50µs). Measured at --shards 2 on one connection, 9 reps, alternating leg order, median; "floor" is the worst within-leg spread, so a delta smaller than its floor resolved nothing:

pipeline shape guard fires before after delta floor
MGET after every 2 SETs 64×/flush 125,885 59,889 −52.4% 5.8%
128 SETs, then one MGET 1×/flush 528,764 512,686 −3.0% 22.2%
SET,SET,GET (guard never fires) never 873,526 867,715 −0.7% 33.9%

A multi-key or keyless command at the END of a pipeline — the shape #507 was filed from, and the shape redis-py's pipeline() produces — costs nothing measurable. One interleaved after every pair of writes halves throughput. redis-benchmark cannot express any of these shapes (it sends a single command type, so the guard never fires), so this came from a purpose-built harness that refuses to report unless the pre-fix binary actually reproduces the bug first.

Recovering that cost means letting multi-key commands participate in the slotted batch instead of executing inline — a cross-shard-coordinator change well outside a correctness fix, filed separately.

  • Writes were paying for a memory measurement on every SET. The maxmemory real-footprint correction (#478) was computed inside evict_to_budget, which runs on the write path, so every write performed open/read/close on /proc/self/statm and an instance-wide accounting sum that takes a read lock on the global replication state. The early-out above it never fired, because Moon auto-sets maxmemory to 75% of RAM, and the correction's own noise-floor guard sat inside the callee with the expensive reader passed as its argument — Rust evaluates arguments eagerly, so the syscalls ran to completion and only then were judged unnecessary.

Measured on the moon-dev VM against v0.8.5, interleaved legs and median-of-5: SET at c=8 P=16 fell 1.36M -> 0.58M ops/s (-57%) and at c=1 P=1 159K -> 134K (-16%), while GET was neutral in both — the shape of a cost paid only by writers. git bisect named the commit.

The measurement now happens once a second, in the shard-0 chore that was already reading /proc/self/statm for the RSS gauge (and already carried a comment about not doing it more often); the write path reads one relaxed atomic. The correction is unchanged in value — a budget divisor in the GB range does not care about a second of staleness — and stays neutral (1.00) until the first sample lands, so an unmeasured process is never over-evicted.

Two new INFO fields make the mechanism observable: maxmemory_footprint_correction is the divisor actually applied, so an operator can tell a cap being honoured from one being silently tightened, and maxmemory_footprint_samples is the sampler's liveness counter. The counter exists because the ratio alone cannot distinguish a healthy small instance from a wedged sampler — both read 1.00 — and because the chore lives in a shard loop with separate monoio and tokio arms, where a correction wired on only one runtime would restore the swap-death bug #478 fixed, invisibly. io20 asserts the counter advances on whichever runtime the test binary was built for. - Five commands were advertised by COMMAND/ACL but dispatched nowhere; they are deregistered. LATENCY, MODULE, DUMP, RESTORE and RECLAMATION sat in metadata.rs while answering unknown command on every dispatch path, so COMMAND, COMMAND COUNT (267 -> 262) and ACL all described a surface Moon cannot serve — a client that introspects before calling was told it could. Three were never top-level commands at all: RECLAMATION is reachable only as DEBUG RECLAMATION, and DUMP/RESTORE only as FUNCTION DUMP / FUNCTION RESTORE, both of which are unaffected. With these gone the registry sweep (cdg1_registry_sweep_no_unknowns) runs with no waiver list at all — every entry in the registry is now proven reachable on both feature legs, which is what the v0.9 exit criterion asked for. - CI now proves reply shapes match across MULTI and pipeline, not just standalone. The compat harness has always supported --contexts standalone,multi,pipeline and CI only ever ran --strict, so the "a command must not change shape by context" rule was asserted by hand and never gated. Wired in as its own step (PASS=201 FAIL=0 across all three contexts). - A test waiver went stale and hid five working commands from the registry sweep. cdg1_registry_sweep_no_unknowns enumerates COMMAND_META and asserts nothing answers unknown command, skipping a list of 10 backlogged-unimplemented names. Five of them had since shipped — WATCH/UNWATCH (+OK), RESET (+RESET), SSUBSCRIBE/SUNSUBSCRIBE (a 3-element confirmation), all delivered by the v0.9 client-compat tasks — and none was removed from the list, so for the whole milestone the sweep reported green over a surface it had quietly stopped covering. The list is now the five that genuinely do not dispatch (LATENCY, MODULE, DUMP, RESTORE, RECLAMATION), and the sweep passes with the other five included on both feature legs. The waiver now also expires by itself: each backlogged command is asserted to STILL answer unknown command, so implementing one fails the test and names it rather than letting the exemption outlive its reason. A skip nobody re-checks is indistinguishable from coverage. - Five Rust SDK helpers sent wire forms Moon rejects on every call; two server-side gaps that hid them are closed. moondb 0.2.1 → 0.3.0 (breaking: five pub methods removed).

Each removed method failed on its first round trip, always, for the whole published lifetime of the crate — so no caller can have depended on its behaviour, only on it compiling. Two named commands Moon does not have, and three named real commands with the wrong arguments, which is why a name-level audit had already cleared them:

Removed Sent Server answered Use instead
MqClient::push_partitioned MQ.PUSH (a command name) unknown command MqClient::push
MqClient::pop_partitioned MQ.POP (a command name) unknown command MqClient::pop
VectorClient::upsert FT.UPSERT unknown FT.* command FT.CREATE, then HSET the index's vector field
TemporalClient::snapshot_at_packed TEMPORAL.SNAPSHOT_AT <hlc> wrong number of arguments snapshot_at (the server captures the timestamp)
TemporalClient::release_snapshot TEMPORAL.INVALIDATE (no args) wrong number of arguments nothing — delete the call

upsert was not reimplemented over the real wire form because there is no faithful one: Moon indexes a vector by HSET-ing a hash whose vector FIELD NAME comes from the index definition, and the signature never carried it. Guessing would trade a loud error for a silent wrong write. release_snapshot has no replacement because its premise was false: TEMPORAL.SNAPSHOT_AT never pinned the connection to a snapshot view — it records a shard-global wall_ms → LSN binding that AS_OF resolves later — so no pin is taken and none can be dropped. Callers can simply delete the call; their reads were already live. snapshot_at's documentation, which described the imaginary pin, is corrected.

TXN was NOT removed — txn_begin / txn_commit / txn_abort are correct and stay. An earlier shell probe appeared to show TXN dead; the probe was wrong (zsh does not word-split an unquoted parameter expansion, so the server received one argument literally named TXN BEGIN).

Server-side, two commands were unreachable through introspection because they are served by intercepts that run before the metadata table: TXN and FT.AGGREGATE are now registered, so COMMAND INFO and COMMAND COUNT (265 → 267) report them. Registration is metadata only and does not reroute either command. Separately, a bare TXN or an unrecognised subcommand answered unknown command 'TXN', which is false — the command exists — and misleads a driver into concluding Moon has no cross-store transactions. It now answers an arity/subcommand error, the shape Redis uses for container commands and the one driver error handling keys on.

The SDK tree had no CI of any kind, which is how all five shipped. Three guards now run: a command-NAME sweep over the SDK sources (tests/sdk_wire_forms.rs), a live round trip through every one of the 52 public Rust helpers (sdk/rust/tests/round_trip.rs), and a Python __version__ derivation check. The round trip is what found release_snapshot, after review had already passed the file — and mutating a helper's argument order fails it while the name sweep stays green, which is the point of having both. - moondb.__version__ reported a release the package had not been for two versions. The Python SDK published as 0.1.1 while __version__ was a hand-maintained literal still answering "0.1.0", and the test covering it asserted the same stale literal — so the suite stayed green while every caller reading __version__, every bug report quoting it, and anything gating on "SDK >= x" got the wrong number. It is now derived from installed distribution metadata (falling back to pyproject.toml for an uninstalled source checkout), so it cannot drift again, and the test asserts the derivation rather than restating the value. - Remote panic on the cluster bus: a truncated v3 gossip header killed the process. The gossip wire v3 (#493) appended a 40-byte sender_master_id, but the deserializer's length guard still admitted any frame of at least the v2 header size so that a genuine v2 peer would still parse — and the v3 branch then read data[2130..2170] unconditionally. Any frame carrying version 3 with a length in 2130..2170 indexed past the end. The cluster bus listener hands this function peer-supplied bytes, so a single unauthenticated 2130-byte frame to the bus port panicked the cluster-ctl thread, which by policy aborts the whole server. A short v3 header is now rejected as malformed. Seeded into fuzz/corpus/gossip_deser — the target was correct but had not synthesised the 4-byte magic plus that 40-byte length window within its PR budget. - A pipelined HELLO no longer re-encodes the replies that came before it. Moon accumulates a read batch's replies and serialized them all at flush time under whichever protocol version was in effect at the END of the batch, so a HELLO 2 sent in the same write as an earlier command retro-downgraded that earlier reply: CONFIG GET maxmemory produced under RESP3 went out as *2 instead of %1. Every reply is now encoded in the protocol that was in effect when that reply was produced, with a switch taking effect from its own index onward — inclusive, so a HELLO's own reply is rendered in the protocol it establishes, which is what redis-server 8.6.1 does. Only the downgrade direction was ever visible: the frame shape is already fixed correctly at dispatch, and a RESP2-flattened array re-serialized as RESP3 still emits *, which is why the upgrade direction looked fine by accident. It is now pinned in both directions. RESET is covered too: it is contracted to return the connection to its default state, RESP2 included, so it moves the protocol exactly as a pipelined HELLO 2 does — and a fix that covered only the two HELLO sites left HELLO 3 + RESET in one write still retro-downgrading. redis-cli cannot express two commands in one write(), which is how this survived — the new suite drives a raw socket, and runs on both runtimes at 1 and 4 shards. Batches without a protocol switch — essentially all of them — keep the previous single-version loop, one branch and no allocation. - CONFIG GET answers every parameter, not just the first. CONFIG GET maxmemory appendonly reported only maxmemory; the rest were silently dropped, which is what redis-py's config_get(*params) and monitoring agents that read several settings per call send. The reply is now the union over all patterns, deduplicated (maxmemory plus maxmemory* reports it once), in the server's own table order rather than the caller's argument order, with unknown patterns skipped rather than erroring — all four properties measured against redis-server 8.6.1. - CLUSTER INFO no longer claims cluster_enabled, and a slotless node no longer claims health. Two integration assertions encoded the pre-fix behaviour and contradicted the measured oracle: redis-server 8.6.1 reports cluster_enabled in INFO only — CLUSTER INFO never carries it — and a node with zero slots assigned answers cluster_state:fail, because state is derived from slot coverage rather than from the --cluster-enabled flag. - Replica routing tests no longer race gossip. A freshly-MEET-ed node holds an incomplete slot map until gossip hands it the rest, and an incomplete map answers CLUSTERDOWN The cluster is down in front of any redirect — matching Redis, where the down-state check precedes MOVED once the slot resolves. The replica tests now wait for the joining node's OWN view to reach cluster_state:ok before asserting on redirects. - Gossip never said WHICH master a replica follows, so replicas were reported as their own shards. The flags word carried a role bit but no master id, and a sender's own role was not propagated at all — a replica was known as one only on the node its CLUSTER REPLICATE ran against, and every other node saw a slotless master. The sender's master_id now rides in the gossip header (wire v3). The header LAYOUT, not just the flags encoding, now depends on the version, so a v1/v2 peer's sections begin 40 bytes earlier and are still parsed. - CLUSTER REPLICATE neither validated its argument nor replicated anything. It answered +OK for a node id it had never heard of, and NodeRole::Replica was read nowhere outside src/cluster/ — so a cluster "replica" held no data and could serve no read. It now rejects with the measured ERR Unknown node <id> / ERR Can't replicate myself, and starts the same replication the REPLICAOF path starts. The validation is not cosmetic: a caller retrying until OK — the only way to wait out gossip convergence — succeeded instantly against an empty node table, leaving the node relabelled but permanently empty. - An inline closing quote followed by \r, vertical tab or form feed was rejected as unbalanced. Redis's sdssplitargs tests isspace(p[1]) after a closing quote; Moon tested only ' ' and '\t', so ECHO "hi"\rx answered -ERR Protocol error: unbalanced quotes in request where Redis splits the line into three arguments. Measured over a raw socket against redis-server 8.6.1, both quote types, one fresh connection per probe: space, TAB, \r, \x0b and \x0c are all accepted as separators, and only a non-whitespace byte (ECHO "hi"y) is unbalanced. The check now reads the same is_inline_space set as the rest of the splitter — the narrow/wide disagreement here was the milder form of the bug that made the splitter unable to advance at all. - cluster_state was a constant, so a cluster that had lost a master still reported itself healthy. The status field was assigned once in ClusterState::new and never written again; cluster_slots_pfail and cluster_slots_fail were the literals 0, and cluster_slots_ok was reported as "however many slots are assigned". A client had no way to learn the cluster was degraded. cluster_state is now derived on read from real coverage — ok only when every one of the 16384 slots is claimed by a node that is not confirmed FAIL — and the three slot counters are counted per slot through their owners' bitmaps rather than by summing per-node counts, so a slot two nodes both claim cannot mask a real coverage hole.

Fixing the reporting exposed that the transition it reports could never happen either: try_mark_fail_with_consensus counted only reports received FROM peers, but Redis's markNodeAsFailingIfNeeded also counts the local node's own suspicion when it is a master (if (nodeIsMaster(myself)) failures++). Without that vote a 3-master cluster is arithmetically unable to promote PFAIL to FAIL — quorum is 2, and each survivor hears about the dead node from exactly one other survivor, so the count tops out at 1 forever. A replica still casts no vote. - A degraded cluster kept serving keys, handing clients a partial view they could not detect. While cluster_state is fail, every keyspace command now answers CLUSTERDOWN The cluster is down — including keys in slots the receiving node owns and could serve. That over-refusal is deliberate and measured: a partially-visible keyspace is worse than a refusal, because a client cannot tell which half of the data it is seeing. Redis's two CLUSTERDOWN messages stay distinct and the order between them is load-bearing — an unclaimed slot answers the per-slot CLUSTERDOWN Hash slot not served even while the cluster is fail, matching a real server (a lone --cluster-enabled node reports cluster_state:fail and yet answers Hash slot not served). A single-node cluster never trips the gate, so bootstrap stays usable. The gate is one bool read on state already under the caller's lock; the coverage scan behind it runs on topology change, never per command. - CLUSTER INFO reported cluster_enabled, which Redis reports only in INFO. Removed from CLUSTER INFO; measured against redis-server 8.6.1, which emits it in INFO's Cluster section and never here. - A multi-node cluster silently served keys it did not own (#485). CLUSTER ADDSLOTS does not bump the config epoch, so a hand-built cluster sits at epoch 0 forever. The gossip merge accepted a peer's slot bitmap only at a strictly higher epoch — 0 > 0 is false — so peer ownership was never merged and every node believed the only slots in existence were its own. route_slot then found no owner for a peer's slot and fell through to a bootstrap fallback that serves everything unclaimed locally. The result was the worst available failure mode: a write went to the wrong node, returned +OK, and was invisible to the node that actually owned the slot — no MOVED, no error, nothing a client could detect. The merge now accepts at equal epoch (a node is authoritative for its own bitmap; a strictly lower epoch is still ignored, preserving Redis's highest-epoch-wins tie-break), and the unclaimed-slot fallback is split: a lone node still serves locally so single-node bootstrap keeps working, while a formed cluster answers CLUSTERDOWN Hash slot not served rather than inventing an answer. That is the per-slot message, measured — Redis's other CLUSTERDOWN, The cluster is down, means cluster_state is fail, and sending it for one unowned slot tells an operator the whole cluster is gone. Direct contact with a node also now clears any suspicion about it — previously nothing ever reset health, so a node that recovered stayed suspected forever and its slots never counted as covered again. - A node's role and its health were the same field, so learning one erased the other. NodeFlags had mutually exclusive Master / Replica / Pfail / Fail variants: marking a node PFAIL destroyed the fact that it was a replica, and with it the master_id recording which shard it belonged to. Redis treats these as orthogonal — a dead master is reported role: master, health: fail, a shape the old type could not represent at all. Split into NodeRole + NodeHealth, both carried together in the gossip flags word (bit 0 role, bits 1-2 health) so the wire header keeps its fixed 2130-byte size; unknown health bits decode as online, so a forward-compatible peer is never treated as failed by accident. CLUSTER NODES and nodes.conf now render both axes in Redis's own spelling: a comma-separated list where PFAIL is fail? and confirmed FAIL is fail, appended to the role (master,fail? → master,fail). Moon previously emitted pfail, a token Redis never produces and no client parses. - The nightly fuzz job never reported, and threw away its corpus every night. The nightly budget was -max_total_time=21600 — exactly GitHub's 6h hosted-job ceiling — so the fuzzer was still running when the platform killed the job. Measured on the 2026-08-14 run: inline_parse, resp_parse and resp_parse_differential each ended at 6h 00m 17s with conclusion cancelled. A cancelled job reports no findings, and because the fuzz step never returned, the Archive corpus step never ran either — so every night's accumulated corpus was discarded and the next night restarted from scratch. Only targets that fail fast enough to finish inside the ceiling (mq_registry_blob, ~7 min) were ever visible, which is why the workflow looked merely "red" rather than broken. The budget is now 5h under an explicit 350-minute job timeout, leaving headroom for setup, build and archiving; an overrun is now a clean job timeout instead of a silent platform cancellation. This is the gap that let a pre-auth inline-parser hang reach main undetected. - RESP3 pub/sub confirmations arrived as Arrays where Redis sends Push frames. Moon's deliveries were already correct — message and pmessage lead with > — but subscribe, unsubscribe, psubscribe and punsubscribe were built by four functions that hardcoded Frame::Array and took no protocol argument. That half-correctness is worse than no RESP3 support at all: a client that separates out-of-band pushes from command replies by the leading byte reads the Array confirmation as the reply to whatever it sends next, and every later reply on that connection is off by one. Confirmations are now built as what they mean — Push — and each protocol's serializer renders it in that protocol's form, so RESP2 clients see byte-for-byte what they saw before. The tokio handler additionally serialized every pub/sub frame through the RESP2 serializer regardless of the connection's protocol, which downgraded Push back to Array; it now picks the serializer the way the codec already does. - RESP3 no longer keeps a subscribed connection in the RESP2 jail. Redis lets a subscribed RESP3 connection run any command — that is a large part of why RESP3 exists — while Moon diverted every subscribed connection into a loop that answers only pub/sub verbs. RESP3 connections now stay in the normal command loop and take a delivery branch alongside it. A reply and a delivery cannot tear: both are whole frames written by the same task. Two further divergences closed themselves as a consequence: a subscribed RESP3 PING now answers +PONG instead of the RESP2 *2 pong "" shape (the shape follows the protocol, not the mode), and a subscribed RESP3 connection can now run GET/SET at all, where Moon used to refuse both. - The subscriber-mode allow-list was stated in three handlers with two different texts and two different behaviours. Only the sharded handler accepted RESET; handler_single advertised HELLO as allowed in its error message while refusing it; and none matched Redis, which names (P|S)SUBSCRIBE / (P|S)UNSUBSCRIBE / PING / QUIT / RESET. The rule now lives in one module. RESET works everywhere, the sharded verbs are admitted, and HELLO stays refused (measured — a RESP2 subscriber genuinely cannot upgrade mid-subscription). - PUBSUB NUMPAT counted subscribers instead of distinct patterns, and double-counted a pattern whose subscribers landed on different shard threads. Two clients on p.* answered 2 where Redis answers 1. It now reuses the same gather INFO uses, so pubsub_patterns and NUMPAT cannot disagree. The bug survived because it is only visible while both subscribers are live — after one leaves, the buggy and correct answers coincide. (#480) - UNSUBSCRIBE with no arguments on a connection subscribed to nothing named an empty channel ($0) where Redis sends a Null channel ($-1); a statically-typed client decodes the two differently. - keyspace_hits and keyspace_misses had been reporting zero for plain GETs on the shipped runtime. try_inline_dispatch — the route a plain GET key actually takes under monoio — frames its reply straight into the write buffer and returns, reaching neither string::get nor string::get_readonly where the recorders live. Every hit-rate dashboard built on those two fields has been dividing by zero-over-zero. The reason it survived CI is worth stating: the counters read CORRECTLY under runtime-tokio, which is what the test matrix ran, so a test that proved the fix under tokio passed while the shipped runtime stayed broken. The inline path now records hit, miss, and the keymiss notification. (#477) - maxmemory now bounds what the process actually costs, not what the allocator says is live. A live instance ran for an hour with used_memory:4.21G against a 10 GB real footprint — a 2.3x gap — so a 19.2 GB cap never engaged, eviction never fired (spill_batches_flushed:0), and the OS reclaimed by swapping 9.9 GB instead: 76M page faults and 53 GB read back off disk. Activity Monitor showing "Real Memory 66 MB" was the symptom, not the reassurance it looks like — that is the residue left after everything else was paged out. The eviction budget is now scaled by how far real footprint exceeds accounted memory, so the cap holds against the number the machine is charged for. Two things this fix needed to get right, both caught before merge: the macOS reader was pulling resident_size (offset 16 of task_vm_info_data_t) rather than phys_footprint (offset 144), which on a swapped-out process reports 91 MB for a 10 GB instance and would have made the whole correction inert; and the ratio must ignore fixed process overhead — binary, thread stacks, arena metadata cost tens of MB before a single key exists — so it prices only the marginal cost of the data over a startup baseline and stays inert below 64 MiB of accounted memory. INFO memory gained mem_fragmentation_ratio and maxmemory, and used_memory_rss now reports footprint rather than a resident figure that understates the truth by two orders of magnitude under swap. - A malformed frame no longer closes the connection silently, and no longer eats the valid commands that arrived with it. Err(_) => break in the read loops discarded two things: the parse error's reason, so a client got a bare FIN and could not tell a bad encoder from a dropped network, and the already-parsed batch — so PING\r\n*-9\r\n in a single write answered nothing at all. All three handlers now execute and flush the valid prefix, then send -ERR Protocol error: <reason> using redis-server 8.6.1's verbatim wording, then close. Also measured and fixed alongside it: *-9 (any negative multibulk count) killed the connection where Redis ignores it; an oversized inline request closed mute even though Moon already built the correct "too big inline request" message; and inline quoting did not exist at all, so SET k "a b" became three arguments containing literal quote bytes and GET "unclosed was silently accepted as a key — the inline parser is now a port of Redis's sdssplitargs. - MULTI is atomic with respect to queue-time faults. Moon had no queue-time validation, so EXEC ran whichever half of a transaction happened to parse: MULTI / NOSUCHCMD / SET k v / EXEC left k set where Redis discards everything — data corruption, not a compatibility nit. A command that could never run (unknown name, impossible arity, SUBSCRIBE, WATCH) is now refused as it is queued and poisons the transaction, and EXEC answers -EXECABORT Transaction discarded because of previous errors. Fixing this exposed a wider defect: Moon decided "am I in a transaction?" hundreds of lines below the INFO / CLIENT / WS / MQ / PUBLISH / SUBSCRIBE intercepts, so every one of those executed for real inside a transaction — SUBSCRIBE ch put the connection into subscriber mode mid-MULTI, and INFO server returned a 3 KB dump where Redis returns +QUEUED. The queue decision now sits directly below each handler's ACL gate, which fixes the whole class at once.

  • HELLO no longer contradicts INFO replication about what the server is. hello_acl built its reply with mode and standalone and role and master as compile-time literals, so a replica told HELLO it was a master while telling INFO it was a slave — on the same connection. Both fields now read ReplicationState/ClusterState, the same source INFO uses. Note that Redis deliberately uses three vocabularies for this one fact (HELLO -> replica, INFO -> slave, ROLE -> slave), verified against a live replica pair on redis-server 8.6.1, and Moon now matches each.
  • The client-compat harness no longer needs a live Moon defect to test itself. Its divergence-classification test borrowed a real bug as its fixture, so it FAILED whenever someone fixed that bug — three fixtures were burned this way (GET-inside-MULTI #457, SISMEMBER RESP3 #463, and COMMAND COUNT here). It now fabricates the divergence through a test-only inject_moon_reply hook, with a guard test asserting the shipped manifest never uses it. The identity_command_count and identity_role waivers were retired as fixed.
  • CLIENT INFO / CLIENT LIST report the real laddr. The field was the literal laddr=127.0.0.1:0 inside the format string, so every client on every listener reported port 0 — a monitoring agent using laddr to tell listeners apart saw nothing.
  • CI now tests the runtime Moon actually ships (check-monoio). Every CI job that executed tests did so under --no-default-features --features runtime-tokio,…, while Moon's default feature set — and what ships on Linux — is runtime-monoio. 26 monoio integration test files and 30 monoio-gated src/ files were unreachable by CI. That is how the v0.8.6 inline-GET ACL bypass shipped green: it was wrong only on the monoio dispatch path, which no CI job could see. The new job runs on the self-hosted Linux runner (the only place monoio's io_uring driver executes at all), uses the default feature set, and invokes cargo nextest run --profile ci so the repo's existing flake policy applies — a bare cargo test has no retries, and an intermittently-red required job gets disabled, which is worse than no job because it still looks like coverage. Measured before landing: 5145 passed, 1 flaky, 244 skipped, exit 0, 80.3s.

Proven by negative control rather than asserted: a deliberate defect on the monoio-only try_inline_dispatch path passes the tokio suite 6/6 and fails the new job 3/6. tests/ci_covers_monoio.rs guards the job against being silently weakened — wrong feature set, continue-on-error, a bare cargo test, or a shared CARGO_TARGET_DIR all fail the suite.

The job's first cut proved the premise but not the driver: MOON_NO_URING: "1" lived in the workflow-level env:, which merges into every job and cannot be unset by one, so the job whose whole point is io_uring ran with io_uring force-disabled. monoio_yield_overhead_is_microscopic caught it at 1.45ms/yield (the sleep(ZERO) timer-park signature). That variable is now per-job, and ci_covers_monoio.rs asserts both scopes. The guard suite itself then failed on Windows — it matched "\n <job>:\n" against a CRLF checkout — so it now normalizes line endings and carries a platform-independent CRLF regression test (Windows is skipped on every PR, so that class is invisible until a main push).

  • Client-compat harness: raw-RESP diff against a real redis-server (scripts/test-client-compat.sh). Moon's existing Redis comparison (scripts/test-commands.sh) goes through redis-cli, which renders replies to text before any assertion can see them — it even strips (integer) — and never sends -3, so the entire RESP3 surface was uncompared. That blindness is why ~22 type-level compatibility defects survived into v0.8.5. The new harness speaks RESP on a raw socket and compares in a fixed order — TYPE, then SHAPE, then VALUE — across the full {RESP2, RESP3} × {standalone, MULTI/EXEC, pipeline} matrix, so a finding says which of the three diverged and whether a reply changes shape by context. A missing redis-server FAILS (ERR_NO_ORACLE) rather than skipping. New CI job client-compat on the self-hosted runner; the manifest carries a recorded baseline of reasoned waivers so the job is a ratchet — a new divergence fails it, and --strict fails the moment a waived divergence is fixed and its waiver goes stale. First full run vs Redis 8.6.1: 152 comparisons, 94 pass, 58 waived, plus 34 named missing INFO fields via --info-manifest.

  • WATCH / UNWATCH now actually guard a transaction on the production dispatch paths. Both commands parsed and answered +OK, and the tokio and embedded handlers re-checked the recorded versions at EXEC — but the two paths clients really reach (handler_monoio, handler_sharded) never consulted the watch set at all. A conflicting write from another client committed anyway, so every check-and-set built on WATCH (inventory decrement, balance transfer, leader election) silently degraded to last-writer-wins. Four defects, all fixed together because a partial fix is indistinguishable from none:

  • The watch set was not consulted at EXEC on the monoio and sharded handlers. EXEC now aborts with a RESP null array on any version mismatch, and clears the watch set on both outcomes by construction (mem::take) rather than by a clear() each exit path has to remember.

  • A watched key on another shard was read from the local slice — a different database entirely, so the version compared was some unrelated key's or zero. WATCH now snapshots versions where the keys live, via a new ShardMessage::ReadVersions (one hop per owning shard, not per key), and a watch set spanning shards is classified and refused CROSSSLOT like a cross-shard MULTI body already was.
  • WATCH inside MULTI was queued as an ordinary command instead of being refused, and WATCH with no arguments answered "unknown command" instead of an arity error.
  • Delete + recreate was invisible (ABA). Versions are per-entry and die with the entry, so every incarnation of a key started at INITIAL_VERSION: DEL k + SET k handed the watcher back the exact token it had recorded and EXEC committed on a key that had been destroyed and rebuilt underneath it — where Redis aborts. Entries are now stamped from a per-database creation ticket, so a recreated key is observably a different incarnation. Restored keys draw tickets too: versions are not persisted, so otherwise the first key created after a restart would collide with the whole restored population.

Residual risk, stated rather than buried: the ticket shares the entry's 24-bit version field and so wraps at 16,777,216 creations (~18s of saturated single-database insert at the measured 914K/s). A miss now needs that wrap to land inside one client's open WATCH..EXEC window and hit the one watched key — ~1 in 16.7M, against the pre-fix certainty. Only a wider incarnation field removes it entirely; WatchToken stays a named struct so adding one later does not churn the call sites.

New suite tests/watch_cas_transactions.rs (10 wire-level tests, raw-byte assertions, two connections, run at both --shards 1 and --shards 4), plus WATCH/CAS entries in scripts/test-consistency.sh and scripts/test-commands.sh.

  • RESP3 reply types now match Redis, and no longer change with the calling context. A live sweep against redis-server 8.6.1 found the conversion table wrong in both directions and, structurally, unable to be right: it keyed only on the command NAME, while WITHSCORES, WITHVALUES and a <count> argument are what actually decide the reply shape. ZRANGE … WITHSCORES arrived as a flat array of bulk strings instead of pair-wrapped [member, Double] — enough to make an unmodified redis-py raise ValueError: not enough values to unpack, so every RESP3 application using sorted-set scores was broken outright. HRANDFIELD WITHVALUES and ZRANDMEMBER WITHSCORES answered a Map where Redis answers an array of pairs. In the other direction the server over-converted: all seven of SISMEMBER, HEXISTS, EXPIRE, PEXPIRE, PERSIST, SETNX and MSETNX returned Boolean where Redis returns Integer, and INCRBYFLOAT/HINCRBYFLOAT returned a lossy Double (,10.6) where Redis returns the exact Bulk "10.59999999999999964".

The conversion is now decided by (command, args) at one policy choke point (Resp3Shape in src/protocol/resp3.rs) instead of by 11 call sites across three handlers. Because the cross-shard reply loop no longer has the command's args by the time its batch returns, the shape is classified at ENQUEUE time and the 1-byte Copy tag travels in RemoteMeta — no per-command allocation on the shard hot path. execute_transaction_sharded now takes the connection's protocol version and converts each inner reply with its own command, so a command answers the same shape standalone, inside MULTI/EXEC and inside a pipeline; previously SMEMBERS was a Set outside a transaction and a flat Array inside one. CONFIG GET (Map) and CLIENT INFO (Verbatim) are fixed at their intercepts, which short-circuit the dispatch exit entirely — the reason CONFIG could never be fixed before. Also added: ZMSCORE and GEOPOS Doubles with Null preserved, SPOP <count> as a Set, ZPOPMIN/ZPOPMAX Double scores (flat with no count, wrapped with one), and XINFO STREAM as a Map.

Emptiness no longer changes the reply type either: HGETALL and CONFIG GET on a miss answered *0 where Redis answers %0, so a client dispatching on the type byte broke on exactly the path it hits most. Every case in the harness populated its key first, which is how that survived a full green run — five miss-path cases were added so it stays diffed against the live oracle.

RESP2 is byte-identical — pinned by a test written before the fix. Harness result vs Redis 8.6.1: 157 pass / 0 fail / 25 waived, up from 98 / 0 / 54, with all 13 now-stale waivers deleted and --strict green. New suite tests/resp3_type_fidelity.rs (13 tests) asserts the wire type byte directly. Known remainder, waived and reasoned: XINFO STREAM is now the right TYPE but still reports 7 fields to Redis's 16 — the missing ones need real stream bookkeeping and are tracked separately rather than fabricated. - A key whose disk-offload spill was in flight was invisible to the whole server — DEL on it was silently undone (#459). evict_one_async_spill freed the hot entry the moment the SpillRequest was queued, and the key was only registered in cold_index when the completion landed. In between it existed in no plane the database consults: spill_inflight held only a request id and was read solely by the completion path. The eviction code documented the window as an acceptable "brief read-miss" backstopped by the AOF — but the AOF backstops durability, not visibility, and nothing considered a write landing inside the window. Measured on 400 × 4 KiB keys against a 512 KiB cap: DBSIZE answered 124 for 400 acked keys and then climbed to 400 on its own; GET/EXISTS denied live keys that returned unaided 250 ms later; and 277 of 400 DELs answered :0 and were then reversed by the completion — which publishes into cold_index unconditionally, so those resurrections reached the manifest and survived restart. A client that deleted data got it back.

spill_inflight is now a real third storage plane carrying the payload (the same refcounted Bytes the queued request already pins — a refcount, not a copy, and no extra peak memory). Reads promote from it with no disk read at all, across all three dispatch paths (promote_cold_if_present for collections and Lua/MULTI, the monoio async GET pre-warm, and the inline GET fast path). EXISTS/DEL count it, DBSIZE/logical_len count it, and KEYS/RANDOMKEY enumerate it. DEL, an overwriting SET, and a promoting read each retire the record, which withdraws the completion's authorization to publish — that is what makes a delete inside the window final. Unpublished completions are counted as spill_completion_superseded in INFO. Non-spilling servers pay one is_empty() load on the affected paths. Known remaining gap, documented at the call site: SCAN's ordered cursor does not merge the unordered in-flight plane, so it may skip a key for the milliseconds its spill is queued — within SCAN's contract, unlike KEYS. - GET inside MULTI was executed instead of queued (monoio). Third defect from the same ungated inline read path, and the one most visible to a working client: MULTI; GET k; EXEC answered +OK, $1 v, *0 where Redis answers +OK, +QUEUED, *1[$1 v]. The client receives a value where it expects +QUEUED, and then an EXEC that silently omits the read — a redis-py/go-redis transaction returns an empty result set for an exchange it believes succeeded. can_inline_writes already carried !conn.in_multi (so SET queued correctly); can_inline_reads now carries it too. MGET, not being inline-eligible, queued correctly throughout, which is what isolated the path. Found by the new scripts/test-client-compat.sh raw-RESP harness on its first run against a real redis-server. New suite: tests/multi_queues_inline_get.rs, including controls pinning that MGET and SET still queue and that a plain GET outside a transaction still takes the fast path (measured: 2000/2000 GETs still local_inline, before and after a completed transaction). - CLIENT TRACKING answered +OK and then never invalidated (monoio). Same root cause: the inline GET path also skips tracking::invalidation::track_read_keys, so a client-side-caching client's own GETs were never registered and nothing ever invalidated them — the cache served stale data indefinitely. At --shards 1 this was total; at --shards 4 it depended on whether the key hashed to the reader's own shard, which also made the existing mset_invalidates_every_second_arg_key test flaky (observed failing 2 of 3 runs on 0.8.5). Reads from a connection with tracking enabled now take the generic path. Deliberately gated on this connection's own tracking_state.enabled, not the process-global tracking_active(): only a connection's own reads populate its invalidation set, so one caching client must not push every other connection off the fast path. - CLIENT TRACKING ON BCAST with no PREFIX never invalidated anything. Redis treats prefix-less BCAST as "invalidate me for every key", but the handlers only registered broadcast interest inside for prefix in &config_parsed.prefixes, so a prefix-less BCAST client registered nothing — at any shard count, and regardless of what it read (BCAST does not depend on reads). parse_tracking_args now normalises prefix-less BCAST to the empty prefix, which TrackingTable's key.starts_with(prefix) match treats as "all keys"; one change fixes all three handlers.

The default (unrestricted, non-tracking) connection keeps the inline fast path byte-for-byte, verified against moon_dispatch_path_total{path="local_inline"}: 2000/2000 GETs still inlined with and without --requirepass, 0/2000 for restricted and tracking connections.

Added

  • INFO Server now reports num_shards (#497), the resolved shard count rather than the configured one, so --shards 0 auto-detection reports the number actually in effect. A client that cannot operate against a multi-shard instance — one committing a cross-key transaction per request — previously had no way to refuse at connect time except a two-key co-location canary, which is slow and still a guess. redis_mode/cluster_enabled do not answer this: a single-process Moon with several shards is standalone and still routes keys across threads. The value is never 0; an embedded harness that never records a count reports the single-shard truth, because a zero reads as a valid answer and is not one.

  • An acceptance suite driven by unmodified redis-py (scripts/client-compat/redis_py/), wired into the client-compat CI job. The raw-RESP differ compares bytes against a real redis-server; it is precise and blind to everything a client library does around the reply — the handshake it opens with, the connection it reuses, the Python type it decodes into, the second command it issues on your behalf. A server can answer every byte correctly and still be unusable from redis-py. The suite therefore drives redis-py's own idioms — connection pools, pipeline(), pubsub(), scan_iter()/hscan_iter(), redis.lock.Lock, from_url, RESP2 and RESP3 handshakes, WATCH optimistic locking — rather than hand-rolled sockets. Stdlib unittest and the distro python3-redis package, because the runner has no pytest and PEP 668 blocks pip.

It found three defects on its first run, each pinned inside the suite so that fixing one breaks the run with an actionable message rather than leaving a stale skip: at --shards >= 2, an MGET in the same pipeline batch as the SETs that wrote its keys returns nulls despite those SETs acking +OK earlier in the batch (a silent read-your-own-writes violation), EVALSHA of a single-key script is rejected with CROSSSLOT (which breaks redis.lock.Lock.release()), and CLIENT INFO reports a literal cmd=NULL for every connection.

The two multi-shard defects fire for roughly half of keys — decided by which shard owns the key relative to the connection's own shard — so each pin runs twenty distinct keys rather than one. A single-trial pin was tried first and made CI flaky, which is how the ~50% rate was found; the amplified form is 0/10 flaky runs on both macOS and Linux, and still fails loudly when the underlying bug is fixed.

  • Nine INFO fields a standard monitoring stack reads. tcp_port, uptime_in_seconds, uptime_in_days, aof_last_write_status, aof_last_bgrewrite_status, rdb_changes_since_last_save, sync_full, sync_partial_ok and sync_partial_err are now emitted, each from a real source rather than a constant: uptime from a start instant captured before the listener binds, rdb_changes_since_last_save from a sharded keyspace-mutation counter reset at save completion, and the three sync_* counters recorded at the one point in the PSYNC handshake where full-vs-partial is still distinguishable (PSYNC ? -1 counts as a full resync the replica ASKED for, not a partial that failed). tcp_port reports the configured listener port, not the port the INFO connection arrived on — behind a container port map the two differ, and the field exists so a client can hand a peer a reachable address. The four fields Moon cannot answer truthfully are recorded as waivers with reasons instead of being emitted: latest_fork_usec (Moon never calls fork(2); BGSAVE snapshots in-process), the two client_recent_max_*_buffer high-water marks (untracked, and tracking them means a counter on every connection read and write), and used_memory_lua — which was implemented and then withdrawn when measurement showed the value the shard can publish is ~2 orders of magnitude below the real VM footprint on the shipped monoio runtime (80 bytes against tokio's 26183, same setup_lua_vm, ruled out as Cargo feature unification). This keeps the INFO emitter's standing rule intact: a wrong number on a dashboard is worse than an absent one.

  • ZRANK/ZREVRANK ... WITHSCORE (#521), the Redis 7.2 option — previously an arity error, so a 7.2-aware client asking for the rank and the score in one round trip got a hard failure. A hit is the two-element array [rank, score]; a miss (absent key or absent member) is the null ARRAY *-1, not the null bulk the option-less form answers, so WITHSCORE changes the null type as well as the hit type — measured against redis-server 8.6.1, and a statically-typed client decodes the two differently. The token is SINGULAR, unlike the WITHSCORES that ZRANGE takes; the plural is a syntax error here, and only a FOURTH argument is an arity error (Redis distinguishes the two and so does Moon). The score is emitted as a RESP double, so RESP3 gets ,1 while RESP2 keeps the identical bulk-string bytes. Both entry points were fixed — zrank and zrank_readonly each carried their own arity check, and which one answers is decided by shard routing, so fixing one would have left the option working on some keys and failing on others. ZRANK/ZREVRANK arity in COMMAND INFO is now -3, matching Redis 7.2+.

  • RPOPLPUSH source destination (#520) — previously "unknown command" even though LMOVE and BRPOPLPUSH both worked. It is deprecated in Redis but never removed, and it is the form baked into a decade of client code: redis-py, jedis, go-redis and node-redis all expose it as a first-class method, so r.rpoplpush(...) failed outright. The name was already in the workspace routing table, in the metrics labels, and was the internal rewrite target of BRPOPLPUSH — missing only from the two places a client reaches, the dispatch table and metadata.rs (COMMAND INFO rpoplpush answered an empty array, so driver feature-detection concluded the command did not exist). Implemented by delegating to LMOVE's body with RIGHT LEFT rather than by duplicating the list logic, and registered with LMOVE's two-key spec (arity 3, write, first_key 1, last_key 2, step 1, ACL @list @slow). Along the way the blocking-wakeup guard — open-coded at eight dispatch sites — moved behind one shared predicate, which fixed a drift the duplication had hidden: the two connection handlers woke LMOVE's SOURCE key, which no blocked reader waits on, while the SPSC handler correctly woke its DESTINATION. A BLPOP dst was therefore woken or left to time out depending on which shard owned the key.

  • Keyspace notifications (notify-keyspace-events). Moon had none: cache-invalidation frameworks and change-data-capture consumers subscribe to __keyspace@<db>__:<key> and __keyevent@<db>__:<event> and got silence. Both channel families are now published, gated by the full Redis flag model — including the parts that are not what the letters suggest, all measured against redis-server 8.6.1 rather than recalled: CONFIG SET KEA reads back as AKE, mn as nm, Km as Km, and A deliberately excludes m (keymiss) and n (newkey) so it stays safe to enable in production. Events wired: set, incrby (INCR publishes incrby, not incr), rename_from/rename_to (RENAME emits BOTH halves, carrying different keys), expired, keymiss. Off by default and genuinely zero-cost when off — one relaxed atomic load. Delivery reaches subscribers on every shard, not just the one that owns the mutated key: a local-only publish would pass at --shards 1 and silently drop roughly (N-1)/N of events at --shards N.

  • ROLE, RESET, and a real COMMAND introspection surface. COMMAND and COMMAND COUNT each returned the OTHER'S RESP TYPE — bare COMMAND replied :0 (an Integer where an Array belongs) and COMMAND COUNT replied *0 (an Array where an Integer belongs); COMMAND INFO/DOCS/LIST/GETKEYS all replied an empty array. A driver that builds its command map at connect time does not read that as "unsupported", it reads it as a protocol violation, and redis-cli renders :0 and *0 identically as "0" so it never showed up by eye. All six now derive from COMMAND_META, so registering a command is what makes it introspectable and there is no second table to drift. ROLE and RESET were unknown commands: RESET was registered in the metadata table with full flags while dispatch rejected it (the same advertise-then-reject class as WATCH/UNWATCH before v0.8.6), and a partial RESET existed only inside handler_sharded's subscribe-mode loop, so it worked if you happened to be subscribed on one runtime and nowhere else. RESET is wired on every production path and on handler_single (the in-process single-shard handler that only tests reach) so the third copy of this surface cannot drift unnoticed the way it just did. ROLE is answered from the shared dispatch table instead, reading the process-global replication handle that INFO already uses — which is what lets a QUEUED ROLE work: EXEC replays the queue through dispatch(), so a connection-layer intercept would have executed ROLE at queue time and dropped it from the EXEC array, shifting every later result index for the client. Caught by the client-compat harness.

Changed

  • One producer→waiter hook instead of eight hand-copied ones (#623). Deciding "is this command a producer, which of its arguments names the key it just made non-empty, and which family's queue should be raised" was open-coded at every dispatch site: both connection handlers and six arms of the SPSC handler. Adding a producer command meant finding all eight, and a new execution path that reached none of them looked healthy in review and in CI — that is how XADD on a locally owned stream (#595) and every write inside MULTI/EXEC (#606) each shipped a lost wakeup. The trio is now wakeup::wake_producer, called once per site; the per-site success gate (is_write && !error, or just !error) stays where it is because it legitimately differs between paths. No behaviour change — the mapping and the gates are the ones that were already there, and the existing blocking suites plus new unit tests over the single copy hold it in place.

  • Cross-shard dispatch no longer clones the command name for every command (#460). RemoteMeta carried a Bytes holding the command name on both handlers. The cross-shard reply loop needed it before RESP3 type fidelity landed, to choose the conversion; since then the shape is classified at ENQUEUE into a 1-byte Copy tag (Resp3Shape) and the reply loop reads only the tag. The field survived as _cmd_name, still paying extract_bytes(&args[0]).unwrap_or_default() — an atomic refcount increment, and a later decrement — once per cross-shard command on the shard hot path.

Removed from both handler_monoio and handler_sharded, along with the corresponding slot in the per-batch metadata tuple, which shrinks by size_of::<Bytes>() = 32 bytes per command (measured in-crate; the issue estimated 16).

No throughput claim is made here. Per-command refcount traffic on the cross-shard path is a structural improvement visible in the diff; a benchmark able to resolve it from noise would have to come from the Linux VM per CLAUDE.md, and was not run. The change's real justification is that the value was dead.

  • compaction.rs and metrics_setup.rs split into directory modules (#479). Both files were over the 1500-line ceiling CLAUDE.md sets (2408 and 2038 lines); each is now a directory module whose largest file is 800 lines. src/vector/segment/compaction/ splits along the pipeline's real seams — graph_build (cell assignment, parallel sub-graph build, cross-cell stitching, builder selection), compact_path (frozen -> immutable), merge (the GraphUnion consolidation), and recall (both verification gates, including the moon#546/#588 distance-scoring rationale). src/admin/metrics_setup/ splits by metric family — init, command_metrics, recorders, memory, globals, publishers. Pure code movement: every moved body is byte-identical to its pre-split text (verified by diffing the extracted ranges against HEAD), the public symbol set is unchanged (8 and 98 items, diffed before/after), and the lib test set is unchanged at 4878 tests with identical names — the count check is what caught the one real hazard, a metrics_setup test that reaches into CachedMetricsHandles private state and would otherwise have silently stopped compiling. pub(super) on the handful of newly cross-file internals reproduces exactly the visibility they had inside the single file; no item became reachable from outside its old scope.
  • Retiring a document from a text index walked the whole field vocabulary, and left the emptied terms behind (T2). TextIndex::remove_field asked every term the field had EVER seen "were you carrying this document?", and the answer was no for almost all of them — so the cost of one removal was the size of the corpus's entire term history, not the size of the document. Worse, a term whose last carrier left kept its now-empty bitmap forever, so the vocabulary only grew: every later removal and every search miss paid for documents that had been deleted long ago. The same unbounded-growth shape as the manifest entries in moon#546 — a cost with no reader.

Each field now keeps a forward map (doc_id -> the terms it contributed) beside the inverted one, sharing the same Arc<str> allocation per term, so the forward side costs one pointer per (document, term) pair and never a second copy of the text. Removal touches only the terms the document actually held, and retires any term it was the last to carry.

Measured on one field, ten distinct terms per document, timing the removal of every document (cargo test --release --lib bench_removal_cost -- --ignored --nocapture, same machine, same build, the only change being remove_field reverted to the old sweep):

docs   vocabulary   before      after     speedup
250    2,500        2.40 ms     0.20 ms   11.8x
500    5,000        7.67 ms     0.37 ms   20.7x
1,000  10,000      21.85 ms     0.80 ms   27.2x
2,000  20,000      68.05 ms     1.87 ms   36.4x

Eight times the documents cost 28.4x more before and 9.2x after: the speedup grows with the corpus because what was removed is a quadratic, not a constant factor.

Not changed here: PayloadIndex::remove_field's tag and numeric sweeps have the same shape, but their per-field value sets are typically bounded (categories, prices) where a text vocabulary is not, so they are a smaller problem and a separate change.

  • Dependency bumps, batched into one verified lockfile change (dependabot #603, #532, #530, #528, #529). Five separate PRs all rewriting Cargo.lock invalidate one another the moment any one of them lands, so they were resolved into a single branch, verified once, and gated once rather than serialised through five matrices. Contents: base64 0.22.1 → 0.23.1 (the only manifest change), num-bigint 0.4.6 → 0.5.1, redis 1.5.0 → 1.6.0, ureq 3.3.0 → 3.4.0, aws-lc-rs 1.17.3 → 1.18.0 (aws-lc-sys 0.43 → 0.44), io-uring 0.7.13 → 0.7.14, roaring 0.11.4 → 0.11.5, futures-util 0.3.32 → 0.3.34, plus bytemuck, cudarc, http-body-util, thiserror, uuid patch bumps and mozilla-actions/sccache-action 0.0.10 → 0.0.11.

Three of these are not routine and were checked deliberately rather than on a green default build. base64 is optional = true and reachable ONLY through the console feature — a build the PR gate does not run — so its green per-PR CI proved nothing about it; verified under --features console, including the admin::auth signature path that decodes with it. num-bigint 0.4 → 0.5 is a breaking change under semver. io-uring backs the SHIPPED monoio runtime, which is likewise dispatch-only. Verified locally on native macOS arm64 across the default, console, and runtime-tokio,jemalloc feature sets, with the full lib suite green (4,877 passed), then through the complete dispatch matrix.

  • Active expiry runs on a deadline-ordered index instead of probabilistic sampling (#541). Every hot entry with a TTL now has an (expires_at_ms, key) pair in a per-database ordered index, maintained O(log n) by the storage-layer writers; the 100ms sweep pops exactly the DUE keys off the front — no 20-key random sample (which went blind when due keys were a small fraction of the volatile population, leaving them resident for many ticks), and no O(N) full-map scans (three per tick on any database with even one TTL'd key). Measured on a 100K-volatile-key database: 1.36ms → 47ns per tick, and the old cost exceeded the sweep's own 1ms budget before it had expired anything. The hash-field-TTL sweep is now latch-gated too, so databases that never touch HEXPIRE skip its scan entirely (the scan itself when the latch is up remains O(N) — #543). INFO's expires count reads the index, O(1). Along the way, two writers that bypassed the expiry machinery were rerouted: GETEX's TTL options now go through set_expiry (previously a GETEX-only TTL never armed the sweep latch, so the key was invisible to active expiry forever), and EXPIRE-family commands hitting an already-expired key now hide and queue it for the emitting drain (#542 semantics) instead of deleting it silently with no expired notification and no dual-plane DEL. Review follow-ups: GETEX EX/EXAT seconds values that overflow the millisecond conversion now answer the range error instead of wrapping (debug builds panicked); the sweep only drops an index pair that is provably stale (entry gone or TTL retargeted) so a backwards wall-clock step can't discard live pairs; the per-key budget clock read is batched to every 64 pops; and sweep 2 lowers the hash latch from its own reap outcomes instead of a second O(N) rescan.
  • volatile-ttl eviction picks the exact nearest-expiry victim (#551). The policy now reads the head of the #541 deadline-ordered expiry index — O(log n), always the globally soonest deadline — instead of sampling maxmemory-samples random volatile keys and taking the sample's minimum, which could evict a key hours from expiring while the one expiring in seconds survived. maxmemory-samples still governs the LRU/LFU/random policies, which remain sampling-based.

  • The INFO field coverage CI step is now a hard gate, no longer continue-on-error. The pinned-field harness gained a waiver syntax (field # WAIVED: <reason>) that refuses an unreasoned waiver at load time (exit 2) and reports a waiver on a field Moon has since started emitting as a failure — so the list cannot go stale unnoticed the way the registry sweep's did.

  • CLUSTER SHARDS, CLUSTER MYSHARDID, READONLY and READWRITE — the four verbs a cluster-aware client needs to bootstrap. CLUSTER SHARDS reports every shard cluster-wide, one entry per master with its replicas, master first; a shard that has lost every live node reports an empty slots array while still listing the dead node as role: master, health: fail. The top level is an Array under both protocols and only the shard and node entries change shape — a Map under RESP3, a flat array under RESP2 — which falls out of building them as Frame::Map and letting each serializer render it, the same approach the RESP3 pub/sub work used. A node entry carries exactly id, port, ip, endpoint, role, replication-offset, health, in that order, because under RESP2 a client may read it positionally. READONLY lets a replica serve reads for slots its own master owns; a WRITE still answers MOVED, the asymmetry a "just return +OK" implementation gets wrong.

  • INFO reports cluster identity honestly. redis_mode and cluster_enabled were the hardcoded literals standalone and 0, under a comment promising the cluster subsystem would say otherwise — nothing ever did. An SDK branches on redis_mode before it ever calls CLUSTER SHARDS, so a server answering SHARDS correctly while reporting standalone stayed undiscoverable as a cluster.
  • MONITOR — the command feed. redis-cli monitor now works against Moon, in Redis's exact line format: +<unix>.<micros> [<db> <addr>] "CMD" "arg" …, arguments quoted and escaped per byte (sdscatrepr semantics — " \, \n \r \t, \a \b, and \xHH for everything outside printable ASCII, so UTF-8 escapes per byte rather than per character). The line is a SimpleString under both RESP2 and RESP3 — measured; Redis does not use a Push frame here, and the reflex to make it one after the RESP3 pub/sub work would have been a new divergence. AUTH's arguments and the credentials in HELLO … AUTH render as (redacted), decided before any argument is written rather than filtered afterwards.

Two behaviours are worth knowing because they are not the obvious implementation. First, administrative commands are hidden at subcommand granularity: CONFIG *, SLOWLOG *, LATENCY *, ACL LIST/SETUSER and CLIENT LIST never reach a monitor, while INFO, DBSIZE, LASTSAVE, CLIENT GETNAME/ID, ACL WHOAMI/CAT and CLUSTER INFO/MYID do. Moon's own CommandFlags::ADMIN is container-granular and could not express that split — using it would have hidden six commands Redis shows — and Redis feeds the entire EVAL family despite flagging it skip_monitor, so neither flag is consulted; the rule is stated explicitly and pinned row-by-row against the measured oracle. MONITOR is absent from its own feed as a consequence of that general rule rather than a self-suppression special case. Second, a monitor that stops reading has its connection dropped: silently skipping lines would leave an operator unable to tell a quiet server from a lossy feed, and blocking would let one slow TCP reader stall every shard.

A monitor connection may not touch the keyspace, matching Redis (-ERR Replica can't interact with the keyspace). The refusal set is measured, not derived from a flag: DBSIZE, KEYS, SCAN, RANDOMKEY, FLUSHALL, FLUSHDB, SWAPDB, EVAL, PUBLISH and MEMORY USAGE name no key yet are all refused, while PING, INFO, TIME, ECHO, COMMAND, LASTSAVE, WAIT, SELECT, CLIENT, ACL, SUBSCRIBE and RESET are served. Neither first_key nor Moon's WRITE/READONLY flags reproduce that split — Moon flags PING and INFO readonly and Redis does not — so the rule is stated explicitly and pinned row by row against the oracle.

Commands issued by a Lua script are fed too, carrying the literal lua in place of a peer address and appearing in execution order after the EVAL line — matching Redis. A script command never passes a connection handler, so it needs its own hook; without it an operator watching a script-driven workload would see every EVAL and none of its effects.

MONITOR requires the admin ACL category, and costs one relaxed atomic load per command when nobody is attached — every other step lives behind that load. While a monitor IS attached the inline fast path stands down, because it answers straight from the read buffer and never sees a peer address; the feed is therefore correct by construction on that path rather than by a hook that must be kept in sync. Fast-path retention when unattached is confirmed by moon_dispatch_path_total{path="local_inline"}, not inferred from latency. - Sharded pub/sub: SSUBSCRIBE, SUNSUBSCRIBE, SPUBLISH, and PUBSUB SHARDCHANNELS / SHARDNUMSUB. Deliveries carry the smessage event name. The sharded namespace is a genuinely separate map from the plain one in both the per-shard registry and the remote-subscriber map, so SPUBLISH ch cannot reach a SUBSCRIBE ch and vice versa — the two may share a channel name while being different destinations, and that separation is structural rather than a filter each call site has to remember. Cross-shard delivery reuses the existing batched fan-out with a sharded marker on each entry, and is proven at --shards 4: a local-only registry would have passed every single-shard test while dropping (N-1)/N of deliveries in production. SPUBLISH was absent from COMMAND_META entirely, while SSUBSCRIBE and SUNSUBSCRIBE were declared there but answered "unknown command" by the dispatcher — so COMMAND COUNT was advertising verbs Moon could not run.

  • CI migration: hosted-only PR gate + local merge bar — the three self-hosted jobs (Check, Check (monoio), Client compat) serialized on the single moon-dev runner on every PR (17–24m wall, with manual dispatches queueing behind them). The PR gate is now entirely GitHub-hosted and parallel (Lint, Check on ubuntu-latest with sccache + rust-cache, MSRV, Memory gate, ~5–8m); the monoio and client-compat legs moved to main-push + workflow_dispatch, and — before every push — to the new scripts/ci-local.sh (host lint gates + both full suites in the moon-dev VM; --full adds the client-compat harness and the macOS host suite). The script captures exit codes directly (no piped gates), keeps VM builds on VM-local target dirs, pins MOON_BIN for the compat harness, and fingerprints the working tree at start/end — a branch switch or edit mid-run marks the run INVALID (exit 3) rather than reporting a false green (both the fail path and the tripwire were attack-tested before landing). tests/ci_covers_monoio.rs still guards the monoio job's integrity; its trigger scope is the documented tradeoff. Console Integration was already path-gated and Integration Tests label-gated; both unchanged.

Removed

  • The dead connection::command stub (#469). COMMAND COUNT replying an empty array (and bare COMMAND replying Integer(0) — each returning the other's RESP type) was fixed in #471, which routed both dispatch sites to introspect::command; verified on the wire, COMMAND COUNT now answers :262 and COMMAND INFO/DOCS WATCH answer real specs. The superseded stub was left behind unreferenced, together with three unit tests that still asserted its wrong replies — a green test pinned to code nothing could reach. Both are deleted, and the issue's three probes are pinned as one test beside the live handler.

[0.8.5] — 2026-08-08

Added

  • Cluster mode now works on the default (monoio) runtime — v0.9 W0/C-1 (#405). The cluster control plane (bus listener, gossip ticker, failover election) runs on a dedicated cluster-ctl std thread hosting a current-thread tokio runtime, on BOTH server runtimes. Before this the monoio startup path never spawned the bus or the ticker: --cluster-enabled accepted CLUSTER MEET but no peer ever learned anything. Under tokio the control plane previously shared the listener runtime with the accept loop and every connection, so gossip starved under load (observed as PFAIL detection stalling); the dedicated thread fixes that too. The unused, untested monoio duplicates of the bus/gossip/election code were deleted — one control-plane implementation on both runtimes. New e2e suite tests/cluster_formation.rs proves a real 3-node cluster forms via MEET + gossip and flags a killed node, per runtime.

Fixed

  • Deep-review wave (2026-08): long-uptime and large-scale correctness. Eight fix groups from a six-dimension architecture review (durability ordering, long-uptime resource growth, concurrency, cluster correctness, long-horizon arithmetic, silent degradation):
  • Eviction/metadata: CompactEntry now stores a full-width u32 last_access (repurposed padding — zero size cost) and a 24-bit entry version starting at 1. Fixes LFU decay that was effectively random (u16/u32 domain mix truncated as u8), LRU inversion for idle gaps beyond ~9.1h, OBJECT IDLETIME wrapping at 18.2h, and WATCH/EXEC's 8-bit version ABA (wrap now needs 16.7M writes; version 0 reliably means "absent" so WATCH detects creation).
  • AOF everysec: tokio deadline-fsync paths no longer swallow errors and record success; all four sites (both runtimes/layouts) latch aof_last_fsync_status:ok|err + aof_fsync_failures into INFO.
  • Spill: a failed background spill pwrite re-inserts the already-evicted key into the hot table (payload carried back in the completion) instead of silently serving nil until restart; counted as spill_failed_reinserted in INFO.
  • AOF rewrite prune: the old generation is deleted only after a new manifest-sync flush barrier (pending deferred ShardManifest commits made durable — they relied on the old incr as their crash backstop) and an explicit durable dir-fsync of the manifest flip.
  • CLIENT TRACKING: the documented max_keys bound is enforced (evict + invalidate instead of unbounded growth).
  • Cluster: inline GET/SET fast path is disabled in cluster mode (was bypassing MOVED/ASK entirely); election acks are now actually received (read back on the request stream), epoch-bound and voter-deduped; the election winner no longer claims an unvoted epoch; graceful CLUSTER FAILOVER errors instead of permanently wedging automatic failover; ClusterState migrated to parking_lot::RwLock (no poisoning, ~25 .unwrap() sites removed from the per-command path).
  • Fail-loud + bounds: vector background-compaction failures now log at every layer; CDC reaps disconnected subscribers on write-idle shards; TemporalRegistry is bounded (262K bindings/shard, oldest evicted).
  • Durability wave 1 (#452, #54): AOF rewrite-window append drops, degraded-state visibility, escalated reason-DEL backpressure, and a fail-loud WAL mid-chain tear policy. (1) While a rewrite fold ran, the writer thread was out of its recv loop, so under sustained pipelined writes the bounded (10k) append channel saturated and acked records were dropped — lost even on a clean restart, and default-on since #433 made rewrites automatic. A per-writer RewriteOverflow spill buffer (aof_rewrite_buf equivalent; 256 MiB cap, strict ordering, all six fold arms on both runtimes) now buffers the overflow and drains it into the committed incr right after the fold; INFO persistence gains aof_rewrite_overflow_spilled. Merge-base A/B (tests/recovery_matrix_w1.rs): main lost/error-failed ~5.6k acked writes per hit; fixed is exact across SIGKILL + recovery. (2) Any dropped acked append now latches sticky aof_last_append_status:err (INFO) via a single accounting helper. (3) Eviction/expiry reason-DELs get a 100× escalated backpressure bound (500ms) plus a dedicated aof_reason_del_dropped counter — a dropped reason-DEL means restart replay resurrects data clients were told was gone. (4) WAL v3 replay now REFUSES to continue past a corrupt record when later segments exist (mid-chain tear = on-disk corruption; applying later segments silently replays operations from after a hole) — MOON_WAL_SALVAGE=1 is the explicit operator override; a torn FINAL segment stays the benign crash-tail it always was.

Adversarial-review hardening round (pre-merge): the rewrite fold's exactly-once contract now extends to the overflow buffer — a snapshot cut (mark_cut, recorded at the fold's atomic snapshot instant) splits spilled entries so a COMMITTED fold discards pre-snapshot spills (their effects are in the new base; replaying them would double-apply INCR/APPEND/LPUSH) while an aborted fold still writes everything. Producer paths are fully gated (spill_first on every enqueue leg, including backpressure/AppendSync parks) so mid-fold channel drains can never invert same-key replay order, and the finish drain is two-phase (channel before buffer, disarm under the buffer lock). Arming is unwind-safe: a panicking fold disarms with drop accounting instead of leaving producers spilling into a dead buffer forever. Boot paths now actually ABORT (exit 70) on a fatal mid-chain tear instead of falling back to legacy recovery, reason-DEL backpressure shares ONE bound per eviction/expiry sweep (was one 500ms bound per victim key), and every cap/shutdown drop path routes through the loss-accounting latch.

  • Cluster formation actually converges (pre-existing, both runtimes). A 3-node cluster could never complete its mesh: (1) CLUSTER MEET's random-id placeholder was never retired when the peer's handshake arrived under its real id, double-counting every met peer (known_nodes 5 in a 3-node cluster); (2) gossip sections were only consumed for PFAIL/FAIL reports, so nodes MEET-ed into a common peer never learned about EACH OTHER; (3) nodes adopted from rumors started with pong_recv_ms = 0, which check_failure_states skips — a rumored node that died before first direct contact could NEVER be marked PFAIL. Placeholders are now retired on handshake (same-address, different-id), healthy rumors are adopted (with self/known-address guards), and adoption stamps a freshness baseline so the staleness clock always runs. Review round: CLUSTER MEET is idempotent by ADDRESS (repeats no longer stack one placeholder per call) and refuses the node's own address; a cluster-bus bind failure now aborts startup loudly instead of leaving a node serving clients while invisible to every peer; --cluster-enabled refuses ports > 55535 (bus port would wrap past 65535); a cluster-ctl thread panic aborts the process via the same hook that guards shard threads.
  • The AOF now compacts itself (#433): Redis-parity automatic rewrite. --auto-aof-rewrite-percentage (default 100, 0 disables) and --auto-aof-rewrite-min-size (default 64mb, size strings accepted) trigger a background rewrite once the AOF has grown the given percentage over its size after the last rewrite. Before this, the AOF grew with write volume rather than dataset size — observed 4.8 GB on disk for a 2.43 GB dataset, ~1 GB/day — until the diskfull guard paused writes. A monitor thread samples the on-disk size once a second and dispatches the same entry point as BGREWRITEAOF; a failed dispatch backs off 60 s instead of hot-retrying. Both knobs appear in CONFIG GET.
  • Multi-shard BGREWRITEAOF is un-gated. The per-shard fan-out rewrite (cooperative snapshot + synchronized manifest commit) is now the default — the historical gate dated from a pre-C4 design that lost ~38% of keys, and the current path holds exact INCR recovery across a rewrite straddling a live write stream plus SIGKILL (crash matrix, 5/5 repeat runs). --experimental-per-shard-rewrite is deprecated (warns, no-op).

Fixed

  • Test-hygiene sweep: pid-only temp dirs + timing-race asserts. Eight spawn sites across six integration suites named their data dirs by pid only, so a crashed run's leftover dir was silently resurrected once the pid was reused — the server then reloaded stale persistence state and the suite failed on ghosts (the documented stale-reload trap). All eight now use tempfile (RAII where a single owner exists; unique keep() dirs where the restart flow shares one dir between two server handles). parked_idle_parity additionally gets deadline-polled asserts: CLIENT KILL's registry removal is polled instead of read once (the CI assert-too-soon race that fired 3/3 on a starved 2-vCPU runner), connects retry against a listening-but-backlogged server, and read deadlines widened 10s→30s (deadline-bound — green runs are unaffected). The client_tracking_invalidation multikey second-key push flake is product-side, not harness-side, and is now tracked as #448.
  • Central accept loop no longer head-of-line blocks on one wedged shard (#438 F3). Every central-listener delivery (tokio plain/TLS, monoio plain/TLS — the monoio sends were synchronous, stalling the whole listener thread) previously blocked on the routed shard's bounded conn channel; one shard wedged at its 4096-conn cap froze accepts for every other shard. Deliveries now try_send and rotate to the next shard with room; only when every shard's channel is full does the loop fall back to the blocking send (server-wide saturation, where back-pressure is correct).
  • Connection-migration fds can no longer leak (#438 F4). MigrateConnectionPayload carried the socket as a raw i32 with no drop semantics: a migration message still queued when a shard shut down — or drained but never spawned — leaked the fd and stranded the client on a connection no task would ever serve. The payload now owns the socket as an OwnedFd end to end (producer → SPSC ring → pending-migrations queue → target-shard spawn), so every undelivered path closes the socket and the client sees a FIN. Three unsafe from_raw_fd blocks became safe ownership conversions in the process. (Resumed-parked connections remain deliberately can_migrate:false; rationale documented at the spawn site.)
  • Migrated connections now carry the real requirepass (#438 F5, sec L3). The target-shard ConnectionContext was built with requirepass: None; the session's auth state was unaffected (it travels in MigratedConnectionState), but a later AUTH on a migrated connection wrongly answered "no password is set", and any future code deriving auth from the context would have failed open.
  • Unauthenticated connections never task-park (#438 F6, sec L2). On an auth-enabled server, a client could open sockets and never authenticate nor speak; each one downshifted and task-parked into the ~3.3 KB watcher state, letting an attacker hold a maxclients-worth of silent connections indefinitely at near-zero cost. Pre-AUTH connections now stay un-parked (full handler task — visible in monitoring, still reaped by timeout N); no-auth servers are unaffected.
  • Idle-park cancel provenance + registry counter hygiene (#438 conn-secondary). A bare ECANCELED (errno 125) from any non-sweep source was indistinguishable from the idle sweep's cancel and re-parked the connection — for a dead fd that is a permanent park→wake→park spin at 100% CPU. Park arms now require the sweep/drain to have marked the slot (was_swept_cancel, consume-once) before treating 125 as a park signal. Separately, a re-register over a still-present client id double-counted TOTAL_CLIENTS/shard gauges permanently (skewing shard_overloaded routing); register now balances against the replaced entry and the kept_registration miss-arm logs the invariant violation loudly. (The third secondary finding — timeout read once at connection setup — was already fixed by the D1 chore-sweep rework, which re-reads the config every second.)
  • Graceful shutdown now drains connections (#438 F1, conn#8). On SIGTERM/SIGINT each shard previously tore its runtime down the moment its event loop observed the cancellation, dropping every pending connection task mid-poll: in-flight replies were truncated and the blocking/subscriber shutdown arms (-ERR server shutting down) never ran — measured 19/50 BLPOP-blocked clients losing their shutdown reply on Linux io_uring (macOS passed only by scheduler luck). Shards now run a bounded drain before persistence teardown: parked stage-1/2 reads are woken through their cancellers (re-fired per tick to close the re-park race), token arms cover the tracking/subscriber/blocking parks, and the loop waits for the shard's live connection tasks to exit through the normal flush+FIN epilogue, up to a 5 s ceiling so a wedged peer cannot hold up shutdown. Writing the drain's red test surfaced a second, tokio-only leak in the same class: connections accepted by the CENTRAL listener were io-bound to the MAIN runtime's driver (tokio io resources bind at creation), so their reads/writes died with main's runtime no matter what the shard drained — with SO_REUSEPORT splitting accepts roughly evenly, about half of all tokio connections failed their final writes with "A Tokio 1.x context was found, but it is being shutdown". Forwarded streams are now re-registered with the owning shard's runtime driver at spawn (the monoio path already did this by forwarding std streams).
  • Parked-connection teardown can no longer close a reused fd out from under a kill scan (#438 F2, conn#9 / sec L1). kill_clients' fd-liveness invariant (registry deregister strictly before fd close) held in handler tasks by local-before-parameter drop order, but a task-parked connection's watcher co-owned guard and stream as future upvars, whose drop order is merely capture order — and the F1 drain makes dropping that future a routine path. Both are now wrapped in a ParkedSession whose hand-written Drop deregisters before closing, with the invariant restated at the kill site.
  • Remotely-triggerable shard-thread crash on pipelined batch tails (#438 follow-on). On any --shards >= 2 deployment, one pipelined write shaped [<remote-key cmd>…, SUBSCRIBE|BLPOP|…] aborted the whole server: the early-flush arm flushed the remote commands' placeholder replies, cleared the response vec, and the batch-end remote-reply drain then indexed into it out of bounds — shard-thread panic, process abort (both runtimes). Early-flush commands (blocking / SUBSCRIBE / PSUBSCRIBE / PSYNC) now defer themselves and the unconsumed batch tail to the next iteration when remote-slotted work is pending, so every prior reply resolves and flushes first.
  • Migration no longer discards MULTI / subscriptions latched mid-batch (#438 D4). The affinity sampler latched migration_target mid-batch, but migration executed at batch end with no re-check — a tail like […GETs, MULTI, SET] migrated with the transaction queued and MigratedConnectionState carries none of that state (queued txn discarded, EXEC answered -ERR EXEC without MULTI; a tail SUBSCRIBE was orphaned). ConnectionState::migration_eligible() (not in MULTI, no cross-store txn, no subscriptions, no CLIENT TRACKING, not a replica) is now evaluated at BOTH the latch and the batch-end execution point; an ineligible batch end keeps the latch and migrates at the first clean one.
  • Migrated connections no longer stall on their carried remainder. A resumed migrated handler received the source's unparsed bytes in its read buffer but awaited a fresh socket read before parsing them — a pipelined tail crossing a migration sat unanswered until the client happened to send more. The resumed handler now parses the carried remainder immediately.
  • The parsed batch tail after SUBSCRIBE/blocking is no longer swallowed. [SUBSCRIBE ch, PING] in one pipelined write silently dropped the PING (the frame iterator discarded the remainder on break); the tail is now re-encoded and carried into the next iteration (subscriber mode answers it in order).
  • INFO persistence reports real AOF state (#432). aof_enabled and aof_rewrite_in_progress were hardcoded 0 even with --appendonly yes (the default) and a rewrite running; they now reflect reality, and new aof_base_size / aof_current_size fields expose the growth the auto-rewrite trigger acts on — an operator can finally see the AOF-vs-dataset ratio the diskfull incident hid.
  • The per-shard BGREWRITEAOF crash matrix no longer reports phantom data loss on nearly-full hosts. The harness lacked --disk-free-min-pct 0 and parsed MOONERR diskfull INCR rejections as silently-dropped writes (the host root volume hovers at ~4% free, making it intermittent). The suite now disables the guard, panics on any non-numeric INCR reply, and is green 5/5 consecutive runs.
  • Replicas now apply streamed SWAPDB (#386), and the record reaches the wire exactly once per client call. Two stacked defects: (1) the replica's apply path had no SWAPDB intercept — generic dispatch hard-errors ("must be issued at the connection handler level") and the error was only logged, so every streamed SWAPDB silently no-op'd and the replica served pre-swap data for both databases until a full resync; (2) a multi-shard master emitted the record once per REMOTE shard leg and never for the coordinator's own leg — against today's single merged replica stream that means N−1 swaps, a net no-op whenever N−1 is even (e.g. --shards 3). The coordinator now emits the replication record exactly once, after the durability gate and the local swap (an aborted SWAPDB can never ship to replicas); remote SPSC legs keep their per-shard AOF/WAL writes but stay off the replication plane; the tokio single-shard handler emits it too. Replicas apply it with the same slice-split swap as WAL replay, skipping (with a warning) indexes outside their own --databases range instead of poisoning the stream. (c10k E1/E3), and cross-shard reply awaits are bounded (E4).** PUBLISH fan-out (immediate, batched, and EXEC-queued) and SCRIPT LOAD propagation used a single try_push — a transiently-full ring lost the message with no log or metric: subscribers on that shard silently missed the publish, or its script cache diverged (NOSCRIPT for a sha the server had just returned). All five sites now retry with the same bounded, shutdown-aware backpressure as the command dispatch path, and a final give-up is loud (moon_xshard_fanout_drop_total). Reply awaits on the slotted dispatch and publish paths — previously unbounded, so one wedged shard could hang a client task forever — now share the coordinator's 30 s bound: publish counts degrade (under-report, counted by moon_xshard_reply_timeout_total), while a slotted-dispatch timeout errors the batch and closes the connection, because the per-connection reply slot cannot be safely reused after an abandoned await. Rehydrated connection handlers (park wake / migration) also start with 512 B I/O buffers instead of 3×8 KiB, shrinking the memory spike of fleet-synchronized wakes; buffers grow back on first real traffic (c10k D3).

Added

  • INFO memory can now explain a used_memory-vs-RSS gap. Build with --features jemalloc-stats and INFO memory reports Redis's allocator_* fields: allocator_allocated, allocator_active, allocator_resident, allocator_retained, allocator_frag_bytes, allocator_frag_ratio, and allocator_unreturned_bytes. active - allocated is fragmentation, resident - active is dirty pages jemalloc holds but has not returned, and anything the OS charges beyond resident belongs to something other than the allocator. Motivated by a real instance reporting used_memory 2.43 GB while the OS charged it 7.3 GB (7.2 GB of that swapped) with no way to tell which. OFF by default — jemalloc's stats add bookkeeping to every allocation. Fields are absent rather than zero-filled when not built in, because a zero would read as "no fragmentation", which is worse than "not measured".
  • INFO clients reports parked_clients. How many connections are currently held only by a task-exit park watcher — real operational visibility into how much of the fleet is actually parked, and a direct signal for tests instead of inferring parking from process RSS.

Known issues

  • An idle moon may not return freed memory to the OS on Apple platforms, and there is no in-process fix. jemalloc's background_thread is compiled out when abi == macho (JEMALLOC_BACKGROUND_THREAD is only defined otherwise), so the background_thread:true moon bakes into its malloc conf is a silent no-op there and decay runs only as a side effect of allocator activity. The same 384 MiB churn-and-free reclaimed to a 3.7 MiB physical footprint on one Apple Silicon machine and retained all 386 MiB indefinitely on a GitHub macOS runner, so the behaviour varies by machine. A --memory-decay-interval-ms timer calling mallctl("arena.4096.decay") — the same call jemalloc's own background thread makes — was implemented and then removed after it failed to reclaim anything on the runner that reproduces the retention, across 30 seconds of driving it and three independent retries, while the ctl itself returned success. Production targets Linux, where the background thread is compiled in and the retention has not been observed; on macOS, build with --features jemalloc-stats and watch allocator_unreturned_bytes to tell whether a given host is affected.

Security

  • One byte made a connection invisible to the idle sweep (c10k D2). The stage-2 park predicate required an EMPTY read_buf, so a single * — the first character of every RESP array — kept it non-empty permanently, made the connection unparkable, and dropped it into the handler's UNREGISTERED plain read. Nothing on that path carries a sweep handle, so the connection held its full stage-2 working set for as long as the attacker left the socket open, with no authentication and no further traffic. One byte and one socket per connection; at 1M connections roughly 10-15 GB that no amount of idle time reclaims. Unparsed input no longer blocks a park: the remainder is carried in read_buf_remainder and re-parsed on resume, exactly as a migrating connection already did. Deliberately NOT capped — read_buf is already bounded by client_query_buffer_limit, and a second threshold would only move the attack past it (send 513 bytes instead of 1). A pending write_buf still blocks the park, because a reply the client is owed is carried nowhere and would be silently dropped.

  • Reply writes were unbounded — a client that stops reading held the whole reply forever (c10k C1). write_all on a socket whose receive window is closed never returns. A client could pipeline a large response and then simply stop reading, parking the handler inside that write while it held the ENTIRE serialized reply — the monoio handler coalesces a whole batch into one Bytes before its single write syscall, so a deep pipeline is hundreds of MB — plus its maxclients slot, for as long as the attacker cared to wait. N such clients is an OOM that costs the attacker nothing: no reads, no CPU, just a TCP window it refuses to open. New --client-write-timeout-ms (default 60000, 0 = the previous wait-forever behaviour) bounds every reply-carrying write in all three handlers — monoio top-level, sharded, and the tokio single-shard Framed path — closing the connection and dropping the reply when a write makes no progress for that long. 60s is far beyond any healthy client's stall and replication does not use this path (PSYNC hijacks the connection before it), but operators streaming very large replies over very slow links should raise it: the budget covers the whole write call, not each byte. Verified at 1 and 4 shards on both runtimes, with the 0 case as a differential control proving the mechanism rather than the harness (tests/write_timeout.rs). Unverified on Windows: the tests need the server's write to actually block, and two attempts to force that under Winsock failed (a 25 MB reply was absorbed by send/receive autotuning, and clamping the victim's SO_RCVBUF to 8 KiB did not change it), so they are skipped there rather than weakened until they pass. The code path is shared and compiles on Windows, but nothing proves the timeout fires; Windows is not a target platform (Linux and macOS are).

  • CLIENT LIST reported obl=0 oll=0 omem=0 unconditionally (c10k C1). The held-output counters were hardcoded, so an operator watching a slow-client output-buffer OOM in progress saw every client reporting zero bytes held — the attack above left no trace anywhere. obl/omem now report the reply bytes a connection has in an in-flight write, and tot-net-out counts reply bytes that actually reached the peer. oll stays 0: moon has one contiguous output buffer, not Redis's static-buffer + reply-list pair, so there is no list whose length it could report. Wired in the two handlers that own a client-registry entry (monoio top-level and sharded); the legacy tokio handler_single path bounds its writes but does not register, so it has nothing to report through.
  • No size ceiling on a reply (c10k C1). Adds --client-output-buffer-limit-normal, Redis's client-output-buffer-limit normal <hard>, and ships it at 256 MiB rather than Redis's unlimited default — that default is exactly why an unread socket can OOM Redis too. A reply exceeding the cap is refused and the connection closed instead of buffering it. Consequence, intended and worth knowing: the cap covers the whole serialized reply, so a single value larger than the cap is undeliverable, not merely a long pipeline. No pub/sub-class knob ships: moon's only subscriber write is already hard-capped at 64 KiB by MAX_COALESCE_BYTES, so such a knob could never fire, and the real pub/sub backpressure is the 4096-slot bounded channel (CONN_CHANNEL_CAPACITY) — which bounds message COUNT, not bytes, and is a genuine follow-up.
  • The write watchdog no longer arms on the hot path. A timer per batch flush would land once per command at pipeline depth 1, the path this project spent a milestone winning against Redis. A reply that fits in the socket buffer cannot block, so only writes ≥ 256 KiB arm the watchdog now (util::arm_write_timeout). Residual, stated rather than hidden: a small write to a genuinely wedged socket is still unbounded — it holds at most 256 KiB instead of hundreds of MB, but keeps its maxclients slot. It is now visible (omem > 0) and killable (CLIENT KILL shuts the fd down), which it was not before.
  • Privileged commands ran BEFORE the ACL permission check (c10k B1). Both connection handlers intercepted EVAL/EVALSHA/SCRIPT, ACL, and CLUSTER above the ACL gate, and each intercept continues on a match, so the gate below it never ran — even though the comment attached to that gate claimed it "must run before any command-specific handlers ... so that low-privilege users cannot reach admin commands". Any authenticated account could escalate to full admin: a user holding nothing but ~app:* +get was correctly refused a plain SET yet could run ACL SETUSER evil on nopass ~* +@all (persisted), CONFIG SET (persisted), arbitrary Lua via EVAL, SCRIPT LOAD, REPLICAOF and BGSAVE. The gate now sits above every privileged intercept in both handlers, with AUTH and HELLO lifted above it as Redis's NO_AUTH commands. Verified end-to-end at 1 and 4 shards (tests/acl_privileged_intercepts.rs).
  • Unknown ACL users failed OPEN (c10k B2). All three permission checks started with users.get(username)?, and None is their "allowed" answer, so a username missing from the table was granted everything. Nothing closes a live session when its account disappears, so ACL DELUSER alice / ACL LOAD silently PROMOTED every connection alice still had open to ~* +@all. Unknown users are now denied; default is guaranteed present after any load (ensure_default_user) so a server whose ACL file omits it is not bricked.
  • CLIENT KILL USER was inert and CLIENT LIST reported user=default for everyone (c10k B3). The client registry captured the username once, at accept time; AUTH/HELLO updated only the connection-local copy. The primary incident-response lever for a compromised credential matched nothing. AUTH and HELLO now publish the adopted identity to the registry, and — matching Redis — ACL DELUSER disconnects the sessions the deleted user still holds.
  • The experimental io_uring bridge served commands with no auth (c10k B4). MOON_URING=1 (tokio, Linux) binds a second SO_REUSEPORT listener on the server's own port whose accept path has no auth gate, no ACL check and no client registry — so on a server with requirepass/aclfile configured the kernel load-balanced roughly half of all new connections onto a listener that skipped authentication entirely. The documented limitation named maxclients and CLIENT LIST/KILL, not this. The bridge now refuses to arm (loudly) when authentication is configured and the shard stays on the tokio path.

Fixed

  • FLUSHALL/FLUSHDB inside MULTI/EXEC cleared one shard and reported success (c10k E2). The transactional executor runs the queued body against the LOCAL slice with no fan-out, so a queued flush emptied a single shard of N while EXEC still answered +OK. Measured at --shards 4: 45 of 64 keys survived a transaction that reported the database emptied — a silent wrong answer to a destructive command, typically noticed much later via a non-zero DBSIZE. The live (non-MULTI) path has broadcast since D-2; the transactional path now does the same. The executor records each flush and the ORIGINATOR fans it out — it cannot fan out itself, being synchronous while the broadcast awaits, and for a routed transaction it runs on the owner shard, where broadcasting from inside that shard's own message loop risks a shard-to-shard wait cycle. A failed leg replaces that entry of the EXEC result array with an explicit partial-flush error, so a +OK for a flush inside a transaction can be trusted exactly as on the live path. Both handlers and both transaction shapes are covered, and a queued SELECT before the flush is honoured. Cross-shard atomicity is unchanged and unchangeable: a concurrent reader can still see shard A flushed before shard B — MULTI bounds the report, not the visibility, in a shared-nothing engine.

  • TLS park veto samples wants_write() after processing, not before (c10k B5). Stream::task_park_safe read wants_write() before process_new_packets(), which can queue outbound bytes and still return Ok — reachable in rustls 0.23 via the TLS 1.2 renegotiation rebuff (NoRenegotiation warning alert) — so the connection could park owing the peer a reply. (The TLS 1.3 KeyUpdate variant is not reachable: rustls defers that reply until the next outbound record; tests/tls_park_keyupdate.rs pins the behaviour so a future rustls change is caught.)

  • A client that disconnects while blocked is now reaped (c10k hardening A1). The infinite-wait select! behind BLPOP/BRPOP/BLMOVE/ BZPOPMIN/BZPOPMAX/BLMPOP/BZMPOP/BRPOPLPUSH had exactly two arms, reply_rx and shutdown — nothing watched the socket. BLPOP key 0 followed by a disconnect therefore leaked the handler task, the WaitEntry, the client-registry entry and the maxclients slot forever: infinite waiters carry deadline: None so the deadline sweep skips them, and timeout exempts blocked clients by design (Redis parity). A few thousand throwaway connections wedged the server until restart, using one unauthenticated command. The wait now also watches the peer, and a vanished client tears down every registration it held (local and remote) before closing. CLIENT KILL of a blocked client works through the same path, since it closes the fd with shutdown(2). Bytes a client legally pipelines behind its blocking command are carried into the parse stream rather than consumed and dropped — they used to wait in the kernel, so the handler now skips exactly one read to drain them. Known gap: TLS connections keep the old behaviour; the vendored monoio-rustls read loops internally until rustls yields plaintext, so a post-readiness read can park inside a partial record, which the cancel-free design must never do.
  • A timed-out cross-shard BLPOP no longer eats the next push to that key (c10k hardening A2/A3). Two defects compounded into silent data loss at --shards > 1. First, the tokio single-key cleanup condition is_remote && !matches!(result, Frame::Null) || matches!(result, Frame::Error(_)) parses as (is_remote && !Null) || Error, so on a timeout — where the result is Frame::Null — the BlockCancel was never sent and the owning shard kept the WaitEntry forever (remote entries carry deadline: None, so the deadline sweep never reaps them). Second, when a later push found that ghost waiter, every wake path popped the element and then did let _ = reply_tx.send(...): the receiver was long gone, so the value was neither delivered, nor requeued, nor offered to the next FIFO waiter — it simply vanished. Wakes now skip waiters whose client has disconnected before touching the datastore, and restore anything popped if the receiver drops in the remaining race window (including BLMOVE's destination push and BLMPOP's multi-element pops, in order). The same boolean bug also sent a local waiter's shutdown cleanup through target_index(shard, shard), which underflows to a panic on shard 0.
  • Blocking registrations are no longer silently dropped or retried forever (c10k hardening A4/A5/A6). Cross-shard BlockRegister / BlockCancel messages were pushed with a bare try_push (tokio) or an uncapped loop { try_push; sleep(10µs) } with no shutdown arm (monoio). A dropped BlockRegister drops the reply sender with it, so BLPOP key 0 returned nil at t=0 — a protocol violation; a dropped BlockCancel leaked the owner-side waiter permanently; and the uncapped retry turned a saturated owner shard into a zombie connection that also blocked graceful shutdown. All four call sites now use the plane-wide bounded, shutdown-aware push, notify the owning shard on success, and fail the command loudly (MOONERR blocking registration failed) rather than blocking on a key nobody is watching.
  • timeout N no longer silently disables the c1M connection park. The idle timeout was enforced by a per-connection select! arm racing the read against sleep(timeout), and that arm sat FIRST in the read loop's if/else chain — so setting timeout made the stage-1 buffer downshift, the stage-2 park and task-exit parking all structurally unreachable. Every connection reverted from the parked footprint to its full working set, silently, in exactly the deployments that use the only slowloris knob moon ships. The same arm also read its config once at connection setup (so CONFIG SET timeout never reached a live connection) and had no exemption for replication links. Enforcement now runs from each shard's existing 1 s chore (client_registry::kill_idle_clients), which needs no per-connection timer and additionally reaches connections whose handler task has exited — matching how Redis enforces timeout from serverCron. The registry is striped by client id rather than by shard, so each shard sweeps a disjoint subset of stripes: total cost stays at one full pass per second regardless of shard count, instead of O(num_shards × total_clients). A parked connection closed by the sweep (or by CLIENT KILL) is now torn down by its watcher directly rather than rehydrating a full handler task purely to observe the kill and exit. Idle clients are closed within one sweep interval of the deadline rather than exactly at it. Blocked clients, subscribers and replication links are exempt, as in Redis.
  • CLIENT LIST's blocked flag was dead code. ClientFlags::blocked was hardcoded false at every call site, so a client parked in BLPOP key 0 never showed flags=b — and would have been treated as idle by the new timeout sweep. The flag is now set for the duration of a blocking command on both runtimes.
  • Warm-tier transitions leaked their staging directory on every failure (#435). transition_to_warm has ~10 fallible steps between creating .segment-{id}.staging and renaming it to its final name, and every early return in that stretch abandoned the whole directory. Found on a live instance: a 27 GB data directory against 2.53 GB of used_memory, 20 GB of which was 14,499 orphaned staging directories versus 174 MB in the 94 real segments — 99% of the vector store was abandoned scratch, written in a single three-minute window and already survived a restart. The dot prefix kept them out of ls and out of every vectors/* glob, so nothing surfaced them. A Drop guard now covers every early return, and sweep_orphan_staging runs at recovery for orphans no in-process guard can catch (a kill -9 between the manifest commit and the rename, or anything left by an older build). The sweep is safe by construction: staging paths are produced in exactly one place and read by nothing — every reader opens the final segment-{id} name. transition_to_warm also removes a stale staging directory for the same id before creating it, so a retry cannot inherit a previous attempt's partial files.

[0.8.4] — 2026-07-29

Added

  • TLS connections task-park too (c1M P1-TLS). Idle TLS connections now exit their handler task after --conn-park-secs, like plain TCP: the parked watcher awaits readability on the raw fd via a vendored io_ref() passthrough, and a vendored task_park_safe() check vetoes any park while the TLS stack holds state the fd cannot signal (wrapper buffer bytes, pending EOF/error status, decrypted plaintext, a received close_notify, or unsent output). Parked footprint keeps the rustls session; the task future and working set are freed.

Fixed

  • Read errors on a parked-capable connection terminate instead of re-parking (park→wake→park spin). The idle-park arms treated every read error as the sweep's cancel; with task-exit parking that spun a dead connection through park→wake→park at 100 % CPU forever (a TLS client vanishing with a raw FIN — no close_notify — surfaces as a read ERROR, and the dead fd stays readable; found live in E11 with 3 000 CLOSE_WAIT connections pinning a shard; an RST'd plain-TCP conn hits the same loop). The arms now match monoio's exact ECANCELED (raw os error 125, both drivers): only the sweep cancel downshifts/parks; real errors tear down promptly. Regression tests: FIN-while-parked (TLS) and RST-while-parked (plain).
  • Resumed parked connections keep their registry identity (c1M P1 follow-up). Waking from task-exit parking previously deregistered and re-registered the client across a task-scheduling boundary: a racing CLIENT LIST could briefly miss the connection, CLIENT KILL ID could return 0, and — worse — the fresh registration silently reset the connection's CLIENT SETNAME and age in CLIENT LIST. The wake now hands the held registration through to the resumed handler (registration handoff), so the entry — name, connected-at, kill state, client counters — persists unbroken across any number of park/wake cycles.

Changed

  • Migrated connections task-park too (c1M P1 follow-up). The migrated-connection spawn path (Linux connection migration) now routes ParkIdle like the primary accept path, so a connection that migrated shards and then went idle parks out of the task model instead of holding its full handler task forever.

[0.8.3] — 2026-07-29

Added

  • Task-exit parking for idle connections (c1M P1, --conn-park-secs, default 60 s). A plain-TCP monoio connection that stays idle past the W11 downshift now has its handler task exit entirely: only a tiny readiness watcher (boxed future holding the stream, ~100 B of session state, and the client-registry guard) remains, reclaiming the ~6 KB task state machine plus the remaining per-task buffers. The watcher wakes on read-readiness (readable(false), race-free on io_uring and epoll/kqueue) or server shutdown and rehydrates a fresh handler through the migration-restore path — wire-invisible across repeated park/wake cycles. Parked connections stay in CLIENT LIST, keep their maxclients slot, and CLIENT KILL still works (its shutdown(2) wakes the watcher; the resumed handler sees EOF). Exclusions: subscriber/tracking/timeout connections (structurally never reach the park arm), MULTI/EXEC or cross-store-txn sessions, partial frames, TLS (keeps W11+P4b buffer downshift), and the tokio runtime. --conn-park-secs 0 disables.

Fixed

  • c10k connection-plane wave (PR #421) — from the 2026-07-29 empirical review (tmp/C10K-REVIEW.md: 10k/25k live-connection ramps, idle-CPU measurement, symbolized flames, pipeline-ratchet A/B on the Linux VM):
  • Pipeline memory ratchet fixed (W1). One 1024-deep pipeline permanently ratcheted a connection from 56.5 KB to ~217 KB RSS until disconnect (measured: 5000 such conns pinned 1.09 GB). Batch scratch vecs now shrink after oversized batches, I/O buffer shrink floor lowered 64 KiB → 16 KiB, tokio bump arena capped, subscriber-mode per-message 8 KiB alloc+zero removed.
  • maxclients rejection is loud (W3). Over-cap connections now receive -ERR max number of clients reached (Redis parity) instead of a silent EOF on the monoio + TLS paths; the gate runs before per-connection context construction; startup now checks RLIMIT_NOFILE against maxclients, raises the soft limit when possible, and warns with the real client ceiling otherwise.
  • Blocked-client sweep is deadline-heap driven (W6) — O(due) instead of walking every blocked waiter at 100 Hz (~2M entry visits/s/shard at 10k BLPOP clients); block-forever waiters are never visited.
  • SPSC drain starvation fixed (W7). The cross-shard drain rotates its start consumer, so the 256-message budget no longer lets a hot low-index peer starve high-index peers indefinitely.
  • IP-affinity funnel load-gated (W8). A shard above 2× the mean connection count no longer receives affinity-funneled connections (falls back to round-robin) — previously one migrated/subscribed client pinned all same-IP connections to one shard with no load feedback.

Added

  • --uring-entries N (c10k P4) — per-shard io_uring submission-queue size (monoio default 1024; also the legacy driver's event-batch capacity). Values below monoio's 256 floor clamp up with a warning; monoio runtime only.

Changed

  • Idle connections downshift to a ~0.5 KB working set (c10k W11). A connection parked in read() ≥1s has its read cancelled by the shard's 1s chore (loss-free cancel-and-await on both io_uring and epoll/kqueue), sheds its 8 KiB rent buffer and empty scratch buffers, and re-parks on a 512 B probe buffer. Idle RSS/conn 43.5 → 19.9 KB (10k-conn VM A/B); GCE-pinned p=1/p=16 A/B perf-neutral. TLS conns and the tokio runtime keep the previous behavior.
  • Idle TLS connections shed their 32 KB wrapper buffers (c10k P4b). monoio-rustls + monoio-io-wrapper are now vendored (moon-patch style, like vendor/monoio): the wrapper's two eagerly-allocated, never-freed 16 KiB per-stream buffers are lazy + releasable, and the TLS stream implements a loss-free cancelable read, so W11's idle downshift now covers TLS connections too — a parked TLS conn drops moon's working set AND both wrapper buffers (they reallocate lazily on the next I/O). Cancel errors are deliberately not stashed for replay in the wrapper (a stashed ECANCELED would resurface on the next read and tear down a healthy connection). Wire parity across downshift/wake cycles is regression-tested on one continuous TLS session (tests/tls_idle_downshift_parity.rs).
  • Cold command futures boxed out of the connection state machine (c10k P3). TXN.ABORT rollback, leaked-txn disconnect teardown, and the FT.* body now Box::pin behind their name-check fast paths; the per-connection monoio task allocation drops 9.3 KB → 5.9 KB plain TCP (13.2 → 9.9 KB TLS) with zero allocation on the KV dispatch path.
  • Client registry striped 16 ways (W5). Accepts/closes no longer serialize on one global lock; CLIENT LIST locks one stripe at a time (output is no longer insertion-ordered — Redis makes no ordering guarantee); CLIENT KILL ID is an O(1) lookup instead of a full scan.
  • active_cross_txn boxed (W2) — ~2.2 KB of inline transaction SmallVecs no longer ride in every connection's task future; ConnectionState is guarded ≤768 B by a regression test.
  • Dead plumbing deleted (W4): the zero-registrant pending_wakers waker relay (per-conn Rc clone + twice-per-iteration sweep) and the never-adopted duplicate server/conn_state.rs module (209 lines).
  • The experimental MOON_URING=1 tokio bridge now logs a loud startup warning that its connections bypass maxclients and CLIENT LIST/KILL (T3; full accounting integration tracked in .planning/rfcs/c1m-connection-plane.md).

[0.8.2] — 2026-07-20

Fixed

  • Millisecond-TTL keys no longer expire up to 999 ms early (W3). CompactEntry stored expiry as seconds (ms / 1000 floor) even though the field was a full u64: a PEXPIRE key 1500 died at the 1000 ms boundary, and PTTL reported the truncated value. Expiry is now stored as absolute Unix milliseconds end-to-end (same 32-byte entry layout), PTTL reads back the exact deadline, and RDB/snapshot round-trips preserve millisecond fidelity (the round-trip test's old 5-second tolerance is now an exact equality).

Changed

  • storage::db is now a directory module. db.rs had grown to 3113 lines against the repo's 1500-line ceiling; the W5 accessor unification made a clean split possible. The file is now db/mod.rs (types, cost constants, ctor/clock, memory accounting, tests) plus db/hash_ttl.rs (HEXPIRE-family per-field TTL primitives), db/kv_ops.rs (core keyspace ops: get/set/remove, lazy expiry, cold-tier promotion, scan) and db/accessors.rs (ValueKind/OwnedKind generics, per-type delegators, read-only refs, blocking-hook helpers, streams). Pure code motion — every method body moved verbatim, public paths unchanged; four private helpers became pub(super) for cross-file visibility within the module.

  • aof_manifest::ShardManifest renamed AofShardManifest (W6). Two unrelated structs shared the name ShardManifest — the page manifest's (persistence::manifest, the durable spill/offload root) and the AOF manifest's per-shard entry — making cross-plane persistence code search and discussion ambiguous. The AOF one (private to its module tree, no external importers) now carries the plane in its name. storage::tier's ResidencyTier gained a doc note disambiguating it from the on-disk manifest::StorageTier (they intersect in concept but not in role and must not be merged); the unified-manifest RFC gained the task-#48 poison-record policy tie-in for its future section decoders.

  • One typed-accessor skeleton in Database (W5). The ~20-method accessor matrix (get_X / get_or_create_X / get_X_ref_if_alive per container type) copy-pasted the same skeleton — expiry check → cold promote/read-through → compact-encoding upgrade → variant projection — once per type. The skeleton now exists once per access shape (get_ref_if_alive::<K> / get_or_create::<K> / get_promoted::<K>); everything genuinely per-type lives in the new storage::db_kind module as ValueKind/OwnedKind marker impls (static dispatch — generated code identical to the hand-written originals). The public per-type methods remain as thin delegators, so command-layer call sites are unchanged. The four caller-less get_X_if_alive variants (compact encodings returned None; superseded by the _ref_if_alive family) are deleted. Pure refactor — behavior pinned by two new tests (hot-WRONGTYPE through the ref accessors; HashWithTtl → HashRef::WithTtl wiring incl. caller-now_ms propagation) plus the existing cold-visibility pins.

  • One eviction entry point (W4). The 13-name try_evict_if_needed* family (every combination of spill-sink / explicit-total / elastic-budget / plain-drop-reporting encoded as a function-name suffix) is replaced by a single evict_to_budget(db, config, EvictionRun) — the victim destination (Plain / SyncSpill / AsyncSpill) and the optional knobs are data on the EvictionRun options struct. The two former core loops (sync + async, five near-identical reclaim-loop copies in total) are now one loop with a per-iteration sink dispatch; the --appendonly durability-ordering branches are unchanged and documented on EvictionSink::AsyncSpill. Pure refactor — no behavior change; behavior pinned by three new entry-point tests plus the full existing eviction suite.

  • Fossil base_ts plumbing removed (W3). is_expired_at / expires_at_ms / set_expires_at_ms / new_string_with_expiry carried a base_ts: u32 parameter that has been ignored since expiry became absolute; the parameter, ~60 dead call-site arguments, dead let base_ts = … bindings, the never-read SnapshotState.base_timestamps field, the dead u32 element in the AOF-rewrite/BGSAVE snapshot tuple (AofFoldSnapshot.dbs and the save_from_snapshot/save_snapshot_to_bytes signatures), and the caller-less merge_shard_snapshots are gone.

Performance

  • Sync eviction spill now batches like the async path (W2). The tick-driven sync spill paths (run_eviction, the memory-pressure cascade's sync fallback) previously wrote one single-entry .mpf file and did one shard-thread-blocking durable manifest commit per victim key; they now route through the same flush_buffer batch writer the background SpillThread and the no-AOF durable path use — shared multi-entry files (up to 256 keys), one manifest commit per file. Evicting N keys under memory pressure costs ~N/256 manifest fsync round-trips instead of N.

Changed

  • One spill pipeline (W2). Victim policy dispatch (select_victim) and victim serialization (build_spill_payload) now exist once instead of being copy-pasted at all three spill entry points; the sync SpillContext path delegates to the one durable batch spiller (evict_batch_durable, formerly evict_batch_durable_no_aof — no longer no-AOF-only). The sync/async divergence class behind the task #34 plain-drop misclassification, the task #45 is_string gate, and the #139 test blindness is structurally gone. kv_spill::spill_to_datafile is no longer a production eviction path (kept as the single-entry primitive test fixtures use).

Fixed

  • Unserializable eviction victims are retained, fail-closed. All three spill sites previously did serialize_collection(..).unwrap_or_default(): a victim whose value failed to serialize (in-memory corruption) was "spilled" as an EMPTY value body and evicted — silent durable data loss on reload. Such victims are now retained hot and skipped, loudly logged.
  • Batch-durable evictions now count in the eviction metric. The durable batch path (--appendonly no cascade) removed keys without calling record_eviction(); single-victim paths did. Both now record.
  • One value codec for RDB snapshots and KV disk-offload spill (W1). The ~150-line value-body encoder/decoder that existed twice — rdb::write_entry's value section and kv_serde::serialize_collection — kept bit-compatible only by discipline, is now a single spec-documented module (src/storage/value_codec.rs) that both call sites delegate to; the RedisValueRef → ValueType mapping likewise exists once (the former inline copies in eviction.rs and kv_spill.rs delegate). Wire format is byte-identical (pinned by hand-built golden byte-spec tests + legacy pre-trailer decode tests); existing RDB files and spill pages load unchanged.

Fixed

  • Spill decode allocation-DoS hardening. The spill value decoder missed the RDB count-validation fix and fed corrupt length fields straight into Vec::with_capacity; the unified codec validates every count against remaining input before allocating, on both planes.
  • Corrupt sorted-set listpack scores are fail-closed. Both former encoder copies silently wrote 0.0 for a listpack zset score that failed to parse; the codec now refuses to encode (ValueCodecError::CorruptScore) — an RDB save fails loudly and a spill aborts instead of persisting a corrupted score.

Documentation

  • New "The Moon Journey" page (docs/journey.md) — the development story toward a production-efficient database, grounded in real evidence. Traces the milestone arc (v0.4 → v0.8.1) and the measured efficiency achievements — throughput (GET 5.11M/s 1.72×, SET 3.50M/s 1.92×, ARM64 2.20×; 1.71× at p=64 on 23× less CPU), the single-connection p=1 conquest, memory (27% less at 1KB values), latency (p50 8–10× lower), the durability write-path campaign (always p16 0.12×→0.91×, everysec p16 1.32×), AI-native vector/graph (12.7K QPS; beats Qdrant 1.6–2.3×, FalkorDB 23×), and the storage-kernel production hardening (46-cell kill-9 matrix, 10× RAM datasets, 24 h replication soak with zero acked-write loss). Includes a "what we measured and threw away" section (prefetch, lock-free Notify, THP-by-default, ef/√G beam split — all rejected on the evidence). Linked from the README (refreshed roadmap table + journey callout) and the mkdocs nav.
  • Visual performance timeline — a step-line chart (docs/assets/journey-efficiency.png) tracing the three efficiency curves that each crossed Redis parity (p=1 latency x86 1.06→1.66×, ARM 0.95→1.21×, fsync=always durable throughput 0.12→0.91×), plus a mermaid milestone timeline (v0.4.x → v0.8.1). Embedded in the journey page and the README.

[0.8.1] — 2026-07-19

Added

  • conf/moon-standalone.conf + a broadened --profile standalone: the best single-shard tuning as one flag or one file, now safe on any host. --profile standalone still fills --shards 1, --io-busy-poll-us 40, and --io-driver epoll, and now additionally drops the jemalloc arena cap to 2 on --features jemalloc builds — a single-shard instance has one hot allocator thread, so the baked-in 8 arenas are oversized and 2 lowers the RSS baseline with no contention cost. The arena cap is resolved in the pre-clap allocator re-spawn scan (the only path that reaches jemalloc before init), so it is CLI-only — --profile standalone on the command line folds it in; a conf-file profile standalone still sets the other three. The new annotated conf/moon-standalone.conf ships the same tuning (durability on) as an editable config file. Paired with the O3 contention governor below, the preset no longer requires pinned cores — the tuning and configuration guides drop the pinned-cores-only caveat throughout and the README/architecture docs are updated to match.

Performance

  • --io-busy-poll-us is now deploy-safe: a per-shard contention governor gates the spin on shared cores (O3). The p=1 busy-poll win (GCE pinned: ARM 0.95→1.21×, x86 1.06→1.66× vs Redis) previously inverted into a regression whenever the shard's core was shared (OrbStack/laptop), making the flag pinned-cores-only operator judgment. Each shard thread's 1s chore now samples its own nonvoluntary_ctxt_switches from /proc/thread-self/status — the kernel's direct "someone else needed this core" signal — and flips a per-thread gate in the vendored monoio legacy driver: one window over 25 preempts/s stops the spin immediately; five consecutive quiet windows re-enable it (asymmetric hysteresis, no flapping). Startup state is ungated, so pinned-core deployments see the win from the first request and a shared-core host pays at most ~one window of spin. Idle cost is unchanged (the existing 10ms idle-disengage already covers quiet threads). MOON_SPIN_ADAPTIVE=0 restores unconditional spinning (same-binary A/B knob); MOON_SPIN_MAX_PREEMPTS_PER_SEC tunes the threshold. Non-Linux has no preemption signal — the governor is inert there (pre-O3 behavior). This also makes --profile standalone (which presets busy-poll 40) safe on non-dedicated hosts.

Documentation

  • SPSC-wake Notify stays on flume — the lock-free replacement was rejected by measurement (O2). A fully-validated AtomicBool token + AtomicWaker implementation (loom lost-wake model, 209/209 consistency at 1/4/12 shards, monoio busy-poll smoke; branch perf/o2-notify-atomic-waker) measured NEUTRAL on GCE t2a shards=4 (SET p1 +0.08%, GET p1 +1.6% across 5 leg-order-alternating rounds): the cachegrind-attributed flume misses are dwarfed by the eventfd+epoll wake delivery, and busy-poll deployments elide flume via the skip-wake gate anyway. A NOTE at the Notify definition records the verdict and the re-attempt bar.
  • --memory-thp is permanently opt-in — the RSS-drift soak disqualified a default flip. 45min mixed-size-churn + idle-decay soak (moon-dev VM, THP leg vs control, AnonHugePages-verified): RSS is flat while hot, but once the workload quiets khugepaged collapses the heap into 2M pages and re-materializes every 4K hole jemalloc had purged — RSS climbed 699→890MB at constant used_memory 539MB, plateaued, and never returned (control: 680MB flat). Same-data RSS settles ~+31% over non-THP, eroding maxmemory eviction headroom. The classic Redis fork-CoW objection does not apply to moon (BGSAVE is an in-process per-shard epoch snapshot; no fork() in the tree). Flag/CLAUDE.md docs updated with the verdict and the enable-when guidance (uniform value sizes + RSS headroom); full data in tmp/THP-SOAK-RESULTS.md.

Fixed

  • Multi-shard SWAPDB is now durable before +OK (#133). Two defects in coordinate_swapdb: (1) no fsync rendezvous — under --appendfsync always the client saw +OK before any shard had fsynced its SWAPDB record; (2) the coordinator shard's own record was written to WAL v3 only, never the per-shard AOF — and for --shards ≥ 2 --appendonly yes the per-shard AOF manifest is the sole recovery authority, so a kill-9 after +OK permanently lost the LOCAL shard's half of the swap while remote shards' halves survived (cross-shard keyspace divergence; reproduced empirically). The local leg now performs a durable AOF group-commit append + fsync barrier BEFORE swapping AND before any remote dispatch (every abort point fires while the whole cluster is still unmutated — the pre-fix order left N−1 shards swapped on a local abort); the coordinator confirms each remote shard's durability with one post-ack fsync_barrier (H1-BARRIER ordering). A remote barrier failure after that shard's swap is reported truthfully as durability-unconfirmed rather than a false +OK. Replica-side SWAPDB application (streamed SWAPDB records currently no-op on replicas — a pre-existing gap on the remote legs too) is deliberately out of scope and tracked separately.

Performance

  • New benches/dashtable_probe.rs: DashTable point-lookup benchmark at cache-exceeding sizes (O1), and a recorded dead end for value-line prefetch. The existing 10K-key bench is cache-resident and cannot see probe miss latency (18.65%/11.09% of cycles on ARM/x86 per real-PMU GCE measurement); the new bench builds 100K/1M-key tables and looks up in fixed-seed shuffled order, with the hit path forcing a real load through the returned entry — black_box alone does not read through a reference (disassembly-verified), a harness trap that initially produced a phantom "3-6% x86 win" for prefetching values[slot] at the first H2 tag match. With the forced load, that prefetch measured ~1-2% SLOWER or noise on pinned x86 (c3) at 1M keys and a consistent regression on aarch64 (t2a prfm variant: get_hit/100K +14%) — so no prefetch ships; the negative result and methodology requirement are documented in segment/mod.rs next to the earlier aarch64 key-prefetch verdict. The probe path keeps its existing one-cache-line ctrl verdicts, directory-level segment prefetch, and x86 key prefetch.
  • --memory-thp opts jemalloc's value heap into transparent huge pages (O4). New opt-in flag layers thp:always onto the baked-in _rjem_malloc_conf (on top of the always-on metadata_thp:auto, which only huge-pages jemalloc's own metadata, not application allocations). Real-PMU measured on GCE vPMU (c4a Axion / c4 Emerald Rapids, tmp/GCE-PMU-RESULTS.md): GET RPS +24.4% (ARM) / +12.1% (x86), dTLB MPKI 7.2→4.7 (ARM, −35%) / 1.10→0.017 (x86, −98.4%), IPC +0.15/+0.09 — ARM pays roughly 10× the x86 dTLB tax at baseline (smaller dTLB, 4K pages) so THP is a disproportionately larger ARM win. Costs RSS +4.2% on both architectures. Implemented via the same _RJEM_MALLOC_CONF execve re-spawn as --memory-arenas-cap (PERF-10) — passing both flags together composes into exactly one re-spawn carrying one conf string (narenas:N,...,thp:always); an operator-set _RJEM_MALLOC_CONF still wins over either flag (warns, no-op). No-op with a warning on non-jemalloc builds, AND on non-Linux jemalloc builds (macOS) — THP is a Linux kernel feature, and jemalloc does not silently ignore an unsupported thp:always under the baked-in abort_conf:true; it aborts the process at init instead (verified experimentally on macOS 2026-07-18). --memory-arenas-cap still composes normally when paired with a downgraded --memory-thp on non-Linux hosts. VM-verified on OrbStack Linux (madvise THP policy): AnonHugePages jumps from ~5% of RSS (baked-in metadata_thp:auto only) to ~96% of RSS with the flag on a 500MB SET load, confirming the opt actually engages. Shipped opt-in only — an RSS-drift soak is still pending before any default flip. Both flags are CLI-only: a moon.conf value cannot reach jemalloc (its config is read at process start, before the conf file is parsed) and now triggers a loud startup warning instead of a silent no-op.
  • Auxiliary threads no longer contend with pinned shard cores (O5). manifest-sync-N, spill-N, and moon-wal-sync-N are spawned from inside the shard's own (pinned) event-loop thread and, on Linux, inherit that thread's exact single-core affinity mask at pthread_create time — they weren't merely sharing the shard's core, they were born wearing its mask and could never leave it (confirmed via ps -eLo psr: manifest-sync-N/spill-N co-located on shard N's core; the main accept/dispatch thread also floated onto a shard core, ~12% CPU steal observed under load). When --shards leaves a remainder of un-pinned cores, every auxiliary thread (main accept/dispatch, manifest-sync-N, spill-N, moon-wal-sync-N, the AOF writer(s), auto-save) now re-pins itself to that non-shard core set (round-robin) as the first act of its closure — except the main thread itself, which re-pins at the LAST moment before entering the accept loop, because children spawned from main inherit its mask and the vector pools must keep the full-machine mask. Adversarial review widened coverage to every OTHER pinned-parent spawn site in the tree: moon-heap-orphan-sweep-N (a ~40s crash-recovery I/O burst previously confined to its shard's own core), the lazily-spawned moon-cold-read-N pool (previously born on whichever ONE shard core served the first cold read and stuck there for the process lifetime, serving all shards), moon-vec-snapshot-N, moon-vec-idx-gc, and the moon-vec-compact-N pool (its old last-core-downward heuristic landed workers on shard cores whenever shards ≈ cores; it now prefers the non-shard set via aux_core_at, keeping the heuristic as fallback). cpuset caveat: core IDs come from the machine-wide online list, so under a restrictive cgroup cpuset the pin can fail — the thread then keeps its inherited mask (debug-logged), never worse than before. No-op when --shards >= cores (no remainder to place them on) or on non-Linux. MOON_NO_AUX_PIN=1 disables it. The vector segment-reload pool (moon-vec-reload-N) is deliberately left scheduler-placed — reloads are latency-critical and the cold-start/crash-recovery reload storm wants every core while shard threads are idle; confining it to the small non-shard remainder risks 3x worse cold-reload latency, the regression PR #237 fixed. jemalloc's background-reclaim threads are unaffected — jemalloc spawns them internally with no Rust-side affinity hook. A/B on GCE t2a-standard-8 (8 dedicated ARM vCPUs, shards=4, AOF on, 4 interleaved rounds vs MOON_NO_AUX_PIN=1): SET median +1.5%, GET median −1.9% (inside noise; GET run-to-run spread tightened 23%→7% under pinning), with ps -eLo psr placement proof — pinned legs show all 12 aux threads on the non-shard cores 4-7, unpinned legs show manifest-sync-N/spill-N sitting exactly on shard cores 0-3.
  • Shard offload paths are precomputed at shard init — the recurring tick paths no longer allocate (#45). The 100ms eviction tick, the memory-pressure cascade, the 10s warm-transition check, and the cold-orphan sweep each rebuilt <offload>/shard-{id} via format! + PathBuf::join (plus a PathBuf clone inside effective_disk_offload_dir()) on every firing, per shard — a CLAUDE.md hot-path-allocation violation. The event loop's existing per-shard disk_offload_dir (built once at init, Some iff disk-offload is enabled) is now threaded into run_eviction_tick and handle_memory_pressure as Option<&Path> and used directly by the warm/orphan tick arms. Cold event-gated builders (save trigger, reclamation-schedule persist, startup/recovery) are unchanged.
  • SCAN cold-plane pages now range-resume from the cursor — no full cold-index filter per page (#368). The in-RAM cold index (spilled keys under disk-offload) is now ordered by the same (hash48, key) SCAN order as the hot plane (BTreeMap keyed by the 48-bit cursor hash), so each SCAN page seeks to the cursor in O(log n) and takes the first COUNT live candidates instead of filtering every spilled key on every page. Completes the two-plane O(COUNT)-per-page walk on the hot-plane change below. Cold lookup/remove pay an O(log n) tree descent instead of an O(1) hash probe — those sit on disk-read-through, promotion, and sweep paths where the descent is noise next to the I/O they front, and the ordered map replaces (not duplicates) the hash map, so per-entry RAM stays comparable. Public ColdIndex API and recovery/rebuild behavior are unchanged.
  • SCAN hot-plane pages are now true O(COUNT) — no full-table walk per page (#368). The SCAN cursor hash is now the DashTable's own fixed-seed key hash truncated to its top 48 bits, which makes cursor ranges line up exactly with the extendible-hashing directory (indexed by top hash bits): segments are range-partitioned in hash space, so a page visits only the segments covering hashes at or after the cursor and stops as soon as COUNT entries are collected (DashTable::hash_page → Database::scan_hot_page). Per-page hot cost is now independent of keyspace size (previously every page walked and hashed all n live entries). Splits, merges, and directory doubling between pages remain safe by construction — the cursor is a position in hash space, and structural churn only changes which segment covers that position, never the set of keys at or above it. The cold plane keeps its filtered in-RAM index walk (bounded by spilled-key count; ordered cold-side paging remains a follow-up in #368). SCAN cursors from before this change are invalidated (cursors are documented as ephemeral; restart scans at 0).

Fixed

  • Cold recovery re-attaches spilled keys to the logical database they were evicted from (#139). Under disk-offload with SELECT N (N>0), the recovery rebuild used to merge every spilled key into db 0's cold index: SELECT>0 keys answered nil after restart (their only durable copy unreachable), and a same-named key could even serve the WRONG database's value from db 0. Spill files are now attributed wholesale via a db_index field in the manifest FileEntry (a byte-compatible repurpose of the always-written-as-zero min_key_hash slot — old manifests read as db 0, which matches their actual provenance), spill flush chunks never cross a db boundary so each file is single-db, and recovery attaches one rebuilt cold index per database before WAL/AOF tail replay (so replayed DEL/FLUSH still tombstone the right plane). A manifest referencing a db beyond the configured --databases count is skipped with a loud warning instead of silently resurrecting keys into a wrong database.
  • SCAN now honors the Redis stable-key guarantee under churn (#368). The cursor was a positional index into a keyspace snapshot re-collected and re-sorted on every page, so any insert/delete (or cold-plane spill/promotion/TTL churn) between pages shifted positions and could skip a key that existed for the entire scan — directly affecting the backup/migration-via-SCAN use case. SCAN pages now iterate in stable 48-bit key-hash order and the cursor is a position in hash space: a key's hash never changes, so churn cannot displace it, and a key present throughout the scan is returned exactly once. Cursors stay numeric (48-bit, fitting the multi-shard composite cursor's per-shard slot unchanged) and remain client-compatible. Page cost drops from collect + sort O(n log n) plus a second full-table lookup pass (and, on the write path, a full-keyspace lazy-expiry probe per page) to one walk with a bounded COUNT-min selection heap. COUNT stays a hint (Redis parity): a full page may defer a trailing equal-hash group to the next page. Follow-up tracked in #368: O(COUNT)-per-page via a DashTable bucket-order cursor.

Testing

  • crash_matrix_cross_plane flake diagnostics for the rare mid-scenario ConnectionReset (#365). The harness's RESP connection now records the last command (or pipeline summary) sent, and includes it — plus the peer address — in the panic when a read/write fails or the connection closes mid-frame, answering "WHICH phase reset". ServerGuard's pkill -9 -f <dir> backstop now logs pgrep -fl survivors before firing (the child is already reaped at that point, so any hit is a leaked respawn or a marker-substring collision about to be collateral-killed — the issue's unproven alternate hypothesis). No behavior change on green runs.

Performance

  • Adaptive idle park (#373 phase 2, monoio runtime). After 64 consecutive provably-no-op 1ms ticks, a shard's event loop stretches its periodic park to 10ms, cutting idle timer wakeups ~10×. Zero chore-cadence changes: every counter-based sub-timer (block-timeout 10ms, expiry/eviction 100ms, WAL-sync 1s, monitors 5s, warm/autovacuum/orphan-sweep) divides by the stretched period and entry is gated on an aligned counter, so each chore fires exactly when it did before — BLPOP timeout precision is unchanged. Eligibility is conservative (no commands counted on the shard thread since the last tick — local dispatch AND replica-applied commands both bump the counter; no cross-shard SPSC message pending at the drain, probed by queue occupancy; WAL buffer empty; no snapshot/checkpoint active; appendfsync != always; no CDC subscribers); any SPSC notify wake exits idle immediately, and the cached clock is refreshed before draining pending cross-shard messages on BOTH race outcomes, so cross-shard commands never observe stretched staleness. Documented trade-off: a command arriving mid-park on an existing LOCAL connection can read a cached clock up to 10ms stale (vs 1ms before) for lazy-expiry checks, only after ≥64ms of complete shard quiet; active expiry keeps its 100ms cadence. Replica-applied commands additionally now count toward total_commands_processed (Redis parity; previously uncounted). Escape hatch: MOON_IDLE_PARK=0 pins the fixed 1ms period (same-binary A/B knob). Tokio runtime unchanged.
  • Idle-tick CPU trimmed (#373 phase 1). Three per-tick costs on the shard event loop's 1ms/100ms cadences were removed without touching any cadence or park timing: (1) PageCache::resident_buffer_bytes — the #1 consumer in an idle-server perf profile (~17% of samples; it walked every frame buffer taking a parking_lot read lock each, every 100ms per shard) — is now an O(1) relaxed atomic load, maintained at the single buffer-grow site (buffers never shrink, so a monotonic counter is exactly equivalent); (2) CheckpointTrigger::should_checkpoint no longer calls Instant::now() every 1ms tick for its whole-seconds timeout — it reads the shard's cached clock; (3) CachedClock::update (the one designated clock read per tick) now makes one SystemTime::now() call instead of two, deriving seconds from milliseconds.

Documentation

  • Persistence guide and Valkey comparison no longer describe the removed WAL v2 (src/persistence/wal.rs); both now document WAL v3 (src/persistence/wal_v3/: segmented files, per-record LSN, lz4 FPI, off-loop fsync agent).

Fixed

  • Test-harness hygiene sweep: hardcoded target/release/moon paths and unguarded child spawns across tests/. 33 integration suites resolved the moon server binary via a bare ./target/release/moon default, a CARGO_MANIFEST_DIR-relative guess, or a local find_moon_binary() copy that never checked CARGO_BIN_EXE_moon — a stale binary of unknown provenance on a shared checkout could silently run instead of the one cargo actually built for the test; migrated to common::find_moon_binary() (5 suites with a documented "skip gracefully when unbuilt" contract keep their own Option<PathBuf> resolver, with the same CARGO_BIN_EXE_moon tier added ahead of the stale-path fallback). Separately, 6 suites (console_gateway_test, scan_fanout_multishard, allocator_mimalloc_smoke, perf_v0112_arenas_cap, replication_readonly_eval, replication_readonly_ws_mq) held a bare Child across many assert!/.expect() calls with cleanup only at the end of the function — a mid-test panic orphaned the server, the same shape behind issue #366's 667%-CPU incident — and now use the kill-on-drop MoonGuard pattern from tests/bgsave_startup_race.rs. Complex multi-restart harnesses (crash-recovery kill-9 cycles, Jepsen, instance-lock, SIGTERM) were left untouched by design.
  • tests/bgsave_startup_race.rs can no longer orphan its server on a mid-test panic — the spawned child is now held by a kill-on-drop guard (the pattern from tests/dir_deleted_degraded.rs). An orphan from this exact suite, its tmpdir later swept, was the 667%-CPU incident behind issue #366.
  • Data-dir deleted under a running server now latches a degraded state instead of error-looping the persistence tick (#366). When --dir vanishes (operator rm -rf, tmp-cleaner sweep, mount loss), the per-shard 1ms WAL/checkpoint tick used to retry its file operations on ENOENT forever — observed in the wild at ~667% CPU on an idle 4-shard server, with a wedged PING. The disk monitor's 5s poll now distinguishes statvfs ENOENT/ENOTDIR on the monitored path and latches dir_lost (one loud ERROR log): write commands are refused with MOONERR dirmissing (reads and PING keep serving, mirroring the diskfull guard), and the persistence tick skips WAL flush / checkpoint / WAL-overflow scan / auto-save entirely — zero per-tick syscalls and zero per-tick log lines while latched. The latch self-heals: the next poll that sees the directory again resumes writes (with an explicit warning that data accepted since the deletion is not guaranteed durable until restart). Only engages when persistence or disk-offload is configured, and works with --disk-free-min-pct 0. Two hardening layers close the shallow-heal hole (adversarial review): WAL segment rotation re-creates a missing parent directory (mkdir -p <dir> after an incident no longer leaves the nested wal-v3/ dir dead), and any tick-flush failure arms a 1s retry backoff so no flush error class can ever loop at the 1ms tick cadence again. The latch is unix-only by design (the non-unix statvfs stub never fails); the latch-behavior unit test is cfg(unix)-gated accordingly, and the e2e suite's latched-CPU check is relative to a pre-deletion baseline window so shared-runner load cannot flake it.

Performance

  • Disk-offload: shard event loop no longer pays manifest-commit fsyncs (task #59 levers 1+2). Under spill load, apply_completion_vec ran one DURABLE manifest commit (up to 2 fsyncs) per spilled file on the shard event-loop thread — strace-measured at 1.0–2.1s cumulative block time per 8s flood window per shard (single fsyncs up to 1.0s), stalling every connection on the shard. ShardManifest now splits into loop-owned RAM state + a per-shard manifest-sync-{id} thread owning the file handle: spill-completion commits are deferred, coalescing snapshot sends (correct because that path only runs under --appendonly yes, where AOF replay + the orphan sweep reconstruct anything a lost manifest commit recorded); commit() keeps its exact blocking durability for the checkpoint protocol and the no-AOF durable batch spill path. The crash-safety write order (overflow pages sync'd before the root that references them; dual-slot atomic root flip) is preserved verbatim. Additionally the spill writer now yields a 4ms quantum after each durable batch flush while a cold read is in flight (COLD_READS_INFLIGHT RAII signal from both the async pool and the synchronous MGET/MULTI/Lua read paths), never pacing when the request backlog is deep (pacing must not push eviction into -OOM) or at shutdown.

Added

  • Data-dir instance lock: a second moon on the same --dir is refused at startup. Two servers writing one directory means two writers on the same per-shard WAL/AOF/spill manifests — silent corruption with no error at either end (observed in dev as dozens of leaked instances accumulating against shared folders). Moon now takes an exclusive non-blocking flock(2) on <dir>/moon.lock before any listener binds or data file opens, held for the process lifetime; a competing start exits fast with the holder's pid. kill -9 releases the lock instantly (kernel advisory lock — no stale-pidfile failure mode; crash-restart is unaffected). A custom --disk-offload-dir is locked too when it differs from --dir. Unix only (Linux + macOS); Windows warns and proceeds. Operational note: concurrent instances on one host must now use distinct --dir values — which was already the only safe configuration.

Fixed

  • BGSAVE issued during shard startup could hang forever (pre-existing, reachable in v0.8.0 and earlier). The listener answers clients as soon as the fastest shard's event loop is up, but each shard seeded its snapshot epoch cursor from the trigger watch channel's current value at loop start — a BGSAVE (or auto-save) broadcast while a slower shard was still initializing was silently swallowed by that shard: its snapshot never ran, BGSAVE_SHARDS_REMAINING never reached zero, and rdb_bgsave_in_progress stuck at 1 with every later BGSAVE refused as already-in-progress until restart. Found as a ~15%/run crash-matrix flake under full-suite CPU contention; reproduced deterministically with a delayed shard start (new test-only MOON_TEST_SLOW_SHARD_START_MS fault injection) and pinned by tests/bgsave_startup_race.rs. The cursor now starts at 0 (epochs are per-process), so a trigger that arrives during startup is honored on the shard's first tick. The same test exposed a second pre-existing gap: the tokio shard loop never called bgsave_shard_done after finalizing its snapshot (only the monoio arm did), so under runtime-tokio every sharded BGSAVE left rdb_bgsave_in_progress stuck at 1 — now decremented in both arms.
  • DBSIZE / INFO # Keyspace count logical keys under disk-offload (issue #355). Both previously reported the resident set only — the 2026-07-16 G2 re-run wrote ~164K distinct keys and DBSIZE answered 24,275 (~86% under-report), breaking operator capacity math. Database:: logical_len() now counts hot + cold keys with hot∩cold overlap (a fresh SET over a cold-only key legitimately leaves its cold shadow until the next touch) counted once via an O(cold) probe pass — same order as the expires_count scan INFO already pays, zero hot-path cost, no counter-drift risk. Wired through all six sites: dbsize, dbsize_readonly, the INFO fallback keyspace section, the KeyspaceStats scatter handler, the embedded/non-sharded handler_single INFO keyspace vector, and coordinate_dbsize's local leg (which inlined a resident-only db.len() while its remote legs dispatched real DBSIZE commands — the two definitions disagreed inside a single reply). Known remaining parity gap, tracked separately: SCAN / KEYS / RANDOMKEY still enumerate the hot plane only.
  • Keyspace enumeration under disk-offload: SCAN/KEYS/RANDOMKEY now see spilled keys (#364). With disk-offload enabled, cold-only keys (spilled by eviction, no in-RAM entry) were readable via GET/EXISTS but invisible to enumeration — a 4-shard instance holding 400 logical keys returned only 116 from redis-cli --scan, silently losing spilled keys for any migration/backup consumer. SCAN, KEYS, and RANDOMKEY (both dispatch tracks) now enumerate the union of the hot plane and the in-RAM cold index, partitioned so a key present in both planes is returned exactly once and TTL-expired cold entries are skipped — pure in-RAM, no disk reads and no promotion. SCAN's TYPE filter judges cold keys from a new ColdLocation::value_type cache (same cached-copy contract as ttl_ms: populated at spill time, re-derived from the on-disk pages by ColdIndex::rebuild_from_manifest after restart; fits existing struct padding, so the cold index does not grow; no on-disk format change).
  • Storage: recovery panic double NeedsSplit after split_segment in DashTable::insert_or_update. The insert-or-update path split an overflowing segment exactly once and declared a second NeedsSplit unreachable — false under hash skew: when the overflowing segment's keys share the next directory bit, the split routes all of them into the same child (still over LOAD_THRESHOLD) and the retry panicked. Deterministically reproduced in production loading a 219k-key shard checkpoint (shard-0.rrdshard) on 0.7.1 recovery — the server crash-looped until the checkpoint was quarantined (code identical back to v0.6.0). Now split-retries in a loop, mirroring insert's recursion; regression test brute-forces 56 keys sharing the top 12 xxh64 bits to force the double split.
  • Test harness: five latently broken integration suites unmasked by the task #59 gate battery. cmd_flush_dbsize_debug_memory still passed the --persistence-dir flag removed in June (server exited 2 before accepting; CI stayed green only because the suite self-skips without a release binary); it and memory_doctor_response ignored MOON_BIN via hardcoded target/release/moon finders (stale host Mach-O via OrbStack proxy → 30s spawn timeouts). mem_watchdog, oom_bypass_closure, and spsc_two_db now pass --disk-free-min-pct 0 so a near-full dev volume's diskfull write pause can't shadow the memory-guard/routing behavior they assert.

[0.8.0] — 2026-07-16

Documentation

  • G2 10×-RAM re-run report (v0.8 close-out evidence). New docs/perf/2026-07-16-g2-10x-ram-rerun.md: re-ran the G2 acceptance benchmark on main @ 4dcfd533 with the full v0.8 storage batch merged. Headlines vs the 2026-07-13 baseline: spill files 236K → 840 (~280×, PR #350 confirmed at scale); used_memory truthful at 1.00× cap steady-state with a ≤5 s post-restart drain to under-cap (task #56, PR #349; demote pass logged 83,788 shadows); cold-GET-during-spill worst tail 1,910 ms → 205 ms (task #59 still open for sub-10 ms); restart now AOF-replay-bound (16.9 s at 3.3 GB unrewritten incr AOF — spill-file count is out of the boot path entirely); 500/500 kill-9 integrity. PRODUCTION-CONTRACT rows CRASH-02 (37→46 cells + scheduled CI) and MEM-10X-01 refreshed with this evidence + audit-trail row. New finding filed as #355: DBSIZE counts only resident keys under disk-offload (~24K reported vs ~164K logical).

Fixed

  • Vector auto-merge CPU livelock: rejected GraphUnion merges now back off exponentially instead of retrying every autovacuum tick. A background merge rejected by the recall gate (or the 512 MiB memory ceiling) keeps its source segments, so every needs_merge trigger condition — segment count

    16, dead fraction > 20% — remains true, and the next tick resubmitted the identical doomed merge. Each recall-gate rejection discards a fully built union HNSW graph (plus the AE-1 self-probe), so an index stuck below tolerance pinned the moon-vec-compact-* worker pool at ~100%/core indefinitely — observed in production at ~300% CPU sustained on an embeddings workload, with no log signal (the rejection was debug!-level). Fixes, all in src/vector/store.rs:

  • poll_install_merge's rejection path now records a MergeBackoff (xxh64 fingerprint of the source-segment Arc identities, consecutive failure count, exponential window 60s→1h) and logs at warn! with the failure count and next-retry delay — naturally rate-limited to once per window.
  • begin_background_merge_due skips dispatch while the same segment set is inside its window; any set change (new compaction, warm-tier transition, installed merge) or a successful install clears the backoff immediately. Manual FT.COMPACT (force-merge) is unaffected, and FT.CONFIG SET <idx> MERGE_RECALL_TOLERANCE clears the backoff so operator intervention retries at once.
  • needs_merge's segment-count trigger now honors the merge memory ceiling (previously only the dead-fraction trigger did): an index whose live vectors exceed 512 MiB was re-dispatching a merge that merge_immutable deterministically refuses — a guaranteed livelock for indexes past the ceiling. The store-level needs_merge now delegates to the index-level check instead of duplicating (and diverging from) it. New unit tests: bg_merge_gate_failure_backs_off (RED against the old behavior), bg_merge_backoff_clears_on_segment_set_change, bg_merge_backoff_clears_on_tolerance_change, bg_merge_backoff_schedule.

Changed

  • CI: heavy vector recall benchmarks moved out of the per-PR gate (~6-24 min tail cut). Duration data from a green main run (2026-07-16) showed the recall family was ~half the whole suite's CPU, with the top case (test_f32_recall_10k_128d) at 182-480s per attempt at the unoptimized test profile — and the two biggest cases were the only tests that flaked (480s nextest timeout × 3 retries = 24 wasted minutes) when the runner VM shared a loaded host. Five recall-quality canaries (test_f32_recall_10k_128d, test_f32_recall_1k_768d, recall_10k_128d_ef128, recall_1k_768d_ef128, test_parallel_recall_parity_with_sequential) are now #[ignore]d, matching the pre-existing convention for the 768d/10K cases; a fast smoke variant (1k/128d, both unit and integration) still runs per-PR, and 27 recall-adjacent tests remain in the PR gate. A new nightly recall-canaries job (.github/workflows/crash-matrix.yml) runs the FULL family — ignored and not — at release profile via --run-ignored all -E 'test(/recall/)', limited to the two binaries that contain them. nextest profile.ci override: test(/recall/) gets retries = 0 (CPU-bound + seed-deterministic — retrying a timeout or a threshold failure only multiplies the waste) and a 120s×5 slow-timeout. Dependabot: new cargo-minor-patch catch-all group collapses the Monday minor/patch flood into one PR — six near-simultaneous dependabot CI runs starved the single self-hosted runner for >20 min on 2026-07-15.

Performance

  • Disk-offload spill files now batch effectively — file count scales as ~keys/batch, not ~keys (v0.8 item 3, task #57 follow-up). The G2 acceptance run (260K × 10KB keys, 256MB cap) produced ~236K heap-*.mpf files despite the format nominally supporting ≤256 keys/file batching. Root cause: SpillThread::flush_buffer (src/storage/tiered/spill_thread.rs) already coalesced up to 256 requests per flush, but partitioned them by value size before writing — every entry over INLINE_MAX_VALUE_BYTES (3500B, e.g. all of G2's 10KB values) was routed to its own dedicated single-entry file via spill_single_entry, defeating the batch entirely. build_kv_spill_batch / write_kv_spill_batch (src/storage/tiered/kv_spill.rs) are rewritten to pack inline AND oversized entries into ONE shared batch file per flush: each oversized entry gets a dedicated leaf (holding an overflow pointer) plus its overflow-page chain placed inline in the same page stream, with the atomic temp-file+sync_all+rename+fsync_directory write sequence unchanged (payload-before-reference durability ordering preserved — a key is only considered cold after its containing segment is durable). Empirical measurement (macOS, matching G2's shape): RED (pre-fix) 3742 files / ~34,908 evicted 10KB keys; GREEN (post-fix) 29 files for the same eviction load — a ~129× reduction, spill_batches_flushed now matches file count 1:1 as designed. Fixing this also uncovered and repaired a latent bug in build_overflow_chain (src/persistence/kv_page.rs): its prev/next link computation silently assumed every chain's start_page_id == 1 (true of every caller before this change, since a single-entry file always put its one leaf at page 0) — a chain-local i+1/i+2 formula that produced dangling/wrong links the moment a caller (this new batching code) passed a variable start_page_id > 1 for the second, third, etc. oversized entry sharing a file. Links are now computed as file-absolute (start_page_id + i), which is also what read_overflow_chain already expected. New unit tests cover oversized-entries-share-one-file, mixed inline/oversized batches, and the recovery-scan/builder agreement for oversized batches (kv_spill.rs::tests); a new end-to-end integration test, tests/crash_recovery_spill_batch_kill9.rs, drives a real server through a sustained eviction+spill burst of 8000-byte values and SIGKILLs it with no settle window, asserting zero acknowledged-write loss and a batched (not ~1-file-per-key) heap-file count survive the crash and restart under both the monoio and tokio runtimes.
  • Adversarial-review follow-ups on the above (same task). Three gaps closed on perf/spill-segment-batching before merge:
  • Genuine post-restart cold-read coverage. The kill-9 test above only proves AOF-replay-derived recovery — every acked key ends up hot via DispatchReplayEngine, never touching cold_read_through, so it cannot tell "the leaf-offset + overflow-chain-with-start_page_id>1 read path works" apart from "AOF replay papered over a broken cold layout". A new spill_batch_shared_file_survives_cold_read_after_kill9 test pre-seeds a 3-entry batch file (1 inline + 2 oversized, ordered so the second oversized entry's chain starts at a file-absolute page_idx > 1) directly via build_kv_spill_batch/write_kv_spill_batch + ShardManifest, boots a real --appendonly no server on top of it, kills it without ever touching the keys, restarts, and GETs all three — the only way any of these bytes can reach the client is a genuine ColdIndex::rebuild_from_manifest + cold_read_through read. Getting this test to actually exercise that path (rather than a hot-RAM hit) required also pulling in src/persistence/recovery.rs's fix from the concurrent task #56 effort (fix/t56-used-memory-offload, commit 689c52e0): v3 recovery Phase 3 used to re-insert every previously- spilled String entry directly into the hot DashTable in ADDITION to rebuilding the ColdIndex stub for the same key — for an entry_flags::OVERFLOW-stubbed entry this hot-preload loop ignored the flag and inserted the raw 4-byte page-pointer as if it were the literal value, silently corrupting every oversized cold key on every restart (independent of this batching fix — it affected the pre-existing single-entry-per-file layout too). That recovery.rs hunk is a clean, self-contained cherry-pick (verified: identical parent-state diff, no dependency on task #56's other, out-of-scope-here, used_memory ledger/RSS/AOF-replay-demotion changes) and is a hard prerequisite for this test's premise, not something this task's rewrite could route around. A companion library-level unit test, test_rebuild_from_manifest_mixed_inline_and_oversized_roundtrip (kv_spill.rs), independently proves the exact same read path correct via direct calls to the production functions, without depending on server boot sequencing.
  • Bounded batch-build memory (BATCH_BYTES_CAP, spill_thread.rs). build_kv_spill_batch materialized an entire flush's pages in RAM before a single byte reached disk — fine for the 256-small-entries case the old comment described (~50 KB), but the whole point of this fix is that oversized entries now share that same unconditional batch, so 256 consistently-large (e.g. near-maxmemory-sized) values could resident hundreds of MB at once, exactly when spills fire under memory pressure. flush_buffer now splits each FLUSH_ENTRY_CAP-sized buffer into sub-batches capped at 4 MiB of cumulative value_bytes (chosen against the OLD per-entry path's peak — one value's pages at a time, no cross-entry amplification — while leaving the G2 acceptance shape, 256 × ~10 KB ≈ 2.5 MiB, as a single file, unchanged), each becoming its own file via that sub-batch's own already-pre-assigned file_id (no new allocation needed — see SpillRequest::file_id's doc on harmless id gaps). New unit test: oversized_flush_splits_into_byte_capped_sub_batches.
  • Hygiene. Fixed two stale comments caught in review: FLUSH_ENTRY_CAP no longer claims to bound in-RAM size by itself (superseded by BATCH_BYTES_CAP above), and entry_flags::OVERFLOW's doc corrected from a fictitious 12-byte file_id:u64 + page_id:u32 layout to the actual 4-byte file-absolute start_page_idx (no file_id is stored — the chain always lives in the same physical file as its leaf stub). Added a one-line comment at write_kv_spill_batch explaining why it hand-rolls temp+fsync+rename instead of calling persistence::atomic::atomic_write_durable: that helper takes one pre-materialized &[u8], and concatenating batch.pages into a single buffer first would undo BATCH_BYTES_CAP's whole point — streaming each already-built page into the temp file keeps this write's own working set at one page at a time.

Added

  • Crash-matrix CI (v0.8 exit-criterion item 4). New .github/workflows/crash-matrix.yml wires tests/crash_matrix_cross_plane.rs (the 46-cell kernel M3/G1 kill-9 durability suite — KV/graph/vector/WS/MQ/cross-store-TXN across appendonly x disk-offload x shards) into scheduled CI on the self-hosted moon-dev runner: a nightly job (03:17 UTC) runs the full matrix at the default single iteration, and a weekly soak job (Saturday 04:41 UTC) runs MOON_CRASH_MATRIX_ITERS=20 on the 10 cells with a probabilistic kill point (mid_checkpoint, graph_drop_survives_repeated_checkpoints, txn_isolated_*) — the cells kernel M3 stage 2 (task #53) proved need repeated sampling, not a single pass, to catch a timing-window regression. Both jobs support workflow_dispatch for on-demand runs. Not wired into the per-PR gate: the self-hosted runner is a single machine and PR-time is already tight (same rationale as integration-tests.yml's ci-full label gate).

Confirmation run on the moon-dev Linux VM (fresh ELF release binary, main @ ec084556, MOON_BIN pinned, --test-threads=1): the full 46-cell matrix passed 2/2 consecutive runs (46 passed; 0 failed, ~79s each), including all former RED cells — the 6 legacy-mode graph cells (task #60 / PR #322) and the MQ generic-DEL resurrection cell (task #46 / PR #301) all now run ungated (no harness::red_guard call sites remain in the suite). A MOON_CRASH_MATRIX_ITERS=5 soak of the 10 probabilistic-kill-point cells also passed 5/5 iterations clean (10 passed; 0 failed, 40.3s) — these are the same cells that, pre-fix, failed at iteration 7/20 and 11/20 in the kernel M3 stage 2 investigation.

Fixed

  • used_memory truthfulness under disk-offload (task #56). The G2 acceptance run (4 shards, --maxmemory 256MB, 2.6GB dataset, disk-offload enabled) reported INFO used_memory at 406-762MB against the 256MB cap, and worse after a kill-9 restart. Three independent bugs, found by instrumenting a 2-shard/8MB-cap/40k-key macOS repro (tests/used_memory_offload_truthful.rs) at load, steady state, and post-restart:
  • used_memory was literally process RSS (src/command/connection.rs), not the logical ledger --maxmemory eviction actually gates on. RSS always carries allocator arena overhead, mmap'd cold-read page-cache frames, thread stacks, the binary image, the (intentionally unbounded) Lua script cache, and the replication backlog ring — none of which the eviction system charges against the cap, so used_memory was guaranteed to read high under any disk-offload workload even when eviction was correctly holding the real ledger under budget. INFO's used_memory now reports the same KV(+ColdIndex)+vector+text+graph "used-term" ShardDatabases::recompute_elastic_budget already gates on (admin::metrics_setup::logical_used_memory_bytes); used_memory_rss / used_memory_peak still report the true OS footprint alongside it, a new moon_used_memory_bytes Prometheus gauge mirrors the ledger figure, and MEMORY DOCTOR gained an explicit "gated vs RSS vs outside-the-cap" header so the two numbers are never conflated again.
  • Restart double-loaded every spilled String key into hot RAM (src/persistence/recovery.rs): v3 recovery Phase 3 re-inserted each previously-spilled ValueType::String entry directly into the hot DashTable (#79-04, predating cold read-through) AND — two days later, unrelated to the first change — rebuilt a ColdIndex stub for the same key pointing at the same on-disk location (#80-02); nobody removed the first loop when the second, correct mechanism landed. Every value type now recovers as a cold-index-only stub, exactly like Hash/List/Set/ZSet/Stream already did; the first GET lazily promotes a key into hot RAM (and only then charges used_memory) via the same Database::promote_cold_if_present path ordinary (non-restart) cold read-through uses.
  • AOF replay re-hydrated already-cold keys back into hot RAM on every restart (src/main.rs, src/storage/db.rs): fixing the bug above exposed a second, independent restart-time re-charge path that only surfaces when --appendonly yes and disk-offload are both active. An evicted-and-spilled key gets no AOF DEL record — a spilled entry stays cold-readable, not deleted, so record_reason_del never fires for it — so the AOF's incr log still holds that key's original SET. AOF replay (main.rs's per-shard and single-shard multi-part branches) has no cold-tier awareness and blindly reapplies every historical SET into the hot DashTable, so a restart transiently re-inflated used_memory to ~6x the steady-state figure (self-healing over several eviction ticks, but proportionally slower to catch up the larger the disk-offloaded dataset — the "got WORSE after restart" symptom on the real 2.6GB G2 dataset). Database::demote_replayed_cold_shadows reconciles the two right after AOF replay finishes and before the server accepts connections: any key still present in the (already crash-consistent) ColdIndex at that point is provably redundant with whatever AOF replay just wrote hot for it (replay cannot re-cold a key, so the last AOF command touching a crash-time-cold key must be exactly the SET the eviction path spilled), so the redundant hot copy is dropped and used_memory no longer spikes at all post-restart. tests/used_memory_offload_truthful.rs drives all three mechanisms end-to-end (spill past --maxmemory, confirm real disk spill, SIGKILL, restart) and asserts used_memory stays within 1.75x --maxmemory at steady state AND immediately post-restart. See docs/guides/monitoring.md for the updated operator story on what counts against --maxmemory and what doesn't.
  • Adversarial-review follow-up on task #56 (three findings, all fixed on the same branch):
  • CRITICAL — stale cold-shadow resurrection on live overwrite. The demote_replayed_cold_shadows reconciliation above assumed any key still present in the recovered ColdIndex after AOF replay must be crash-time-cold and untouched. False whenever a key is spilled at v1 and then live-overwritten to v2 with no further eviction before a crash: the manifest still shows it spilled at v1, so a restart rebuilds ColdIndex[key] = v1, AOF replay reconstructs hot[key] = v2, and the (pre-fix) demote pass wrongly dropped the hot v2, resurrecting stale v1 on the next GET. Database::set (src/storage/db.rs) now clears the key's ColdIndex shadow the instant a SECOND write to that key is observed (InsertOrUpdate::Updated, not Inserted) — provably safe because AOF replay always starts from an empty DashTable (db.clear() runs before every replay branch), so the first replayed write to a key is always Inserted (left alone; it may be exactly the write that later got spilled) and any subsequent write to the same key during that same replay can only happen if the AOF recorded a write after the one that got spilled — proving the cold copy stale. No new record type or spill-file format change needed. Covered by the new unit test storage::db::tests::test_second_write_invalidates_cold_shadow and the new end-to-end kill-9 test tests/cold_shadow_overwrite_resurrection.rs.
  • HIGH — demotion never wired into the no-manifest (tokio + --shards 1) recovery branch. demote_replayed_cold_shadows was only invoked from the manifest-based AOF replay branches in src/main.rs; under runtime-tokio + --shards 1 no manifest is ever created (AofManifest::initialize is monoio-only there), so the only KV replay for that runtime/shard combination — the Phase 4b appendonly.aof fallback inside recover_shard_v3_pitr (src/persistence/recovery.rs) — never demoted anything, leaving that config just as exposed to the AOF-replay-rehydration bug as the manifest-based paths were before task #56. The demote call is now also invoked at the end of Phase 4b. Covered by the new tokio+jemalloc-feature integration test tests/cold_shadow_single_shard_tokio.rs (single_shard_overwritten_cold_key_returns_new_value_after_crash), using SETEX rather than a bare SET per the documented monoio-write-gate/inline-SET-path gotcha.
  • MEDIUM — used_memory parity gap: Lua script cache and replication backlog. Real Redis's used_memory is "total allocator-attributed memory", not "memory eviction can reclaim" — it counts the Lua script cache and replication backlog even though neither is evictable. Decision: include both in the reported figure rather than document an exclusion list, since both were already tracked as separate moon_memory_bytes{kind="lua_scripts"} / {kind="replication_backlog"} gauges — folding them into logical_used_memory_bytes() (src/admin/metrics_setup.rs) and the moon_used_memory_bytes gauge cost no new instrumentation. The actual --maxmemory eviction gate (ShardDatabases::recompute_elastic_budget) is deliberately left unchanged and narrower — eviction has no mechanism to reclaim Lua bytecode or backlog bytes, so gating on them would free nothing. MEMORY DOCTOR (src/command/server_admin.rs) now prints three distinct figures (elastic budget / used_memory reported / RSS) instead of conflating the first two; see the expanded docs/guides/monitoring.md "used_memory vs RSS under disk-offload" section.

Documentation

  • Roadmap Rev 2 (task #68, doc half). docs/roadmap/ROADMAP.md updated to post-v0.7.1 actuals: v0.6.1/v0.7.0 marked shipped (with the R5/R6 slips and the single-shard-replica disclosure recorded), v0.8 re-slotted to One Storage Kernel GA (close-out of tasks #49/#56, spill-file batching, crash-matrix CI, 10×-RAM benchmark publication), cluster hardening + multi-shard replicas moved to v0.9, enterprise foundation to v0.10; debt register refreshed.
  • Task #49 v0.8 close-out audit: no code change needed, roadmap corrected instead. Re-verified every bare-write persistence site named in the kernel review (ACL SAVE, cluster nodes.conf, CONFIG REWRITE, replication state, native BGSAVE/RDB, clog, kv_page) against current HEAD: all 7 already route through atomic_write_durable (temp file → sync_all → rename → dir-fsync), shipped in PR #304 (merged 2026-07-13) and released as part of v0.7.0 — git log 4e0688e6..HEAD on every touched file confirms none of the 9 converted call sites (ACL SAVE has two: acl::io::acl_save + the command handler routing through it; native BGSAVE has three: rdb::save, rdb::save_from_snapshot, redis_rdb::save) were reverted or bypassed since. The Rev 2 roadmap pass (task #68, above) had re-listed task #49 as an open v0.8 gap without checking it against the already-shipped v0.7.0 CHANGELOG entry; docs/roadmap/ROADMAP.md §1 gap table, §4 v0.8.0 item 1, and §5 debt register are corrected to reflect this (struck through / removed, not deleted from history). Also audited adjacent hand-rolled writers for regression risk: storage/tiered/warm_tier.rs's staging-dir → final-dir rename (own documented atomicity protocol, per-file .mpf writes inside the staging dir are covered by the directory-level rename, not a bare-write gap) and storage/tiered/kv_spill.rs::write_kv_spill_batch (already hand-rolled tmp+fsync+rename+dir-fsync correctly; not one of the 7 named sites) both remain correctly out of scope — no change made.

Changed

  • TLS: migrated off the unmaintained rustls-pemfile onto rustls-pki-types's PemObject trait (task #66). build_tls_config (src/tls.rs) now parses certificate chains via CertificateDer::pem_reader_iter and private keys via PrivateKeyDer::from_pem_reader — both provided directly by rustls-pki-types (already in the dependency graph as rustls's own pki_types re-export), so no new supply-chain surface is added. Behavior is unchanged: same fail-loud io::Error wrapping per stage (TLS cert file / TLS cert parse / TLS key file / TLS key parse / CA cert parse), same SIGHUP hot-reload path. Two new unit tests (test_build_tls_config_garbage_cert_parse_error, test_build_tls_config_garbage_key_parse_error) exercise the parser seam directly with corrupt-but-well-formed PEM bodies — the existing suite only covered missing-file paths, not actual DER decode failures. rustls-pemfile is fully removed from Cargo.toml/Cargo.lock; the RUSTSEC-2025-0134 ignore entries in deny.toml and .cargo/audit.toml are removed since the advisory no longer applies. cargo deny check advisories licenses bans sources and cargo audit both pass clean with no ignore needed.
  • cargo clippy --tests -- -D warnings is now clean on both feature configurations (task #39). CI previously only gated non-test code; ~170 test-target warnings (mostly collapsible_if, doc_lazy_continuation, field_reassign_with_default, needless_range_loop) are fixed mechanically, plus a handful of real bugs surfaced along the way: a #[deny]-level approx_constant false positive that was silently blocking --tests compilation entirely; three test-only helpers whose #[cfg] was looser than their actual (feature-gated) callers, making them dead code under --no-default-features --features runtime-tokio,jemalloc (src/runtime/mod.rs, src/text/store.rs, tests/vector_db_isolation.rs); and a cargo clippy --fix autofix that would have deleted a key_to_shard import still required under the default graph feature (tests/sharded_multi_exec_locality.rs) — caught by cross-checking against a default-features build before committing. cargo clippy --all-targets -- -D warnings (default features) is also clean — no residue in benches. No test assertions were altered.

[0.7.1] — 2026-07-15

Patch release closing the two follow-ups disclosed in the v0.7.0 tag notes: the SQ8 vector CPU error-storm and replica TTL determinism.

Fixed

  • Vector: SQ8/TQ code-size mis-dispatch CPU error-storm. SQ8's bits() returns 8, which falls outside TurboQuant's supported 1..=4 range; the free code_bytes_per_vector(padded, bits) helper hit its _ => arm and returned 0 while logging a tracing::error! on every call. On the hot memory-accounting path (store.rs via bytes_per_code_per_vector()) this produced a CPU-pegging error storm and wrong resident-byte accounting for SQ8 indexes. Dispatch now computes the true SQ8 layout (dim unpadded u8 codes + 8-byte affine (min,scale) trailer) directly, and the free-fn logs the unsupported-bit-width error at most once via an AtomicBool latch.
  • Replica TTL semantics — deterministic cross-node expiry (#71). Two changes close the replica-TTL caveat disclosed in the v0.7.0 tag:
  • #71a — master-side absolute rewrite. Relative-expiry commands are now rewritten to absolute deadlines before they enter the durable log and the replication stream: EXPIRE/PEXPIRE → PEXPIREAT, SETEX/PSETEX → SET … PXAT, SET … EX/PX → SET … PXAT, GETEX … EX/PX → PEXPIREAT. The absolute deadline is computed from the master's per-tick cached clock — the exact value the command handler stored — so a replica (or an AOF replay after a restart) reproduces the master's expiry instant instead of restarting the countdown at apply/replay time. Already-absolute forms (PEXPIREAT, EXPIREAT, EXAT, PXAT, PERSIST) and past-time deletes propagate verbatim.
  • #71b — role-gated active expiry. A replica no longer runs its own active-expiry deletion sweep (both the monoio shard tick and the tokio background task); it keeps a logically-expired key resident (reads still see it as gone) until the master streams the authoritative removal, so both nodes delete a key at the same point in the stream instead of racing independent TTL sweeps.

[0.7.0] — 2026-07-15

Replication GA for multi-shard masters. Moon now supports real Redis-compatible asynchronous replication with WAIT/ACK acknowledgement semantics: a multi-shard master (--shards N) streams a merged, exactly-once command feed to a single-shard streaming replica across all data planes — KV, vector/text index, graph, workspace (WS), message-queue (MQ), and temporal. Every write plane is crash-durable and survives kill -9 on either side, validated by a 24h continuous-load kill-9 soak (alternating master/replica restarts every 12 min) that asserts zero loss of any WAIT-acknowledged write. Also folds in the full v0.6.1 hardening scope (WAL v3 storage-kernel M1–M4: cross-plane crash matrix, unified per-shard WAL-recycle floor, atomic durable writes, FTS term-dict durability) and a supply-chain CI gate (cargo audit + cargo deny).

Soak evidence (release gate REPL-SOAK-01): SOAK-PASS duration=86400s cycles=114 acked=82044 inflight=7 master_kills=57 replica_kills=57 — 82,044 WAIT-acked writes preserved across 114 alternating kill-9 cycles, zero acked-write loss (2026-07-15, run dir moon-soak/runs/20260714-141946, RC e2d87893).

Known limitation: the streaming replica is single-shard only (--shards 1) — the multi-shard work in this release is master-side (merged N-shard PSYNC feed). Multi-shard replicas are roadmapped for v0.8/v0.9.

Replica TTL semantics (disclosure): relative-expire commands (EXPIRE, SETEX, PEXPIRE, GETEX with a relative TTL) currently replicate verbatim rather than being rewritten to absolute PEXPIREAT on the master, and replicas run their own active-expiry cycle regardless of role. In practice keys expire correctly on both sides under normal clock sync, but a master/replica clock skew can shift a relative-TTL key's expiry moment between the two by up to that skew. Absolute-expiry rewrite + role-gated passive expiry land in v0.7.1 (task #71b). Applications needing exact cross-node expiry parity should set absolute deadlines with PEXPIREAT until then.

Documentation — replication configuration & tuning

Documented the v0.7 replication surface: corrected the stale guides/clustering.md replica walkthrough (it still showed the pre-GA topology — single-shard leader, multi-shard replica — the reverse of the shipped shape) and its outdated "WS/MQ not replicated yet" note (all six planes replicate as of Wave B). Added the WAIT-durability ladder, replica promotion (REPLICAOF NO ONE), INFO replication monitoring, a Read/write splitting subsection (Moon has no built-in R/W-split proxy — client-side vs command-aware external-proxy patterns), and the replica-TTL caveat. New Replication section in configuration.md and a Replication durability section in guides/tuning.md (RPO/latency trade-offs, read-scaling, lag alerting).

Fixed — segment_plane_scan missed v0.6.0's nested-Command plane framing, risking WS/MQ/temporal data loss on upgrade (task #69)

WalWriterV3::recycle_aggressive/recycle_segments_before gate deletion of sealed WAL v3 segments on segment_plane_scan's blocks_recycle verdict, which — until this fix — classified a segment purely from its OUTER record-type byte. v0.6.0's shard event-loop wal_append drain hardcoded WalRecordType::Command as the outer wrapper for EVERY plane record (WorkspaceCreate/Drop, MqCreate/Push/Pop/Ack/Trigger/Drop, TemporalUpsert, GraphTemporal) — verified against the v0.6.0 tag's event_loop.rs (lines 1452/2070). So on a v0.6.0→v0.7.0 upgrade, every plane record in a pre-upgrade segment was invisible to the scan: it saw only Command at the outer level and classified the whole segment as recyclable pure-KV history. WS/MQ/temporal-upsert have no snapshot or checkpoint format in ANY mode — their replay is WAL-only (task #43) — so recycling one of these segments permanently destroyed that history, with no way to recover it.

segment_plane_scan now peeks a Command record's payload for a nested v3 record frame (length-prefix + CRC32C validated, matching read_wal_v3_record's own framing) and classifies the INNER type against the same blocking set, mirroring the unwrap shared_databases.rs::replay_workspace_wal/replay_mq_wal/ replay_temporal_wal already perform via read_wal_v3_record(&record.payload) at replay time — scan and replay now agree on what counts as a nested record. Fails closed: a structurally valid nested frame with an unrecognized inner type byte (a future plane type this build predates) still blocks recycling. An ordinary (non-nested) Command payload — the overwhelmingly common case — is unaffected and stays recyclable.

Byte-slice peek only, no allocation and no LZ4 decompression (Command payloads are never LZ4-compressed by write_wal_v3_record) — off the hot path (recycle passes only) but kept allocation-free per repo convention.

Fixed — Replication fanout gate perf debt: parking_lot migration, FANOUT_HINT mis-activation, single-guard write path (task #70); deleted dead master-side PSYNC path (task #72)

Two pre-soak perf/hygiene defects blocking the v0.7.0 tag, fixed together since both touch ReplicationState locking on the per-write hot path.

Task #70 (4 sub-fixes):

  1. ReplicationState's outer lock (ConnectionContext::repl_state and every Arc<RwLock<ReplicationState>> holder — 27 call sites across admin/metrics_setup.rs, cluster/gossip.rs, command/server_admin.rs, main.rs, persistence/aof/pool.rs, replication/{reason_del,replica, master,state}.rs, scripting/bridge.rs, server/conn/*, shard/*) was std::sync::RwLock while the inner per_shard_backlogs was already parking_lot::Mutex — an inconsistent locking policy and a source of .read().unwrap() / if let Ok(g) = rs.read() poisoning dances on the write path. Migrated the outer lock to parking_lot::RwLock throughout; removed all poisoning-related Result unwrapping (parking_lot's read()/write() return the guard directly, try_read()/try_write() return Option, not Result).
  2. ensure_backlogs_allocated() (called from try_handle_replconf on ANY bare REPLCONF, including monitoring probes and failed handshakes) used to call mark_fanout_active() unconditionally. Since FANOUT_HINT is a sticky, never-cleared process-global flag, a single stray REPLCONF permanently taxed every subsequent write with the fanout-active gate check — even on a server that never completes a PSYNC. Split allocation from hint-activation: ensure_backlogs_allocated() still allocates backlogs (load-bearing for the handshake) but no longer touches FANOUT_HINT; mark_fanout_active() now fires only from try_handle_psync, on an actual PSYNC/replica registration.
  3. record_local_write / record_local_write_db took up to 3-4 separate repl_state.read() acquisitions per write (SELECT-needed check, backlog append, offset advance, ...). Collapsed the internal chain to ONE guard acquired at the top of record_local_write_db and threaded through to record_local_write (now fn record_local_write(g: &ReplicationState, shard_id, bytes), no longer self-locking). replication_fanout_active stays a separate, short-lived gate lock — deliberately NOT folded into the same guard, because some external call sites (handler_monoio/mod.rs's MOVE/COPY handling) reach an .await between the gate check and the next repl-state touch, and Rust drops lock guards at end of lexical scope, not at last use — holding one guard across that span would violate the never-lock-across-.await rule. Net effect: at most 2 lock acquisitions per replicated write (1 gate + 1 write guard), down from up to 4-5.
  4. Added #[inline] to record_local_write, record_local_write_db, and replication_fanout_active, matching the sibling try_handle_* convention in dispatch.rs.
  5. Correctness follow-up (review-caught regression on #2): splitting backlog allocation from hint activation opened a REPLCONF→PSYNC window where a bare REPLCONF could allocate + seed a shard's ReplicationBacklog at its offset at that instant, then — with FANOUT_HINT still false — every subsequent local write advanced the shard's real offset counter via the bare-issue_lsn branch with NO matching backlog append. The backlog's end_offset silently skewed stale relative to the real counter for as long as the window stayed open (an AOF-enabled master under continuous write load with periodic replica kill-9/reconnect — exactly the v0.7.0 24h soak's shape). Any snapshot/cut/push-offset captured from that backlog after activation would then be range-inconsistent with it, corrupting partial-resync catch-up reads (wrong bytes or a silently-skipped catch-up) and acked-write accounting. Fixed by adding ReplicationBacklog::realign_to(offset) (resets the buffer to empty, reseeded at offset — a skewed buffer's contents are already useless, so dropping them is correct; an offset request that falls below the new start_offset fails safe into a full resync) and ReplicationState::realign_backlog(shard_id), called at every activation site — try_handle_psync (dispatch.rs, right after mark_fanout_active) and the RegisterReplica/PrepareReplicaSync arms (shard/spsc_handler.rs) — BEFORE any snapshot/cut/push offset is captured. All three sites run on the shard thread that owns that shard's own offset advances, so the offset read + realign is race-free with the shard's own append/advance sequence.

Task #72: handle_psync_on_master (both runtime-tokio and runtime-monoio variants) and register_replica_with_shards in src/replication/master.rs were unreachable — every PSYNC handler actually wired from shard/conn_accept.rs goes through handle_psync_inline_single_shard / handle_psync_inline_multi_shard instead. The dead pair also carried a latent broken WAIT implementation: it initialized ack_offsets but never spawned the ack_read_loop that drains replica ACKs into them, so wait_for_replicas against a registration made through that path would have blocked forever. Deleted both variants plus evaluate_psync_shared (only reachable from the deleted code); kept backlog_bytes_from, still used by the live send_backlog_range.

Two new unit tests (replication::state::tests) cover the FANOUT_HINT fix with delta-based assertions (robust against shared-binary test-order contamination): test_ensure_backlogs_allocated_does_not_activate_fanout_hint (RED on the pre-fix code — verified by temporarily reintroducing the bare mark_fanout_active() call and observing the assertion fail) and test_mark_fanout_active_sets_hint. Four more cover the backlog-realign correctness fix (#70.5): test_realign_to_resets_skewed_backlog and test_realign_to_noop_when_already_aligned (replication::backlog::tests, the low-level buffer reset behavior) plus test_realign_backlog_fixes_skew_after_hint_false_window and test_realign_backlog_then_append_reads_back_exact_record (replication::state::tests, reproducing the exact skew scenario — allocate, advance the offset via issue_lsn with no append, then activate — and proving the post-fix backlog is byte-position-consistent with the real offset again). RED verified by temporarily neutering realign_backlog's body and observing both state::tests cases fail (end_offset stuck at 0 instead of matching the real offset 300; the post-append record unreadable at its own pre-append offset), then reverting.

No wire-protocol or observable behavior change beyond the hint gating: the REPLCONF → PSYNC handshake sequence is unchanged and covered end-to-end by the replication_hardening / replication_multishard / replication_planes integration suites (all green, monoio — master-side PSYNC is monoio-only by pre-existing design, try_handle_psync_unsupported under tokio, untouched by this change) plus the full replication:: unit-test module under both runtime-monoio (default) and runtime-tokio,jemalloc.

Fixed — crash_recovery_disk_offload_no_aof harness assumed eviction-throughput durability the write path no longer provides (task #44)

cold_keys_recover_after_crash_without_aof forced LRU eviction with a 16,000-key pipelined SET burst and asserted post >= 65% of PROBE_COUNT, assuming most evicted probes land durably in a heap-*.mpf file. That predates disk_offload_spill_inert() / PR #273 (policy-aware eviction fail-close): under --appendonly no, the per-connection write-path eviction gate (run_write_eviction_gate, src/server/conn/handler_monoio/mod.rs) — the path every ordinary SET/HSET past maxmemory actually takes — has no ShardManifest handle and PLAIN-DROPS victims (documented Redis "pure cache, no durability" semantics; see docs/PRODUCTION-CONTRACT.md's CRASH-02 note on task #57, which made this drop-not-OOM behavior the intended fix). Only the periodic memory-pressure tick (shard/persistence_tick.rs) durably spills under appendonly no, and a busy connection's synchronous per-write eviction always wins the race before that tick observes a sustained over-budget shard — so the pipelined-SET burst durably spilled 0-1/200 probes (ground-truth-verified via direct manifest/heap-*.mpf inspection), making the 65% floor permanently unreachable. Not a data-loss regression: the dropped keys were never claimed durable (the boot-time WARN says so explicitly for this exact config).

Two changes, no src/ change: (1) write_filler now sends the filler as ONE MSET instead of 16,000 pipelined SETs — MSET's handler applies all pairs with no per-key eviction check, and the write-path gate that does run checks memory once, pre-MSET, so a shard can end up far over budget with no further write pending to re-trigger the plain-drop path; the periodic tick is then the only thing left to reclaim it, durably spilling the LRU-oldest keys (the probes) via the manifest. (2) The fixed RECOVERY_FLOOR is replaced by count_durable_probe_entries, which reads each shard's manifest + Active KvLeaf DataFiles directly (the same way recover_shard_v3_pitr does at boot) right before the kill to establish ground truth, and asserts recovery returns AT LEAST that many probes — the actual #22 regression signal, decoupled from eviction throughput. Verified the rewritten test still red-flags a reintroduced #22 regression (manually reverted the persistence_dir.is_some() || disk_offload_base.is_some() gate in main.rs, confirmed RED, reverted back). Green 10×+ consecutive on macOS (monoio + tokio runtimes); crash_recovery_cold_del_resurrection and the crash_matrix_cross_plane no-durability-contract spot check show no regression.

Fixed — crash_recovery_graph_durability g1/g2/g3 harness polled the deleted WAL v2 flat file (task #31)

The G1–G3 legacy-mode graph crash tests never actually ran their kill -9 scenario: wait_for_wal_bytes (the "my last write reached the WAL fd" probe that makes the crash point deterministic) and G3's replay-idempotency length check still read the flat <dir>/shard-0.wal WAL v2 file, which was removed in PR #236 — one day before this test file was added (#237). The probe waited 20s for a file that is never created and panicked, on every platform (verified identical on the Linux dev VM — not a macOS durability gap). Masked because the suite is #[ignore]d. Both helpers now scan the real <dir>/shard-0/wal-v3/*.wal segment directory; the G3 length check excludes each segment's fixed 64-byte header so normal per-boot segment rotation isn't misread as replay double-append. With the task #60 replay fix already on main, g1–g5 are green 3× consecutively on macOS. The underlying replay_graph_wal RESP-parse bug these tests expose was fixed separately in PR #322.

Changed — PRODUCTION-CONTRACT.md reconciled to post-replication-GA main (v0.7.0 prep)

2026-07-14 owner reconciliation of the GA Exit Ledger against the tree: re-ticked FUZZ-01 (12th fuzz target restored), ACL-REG-01 (PR #258), COLD-TTL-01 (ColdIndex::sweep_expired + reclaim counters), SEC-07 (release-agnostic supported-versions policy); FT-PARITY-01 updated (FT.AGGREGATE shipped, FT.ALTER open). CRASH-01 gained honesty caveats (g1–g3 harness never ran before PRs #322/#324; crash_recovery_disk_offload_no_aof residual red, task #44). New rows for shipped-but-untracked guarantees: CRASH-02 (37-cell cross-plane kill-9 matrix), MEM-10X-01 (10× RAM G2 acceptance), REPL-PLANES-01 (all-plane replication); new REPL-SOAK-01 row gates the v0.7.0 tag on the 24h replication soak. GA-blocking gap now 17 rows.

Added — supply-chain security CI gate: cargo audit + cargo deny check (task #63, SUPPLY-01)

deny.toml existed in the tree but was never wired into CI (its own header comment said so). Added .github/workflows/supply-chain.yml: two ubuntu-latest jobs, audit (cargo audit, blocking on RUSTSEC vulnerability-class advisories) and deny (cargo deny check advisories licenses bans sources, blocking on any deny.toml violation), triggered on PRs touching Cargo.toml/Cargo.lock/deny.toml/.cargo/audit.toml, push to main, and a weekly schedule (advisories publish independent of code changes). Runs on the hosted runner, not the self-hosted moon-dev box — no build is required, just dependency-graph inspection.

Fixed three real advisories to get to green: memmap2 0.9.10 → 0.9.11 (RUSTSEC-2026-0186, unsound pointer-offset validation), crossbeam-epoch 0.9.18 → 0.9.20 (RUSTSEC-2026-0204, invalid pointer dereference in Display), spin 0.9.8 → 0.9.9 (yanked), and dropped the core2/proc-macro-error2 unmaintained transitives by bumping rust-embed 8.11.0 → 8.12.0 (console feature). Three unmaintained transitive advisories with no available safe upgrade (fxhash via monoio, paste via tikv-jemalloc-ctl, rustls-pemfile pending a rustls-pki-types::PemObject migration) are explicitly ignore-listed with reasons in deny.toml / .cargo/audit.toml rather than left to silently pass — an always-red gate is worse than none, but a silently-permissive one is worse still.

Added — 24h replication kill-9 soak harness gating the v0.7.0 release tag (task #61)

scripts/soak-replication-24h.sh + scripts/soak_replication_driver.py: a machine-verifiable zero-data-loss soak for the "Replication GA for multi-shard masters" headline. A master (--shards 4 --appendonly yes --appendfsync always) and replica (--shards 2) run under continuous SET soak:{seq} {seq}:{ts} + WAIT 1 <timeout> load; only a WAIT>=1 reply appends the seq to an fsync'd acked-write ledger (a WAIT timeout is recorded separately as "in-flight" — allowed to be lost or present, never a failure). Every ~12 minutes the harness alternates kill -9 on master/replica, restarts the killed side, waits for a data-driven resync gate (polls the last few acked seqs back from the replica — INFO replication's master_link_status:up only proves the TCP link is back, not that the backlog/RDB replay has landed), then samples ≥1000 random + the last 200 acked seqs and asserts exact value parity on both master and replica. Any mismatch prints SOAK-FAIL seq=<n> side=<m|r> cycle=<k> and exits 1 immediately; a full-ledger sweep runs at soak end. Hourly progress: SOAK-OK hour=<h> acked=<n> cycles=<k> master_kills=<a> replica_kills=<b>. --smoke runs a 30-minute validation (3 chaos cycles) instead of the full 24h. Builds from a VM-local clone (~/moon-soak/repo) per the OrbStack diskfull-guard/tmpfs rules; kill/restart is strictly PID-targeted kill -9 (never a broad pkill pattern) per the SO_REUSEPORT hang trap.

Fixed — WAIT wedged forever on a restarted multi-shard master after kill-9 (task #67)

After a multi-shard master (--shards >= 2) was kill -9'd and restarted with prior write history, the surviving replica kept streaming and applying writes correctly, but WAIT 1 <timeout> on the restarted master timed out indefinitely — the replica's periodic REPLCONF ACK never registered as "caught up" even though it was arriving every second. Root cause: AOF recovery's ReplicationState::seed_master_offset seeded ONLY the process-wide master_repl_offset (total_offset()) from the recovered max LSN, leaving every per-shard shard_offsets[i] at the fresh-boot 0. handle_psync_inline_multi_shard's full-resync handshake advertises Σ shard_offset(i) — not total_offset() — as a reconnecting replica's new baseline (each shard captures its own offset atomically with its RDB body, which the exactly-once live-fanout cut gate depends on), so a replica reconnecting after the restart adopted a near-zero baseline while wait_for_replicas kept comparing ACKs against the correctly-seeded (large) total_offset() — a gap the replica could never close. seed_master_offset now also seeds shard 0 to the same recovered value, restoring the Σ shard_offsets == total_offset() invariant every write already maintains going forward. Reproduced 2x by the v0.7.0 replication soak (scripts/soak-replication-24h.sh); regression test: tests/replication_hardening.rs::master_kill_restart_wait_acks.

Fixed — legacy-mode (--disk-offload disable) graph WAL replay silently dropped the entire graph plane on kill-9 restart (task #60)

replay_graph_wal (src/shard/shared_databases.rs, the legacy-mode-only replay path taken when persistence_dir is set but disk-offload is disabled) passed the raw RESP-encoded WalRecord::payload directly as the cmd argument to CommandReplayEngine::replay_command, instead of parsing it into frames and passing the bare command name + &[Frame] args as the sibling replay_graph_wal_v3 (disk-offload-enabled path) and recovery.rs's KV Command replay both already did. GraphReplayCollector::is_graph_command compares against literal names like b"GRAPH.CREATE", so it never matched the multi-line RESP blob — graph_command_count() stayed 0, the replay_graph_commands guard never fired, and every graph mutation was silently lost on restart. Regression from PR #236 (WAL v2 → v3 port). Un-gates 6 previously-RED crash-matrix cells (tests/crash_matrix_cross_plane/{tests_legacy.rs,tests_spot.rs}): legacy s1 graph_isolated, txn_isolated_committed, txn_isolated_atomicity, mixed_all_planes_synced, mixed_all_planes_mid_pass_c, and spot-check cross_plane_spot_legacy_s4_graph_isolated.

Fixed — async-park cold GET off the shard event loop, no timeout (task #59)

Database::get()'s cold-tier fallback (promote_cold_if_present -> cold_read::read_cold_entry) does a blocking pread inline on the single-threaded shard event loop; under spill/AOF write backlog on the same disk that read can block for up to ~1.9s, stalling every connection on the shard for the duration. The monoio connection handler's GET fast path (src/server/conn/handler_monoio/mod.rs) now peeks whether the key needs a disk read and, if so, .awaits the real result via a new off-thread worker pool (storage::tiered::cold_read_pool::read_cold_entry_async, 2 threads by default, MOON_COLD_READ_POOL_THREADS override) — no timeout: the awaiting connection task yields, letting sibling connections on the same shard thread keep running while the real disk I/O completes, and the resolved outcome (never a placeholder) is promoted into hot RAM (Database::promote_cold_outcome) before the normal synchronous dispatch path answers. An earlier revision of this fix used a bounded/timeout pool that could answer Miss for an existing spilled key under backlog — a silent correctness regression — and has been fully removed; there is no timeout anywhere in the new design. A second review round caught a DEL-during-await TOCTOU: a DEL/FLUSHDB on the same key running on the same shard thread while the GET's .await was suspended could resurrect the deleted key once the stale Hit outcome resolved. Fixed by revalidating, synchronously and without any intervening .await, that the cold index still maps the key to the SAME ColdLocation (now PartialEq/Eq) the outcome was read from before promoting anything — otherwise the outcome is discarded and the caller's normal dispatch answers from current state. MULTI/EXEC bodies, Lua redis.call, and every non-GET command (MGET/HGET/etc.) are architecturally unable to .await from their current synchronous call sites and continue to use the original, unmodified synchronous blocking read — slower under backlog, but never wrong. A third review round caught that the round-1/round-2 fix, though correct in isolation, was landed on a branch plain GET key never actually reaches: src/server/conn/blocking.rs's synchronous inline fast path (try_inline_dispatch, which handles plain GET UNCONDITIONALLY — unlike SET, GET inlining isn't gated by disk-offload) still called the blocking read_cold_entry_at directly on a cold miss, before the async pre-warm hook was ever reached — so redis-benchmark GET and the task's own repro were completely unaffected by rounds 1-2. Fixed: the inline GET path now distinguishes "genuinely absent" (still answered inline, fast $-1) from "a real cold entry exists" (declines to inline — returns the pre-existing "not inlined" sentinel with the command bytes unconsumed, exactly like every other early-return in that function) and lets the already-correct fall-through hand it to the async pre-warm hook instead. See tmp/task59-design.md for the full call-site audit and the scope this pass covers vs. defers.

Added — allocator-overhead and PageCache observability in INFO memory / Prometheus (task #58)

INFO memory and /metrics previously had no way to see the gap between process RSS and what Moon could account for (DashTable+entries, vector/text/ graph/lua planes, replication backlog) — the only place that gap was ever computed was MEMORY DOCTOR's on-demand snapshot. Similarly, the disk-offload PageCache's resident 4KB/64KB frame buffers were tracked internally but never surfaced anywhere.

Added allocator_overhead_bytes (= RSS - tracked_sum, now including PageCache and the replication backlog) as a continuously-updated metric, sampled once per 100ms tick by shard 0 (persistence_tick::run_eviction_tick) rather than recomputed per-request, and published through a single atomic (admin::metrics_setup::{update,get}_allocator_overhead_bytes) that both INFO memory and the Prometheus moon_memory_bytes{kind="allocator_overhead"} gauge now read — the 15s Prometheus updater no longer independently recomputes it against a separately-timed RSS read, closing a drift source. Added pagecache_bytes (PageCache::resident_buffer_bytes(), published into a new ShardStoreMemory.pagecache atomic alongside vector/text/graph/lua on the same tick) as its own INFO/Prometheus line (moon_memory_bytes{kind="pagecache"}, memory_prometheus_kinds test now expects 10 kinds). Both figures are observability-only — neither feeds eviction or elastic-budget gating. See src/shard/shared_databases.rs (ShardStoreMemory.pagecache), src/shard/persistence_tick.rs (compute_allocator_overhead, publish sites), src/admin/metrics_setup.rs, src/command/connection.rs (INFO memory). New test: tests/info_memory_allocator_pagecache.rs.

Fixed — startup blocked ~153s on the crash-orphan heap-file sweep at scale (task #55)

recover_shard_v3_pitr called kv_spill::sweep_orphan_heap_files synchronously — classify AND delete — for every shard on the main thread, before any listener/event loop started. At production scale (G2 crash-matrix bench: ~236K spilled files) this stalled restart-to-first-served-command by ~153s (~40s/shard × 4 shards), even though the manifests themselves recovered in ~4s: the deletion I/O, not the manifest rebuild, was the bottleneck.

Split the sweep into a cheap synchronous classify (read_dir + HashSet membership, no remove_file syscalls — safe to run during recovery because it observes the exact manifest state recovery just rebuilt, before the shard serves a single command) and a deferred delete (kv_spill::remove_orphan_heap_file) that now runs on a plain std::thread spawned once the shard's own event loop starts, fully decoupled from the accept path. Orphan files are by definition unreferenced by the manifest or any in-memory index, so reclaiming them needs no shard-state synchronization, and the file-id namespace is monotonic — nothing a shard writes after this snapshot can retroactively register one of the already-classified paths, so no epoch fence beyond "classify before serving" is required. Restart-to-ready is now seconds-scale regardless of cold-plane size, and the orphans are still fully reclaimed shortly after. See src/persistence/recovery.rs, src/storage/tiered/kv_spill.rs (classify_orphan_heap_files / remove_orphan_heap_file), src/shard/mod.rs (pending_heap_orphans staging, mirroring the existing recovered_warm_segments pattern), and src/shard/event_loop.rs. New test: tests/crash_recovery_orphan_sweep_readiness.rs.

Fixed — read-only replicas silently executed writing EVAL/EVALSHA (task #38)

A client-issued EVAL/EVALSHA on a read-only replica ran unconditionally: EVAL/EVALSHA deliberately carry no WRITE command-metadata flag (a script's writes replicate as individually-emitted effect records, not as the literal EVAL), so the connection-layer try_enforce_readonly gate that blocks every other write command on a replica never saw them — a script's redis.call('SET', ...) mutated the replica's local state with no rejection at all, silently diverging it from the master.

Fixed at the actual write boundary — scripting::bridge::make_redis_call_fn (the redis.call/redis.pcall bridge) now rejects a WRITE-flagged inner command with -READONLY You can't write against a read only replica. the moment the shard's ReplicationState reports the replica role, mirroring upstream Redis semantics: a script that never writes still runs to completion on a replica (EVAL_RO/FCALL_RO and plain read scripts are unaffected), while one that does write aborts at the first offending call. Reuses the same WS/MQ/GRAPH.QUERY read-only-subcommand carve-outs as the connection-level gate, and reads a lock-free AtomicBool mirror of ReplicationState::is_replica_mirror (same pattern as ConnectionContext::is_replica_mirror) so a tight script write-loop never takes the ReplicationState lock.

Verified this does not regress master→replica Lua effect replication (Wave A part 2, task #34): replayed effects apply via replication::apply::apply_local, which dispatches the inner write command directly against storage and never touches the Lua bridge, so the new gate cannot see — and cannot block — replica apply of a master's script effects (eval_effects_parity_shards1/_shards4, eval_incr_no_double_apply, eval_writes_survive_restart all still green).

Multi-db note: redis.call('SELECT', ...) was already rejected inside scripts (pre-existing fail-loud guard — the CURRENT_DB thread-local is pinned to one Database for the whole script execution), so there is no per-script db-switching path that could race or diverge from this new readonly check.

Fixed — MqPop WAL/replication payload v2 carries real PEL idle metadata (task #47)

apply_mq_pop (boot-time WAL replay AND live replica apply — same function, src/shard/shared_databases.rs) hardcoded a claimed PEL entry's delivery_time/seen_time to 0 regardless of when the original MQ.POP actually claimed it. Since XPENDING's idle time is now - delivery_time, a kill-9 restart or a promoted replica reported "idle since the Unix epoch" (an astronomically large idle time) instead of the real elapsed duration — XPENDING/XPENDING ... IDLE consumers on a recovered node saw wrong data for every pre-crash claim.

Fixed by bumping the MqPop WAL record to its own v2 payload (MQ_POP_WAL_V1/MQ_POP_WAL_V2 in src/mq/wal.rs — MqPop now versions independently of the shared MQ_WAL_VERSION used by every other MQ record kind) that appends [delivery_time_ms:u64][seen_time_ms:u64] captured once per POP batch (every entry one MQ.POP claims shares a single now, per Stream::read_group_new). src/shard/mq_exec.rs's POP handler now reads these back out of the PEL/consumer state it just wrote instead of discarding them. Decode is version-led and fail-closed: v1 payloads (no trailing timing fields) still decode, filling an explicit 0 sentinel that apply_mq_pop resolves to current_time_ms() at apply time — a strict improvement over the old hardcoded-0 behavior, not a behavior match to it. Any version byte other than 1 or 2 is rejected (None), same fail-closed posture as every other MQ WAL decoder.

Producer (src/shard/mq_exec.rs), boot-time replay (src/shard/shared_databases.rs::apply_mq_pop), and live replica apply (src/replication/apply.rs::apply_mq) all updated; the MQ WAL fuzz target (fuzz/fuzz_targets/mq_wal_record.rs) needed no changes since it fuzzes decode_mq_pop generically. New coverage: codec round-trip tests for v1 (sentinel-filled) and v2 (full round-trip) in src/mq/wal.rs, a kill-9 integration test asserting XPENDING idle survives restart (tests/crash_recovery_mq_effects.rs::xpending_idle_metadata_survives_kill9), and a replica-promotion test asserting the promoted node's XPENDING idle reflects the master's original claim time, not the replica's own apply time (tests/replication_mq.rs::promoted_replica_reports_real_pending_idle_time).

Fixed — unified replica poison-record policy across graph/MQ/WS/temporal planes (task #48)

Each replica-apply plane (graph WAL replay, MQ effect records, workspace create/drop, temporal invalidate, RESP framing) previously handled a malformed/undecodable record ad hoc — mostly tracing::warn! + silently skip-and-continue, one path (WS.CREATE.APPLY/WS.DROP.APPLY) even swallowed the failure entirely (try_with_shard(...).is_some() returned true regardless of the inner decode outcome). A silent skip means every subsequent record in the stream keeps applying against state that has already diverged from the master — invisible until an operator notices missing data.

Unified policy (replication::apply, see its "Unified poison-record policy" module docs): any decode/parse failure on a live-stream record is now a ApplyOutcome::Poisoned — logged loud (rate-limited to 1/sec), counted in a new INFO counter, and propagated up so the replica connection task drops the link and lets the existing reconnect/resync loop renegotiate PSYNC (never silently skip, never panic). PSYNC snapshot install already fail-closed correctly (a malformed aux blob fails the whole install); it now also increments the same counter. Semantic apply errors on well-formed records (e.g. WRONGTYPE, "entity not found") keep the existing warn-and-continue posture — that is data-level divergence from a command that legitimately executed differently, not stream corruption.

  • New INFO replication field: replication_poison_records_total.
  • src/replication/apply.rs: ApplyOutcome enum, poison() helper, apply_graph/apply_mq/apply_ws_create/apply_ws_drop/ apply_temporal_invalidate now return bool (poisoned vs. applied) instead of logging-and-swallowing.
  • src/replication/replica.rs: both tokio and monoio stream-apply loops match on ApplyOutcome and drop the connection on Poisoned, same as the existing NoShardSlice / DrainResult::fatal paths.

Added — real SHUTDOWN [NOSAVE|SAVE] (task #27)

SHUTDOWN was previously a stub that always replied ERR Errors trying to SHUTDOWN. Check logs. and never terminated the process. It is now handled at the connection-handler level (like BGSAVE/ACL, alongside all three dispatch paths: handler_single, handler_sharded, handler_monoio):

  • Parses the optional NOSAVE / SAVE modifier (Redis parity: bare SHUTDOWN forces a save iff RDB save points are configured; NOSAVE always skips it; SAVE always forces it). ABORT is rejected (Moon's SHUTDOWN runs synchronously to completion, so there is never an in-progress shutdown); FORCE/NOW are accepted no-ops.
  • A forced save that fails replies an error and the server stays up (single-shard: synchronous SAVE; sharded/monoio: the cooperative per-shard BGSAVE snapshot, polled with a bounded timeout).
  • On success, triggers the exact same CancellationToken-driven graceful shutdown sequence already used for SIGTERM — per-shard WAL/AOF flush + fsync across every plane, accept-loop stop, and clean connection teardown — rather than reimplementing it. No reply is sent (Redis parity: the client observes the connection close).
  • ACL category was already admin/dangerous (DNG) in the command registry; unchanged.

Coverage: tests/shutdown_integration.rs (NOSAVE prompt exit + waitpid status, AOF durability across a SHUTDOWN→restart round trip, a failed forced SAVE keeps the server up, syntax errors keep the server up) plus a cross-shard durability smoke section in scripts/test-consistency.sh.

Fixed — AOF writer manifest-wait/cancellation data-loss race (found via task #27)

Writing the SHUTDOWN durability test surfaced a real, pre-existing bug: on first boot, main.rs's recovery creates the AOF manifest on a separate thread/task from the one that starts accepting connections, so a client can already be writing (queuing Append messages) while the AOF writer thread is still polling for that manifest to appear (src/persistence/aof/writer_task.rs, three writer-loop variants). That polling loop checked cancel.is_cancelled() before attempting AofManifest::load each iteration — a graceful shutdown landing in the same ~50ms poll window made the writer return immediately without ever loading the (by-then-current) manifest, silently discarding every already-queued Append. Reproduced reliably (>80% of runs) by sending SHUTDOWN within tens of milliseconds of boot: AOF writer: cancelled while waiting for manifest in the log, and a 0-byte incr AOF file. Fixed by trying the load first in all three affected loops (TopLevel-monoio, PerShard-tokio, PerShard-monoio); a manifest that exists by the time the writer gets there is never missed just because a shutdown signal happened to land in the same instant. In practice this window was rarely hit via SIGTERM (real signal delivery has enough latency to clear it), which is presumably why it went unnoticed until a command that can call cancel() synchronously, milliseconds after the connection that issued the write, existed.

Changed — CI pipeline optimization (round 2: cargo-nextest)

  • CI test steps (Linux/macOS/Windows) now run under cargo nextest run --profile ci: per-test-binary parallelism and two automatic retries for the known flaky classes (fixed-port listeners, kill-9 timing under full-suite load) — a pass-on-retry is reported as FLAKY, keeping the signal while ending manual reroll round-trips. fail-fast = false surfaces every failure in one run. Doctests keep a dedicated cargo test --doc step (nextest does not run them). Config in .config/nextest.toml; local cargo test is unaffected.

Changed — CI pipeline optimization (round 1)

  • CodeQL no longer runs on every PR (main-push + weekly schedule only) — it was a full instrumented 20-40 min build and the slowest job on every PR, while serving as audit tooling rather than a per-PR review gate.
  • Swatinem/rust-cache shared-keys no longer embed the Cargo.lock hash: the action already invalidates changed dependencies via its own lockfile-aware restore, and hashing the lock into the key cold-started every job's cache on each dependency bump.
  • CI builds/tests run with debug = 0 (no debuginfo) via CARGO_PROFILE_*_DEBUG env — smaller artifacts (faster cache save/restore) and faster linking; local builds unaffected.

Fixed — FTS term-dict + FST sidecar durability, ends restart-rescan-only recovery (kernel M4, task #50)

The full-text-search inverted index was the last plane not kill-9-lossless: every restart rebuilt every text index by rescanning the keyspace and reassigning term ids by DashTable hash-iteration first-encounter order — not reproducible across restarts, which is why the pre-existing .fst sidecar write path had its load path (load_fst_sidecars) deliberately left uncalled in production (wiring it once corrupted FUZZY/PREFIX results: the sidecar's baked-in term ids silently collided with a freshly-rescanned dictionary's differently-assigned ids).

Fixed by persisting the term dictionary itself alongside the FST, in one atomic sidecar ({shard_dir}/{index}.tfst, magic TFS2, version-stamped, atomic_write_durable): per TEXT field, next_id, fst_high_water_mark, every (term, id) pair, and the optional FST bytes. TermDictionary::from_pairs reconstructs a dictionary whose ids are taken verbatim from the sidecar (never reassigned) and whose next_id continues the persisted high-water mark. Wired into shard boot (src/shard/event_loop.rs) so TextStore::load_term_fst_sidecars runs AFTER text index schemas are restored but BEFORE the keyspace auto-reindex rescan — seeding the term dicts first makes the rescan's get_or_insert calls resolve known terms to their persisted ids and assign fresh, non-colliding ids only to genuinely new terms, which is what makes loading the FST alongside it safe (FST ids and live dict ids are the same id-space by construction). Fails closed per index on any missing/truncated/corrupt/version-mismatched/field-count- mismatched sidecar — falls back to today's full rescan, never partially applies a sidecar. FT.COMPACT now calls the combined saver (save_term_fst_sidecar_for_index) instead of the old FST-only one. New FT.INFO counters sidecar_recovered_indexes / text_indexes_total (additive across shards) surface fast-boot coverage. New default-GREEN crash-matrix cells cross_plane_prod_{s1,s4}_text_fts_sidecar_isolated verify a FUZZY query survives kill-9 identically to a from-scratch rebuild. New fuzz target term_fst_sidecar covers the sidecar decoder.

Security — clear dependency vulnerability backlog (task #51)

sdk/python (uv.lock, 36 open Dependabot alerts incl. 1 CRITICAL) and console (pnpm-lock.yaml, 17 open alerts incl. 2 HIGH react-router):

  • sdk/python: bumped requires-python floor >=3.9 → >=3.10 (Python 3.9 reached EOL 2025-10). This was required, not cosmetic — nltk (pulled transitively via the llama-index extra) has no python-3.9-compatible release past 3.9.2, which carries the CRITICAL zip-slip advisory (GHSA, nltk.corpus.util.LazyCorpusLoader zip extraction) plus 4 more high/medium path-traversal and XSS CVEs. Collapsing the py3.9 resolution branch lets uv lock --upgrade land on nltk 3.10.0 everywhere. Also picked up pillow 12.3.0, aiohttp 3.14.1, urllib3 2.7.0, requests 2.34.2, orjson 3.11.9, langsmith 0.10.2, langchain-core 1.4.9, pytest 9.1.1 (all natural resolutions within existing >= floors — no manifest ceiling changes needed beyond the python floor). 263/263 SDK tests green post-bump.
  • console: targeted pnpm update (no package.json range changes needed — all fixes were already inside existing ^/transitive ranges) for react-router-dom 7.14.0→7.18.1 (clears the react-router vendored turbo-stream RCE + 3 more), dompurify 3.4.0→3.4.12 (transitive via @cosmos.gl/graph, clears 8 XSS/pollution advisories), vite 7.3.5→7.3.6 (needed first — 7.3.5 hard-pins esbuild@^0.27.0, below the fixed 0.28.1), esbuild 0.28.1, form-data 4.0.6, ws 8.21.0, js-yaml 4.3.0, @babel/core 7.29.7. pnpm audit --audit-level high clean. pnpm build and pnpm install --frozen-lockfile verified; the pre-existing console.test.ts failures (7/56 tests — zustand persist middleware vs. mocked storage) are unrelated to this change, reproduced identically on the pre-bump lockfile, left untouched.

Fixed — cross-shard MULTI/EXEC graph-leg misrouting (kernel M3 stage 3 / task #52, review round 3, P1)

CodeRabbit finding on PR #300: analyze_txn_locality (the pre-EXEC classifier that decides whether a queued body runs locally, hops to a remote owner shard, or is rejected CROSSSLOT) only scans KV keys via command_keys. GRAPH.* commands declare NO keys in command metadata (first_key==last_key==0), so a graph-bearing body was misclassified — Keyless for a graph-only body, or by its KV keys alone for a mixed body — completely ignoring the graph name's REAL owner shard (the standalone GRAPH.* path routes by graph_to_shard(name, num_shards), identical hash-tag-aware xxh64 to key_to_shard). At num_shards > 1 a transaction whose true graph owner differed from the classified owner would durably apply the GRAPH.* write to the WRONG shard's graph_store — invisible to subsequent normally-routed single-command reads. The task #52 crash cells never caught this because their KV key and graph name deliberately share one {txniso} hash tag (co-location by construction), so they always agree on owner regardless of the bug.

Fixed by folding the graph name into analyze_txn_locality's existing visit closure (the same one KV keys and the SORT/GEORADIUS STORE-dest already use): a body whose KV keys and graph name disagree on owner becomes CrossShard (rejected CROSSSLOT, same posture as any other genuine cross-shard body); a graph-only body becomes SingleShard(graph_owner), which routes through the SAME whole-body-atomic owner hop (execute_txn_on_owner / ShardMessage::TxnExecute) a KV-only remote-owner body already uses — no new hop mechanism, so the "body runs on ONE local shard slice" invariant is preserved.

New tests in tests/sharded_multi_exec_locality.rs: graph_leg_cross_shard_rejected (a KV key and graph name deterministically computed, not guessed, to hash to different shards at shards=4 → CROSSSLOT, neither leg applied) and graph_only_txn_routes_to_true_owner (a graph-only body commits and is visible via a fresh, normally-routed connection — proving owner placement, not "whatever shard the connection landed on"). Confirmed RED without the analyze_txn_locality fix (both new tests fail: the rejected-body test instead applies the KV leg and errors "graph not found" on the misrouted GRAPH.ADDNODE; the owner-routing test reads back 0 rows), GREEN with it, 3× consecutive. no_silent_divergence_across_shards / multi_shard_span_rejected / single_shard_unaffected (pre-existing KV-only locality tests) and the task #52 crash-matrix txn cells are unaffected.

Fixed — cross-store MULTI/EXEC graph leg replication (kernel M3 stage 3 / task #52, review round 2, P1)

Follow-up to the durability fix below: execute_transaction_sharded's new GRAPH.* branch originally wal_appended the graph-leg records itself, inside the function — but that function has no ConnectionContext / replication access, so the graph leg applied and persisted on the master but never reached a replica, a NEW master/replica divergence the durability fix introduced (pre-fix the in-txn GRAPH.* applied nowhere, so there was no divergence to have). execute_transaction_sharded now returns the collected graph records (bound to the db each command executed in, same per-entry-db rule as the AOF entries — a queued SELECT redirects later commands) instead of appending them itself. The monoio EXEC caller (handler_monoio/write.rs) now replicates each record via record_local_write_db — gated ctx.num_shards == 1 && replication_fanout_active(ctx), matching the live single-command graph path's scope exactly (graph replication is single-shard only; multi-shard graph replication rides the R2 broadcast redesign) — in the same synchronous stretch as the mutation, THEN wal_appends, same replicate-then-append order as the live path. The tokio/sharded caller and the cross-shard TxnExecute shard-message arm (src/shard/spsc_handler.rs) only wal_append (no replication plane in scope there; the TxnExecute hop only fires at num_shards > 1, where graph replication is out of scope by design regardless). New regression test replica_syncs_multi_exec_graph_leg (tests/replication_graph.rs): shards=1 master+replica, MULTI/SET+GRAPH.ADDNODE/EXEC on the master, asserts both legs land on the replica over one live stream (no reconnect/full-resync masking) — confirmed RED (graph leg missing) with just the new replication call disabled, GREEN 3× consecutive with the fix.

Fixed — cross-store MULTI/EXEC graph leg durability (kernel M3 stage 3 / task #52)

A committed cross-store transaction (MULTI; SET k v; GRAPH.ADDNODE g Label id 1; EXEC) survived kill-9 on the KV leg (PR #247's persist_txn_aof) but silently lost the graph leg — worse than a missed durability leg, the GRAPH.ADDNODE never even applied in memory. Root cause: execute_transaction_sharded (src/server/conn/shared.rs), the executor both the sharded and monoio connection handlers use for EXEC, had no GRAPH.* branch at all — a queued GRAPH.ADDNODE fell through to the generic KV dispatch() table (which only knows keyspace commands) and errored ERR unknown command, an error MULTI/EXEC's per-command tolerance swallowed without surfacing it to the client. Fixed by giving the txn executor a GRAPH.* branch that mirrors the single-command path (try_handle_graph_command): dispatches writes/reads against ShardSlice::graph_store, and flushes every command's drained wal-v3 records via wal_append after the whole transaction body runs (same per-shard append-order contract as the live path). Verified 3× consecutive GREEN at shards=1 and shards=4; the sibling atomicity claim (a transaction queued but killed before EXEC applies nothing) was unaffected and stays GREEN. Crash-matrix cells cross_plane_prod_s1_txn_isolated_committed / cross_plane_prod_s4_txn_isolated_committed (tests/crash_matrix_cross_plane/) flip from red_guard-gated RED to default-GREEN tripwires (42 cells: 33→35 GREEN by default, 9→7 RED).

Fixed — MQ durable-stream tombstone: DEL/UNLINK/FLUSHALL/FLUSHDB no longer resurrect on kill-9 (kernel M3 stage 3 / task #46)

Root cause: MQ durable streams live as ordinary keys in the shard keyspace, but replay_mq_wal had no way to represent "this queue was deleted after these pushes" — a generic DEL/UNLINK/FLUSHDB/FLUSHALL removed the stream from the live db and the DurableQueueRegistry, but a kill-9 afterward re-materialized the full pre-delete content on the next restart (replay reapplies every MqCreate/MqPush/... record regardless). Same bug class as the already-fixed KV/vector cold-plane resurrection (PR #257), now closed for MQ. Proven by the former RED crash-matrix cell cross_plane_seeded_red_mq_generic_del_resurrection (tests/crash_matrix_cross_plane/tests_seeded_red.rs), now un-gated (no more harness::red_guard) and green 3/3 consecutive runs.

Fix: a new MqDrop WAL v3 record (discriminant 0x75, src/persistence/wal_v3/record.rs; versioned encode/decode in src/mq/wal.rs — [version:u8][db_index:u32][key_len:u32][key:N], fails closed on any malformed/future-version payload like every other MQ record) is emitted whenever a durable MQ stream is removed via generic DEL/UNLINK/FLUSHDB/FLUSHALL — new hooks mq_exec::auto_drop_mq_streams / auto_drop_mq_streams_on_flush, wired into every connection-layer write path that already runs the equivalent vector/text index-parity hooks (handler_monoio/mod.rs, handler_sharded/mod.rs, the sharded MULTI/EXEC helper in server/conn/shared.rs, and the replica-side apply_index_parity_hooks in replication/apply.rs — a replica must tombstone its OWN WAL too, or it resurrects the stream on its own restart even though its live copy stayed correctly deleted). Boot-time replay applies apply_mq_drop (src/shard/shared_databases.rs) strictly in WAL order alongside every other MQ record, so a Drop only kills records that PRECEDE it for that key — a later MqCreate/MqPush for the same key survives intact (create -> drop -> create round-trips a kill-9 with the second incarnation whole; new regression tests test_replay_mq_wal_drop_only_kills_prior_records and the crash-matrix cross_plane_mq_create_drop_create_survives). segment_plane_scan's plane-history block set (src/persistence/wal_v3/segment.rs) gained MqDrop alongside the other MQ discriminants so autovacuum/recycle never deletes a sealed segment still holding an unfloored tombstone. The MQ WAL fuzz target (fuzz/fuzz_targets/mq_wal_record.rs) now also fuzzes decode_mq_drop.

Replication decision: MQ effect records replicate live only at num_shards == 1 (the pre-existing gate, matching MqCreate/etc — see mq_exec::replicate_mq_record's docs). MqDrop follows the same posture: MQ._REPL.DROP is emitted and applied by replication::apply::apply_mq exactly like the other MQ replay commands. Separately — and regardless of that gate — a REPLICATED generic DEL/UNLINK/FLUSHDB/FLUSHALL already removes the stream from the replica's live keyspace via normal command replication; what it did NOT do before this fix is tombstone the replica's OWN wal-v3 MQ plane, so the replica's live state was correct but its restart-durability was not. apply_index_parity_hooks now closes that gap identically on the replica side.

Scope decision (documented in code, not re-litigated per-callsite): DurableQueueRegistry entries are NOT db-indexed (pre-existing limitation, same as MqCreate's registry) — a key match tombstones regardless of which db the deleting command ran in, and FLUSHDB drops every registered durable queue exactly like FLUSHALL (the registry has no way to scope to one db). Two different dbs sharing an MQ queue NAME is not a supported configuration.

Test-gotcha fixed along the way: the crash-matrix DEL/FLUSHALL scenarios' original sync-marker strategy waited on an AOF-family write to prove the delete was durable — correct pre-fix (DEL only touched the AOF), but the new MqDrop record lands on the wal-v3 MQ plane via a separate fire-and-forget channel drained on its own 1ms tick, so an AOF-only marker no longer proves anything about the tombstone's own durability. Both tests now also sync a throwaway durable queue's MQ.PUSH (wal-v3-family) after the delete/flush before crashing.

Fixed — adopt shared atomic_write_durable at remaining bare-write persistence sites (kernel M4 prep / task #49)

Six durable-state writers still hand-rolled their own (often incomplete) tmp-write/rename sequence instead of the shared K3 primitive (src/persistence/atomic.rs, PR #296): acl::io::acl_save and the ACL SAVE command handler (src/command/acl.rs) had zero atomicity at all (bare fs::write + rename, no fsync); CONFIG REWRITE (src/command/config.rs), cluster::migration::save_nodes_conf (nodes.conf), and replication::save_replication_state all did write+rename with no fsync, so a kill-9 immediately after rename() returns could still revert the file to stale/empty contents on data=ordered ext4/xfs; persistence::clog::write_clog_page and persistence::kv_page::write_datafile/write_datafile_mixed wrote directly to the FINAL path with no temp file or rename at all (only a trailing sync_all/fsync_file), so a crash mid-write could leave a torn CLOG page or KV data file. The native RDB save paths (persistence::rdb::save/save_from_snapshot, persistence::redis_rdb::save, used by BGSAVE) also lacked the parent-directory fsync step. All eight sites now route through atomic_write_durable (temp file, sync_all, rename, dir-fsync); each conversion is paired with a "no leftover temp file after a successful write" regression test. AOF appends, WAL segments, and the legacy #[allow(dead_code)] rewrite_aof_sync RDB-preamble path are out of scope (different framing/fsync mechanisms, per task brief).

Fixed — task #45: tick-path eviction now spills collections instead of plain-dropping them

Root cause of the intermittent tests/cold_collection_visibility.rs CI flake (stable GHOST: EXISTS == 1, LRANGE/ZRANGE/etc. permanently empty — only observed on shared/slow Linux runners, never test-induced). Two compounding defects in src/storage/eviction.rs's synchronous spill path (evict_one_with_spill):

  1. Collection victims (Hash/List/Set/ZSet/Stream) were gated behind a stale is_string check that predates kv_spill::spill_to_datafile gaining full collection support (it already serializes any RedisValueRef via kv_serde) — so a collection picked as a victim while a SpillContext WAS present still fell through to a silent plain-drop, indistinguishable from spill: None.
  2. Even for strings, a successful sync spill never registered a ColdIndex entry (spill_to_datafile was always called with cold_index: None) — the durable .mpf file existed and was manifest-registered, but nothing could ever read it back via promote_cold_if_present.

Separately, src/shard/timers::run_eviction (the periodic 100ms tick, independent of the memory-pressure cascade) never received a SpillContext at all — every victim it picked was plain-dropped even with --disk-offload enable and a durability backstop (--appendonly yes) configured. src/shard/persistence_tick.rs's run_eviction_tick now builds a real SpillContext (from the shard's ShardManifest + shard data dir + next_file_id) whenever disk-offload is enabled and a manifest is present, and threads it through — falling back to the pre-existing fail-close plain-drop (PR #273 policy-aware discipline: noeviction still OOMs, an evicting policy still frees RAM) when no durability backstop exists, matching the already-documented "spill is inert without one" rule.

New deterministic regression tests (no server process, no timing race): storage::eviction::tests::sync_spill_non_string_victim_is_durably_spilled_not_plain_dropped and shard::timers::tests::test_run_eviction_spills_collection_victim_when_spill_context_given. tests/cold_collection_visibility.rs (10x local reruns, both runtimes) remains green and unmodified — its GHOST assertions still fail the suite on any regression.

Added — unified per-shard floor register + min-across-planes WAL recycle (kernel M3 stage 2 / K2)

ShardControlFile (src/persistence/control.rs) gains graph_floor_lsn, ws_floor_lsn, mq_floor_lsn: u64 fields (control payload 57→81 bytes, backward-compatible: fields default to 0 when reading a control file written by an older binary, keyed off the on-disk payload_bytes header). Both WAL-recycle call sites — checkpoint Finalize (src/shard/persistence_tick.rs) and autovacuum Pass C (src/shard/autovacuum.rs, wired via a new control_file parameter on run_tick) — now recycle up to min(kv_floor, graph_floor) instead of the KV-only floor, so a WAL segment is never recycled while it still holds the only copy of an unflushed graph record. WS/MQ floors are tracked but deliberately excluded from the min() (brief's Risk #2): both planes have no snapshot format in any mode, so folding their sentinel 0 floor into the minimum would permanently freeze recycling on any shard that ever saw a WS/MQ record. Their content stays protected by the existing, orthogonal segment_plane_scan content-scan (renamed from segment_holds_plane_history; GraphTemporal removed from its blocking match arm now that the LSN floor covers that record type, with a new RECL_WAL_RECYCLE_GRAPH_TEMPORAL_FREED_TOTAL counter in the # Reclamation INFO section marking segments it stops blocking).

Task #53 root-caused and fixed in the same stage (required by this stage's mandate: fix in-scope if the root cause is the Finalize floor/durability-ordering invariant K2 formalizes). The checkpoint-Finalize-window graph-total-loss finding (tests/crash_matrix_cross_plane.rs's former RED cell 4) was exactly that invariant: save_graph_store (src/graph/recovery.rs) wrote the reference/floor (graph_metadata.json, via GraphStore::save_metadata) BEFORE the payload it claims durable (CSR segments + manifest.json) — an ARIES-inverted write order. A kill-9 between the two writes left a fully-advanced floor pointing at a payload that never reached disk, so recovery trusted the floor and skipped WAL replay for records it claimed were already covered, total-losing the graph batch. Fixed by reordering save_graph_store to write CSR segments + manifest.json first, store.save_metadata last, and making save_metadata itself atomic (temp+fsync+rename+dir-fsync via the shared persistence::atomic::atomic_write_durable helper) so the floor write can never itself be torn. Safe against double-replay: graph::replay::node_present checks both write_buf and loaded CSR segments before re-inserting a WAL-logged node/edge, so replaying a record whose payload actually made it to disk before the kill is a no-op, not a duplicate. Confirmed via MOON_CRASH_MATRIX_RED=1 MOON_CRASH_MATRIX_ITERS=20 soak: 20/20 clean at both prod_s1 and prod_s4 (pre-fix baseline: hit at iteration 11/20 and 7/20 respectively). cross_plane_prod_s1_mixed_all_planes_mid_checkpoint / cross_plane_prod_s4_mixed_all_planes_mid_checkpoint now run ungated — the crash-matrix suite's default-GREEN count moves from 29/40 to 31/40 (9 RED cells remain: task #52's cross-store TXN graph leg and the legacy-mode graph-reconstruction/MQ-resurrection findings above, all still tracked separately and explicitly out of scope for this stage). A subsequent adversarial review round found a second graph-durability P0 plus a VACUUM floor bypass on top of this fix — see the next section — which add 2 more GREEN cells, moving the suite to 42 cells total (33 GREEN by default).

src/persistence/recovery.rs's PITR path threads the three new floors through unchanged on replay. New unit tests: control-file round-trip + backward-compat (old-format read defaults new floors to 0), an inverted GraphTemporal recycle test (segment recycles once the graph floor covers it even though the plane-scan no longer blocks it), and a Risk #2 regression test (a pure-KV write after a WS/MQ record on the same shard still recycles — the WS/MQ sentinel floor must never collapse recycling to zero).

No per-write hot path touched — ShardControlFile is only written at checkpoint Finalize and read at Pass C/recovery, both off the command dispatch path; a VM hot-path A/B bench was judged unnecessary for this change and not run.

Fixed — K2 adversarial review round: graph drop-resurrection + VACUUM floor bypass (kernel M3 stage 2)

An adversarial review of the floor-register work above (SHIP-WITH-FIXES verdict) found two P0s and a P1 in the same area before the stage could close:

P0 — GRAPH.DELETE'd graphs could resurrect across repeated checkpoints. persist_graph_at_checkpoint's short-circuit (!store.is_dirty() || store.graph_count() == 0 -> return true) skipped save_graph_store whenever a delete emptied the graph map — even though the delete itself marks the store dirty via the WAL drain. graph_count() == 0 and "nothing to persist" are independent conditions: an empty-but-dirty store still needs graph_metadata.json rewritten to reflect zero graphs at the new snapshot_lsn. With the skip in place, metadata stayed stale forever (dirty never clears without a real save) while every later checkpoint kept advancing control.graph_floor_lsn past the WAL record holding the DELETE — once that segment was recycled, a crash+restart loaded the stale metadata with nothing left in the WAL to replay the deletion. Fixed by dropping graph_count() == 0 from the short-circuit; dirty alone gates it now, and save_graph_store's per-graph loop correctly no-ops on zero graphs while store.save_metadata still durably rewrites the (now-empty) graph list. New regression cells cross_plane_prod_s1_graph_drop_survives_repeated_checkpoints / _s4_... (tests/crash_matrix_cross_plane/scenarios.rs), RED-first verified against the pre-fix binary at both shard counts (two earlier padding designs — a live second graph, then MQ pushes — were each proven NOT to reproduce the bug before a third, using throwaway create+delete graph cycles as padding, did).

P0 — manual VACUUM bypassed the unified floor register. run_vacuum_passes (src/command/server_admin.rs) recycled WAL to w.current_lsn() unconditionally — a client-reachable path around the min(kv_floor, graph_floor) invariant K2 established for the checkpoint and autovacuum recycle sites. Fixed by threading the control-file floors into VACUUM: in checkpoint-backed (disk-offload) mode it now recycles to min(last_checkpoint_lsn, graph_floor_lsn) read from the shard's ShardControlFile, mirroring autovacuum Pass C; in legacy mode (no disk-offload dir — no snapshot format to bound WS/MQ recycling against) it now REFUSES to recycle WAL at all and increments the shared RECL_WAL_RECYCLE_BLOCKED_NO_CHECKPOINT_TOTAL counter (surfaced in INFO/DEBUG RECLAMATION's # Reclamation section), the same skip-and-warn shape Pass C already uses for the identical precondition.

P1 — the WAL-overflow emergency recycle path (src/shard/ persistence_tick.rs) used last_checkpoint_lsn alone; now .min(control.graph_floor_lsn), closing the same gap in the emergency code path.

Plus two P2s: GraphManifest::save (src/graph/manifest.rs) now routes through the shared atomic_write_durable helper instead of a hand-rolled temp+fsync+rename+dir-fsync sequence (behavior-equivalent); and a stale doc comment on scenarios::mixed_mid_checkpoint that still described the Finalize-ordering bug as unfixed was corrected.

Re-gated after these fixes: the crash-matrix suite (42 cells, 33 GREEN by default, including the 2 new regression cells), mixed_mid_checkpoint MOON_CRASH_MATRIX_RED=1 MOON_CRASH_MATRIX_ITERS=20 soak at both shard counts (re-verified since the P0 fix touches the same Finalize path), g4/g5 graph durability cells, cargo fmt --check, and both clippy matrices (default features; runtime-tokio,jemalloc).

Added — cross-plane kill-9 crash-matrix suite (kernel M3 stage 1 / G1)

New tests/crash_matrix_cross_plane.rs + tests/crash_matrix_cross_plane/ harness: 40 kill-9 crash-recovery cells spanning every persistence plane (KV/AOF, KV disk-offload, graph/vector/WS/MQ WAL v3, cross-store MULTI/EXEC, mixed-plane concurrent workloads, and a checkpoint-Finalize-window kill) across the {shards=1, shards=4} x {appendonly, disk-offload} matrix. This is a RED-first tripwire suite — 29 cells run GREEN by default; 11 known-RED cells are gated behind a runtime guard (harness::red_guard, checked via MOON_CRASH_MATRIX_RED=1) rather than #[ignore = "RED: ..."], because the suite's own --ignored invocation convention (every cell needs a release binary) makes an ignore-reason string alone non-functional as a gate. A MOON_CRASH_MATRIX_ITERS env var (default 1) enables ad-hoc soak runs for probabilistic findings.

A second review round caught a vacuous-precondition bug and a gating granularity bug before this suite could be trusted as a tripwire: kv_spilled_isolated's filler (3000×512B against a 4 MiB cap) never actually forced a cold-tier spill at either shard count, silently degenerating to the same coverage as kv_isolated — fixed by mirroring crash_recovery_cold_del_resurrection.rs's proven filler ratios (16,000×600B against 8 MiB) plus a hard heap-*.mpf precondition assert so a future load-parameter regression fails loud instead of vacuously passing. txn_isolated was split into txn_isolated_committed (still RED, task #52 below) and txn_isolated_atomicity (a transaction queued but killed before EXEC must apply nothing) — the two halves have unrelated root causes and sharing one red_guard hid a working, GREEN regression tripwire on prod_s1/prod_s4 by default; the atomicity half now runs ungated on those two configs (still RED on legacy_yes_s1, same pre-existing legacy-graph gap as everything else there).

Two real, unresolved findings surfaced and are intentionally left RED (not fixed — that is out of scope for this stage) for the kernel M3 backlog:

  • Cross-store TXN graph leg not durable (task #52): a committed MULTI/EXEC mixing SET + GRAPH.ADDNODE survives kill-9 on the KV leg (PR #247's persist_txn_aof) but the graph leg is lost. Deterministic repro (3x consecutive at shards=1 and shards=4): kill immediately after the post-EXEC WAL-v3 sync wait, with --checkpoint-timeout 3600 so no periodic checkpoint can incidentally rescue the graph leg. See scenarios::txn_isolated_committed's doc comment for the minimal repro.
  • Checkpoint-Finalize window can total-loss the graph plane (new): some kill offsets in the 0-150ms window after BGSAVE lose the entire already-wal-v3-synced graph batch (not just an unsynced tail) while KV/vector/WS/MQ all survive in the same run. Probabilistic (~1-in-7 to 1-in-12 per iteration) — a single default run can report false-green even with MOON_CRASH_MATRIX_RED=1; reproduce reliably with MOON_CRASH_MATRIX_RED=1 MOON_CRASH_MATRIX_ITERS=20. Strong P0 candidate for kernel M3 stage 2 (K2). See scenarios::mixed_mid_checkpoint's doc comment.

Also confirmed GREEN (not RED, contrary to the brief's grouping at the design-doc level): WS DROP durability under kill-9 — WorkspaceDrop WAL records replay correctly in order. The sibling MQ generic-DEL resurrection gap (same bug class as the already-fixed KV/vector cold-plane resurrection, PR #257) was RED at the time of this entry; fixed in kernel M3 stage 3 (task #46) — see the MqDrop entry below.

tests/common/mod.rs gains find_moon_binary()/sigkill()/ wait_for_port_down() shared helpers this suite depends on (the latter now panics on loop exhaustion instead of silently returning — a tripwire codebase should not have silent-pass verification helpers). Shared helpers added; migration of the 5 existing crash_recovery_*.rs suites' own local copies to the shared versions is deferred, not done in this stage.

Test-only; no production code changed in this stage.

Added — memory + tier accounting spine (kernel M2 stage 2 / K4)

resident_bytes() is now implemented by every storage plane, and the elastic memory budget's used-term (ShardDatabases::recompute_elastic_budget, GAP-1/PR #170) folds in all of them — previously kv+vector only.

  • TextStore/TextIndex::resident_bytes(): posting lists, term dictionaries, FST fuzzy/prefix sidecars, per-document bookkeeping maps, and TAG/NUMERIC secondary indexes. FTS memory was hard-coded 0 everywhere it was published (elastic budget, MEMORY DOCTOR, Prometheus) until this change. Data-size-independent incremental accumulator (the publish-site read sums cached per-structure totals — O(schema field count), bounded by FT.CREATE definitions, never corpus size; same contract as ColdIndex/graph below) — an initial version was an O(doc-count + vocabulary) full-recompute walk invoked unconditionally every 100ms from the shard eviction tick regardless of maxmemory, measured 6.4–21.3ms/call at 50K–200K docs (>20% of the tick budget, recurring P99 spikes for every command on that shard). Fixed before merge: PostingStore/ TermDictionary/TextIndex each carry a cached total maintained incrementally at every mutation site (index/delete/upsert/TAG/NUMERIC update/FST rebuild), verified against a #[cfg(test)] ground-truth full-walk after a mixed mutation sequence.
  • ColdIndex::resident_bytes() (KV disk-offload bookkeeping): an O(1) incremental accumulator (not a per-tick walk — sized for G2's "10x RAM" scale target), charged into the shard's published KV memory.
  • Graph resident bytes (already computed) now also feed the elastic budget's used-term, not just the observability atomic.
  • moon_memory_bytes{kind="text"} Prometheus gauge + Text (FTS): line in MEMORY DOCTOR. Also fixes a pre-existing gap where moon_memory_bytes{kind="lua_scripts"} was emitted but never primed.
  • New src/storage/tier.rs: ResidencyTier (Hot/WarmReloadable/ColdStub)
  • TierPolicy trait skeleton — types only, no plane adoption in this milestone (that is M4).

No eviction policy semantics changed: this widens what the existing donor/hot formula sees, not how it decides. Verified against eviction_parity/eviction_parity_hash_disk_offload (shards 1 and 4, including the disk-offload ColdIndex path) with no behavior change.

Fixed — Windows CI: replication_planes used un-gated libc::kill

tests/replication_planes.rs's sigkill helper called libc::kill unconditionally — libc is not linked on Windows, so the Check (Windows) job failed to compile the suite on every main push since the Wave A merge. Now cfg-gated exactly like tests/aof_multidb_kill9.rs (Child::kill on non-unix). Test-only.

Fixed — graph CSR segments and text/vector sidecars were not crash-durable (K3, storage-kernel M2 stage 1)

Extracted the vector engine's Stack-B temp+fsync+rename+dir-fsync sequence into one shared atomic_write_durable(path, bytes) (src/persistence/atomic.rs) and adopted it at every durable-artifact write site the K3 audit flagged as missing part of the sequence (storage-audit-2026-07-12-graph-fts.md):

  • Graph CSR segments (CsrSegment::write_to_file) used a bare std::fs::write — no fsync at all, and the final path was overwritten in place instead of replaced via rename. A crash mid-write left a torn segment directly at the path GraphManifest already references. Proven with a concurrent-writer regression test that reliably caught a torn read (InvalidData("data shorter than header")) against the pre-fix code.
  • Text sidecars (save_text_index_metadata, save_fst_sidecar) and the vector index metadata sidecar (save_index_metadata_v3) already did temp-write + fsync + rename, but never fsync'd the parent directory — on data=ordered-journaled filesystems a crash between rename and dir-fsync can revert the directory entry to the old name even though rename() returned success.
  • FST term-dict sidecar deliberately left unwired (and now documented as such on TextStore::load_fst_sidecars): wiring the load was attempted and REVERTED after adversarial review proved a loaded FST's baked-in term-ids are stale id-space garbage against the term dictionary the restart rescan rebuilds (first-encounter- order ids over non-reproducible hash iteration) — merging them silently corrupts FT.SEARCH FUZZY/PREFIX results. Restart fuzzy/ prefix queries keep the slow-but-correct brute-force path; making the sidecar loadable requires persisting the term dictionary itself (task #50, kernel M4 FTS persistence).

Ref: .planning/reviews/kernel-m2-brief-2026-07-12.md (K3, stage 1).

Fixed — read-only replicas accepted client-issued WS/MQ/TEMPORAL.* writes (Wave B readonly-enforcement, task #34 follow-up)

WS and MQ (dispatched as WS <SUB> ... / MQ <SUB> ...) and the dotted TEMPORAL.SNAPSHOT_AT / TEMPORAL.INVALIDATE commands were entirely absent from command::metadata::COMMAND_META, so metadata::is_write silently returned false for all of them. try_enforce_readonly (handler_monoio::dispatch / handler_sharded::dispatch) is driven purely by is_write, so a client could issue WS CREATE/WS DROP, MQ CREATE/PUSH/POP/ACK/TRIGGER/PUBLISH, and TEMPORAL.SNAPSHOT_AT/TEMPORAL.INVALIDATE directly against a read-only replica — mutating its workspace registry, message queues, or graph temporal state (TEMPORAL.INVALIDATE sets valid_to) with no path back to the master. A silent divergence source, same class of bug fixed for ACL in PR #258.

WS and MQ now carry the blanket WRITE flag (their read-only subcommands — WS LIST/INFO/AUTH, MQ DLQLEN — are carved out at the try_enforce_readonly call site via the new command::workspace::is_ws_readonly_subcommand / command::mq::is_mq_readonly_subcommand classifiers, mirroring the existing SELECT/GRAPH.QUERY blanket-write-with-carve-out pattern). Both TEMPORAL.* commands are unconditionally WRITE (no subcommand token, so no carve-out is needed). The replica apply path is unaffected — it never calls try_enforce_readonly by construction (see the module doc on replication::apply), so replaying a master's own WS/MQ/TEMPORAL records on a replica still works.

Covered by unit tests in command::metadata, command::workspace, and command::mq, plus a new black-box integration test, tests/replication_readonly_ws_mq.rs (--ignored, spawns a real moon binary — WS/MQ are only wired in the sharded handlers, so the synthetic single-shard harness used by tests/replication_test.rs can't exercise this fix). Verified against both a monoio (default-feature) build and a runtime-tokio,jemalloc build.

Added — WS.CREATE/DROP replication (Wave B ws-plane)

WS.CREATE/WS.DROP mutated the process-global WorkspaceRegistry from whichever connection thread received the command and pinned their WAL records to shard 0 without actually running on shard 0's thread — the replication offset advance (which must happen synchronously with the mutation on shard 0, matching the snapshot-capture atomicity argument used elsewhere) had no such guarantee, and the plane was not replicated at all (round-2 finding A / task #34).

Routed the whole mutation + WAL-append + replication-record sequence through a dedicated shard-0 hop (ShardMessage::WsRegistryCreate / WsRegistryDrop, mirroring the existing WsDropCleanup hop): a connection already on shard 0 runs the sequence inline, every other connection hops over via spsc_send + a bounded oneshot reply. After the shard-0 execution, the WorkspaceCreate/WorkspaceDrop record (same payload as the WAL record — ws_id + name + created_at_ms) is pushed into the replication stream via the graph-plane's record_local_write_db pattern (offset advances IFF the backlog append happens).

Replica apply arms (WS.CREATE.APPLY / WS.DROP.APPLY internal pseudo-commands, replication::apply::apply_local) install the MASTER's ws_id + created_at VERBATIM — UUIDv7 is nondeterministic, so a replica must never mint its own id. WS.DROP.APPLY reruns the master's best-effort {ws_hex}:-prefix key sweep on the replica's single shard (R0 replication is single-shard-only). Full-resync backfill: a new versioned MOON_AUX_WORKSPACE_REGISTRY RDB aux blob (replication::ws_sync, shard-0-authoritative — captured only in shard 0's PrepareReplicaSync leg on the multi-shard master path) is installed authoritatively at load_snapshot, same convention as the graph/vector/text aux blobs. New ws_registry_record cargo-fuzz target covers the blob decoder.

The former combined warn_unreplicated_plane fail-loud marker (handler_monoio/ft.rs) is fully retired: its WS half by this WS-plane work, and its MQ half by the MQ-plane replication entry below.

New tests/replication_ws.rs: live-stream parity (id + created_at round-trip through WS.INFO/WS.AUTH on the replica), snapshot-leg backfill, WS.DROP propagation, and a --shards 4 leg exercising the shard-0 hop from connections that may land on any shard.

Added — MQ-plane replication (Wave B stage 2b)

Durable-queue MQ. mutations now replicate to attached replicas, closing the plane gap warn_unreplicated_plane previously fail-loud-warned about for MQ.. Builds on the stage-2a MQ WAL effect records (PR #291):

  • Live stream: shard::mq_exec::replicate_mq_record emits the SAME encoded MqCreate/MqPush/MqPop/MqAck/MqTrigger payload bytes already built for the WAL as one of five synthetic MQ._REPL.* pseudo-commands (never dispatched to real clients), gated single-shard-only (num_shards == 1) — the same posture graph's own live path takes for its multi-shard gap. MQ.PUBLISH's TXN-materialization leg (server::conn::handler_monoio::txn.rs) replicates the same way at commit. replication::apply::apply_mq decodes and applies through the SAME apply_mq_* engine boot-time WAL replay uses (shard::shared_databases, generalized over a new MqApplyTarget trait so ShardSliceInit boot replay and ShardSlice live apply share one codec instead of two).
  • FULLRESYNC: a per-shard MOON_AUX_MQ_REGISTRY aux blob (replication::mq_sync) carries the durable-queue registry and trigger registry — the shard-level bookkeeping that lives outside the keyspace. Installed additively into every replica shard (mq_sync::install_mq_registry_many), mirroring graph_sync::install_graph_store_many.
  • Fixed a pre-existing FULLRESYNC data-loss bug found along the way: the PSYNC RDB codec (persistence::redis_rdb, distinct from the persistence::rdb SAVE/BGSAVE codec) serialized Stream values as a placeholder "__stream__:<len>" string, discarding all entries/PEL/ consumer-group state. A full RDB_TYPE_STREAM_MOON (0xC8) codec — ported from persistence::rdb's battle-tested Stream serializer — now carries entries, last_id, consumer groups (PEL + per-consumer pending), the durable flag, and max_delivery_count. Without this fix, MQ FULLRESYNC backfill would silently lose every durable queue's message content. Operator note — upgrade masters and replicas together if any durable Streams exist: a pre-fix replica syncing from a fixed master hard-fails FULLRESYNC on the new RDB_TYPE_STREAM_MOON tag (and retries on its reconnect loop until upgraded), while a fixed replica syncing from a pre-fix master absorbs the old placeholder as a corrupted string value — both inherent to fixing a wire-format bug, neither loses data already durable in the master's WAL.
  • Two new cargo-fuzz targets: mq_registry_blob (the new install_mq_registry_many decoder) and redis_rdb_load (the FULLRESYNC RDB loader as a whole, including the new Stream type tag) — wired into both PR and nightly fuzz matrices.
  • tests/replication_mq.rs: live-stream, snapshot-backfill, and a pinned multi-shard-live-write-not-streamed limitation test (shards=1 scope, matching graph's own precedent).

A multi-shard master still does not live-stream MQ writes (durability is unaffected — WAL + FULLRESYNC still cover it); tracked as a known follow-up alongside graph's own multi-shard live-stream gap.

docs/guides/tuning.md (PR #245) and docs/PRODUCTION-CONTRACT.md (PR #263) linked repo-root files (BENCHMARK.md, RELEASES.md, .github/workflows/release.yml) via relative paths that escape the MkDocs docs/ tree — mkdocs build --strict aborts on the 5 resulting warnings, so the GitHub Pages deploy had failed on every main push since 2026-07-08. Converted the 5 links to absolute GitHub blob URLs; strict build verified clean locally. Docs-only.

Fixed — MQ effect records now survive kill-9 (Wave B stage 2a, task #34)

MQ.CREATE/PUSH/POP/ACK/TRIGGER intercept before the generic AOF-logging dispatch path (they route via execute_mq_on_owner in src/shard/mq_exec.rs), so under --appendonly yes the AOF-authority recovery (db.clear() on every shard, rebuild solely from the AOF manifest) silently discarded every durable Stream, consumer-group PEL, DLQ routing decision, and trigger registration on every restart — regardless of whether replay_mq_wal itself was correct. Two additional pre-existing defects in replay_mq_wal made it dangerous even when reached: MqAck records were applied by COUNT (rolling the whole PEL back to a snapshot cursor) instead of by ID, and every MQ payload hardcoded db index 0.

Fixed by giving each MQ mutation its own versioned WAL v3 effect record, emitted at the owner-shard execution site: MqPush (0x72), MqPop (0x73), and MqTrigger (0x74) are new discriminants; MqCreate (0x70) and MqAck (0x71) keep their existing discriminants but move to a versioned, db-index-carrying payload (precedent: the XactCommit 0x51→0x53 format freeze — WAL v3 segments are short-lived with no cross-version contract). Every decoder in the new src/mq/wal.rs module returns None on a malformed OR unsupported-version payload; replay_mq_wal skip-and-warns (tracing::warn!) rather than aborting the scan. MqPop now carries the full claim set (id + delivery_count per claimed message) plus any DLQ routing decisions (source id → assigned DLQ id), so replay reconstructs the consumer group's PEL and last_delivered_id exactly instead of guessing. MqAck now applies via Stream::xack by id — idempotent, and immune to the old count-based rollback bug. MQ.PUBLISH's TXN materialization hop also emits MqPush records, at both the self-fold and foreign-shard legs, on both the monoio and tokio connection handlers.

Trigger registrations are durable/replayed as opaque data — replay never fires a trigger; only a live MQ.PUSH's debounce arming does (src/shard/timers.rs::fire_pending_mq_triggers).

New RED/GREEN kill-9 crash test: tests/crash_recovery_mq_effects.rs (--ignored, needs a built binary) exercises the full lifecycle — CREATE → PUSH×5 → POP(3) → ACK(2 of 3) → a second CREATE/PUSH/POP pair that forces immediate DLQ routing → a TRIGGER registration — kill -9, restart on the same --dir, and asserts stream content, delivery cursor, PEL-by-id, DLQ routing, and trigger re-arming all survive. New fuzz target mq_wal_record covers the five new op-blob decoders (fuzz/fuzz_targets/, registered in both fuzz-pr and fuzz-nightly CI matrices).

Measured WAL footprint (release, single shard): CREATE + 5×PUSH + POP(3) = 7 records, 576 bytes on disk including the 64-byte segment header — ~68 bytes per small PUSH (48-byte payload + 20-byte framing/CRC), ~132 bytes for a 3-claim POP.

WAL recycle plane guard (adversarial-review fix): recycle_aggressive and recycle_segments_before now refuse to delete any sealed segment holding workspace/MQ/temporal records — those planes have no snapshot format in ANY mode, so the WAL is their sole durable copy and no caller's LSN floor (autovacuum Pass C in disk-offload mode, the checkpoint protocol, admin VACUUM) makes such a segment safe. Previously, disk-offload deployments under --max-wal-size pressure would silently and permanently lose MQ/WS/temporal history during normal operation, no crash required. Kept segments are counted in reclamation_wal_recycle_blocked_no_checkpoint_total and warned (rate-limited); WAL size may exceed --max-wal-size for plane-heavy workloads until plane checkpointing lands (storage-kernel M3). The guard fails closed: unreadable/torn/unknown-type segments are kept.

Known limitation (pre-existing, tracked as task #46): durable MQ streams deleted via generic DEL/UNLINK/FLUSHALL/FLUSHDB are resurrected — now with full content — by MQ WAL replay after a restart; there is no MQ tombstone record yet (same bug class as the fixed vector/KV cold-plane resurrection). Fixed in kernel M3 stage 3 — see the "MQ durable-stream tombstone" entry below.

Out of scope (stage 2b+): replication emission/apply for the MQ plane.

Changed — WAL v3 wal_append channel now preserves the caller's REAL record type end-to-end (K1a, storage-kernel M1 stage 1)

ShardDatabases::wal_append / try_wal_append_required and mq_exec::wal_append_on_slice took a plain Bytes payload; the shard event-loop's 1ms-tick drain (event_loop.rs, two sites) unconditionally re-wrapped every message as an outer WalRecordType::Command record. Every non-Command producer (XactCommit, WorkspaceCreate/WorkspaceDrop, MqCreate/MqAck, GraphTemporal) worked around this by pre-framing its own record with write_wal_v3_record and sending the ALREADY-FRAMED bytes through — a second, nested WAL frame inside the outer Command frame, recovered on replay via a bespoke read_wal_v3_record-on-payload unwrap per record type. One nested-framing call (handler_monoio/txn.rs / handler_sharded/txn.rs XactCommit) additionally passed the transaction's txn_id into the inner frame's lsn field, mislabeling it.

Fixed structurally: the channel item is now (WalRecordType, Bytes) — the producer's real type plus the UNFRAMED payload — threaded through ShardSlice::wal_append_tx / ShardDatabases::wal_append_txs end-to-end. The drain calls wal.append(record_type, &payload) directly, so the WAL writer assigns the real LSN and does the single framing. Every producer's pre-framing (write_wal_v3_record call) was deleted: handler_monoio/handler_sharded write.rs (WS.CREATE/DROP) and txn.rs (XactCommit — the txn_id-as-lsn mislabel is gone with it), shard/uring_handler.rs's WS batch path, shard/mq_exec.rs (MqCreate/MqAck), and command::temporal::apply_invalidate (GraphTemporal — PR #286 had just added this record's pre-framing; it is deleted in favor of the typed channel, and the function now RETURNS the raw payload instead of pushing it into GraphStore::wal_pending, which stays exclusively Command-typed RESP bytes). The one caller that cannot change its Frame-only return signature without a ~90-call-site test ripple (command::graph::dispatch_graph_command, the cross-shard GraphCommand entry point) stashes the returned payload in a new GraphStore::temporal_wal_pending side-channel instead, which its two real callers (shard/spsc_handler.rs) .take() right after dispatch.

Replay is fully backward compatible, no format bump: every existing direct-type match arm (replay_workspace_wal, replay_mq_wal, replay_temporal_wal in shared_databases.rs) already had a nested-Command unwrap fallback (added when the records were still nested) plus, for GraphTemporal, a legacy-raw fallback for even older un-nested records — both are UNTOUCHED, so segments written before this fix keep replaying exactly as before. New records simply hit the direct-type arm for the first time instead of falling through the unwrap.

persistence::recovery.rs Phase 4's on_command closure had no XactCommit arm at all (silent _ => {}) — now reachable for the first time because the outer type used to always be Command. Decision: an EXPLICIT documented no-op (counts toward commands_replayed, never dispatches). The forward-image KV payload (encode_xact_commit_payload) is redundant in every reachable config — the transaction's individual SET/DEL ops already ride either this same Phase-4 WAL replay as ordinary Command records (--wal-kv-log on) or the AOF, the KV recovery authority in every config (Phase 4b falls back to it whenever kv_commands_replayed == 0). Decoding it would be a no-op at best and risks double-applying a non-idempotent op at worst — the same overlap risk the adjacent Phase 4b comment already flags. wal_v3::replay::replay_wal_v3_dir_commands (the legacy last-resort-fallback path, used only when no AOF exists) already had a correct replay_xact_commit call for the real outer type — untouched, and now reachable for the first time too since it stops receiving records nested inside Command.

Fixed — WS.CREATE's created_at was silently dropped by the WAL, restored as 0 after every restart (K1b, storage-kernel M1 stage 1)

encode_workspace_create/decode_workspace_create (src/workspace/wal.rs) only ever serialized [ws_id][name_len][name] — the created_at computed at WS.CREATE time never reached the payload, so replay_workspace_wal (shared_databases.rs) had no choice but to hardcode created_at: 0 on every restart, silently losing the real creation time. Fixed with a versioned, backward-compatible layout: new records append a trailing created_at_ms: i64 LE; the decoder accepts BOTH the old (20 + name_len-byte) and new (20 + name_len + 8-byte) lengths, returning created_at = 0 for old records (matching prior restart behavior exactly — no format bump, mixed-segment compatible). All three producers (handler_monoio/handler_sharded write.rs, shard/uring_handler.rs's batch path) now pass the already-computed created_at through; the replay decoder threads the real value into WorkspaceMetadata instead of the hardcoded 0.

Fixed — autovacuum Pass C recycled the sole durable copy of graph/WS/MQ/temporal WAL history in legacy mode (task #43, P1)

AutovacuumDaemon::run_tick's Pass C (src/shard/autovacuum.rs) called WalWriterV3::recycle_aggressive(redo_lsn = wal.current_lsn()) — the LSN about to be assigned, i.e. "everything sealed is durable elsewhere" — whenever total on-disk WAL bytes exceeded --max-wal-size (default 256 MiB), regardless of whether disk-offload was enabled. That floor only holds in disk-offload mode, where the checkpoint protocol maintains a real redo_lsn and persist_graph_at_checkpoint snapshots the graph write-buffer before the floor advances. In legacy (non-disk-offload) mode there is no periodic plane snapshot at all: save_graph_store runs only at graceful shutdown, and the workspace/MQ/temporal registries have no snapshot format whatsoever — replay_workspace_wal / replay_mq_wal / replay_temporal_wal rebuild them purely by re-scanning the WAL. Once the ceiling was breached, Pass C deleted every sealed segment unconditionally, including segments holding the sole durable copy of that history; a kill-9 afterwards lost it permanently, even though replay itself (task #42, PR #286) works correctly for whatever records survive.

Fixed per the "never trade data loss for disk space" principle: run_tick now takes a disk_offload_enabled flag. When true, behavior is byte-for- byte unchanged (recycle_aggressive against current_lsn(), as before). When false, Pass C no longer recycles: it skips the delete, emits a rate-limited (5 min) tracing::warn!, and increments the new reclamation_wal_recycle_blocked_no_checkpoint_total INFO counter (RECL_WAL_RECYCLE_BLOCKED_NO_CHECKPOINT_TOTAL) so operators can see WAL growth is unbounded in this mode. A snapshot-then-floor design (option A) was considered and rejected for the workspace/MQ/temporal planes — they have no snapshot format to floor against, and inventing one was explicitly out of scope for this fix; only the graph plane has a callable snapshot (persist_graph_at_checkpoint), so a partial-A fix would still have left WS/MQ/temporal unprotected. Operators needing bounded WAL growth without this trade-off should enable --disk-offload enable.

New integration test tests/crash_recovery_wal_recycle_legacy.rs (t43_legacy_mode_pass_c_must_not_lose_ws_and_graph_history, live kill -9 + restart against tiny --wal-segment-size/--max-wal-size to force several Pass C ticks within seconds) proves two things: (1) byte-level survival — after Pass C has had several chances to run, the earliest workspace AND graph WAL records are still present on disk; and (2) round-trip survival — after a real kill -9 + restart, WS LIST still returns the workspace (replay-only plane, no snapshot format, the exact risk called out above). Full GRAPH.QUERY reconstruction after restart is deliberately not asserted in legacy mode — independent probing found graph command replay from WAL v3 does not reconstruct the graph in that mode even with zero Pass C interference (default --max-wal-size), a separate, pre-existing gap unrelated to WAL recycling; assertion (1) above still proves task #43 is fixed at the layer Pass C actually operates on. Disk-offload mode is unchanged, verified by the existing crash_recovery_graph_durability.rs G4/G5 suite (checkpoint-driven recycle path) passing unmodified.

Fixed — cold-collection visibility: silent data loss on evicted Hash/List/Set/ZSet/Stream (task #41, P0)

With disk-offload enabled (the default), the production eviction paths (evict_one_async_spill, evict_batch_durable_no_aof in storage/eviction.rs) spill Hash/List/Set/ZSet/Stream values to the cold tier via kv_serde::serialize_collection — but the type-specific accessors never consulted the cold index before this fix. Consequences: HGET/ HGETALL/LRANGE/SMEMBERS/ZRANGE/EXISTS reported the key absent after eviction even though it was durably spilled, and HSET/LPUSH/ SADD/ZADD silently fabricated a new empty container, permanently shadowing (destroying) the cold copy on the next flush — a genuine, silent user-data-loss bug for any collection key subject to eviction.

Fixed with a single promote-if-cold hook rather than rewriting every accessor:

  • Database::promote_cold_if_present — on a hot miss, does one ColdIndex lookup; on a hit it decodes the cold value, installs it hot, and removes the cold-index entry (single owning copy, no dual reference). All 8 get_or_create_*/get_* mutable accessors (get_or_create_hash[_listpack], get_or_create_list[_listpack], get_or_create_set/get_or_create_intset, get_or_create_sorted_set, get_or_create_stream, get_hash, get_list, get_set, get_sorted_set, get_stream[_mut]) now call this hook before falling through to their existing "fabricate new" path — so HSET/LPUSH/SADD/ZADD/ XADD merge into the promoted collection instead of shadowing it.
  • The four &self-typed "*_ref_if_alive" shared-read accessors (get_hash_ref_if_alive, get_list_ref_if_alive, get_set_ref_if_alive, get_sorted_set_ref_if_alive, get_stream_if_alive) cannot promote (they're also called from the _readonly dispatch path holding only a shared &Database), so they instead do a non-promoting cold read-through: new Owned/Borrowed variants on HashRef/ListRef/ SetRef/SortedSetRef (and a new StreamRef<'a> Deref wrapper replacing the old &StreamData return type) let a cold hit return a freshly-decoded owned value without a backing hot entry, at the cost of exactly one extra branch + one ColdIndex lookup on a hot miss. Hot-hit cost is unchanged (single probe, same as before).
  • exists/exists_if_alive now fall back to a cheap cold_contains_alive check (ColdIndex presence + TTL only, no disk I/O, no promotion) instead of unconditionally returning false on a hot miss.
  • zrank_readonly/zrevrank_readonly (command/sorted_set/sorted_set_read.rs) updated to route the new SortedSetRef::Owned variant through the same O(n) fallback as the existing Listpack variant (previously any unmatched variant silently fell into a Frame::Null wildcard arm — would have reported a promoted member as not-ranked).

New tests: 13 unit tests in storage::db::tests (spill-then-promote for all 5 collection types + EXISTS-without-promoting) plus a new black-box suite tests/cold_collection_visibility.rs (real server, --disk-offload enable, sampled-LRU-driven eviction, asserts EXISTS, read-after-evict, and write-merges-with-promoted-data for Hash/List/Set/ ZSet). Explicitly out of scope for this fix (unchanged): recovery.rs rehydration, replication, and the eviction gates' Wave A reporting sinks.

Fixed — TEMPORAL/MQ WAL v3 replay data loss on kill-9 (task #42)

TEMPORAL.INVALIDATE and MQ.CREATE/MQ.ACK durable state was silently lost on every crash restart. replay_mq_wal / replay_temporal_wal (src/shard/shared_databases.rs) had two bugs: (a) they scanned persistence_dir/shard-{id}/ instead of .../shard-{id}/wal-v3/ — std::fs::read_dir is non-recursive, so every boot replayed zero records from an empty directory; (b) even with the directory fixed, they matched record.record_type directly, but every cross-thread wal_append blob is re-wrapped as WalRecordType::Command by the shard event-loop drain — the typed inner record (MqCreate/MqAck/GraphTemporal) is nested in the Command's payload and was never unwrapped (replay_workspace_wal was the only replay function that already did this). GraphTemporal records also turned out to be pushed as raw, un-framed bytes (not nested-record-framed) by command::temporal::apply_invalidate, requiring a direct-decode branch alongside the nested-unwrap for replay_temporal_wal. Also fixed a related dead-channel bug found while verifying the MQ path: ShardSlice::wal_append_tx (used by mq_exec.rs's owner-shard MQ.CREATE / MQ.ACK path) was never assigned anywhere — ShardDatabases::wal_append_txs (a separate field, used by TEMPORAL/WS) was wired in event_loop.rs, but the per-slice field was not, so MQ.CREATE/MQ.ACK never reached the WAL v3 writer at all regardless of the replay-side fix. Wired it in event_loop.rs alongside the existing set_wal_append_tx call. New crash-recovery integration test tests/crash_recovery_temporal_mq.rs (temporal_invalidate_survives_kill9, live kill -9 + restart against GRAPH.QUERY ... VALID_AT) and two white-box unit tests in shared_databases.rs (test_replay_mq_wal_restores_registry_and_rolls_back_cursor, test_replay_mq_wal_missing_wal_v3_subdir_is_a_noop) proving the MQ replay fix directly against real WAL v3 bytes — a live MQ.CREATE round trip is architecturally blocked by a separate, deeper, out-of-scope gap (MQ has zero AOF durability, so --appendonly yes's unconditional AOF-authority recovery wipes the Stream on every restart independent of this fix); that gap is left for follow-up.

CodeRabbit review follow-ups on the above (same PR): - Dead-channel sender wiring under appendonly=no + disk-offload=on. event_loop.rs wired ShardSlice::wal_append_tx / ShardDatabases::set_wal_append_tx whenever appendonly_enabled || disk_offload_enabled(), but wal_writer itself is only ever created when appendonly_enabled is true (wal_shard_dir requires it regardless of disk-offload) — so with disk-offload on and AOF off, a live sender was wired to a channel with no writer behind it, and the 1ms-tick drain silently discarded every record instead of the None-sender no-op wal_append/try_wal_append_required are documented to degrade to. Gated the wiring on wal_writer.is_some() instead. - replay_temporal_wal's legacy raw-GraphTemporal discriminator could drop a real invalidation. The un-framed-record fallback rejected a payload whenever its first byte was b'*', to distinguish it from a RESP command payload — but that byte is the low byte of entity_id, so any invalidation on an entity whose id happened to end in 0x2A was silently skipped on replay. Fixed both ends: command::temporal::apply_invalidate now pre-frames its GraphTemporal WAL payload with write_wal_v3_record (matching how MQ's MqCreate/MqAck are already pre-framed), so NEW records replay unambiguously via the existing nested-Command unwrap (CRC-validated, no byte-pattern guessing); the raw-legacy fallback (still needed for WAL segments written before this fix) now runs only when the nested unwrap finds nothing, gated by a new decode_graph_temporal_legacy_raw sanity check (is_node byte must be literally 0/1, both timestamps plausible) instead of the single leading-byte guess. Added a red/green regression test reproducing the exact 0x2A-collision entity id, a round-trip test for the new pre-framed producer format, and an adversarial test proving a RESP-shaped 25-byte payload is not misdecoded into a spurious invalidation. - temporal_invalidate_survives_kill9 only asserted absence. The test checked that the node was invisible at a far-future VALID_AT post-restart, which would pass vacuously if GRAPH.ADDNODE replay had failed entirely (no node at all is also "not visible"). Added a positive assertion that the node is visible at a VALID_AT preceding its own invalidation, proving the node itself actually survived the restart. - MQ.ACK's PEL/queue-key drop on replay is deliberately NOT addressed here — deferred to the same follow-up as MQ's zero-AOF-durability gap above, since PEL replay is only meaningful once MQ stream content itself survives a restart.

Fixed — Wave A adversarial-review fixes (task #34)

Three defects found reviewing Wave A (plane replication) before merge, all fixed on top of the branch's part 1/2/3 work below.

  • Non-string eviction victims under a SpillContext were dropped with no cold copy AND no DEL record. storage::eviction::evict_one_with_spill snapshotted is_plain_drop = spill.is_none() BEFORE its spill body ran, but that body only ever spills RedisValueRef::String values — a Hash/List/Set/ZSet victim picked while a SpillContext was live took neither branch of the is_string check (no bytes written anywhere) yet still fell through to the unconditional db.remove, and because is_plain_drop was already false, on_plain_drop never fired: a genuine, silent, unreported data loss. Fixed by tracking whether THIS victim was actually spilled (spilled, set only inside the branch that performs the write) instead of snapshotting the precondition. Also threaded the real record_reason_del sink into persistence_tick::handle_memory_pressure's "durable spill, manifest reachable" branch, which previously passed a hardcoded no-op sink on the (incorrect, for non-strings) assumption that this branch could never plain-drop. New unit test: eviction::tests::sync_spill_non_string_victim_reports_plain_drop.
  • record_reason_del/record_reason_del_conn/record_effect_write allocated a fully serialized RESP record BEFORE checking whether any replica or AOF pool was even wired. The single most common deployment shape (standalone server, no replica ever attached, --appendonly no) paid a Bytes/Frame allocation for every background-tick DEL or script write effect, only to discard it a few instructions later inside wal_append_and_fanout's own no-op fast path. Fixed by hoisting the exact same has-work predicate (wal_fanout_has_work for the shard-loop context; a new conn_has_work — fanout_hint_active() || aof_pool.is_some() — for the connection context) to the top of all three entry points, before any serialization. Behavior when there IS work is byte-for-byte unchanged (same predicate, same downstream call). This is a perf-only fix with no allocation-counting harness in this repo to red/ green against; verified instead by regression tests confirming both the no-op path (record_reason_del_noop_when_nothing_wired) and the has-work path still emits correctly (record_reason_del_still_emits_to_aof_when_wired, conn_has_work_true_when_aof_pool_wired).
  • Lua's bystander-eviction gate was unwired from both durability planes. scripting::bridge::LuaEvictionCtx::gate — the OOM check run before every WRITE-flagged redis.call/redis.pcall inside a script — called the non-reporting try_evict_if_needed_budget/ try_evict_if_needed_async_spill_budget, which hardcode a no-op sink. When a script's write pushed the shard over maxmemory, whatever BYSTANDER key the policy sampled (never anything the script itself touched) was plain-dropped with no AOF/replication record — invisible eviction, same class of bug as every other Wave-A plain-drop site this branch fixed elsewhere. New eviction::try_evict_if_needed_async_spill_budget_reporting wrapper (mirroring the existing try_evict_if_needed_budget_reporting); gate now threads a real sink through to record_reason_del_conn with the gate's own db_index. Unit-level regression test (not a full black-box EVAL+replica repro, to avoid a flaky multi-process harness for a pure wiring check): scripting::bridge::tests::gate_reports_bystander_eviction_to_aof.
  • Same unwired-reporting bug at the two call sites ordinary writes actually reach. Black-box testing defect 1 (disk-offload enabled, Hash victims, real replica) surfaced that the narrowly-described evict_one_with_spill/SpillContext path above is effectively dead code via any CLI configuration today — spill_thread and shard_manifest are co-gated on the identical disk_offload_enabled() check, so the sync SpillContext::Some branch that persistence_tick wires is never actually constructed. The real reachable paths for an ordinary write under --disk-offload enable are server::conn::handler_monoio:: run_write_eviction_gate (monoio) and shard::spsc_handler:: spsc_eviction_gate (cross-shard dispatch), and both had the identical defect-3-class bug: their spill-sender branch called the non-reporting try_evict_if_needed_async_spill_budget with a hardcoded no-op sink (spsc_eviction_gate additionally had an already-plumbed on_plain_drop parameter that this one branch silently discarded). Fixed both to call try_evict_if_needed_async_spill_budget_reporting with the real record_reason_del_conn sink, mirroring the sibling no-spill-sender branch each function already had correct. New black-box regression: tests/replication_planes.rs eviction_parity_hash_disk_offload_shards1/_shards4 (disk-offload enabled, HSET-driven eviction, real replica attached). Note: full dbsize equality is not asserted — the tick-driven memory-pressure cascade's evict_batch_durable_no_aof batch spill is a real, correctly unreported spill (durable cold copy + cold_index registration) for non-string values too, so a residual share of evictions legitimately never reach the replica. That cold copy is unreadable today for Hash/ List/Set/ZSet (get_hash_ref_if_alive and siblings never consult cold_index, unlike the generic Database::get() path GET uses) — a separate, pre-existing, out-of-scope gap — which is why the test gates on a large-majority convergence ratio (0.70, matching this repo's own MERGE_RECALL_TOLERANCE-style "good enough" convention) rather than exact equality; before this fix the ratio was 0.0 (the replica retained every hash key, full stop).

Known limitations (documented, not fixed here): - redis.pcall('SELECT', ...) inside a script swallows the new SELECT-rejection error into a Lua table instead of raising — the script may continue running against the ORIGINAL db, silently, exactly as before Wave A's SELECT lockout. Only redis.call fails loud. A real multi-db-scripts feature (tracked as a follow-up) is the actual fix; until then, avoid redis.pcall('SELECT', ...) in scripts. - Nondeterministic write commands issued inside a script (SPOP, SRANDMEMBER- derived writes, etc.) replicate verbatim via record_effect_write (the exact cmd + args the script invoked) and can diverge from the master on a replica, exactly like any other verbatim-replicated nondeterministic command in this codebase (pre-existing class, first exposed for Lua by Wave A). A real fix requires replicating the script's effects in a deterministic form, not the nondeterministic call itself — out of scope for task #34. - replication::state::fanout_hint_active() is sticky: once ANY replica has attached, it stays true for the rest of the process's lifetime (see FANOUT_HINT, a one-way latch, not a live "is a replica currently attached" flag). Consequence for operators: after the first-ever replica attaches and later detaches, the monoio inline-SET fast path (server::conn::blocking::try_inline_dispatch) stays permanently disabled for that process — every subsequent SET falls back to the generic dispatch path for the rest of the process's life, even with zero replicas currently attached. A process restart is the only way back to the fast path.

Added — Wave A part 1: record_reason_del dual-plane DEL emission (task #34)

Master-side key removals for a reason OTHER than a client write command (active TTL expiry, --maxmemory eviction plain-drops) previously reached NEITHER the AOF plane nor the replication plane: the key vanished from the master's own keyspace, but an attached replica kept serving it forever, and a kill -9 + restart against --appendonly yes resurrected it from the AOF replay (the AOF never recorded the DEL, only the original SET/write).

  • New replication::reason_del::record_reason_del (shard-event-loop context, reuses wal_append_and_fanout — the same mechanism the ShardMessage::SwapDb synthetic-command record already uses) and record_reason_del_conn (connection-handler context, mirrors handler_monoio::ft::record_local_write_db's backlog/offset/fan-out mechanics and adds the AOF leg that helper omits). Both emit a real DEL <key> RESP record, respecting the fused-SELECT multi-db framing every ordinary write already uses.
  • Wired into: active expiry's whole-key sweep (expire_cycle), background eviction's plain-drop path (timers::run_eviction, persistence_tick::handle_memory_pressure's no-manifest fallback), the inline fast-path SET eviction gate, the generic per-command write-eviction gate, and the SPSC cross-shard write-eviction gate. Every call site is restricted to spill.is_none() plain drops — a spilled/cold-tiered key is NOT a delete and must never be reported here.
  • Fixed a related pre-existing gap found while wiring this: with --disk-offload disable, server::conn::blocking::try_inline_dispatch (the monoio inline SET fast path) fed the AOF but never the replication backlog/fan-out at all — a plain SET on such a master silently never reached an attached replica. can_inline_writes now also gates on !replication::state::fanout_hint_active(), falling back to the generic dispatch path (which replicates correctly) once any replica has ever attached.
  • Known Wave-A-scoped gaps (documented, not silent): hash-field TTL reaps, Lua redis.call write effects, per-db quota (--db-maxmemory) eviction, and cross-db COPY ... DB n destination eviction do not yet emit — tracked as follow-ups.

Added — Wave A part 2: Lua script write-effect replication + SELECT lockout (task #34)

Continues part 1: a Lua script's redis.call/redis.pcall write effects previously reached NEITHER durability plane — EVAL/EVALSHA carry no WRITE command-metadata flag (by design, see below), so the generic per-command AOF/replication gate never saw them. A script's writes vanished on kill -9 + restart against --appendonly yes, and an attached replica never observed them at all. redis.call('SELECT', ...) inside a script also silently corrupted state: dispatch's SELECT handler mutated only a local variable, so every subsequent write in the script kept landing in the ORIGINAL db while looking like it had switched.

  • scripting::bridge::make_redis_call_fn now records every successfully- executed, WRITE-flagged inner command to both the AOF and replication planes itself, immediately after each redis.call/redis.pcall returns (not batched to script end) — a script that writes two keys and then errors on a third still durably records the first two. New replication::reason_del::record_effect_write (+ shared record_bytes_conn core, refactored out of part 1's record_reason_del_conn) does the emission: same fused-SELECT/backlog/ offset/fan-out mechanics as every other write path, recording the verbatim cmd + args the script invoked.
  • LuaEvictionCtx (built once per shard at Lua-VM setup, already caching the OOM eviction handles) now also carries num_shards/repl_state/ aof_pool for this emission — one addition at each of the 5 production construction sites (4 in shard::conn_accept, 1 in server::conn::core::ConnectionContext::build_lua_eviction_ctx, shared by EVAL/EVALSHA and FCALL/FCALL_RO, which route through the identical bridge closure).
  • redis.call('SELECT', ...) inside any script now fails loud with ERR SELECT inside scripts is not supported by moon yet instead of silently corrupting state — intercepted before dispatch, so no partial write from the same call ever lands. A real multi-db-scripts feature is a follow-up.
  • Deliberately did NOT flip EVAL/EVALSHA to WRITE in the command metadata table — that would make the generic per-command AOF/replication gate ALSO record the literal EVAL <script> ... invocation on top of the effect records above, double-applying every write the script made (e.g. an INCR landing as 2 instead of 1). Guarded by a new unit test (command::metadata::eval_evalsha_never_write_flagged) and a black-box regression test (eval_incr_no_double_apply). FCALL (unlike EVAL) IS WRITE-flagged — mirrors upstream Redis Functions and only feeds ACL / READONLY-replica gating, since try_handle_functions always consumes FCALL before the generic AOF/replication block runs; it rides the same single-emission bridge path.
  • Turns the 4 remaining RED tests in tests/replication_planes.rs green: eval_effects_parity_shards1, eval_effects_parity_shards4, eval_writes_survive_restart, select_in_script_errors. All 9 tests in the suite (the 5 from part 1 plus the new eval_incr_no_double_apply) pass.
  • Known gap, unchanged from today: a write-EVAL issued directly against a read-only replica is not rejected (try_enforce_readonly only gates on the WRITE metadata flag, which EVAL intentionally lacks) — tracked as a follow-up; replicas never receive EVAL itself via replication (only the effect records), so this cannot cause replication divergence, only local replica-side corruption if a client is misdirected to write against a replica directly.

Fixed — test-harness port-flake sweep (task #18)

33 integration suites that spawn a real moon process shared a copy-pasted free_port() that binds :0, reads the port, and drops the listener before the server spawns. Two CI-observed failure modes shipped with that pattern: a port TOCTOU (between probe drop and moon's bind, a concurrent test's probe — or an outbound connection's ephemeral source port — takes the port, so moon exits with EADDRINUSE) and a dead-server blind poll (harnesses polled connect() for up to 30s without checking child liveness, reporting "server never accepted: Connection refused" while the real bind error sat unread in the server's stderr log).

  • New shared tests/common/mod.rs: reserve_port() (process-wide dedup set over kernel-chosen probe ports) and spawn_listening() (spawns via a caller closure, polls TCP accept while watching child.try_wait(), and respawns on a fresh port the moment a child dies — external ephemeral-port steals can't be prevented, only recovered from).
  • All 33 suites converted; protocol-level readiness (PING/AUTH) stays with each suite. Kill-9/SIGTERM/restart tests keep their deliberate same-port+same-dir restart legs untouched — only the first spawn of each server lifecycle goes through spawn_listening. Expected-startup-failure tests (CLI validation) keep direct spawns on reserve_port() ports.
  • Verified: all suites green plus a 10-rep stress running 14 port-hungry suites concurrently (the CI contention pattern that produced the original flakes).

Added — R2: multi-shard master PSYNC (task #20, RFC 1B)

A master running --shards N now serves full replication to a single-shard replica — previously PSYNC was rejected with -ERR PSYNC across multiple shards is not yet supported.

  • Per-shard atomic snapshot legs. A new ShardMessage::PrepareReplicaSync fans out to every shard (own shard via the self queue, the rest over the SPSC mesh). Each shard serializes its keyspace slice to an RDB body, captures its replication offset, and registers the replica's live channel in ONE synchronous stretch on its own thread — per shard, nothing can land between "inside the snapshot" and "streamed live", so there is no backlog catch-up leg and non-idempotent commands (INCR) can never double-apply.
  • One merged Redis-format RDB. The PSYNC task stitches the per-shard bodies into a single valid RDB (redis_rdb::write_rdb_merged: header + bodies + EOF/CRC64) and answers +FULLRESYNC <replid> <Σ shard offsets> with one $<len> bulk — the replica's existing R0 loader needs no changes. Index definitions ride once; graph content is sharded, so the snapshot carries one moon-graph-store aux entry per shard and the replica imports all of them (install_graph_store_many, read_moon_aux_all).
  • Per-record SELECT framing on the merged wire. N shard threads feed one replica socket, so a shared "current db" context cannot exist: on multi-shard masters every db-scoped record is fused with its own SELECT <db> prefix (single channel send / backlog append pair / offset advance) — no cross-shard interleave can split a SELECT from the write it frames. Gated on the replica-attach hint; single-shard masters keep the cheaper emit-on-change tracking.
  • Partial resync degrades to full. A replica's single scalar offset cannot be mapped back onto N per-shard backlogs, so a multi-shard master answers every PSYNC (any replid/offset) with +FULLRESYNC.
  • Overflow-kick (task #35), REPLCONF ACK, and WAIT all carry over: the summed snapshot offset keeps total_offset - base == bytes on wire, so WAIT/ACK math stays exact on multi-shard masters.
  • New e2e suite tests/replication_multishard.rs: 2/4/8-shard full resync + live-stream convergence with INCR exactness, interleaved multi-db writers with db-leak asserts, per-shard graph snapshot import, and the partial→full degradation handshake.
  • Known limitation (unchanged from R1): master-side PSYNC requires runtime-monoio (the default). A runtime-tokio master now answers PSYNC with a clear -ERR PSYNC requires runtime-monoio on the master instead of an unknown-command reply (RFC R3/2A). Multi-shard replicas remain unsupported (--shards 1).
  • Known limitation: during snapshot preparation + transfer the live stream buffers in the replica's 16,384-record channel. A very large keyspace under sustained heavy write load can overflow it mid-attach — the replica is then KICKED (loud) and retries the sync; it never diverges silently. Attach such deployments during a write lull, or raise the buffer if this becomes a practical constraint.
  • Docs refreshed for the new topology: docs/guides/clustering.md deployment shape, README.md replication bullets, and docs/PRODUCTION-CONTRACT.md rows REPL-MULTISHARD-01 + WAIT-01 flipped to ✅ with evidence.

Fixed — live-fanout exactly-once redesign + replica-task leak (task #20 follow-up)

Attach-under-write stress testing of R2 surfaced three defects; all fixed before release (none shipped):

  • REPLICAOF leaked the previous replica task — every re-attach stacked one more live applier. REPLICAOF host port spawned a fresh run_replica_task without stopping the old one, and REPLICAOF NO ONE only flipped the role state: the old task kept its master link open and kept APPLYING the stream. After NO-ONE → re-attach cycles the replica ran INCR counters ~25-35% ABOVE the master (reproduced at shards=1 AND shards=4 — pre-existing, not an R2 defect). Replica tasks now carry a process-global epoch ticket; a new REPLICAOF target or NO ONE bumps the generation and superseded tasks exit before their next connect, before loading a snapshot, and before applying any parsed chunk.
  • Snapshot-vs-live exactly-once is now offset-cut based, not FIFO-placement based. Two adversarial-review rounds found opposite failure modes for placement schemes (a queued self-shard snapshot leg double-delivered local writes; an inline-captured one lost same-cycle cross-shard writes — neither in the body nor live-sent, with the offset advanced: permanent replica lag). Every fan-out entry now records cut = <shard offset at body capture> and every live record carries its per-shard end_offset; delivery requires end_offset > cut, making correctness independent of where the registration lands in the drain FIFO. All shards (including the PSYNC connection's own) use the same PrepareReplicaSync arm.
  • Same-key wire ordering. Cross-shard (SPSC-dispatched) writes used to send to replicas directly from the execute arm while local handler writes deferred through the self queue — a later-offset write could reach the wire before an earlier-offset write to the same key, replaying same-key writes out of the master's order on the replica (found by analysis; not reproduced in ~10 black-box runs). ALL live replica sends now flow through the self-queue ReplicaLiveFanout arm, so per-shard wire order equals offset order by construction.
  • New e2e regressions: attach_under_write_no_double_apply (4-shard + single-shard control, 5× detach/re-attach under pipelined INCR load, exact per-counter parity) and same_key_write_order_parity (12 connections APPEND-race the same 32 keys through both write paths; replica strings must byte-equal the master).

Fixed — consistency/durability defects caught by the R1 gates (task #35)

Three pre-existing data-integrity bugs surfaced by the new load/kill-9 gates run for the R1 WAIT/ACK work (all reproduced on main before the fix):

  • Replication: silent record drop for lagging replicas. The per-replica fan-out used try_send and skipped the record when the bounded channel was full — under one pipelined burst a replica received ~2k of 40k keys, stayed master_link_status:up, and diverged forever. A replica whose channel overflows is now KICKED (shared kicked flag → drain loop closes the socket) so it reconnects and resyncs from the backlog — Redis's output-buffer-limit policy. Channel capacity raised 1024 → 16384 records so bursts kick rarely. New e2e replica_converges_under_interleaved_multidb_load locks in exact post-load parity.
  • Client SELECT commands leaked into the AOF and replication stream. SELECT is W-flagged (routing), so every handler persisted the literal client SELECT while bare writes carried no db context — two interleaved connections on different dbs corrupted BOTH planes' db attribution (observed: all 40k keys recovered into db2 after kill-9). SELECT is now connection-state only (metadata::is_persisted_write), never persisted or replicated.
  • AOF had no per-record db attribution. The AOF writer now threads the executing db through AofMessage and tracks the stream's current db, prepending a SELECT <db> record exactly when it changes (Redis's aof_selected_db), including through BGREWRITEAOF fold drains (context resets per fresh incr segment; ambiguous drains use a force-emit sentinel). New e2e aof_multidb_kill9 (TopLevel + PerShard) proves interleaved multi-db writes recover into the correct databases after kill-9 on both runtimes.

Added — R1: real WAIT/ACK plumbing (replica acknowledgements)

  • WAIT <numreplicas> <timeout> now works on the production (monoio) runtime. It previously answered :0 unconditionally from the synchronous dispatch table; a new connection-layer intercept awaits wait_for_replicas (10ms poll, early exit; timeout 0 = block until satisfied, capped at one year). Wired on the tokio sharded path too.
  • Replicas acknowledge their applied offset: a dedicated 1s ticker task owns the write half of the (split) replication socket and sends REPLCONF ACK <offset> — Redis's replicationCron cadence, doubling as an idle keepalive for master-side lag detection. A timeout-wrapped read was rejected: cancelling an in-flight io_uring read whose completion already landed DISCARDS those bytes (silent stream corruption).
  • The master reads ACKs off the hijacked PSYNC socket: the inline drain loop splits the stream; a same-thread reader task parses REPLCONF ACK <offset> frames (dedicated parser — the shared replication drainer deliberately drops REPLCONF as chatter) and records them into the replica's ack_offsets/last_ack_time via fetch_max (reordered or duplicate ACKs can never regress the recorded offset).
  • New e2e wait_returns_acked_replica_count: WAIT with no replicas → 0 fast; WAIT 1 after a write → 1 within the ACK cadence; WAIT 2 with one replica → times out reporting 1 (RED before this change).

Fixed — multi-db replication: master streams SELECT, replicas serve it (HIGH-2)

  • Writes outside db 0 landed in db 0 on the replica: the replica-side drain already understood in-stream SELECT context, but the master never emitted it — every replicated command applied to the replica's db 0 regardless of the db it executed in on the master. record_local_write_db now prepends a SELECT <db> record whenever the writing connection's db differs from the stream's per-shard db context (ReplicationState:: stream_db); the context is reset in the same synchronous stretch as every FULLRESYNC snapshot capture (Redis's slaveseldb = -1 idiom), so a fresh replica always sees an explicit context before its first non-0-db write.
  • +CONTINUE keeps the db context across reconnects: resumed backlog bytes only carry SELECT at db CHANGES, so the replica now preserves its drain-side db in ReplicaTaskConfig::stream_db across link drops (reset to 0 on FULLRESYNC). In-memory only — a replica process restart starts at offset 0 and always full-resyncs.
  • SELECT was rejected on read-only replicas (task #23): flagged W in the metadata table, so the READONLY guard blocked it — a client could never read a replica's non-zero dbs. All three dispatch paths now serve SELECT on replicas (connection-state only, Redis parity).

Fixed — replication round-2 hardening: TEMPORAL.INVALIDATE, replay liveness, blob endpoint checks

  • TEMPORAL.INVALIDATE never replicated (round-2 finding B): the handler drained the graph WAL — the same record mechanism GRAPH.* replication uses — but only fed the local WAL, never the replication plane; a replica silently kept valid_to = ∞ for entities the master had invalidated. The master now streams a deterministic, wall-clock-pinned internal form (TEMPORAL.INVALIDATE-AT <graph> <N|E> <entity_id> <wall_ms>, single-shard scope like every replication leg) so master and replica agree on the exact valid_to; the replica applies it through the same apply_invalidate the master ran. New e2e REPL-GRAPH-03 proves temporal visibility converges (red without the master leg, green with it).
  • Streamed replay could resurrect a tombstoned node (round-2 finding F, regression from the P1-4 lazy-resolver rewrite): the lazy node_exists accepted DEAD write-buffer entries (get_node does not filter deleted_lsn), so a stray SETPROP for a node removed in an earlier streamed replay call re-registered it into the live property index. Split into node_present (AddNode dedup — any record of the id, matching the never-reuse slotmap id contract) and node_alive (edge endpoints / SETPROP / SETLABEL / REMOVENODE — write-buffer entry is authoritative, live-only, matching the old pre-seeded map's iter_nodes() semantics).
  • Graph snapshot install now rejects delta edges with unknown endpoints (round-2 finding E, defense-in-depth): add_edge_across_tiers_with_id's aliveness check only fires for resident endpoints — a corrupted blob referencing a nonexistent node installed silently. The install loop now verifies both endpoints against the just-installed segments and drops the edge LOUD (tracing::warn!) otherwise.
  • WS.*/MQ.* writes are NOT replicated in v0.7 — now fail-loud (round-2 finding A, known limitation): WS.CREATE/WS.DROP and MQ mutations persist durably on the master (WAL) but have no deterministic replication record form yet (WS.CREATE mints a fresh UUIDv7 per execution — verbatim streaming would diverge). A one-time tracing::warn! now fires when such a write executes while a replica is attached, instead of silent divergence discovered at failover. Full support (id-pinned record forms + replica apply arms + snapshot coverage) is tracked as follow-up work, alongside the pre-existing Lua-EVAL and expiry/eviction propagation gaps.

Fixed — listing parity: INFO # Keyspace all dbs × all shards, GRAPH.LIST all shards

  • INFO's # Keyspace section always printed a single db0: line holding the SELECTED db's LOCAL-shard key count — SELECT 2; SET k v; INFO reported the db-2 count as db0, every other db was invisible, and at --shards N the other shards' keys were uncounted. It now lists every NON-EMPTY logical db (Redis semantics) with (keys, expires) summed across all shards via a new KeyspaceStats scatter (O(#dbs) counter reads per shard, no key iteration). All three dispatch paths (monoio / tokio sharded / single).
  • GRAPH.LIST at --shards N listed only the connection shard's graphs (~1/N of them, since a graph lives on the shard that owns its name). It now scatters to every shard and returns the sorted, deduplicated union.

Fixed — CLIENT TRACKING dead on the monoio runtime (H-3 reorder regression)

  • Since the H-3 ACL reorder (#258), CLIENT TRACKING ON|OFF answered ERR unknown subcommand 'TRACKING' on the monoio runtime (the production default) at every shard count: try_handle_client_admin runs before try_handle_client_tracking in the frame loop and its unknown-subcommand fallback consumed TRACKING before the dedicated handler could see it. RESP3 invalidation push was therefore entirely unavailable. Invisible to CI because the test matrix exercises the tokio handler (handler_sharded), whose intercept ordering differs.
  • try_handle_client_admin now falls through for TRACKING (both handlers are post-ACL, so the H-3 deniability guarantee is unchanged). Re-greens all 5 client_tracking_invalidation black-box tests on monoio.

Fixed — replication exactly-once + txn/graph fidelity (adversarial-review P0/P1)

  • MULTI/EXEC bodies never replicated at --shards 1 (P0-1): EXEC persisted its body through persist_txn_aof's AOF-only leg — the one local write path that skipped the replication plane. A MULTI/SET/EXEC committed durably on the master and never reached the replica: silent deterministic divergence for every application using transactions. The txn body now records each entry through the same record_local_write leg as single-command writes (AOF lsn = 0, no double-advance). New e2e replica_applies_multi_exec_bodies (INCR doubles as a double-apply canary).
  • FULLRESYNC snapshot capture raced undrained local writes (P0-2): the original design queued backlog append + offset advance + live fan-out as ONE deferred event-loop message, so a mutation could sit inside the RDB while still below the advertised snapshot offset — re-delivered via backlog catch-up and double-applied (INCR/LPUSH divergence). record_local_write now appends the backlog bytes and advances the shard offset SYNCHRONOUSLY at write time (atomic with the mutation w.r.t. the inline PSYNC capture); only the live replica try_send is deferred (ReplicaLiveFanout). RegisterReplica correspondingly carries a push-time offset so catch-up and live delivery stay disjoint for every write/attach interleave.
  • Graph snapshot lost soft state that lives outside CSR segments (P1-5 + two adjacent gaps): the CSR byte format has no validity section, freeze() RETAINS cross-tier delta edges in the write buffer, and copy-up node tombstones never freeze — the master recovers all three from its WAL on restart, but a replica has no WAL, so it resurrected deleted edges/nodes and silently lost every cross-tier edge. Blob format v2 ships a per-segment deleted-edge sidecar, the retained delta edges (original edge ids), and the dead-shadow list; install re-applies all three.
  • Streamed graph replay was O(N²) (P1-4): each replicated GRAPH.* record re-scanned every write-buffer node and every segment row to pre-seed the replay id map — unbounded replication lag on bulk graph loads. Node existence is now resolved lazily (O(1) write-buf probe + MPH segment lookup), and the id-allocation floor is raised per segment header max instead of per row.
  • Also: FULLRESYNC graph export skips freezing untouched write buffers (P1-6, repeated-resync latency), and the shard self-queue is drained unbounded per cycle (P2-7 — entries are cheap try_sends; a cycle cap could strand a replica's live bytes by a full tick).

Fixed — single-shard live replication stream was DEAD (self-SPSC gap)

  • The R0 live stream never actually flowed at --shards 1. The SPSC mesh is N·(N−1) with skip-self mapping, so a task on a shard's own thread had NO producer to that shard: the inline PSYNC task's RegisterReplica failed every attach ("shard 0 producer missing") and the replica fell into a 0.5s reconnect/full-resync loop. Tests stayed green because each resync's RDB carried the latest keyspace + FT defs — data crawled across via snapshot polling, masking the dead stream. Fixed with a thread-local self-message queue (shard::self_msg) drained by the event loop alongside its SPSC consumers; PSYNC registration, FT.*/graph fan-out, and local-write fan-out all route through it.
  • Local (same-shard) writes now feed the replication plane. Successful local writes push their wire bytes as ReplicateVerbatim before any await (mutation + record are one synchronous stretch, atomic w.r.t. snapshot capture); the drained message does backlog + offset + replica fan-out together, and the AOF leg no longer double-advances the offset (lsn = 0 when fan-out owns the advance).
  • Replication backlog now seeds at the current shard offset on lazy allocation (ReplicationBacklog::new_at) — an unseeded backlog made every catch-up range read on a pre-written master fail as "evicted".
  • New process-global fanout_hint_active() (one Relaxed load, set on first replica attach, never cleared) gates all fan-out serialization so non-replicating servers pay nothing on the hot path.

Added — v0.7 graph-plane replication (live stream + snapshot backfill)

  • Live leg: graph mutations (GRAPH.* + Cypher writes) stream to replicas as their deterministic, id-pinned WAL records (GRAPH.ADDNODE <g> <id> …; label/prop ids are a stateless FNV hash, identical on both sides). The replica applies them through the same GraphReplayCollector restart recovery uses — no id re-allocation, no divergence. Replay's edge/SET resolution now also seeds from write-buffer-resident nodes so one-record-at- a-time streaming replay resolves endpoints applied by earlier records.
  • Snapshot leg: the FULLRESYNC RDB carries a moon-graph-store aux blob — every graph's write buffer is frozen to CSR segments (the checkpoint's own "freeze is the only serialization path" contract) and shipped as to_bytes() encodings + id cursors; the replica installs them exactly like restart recovery (replication::graph_sync). Mmap (restart-loaded) segments export their mapped bytes verbatim.
  • READONLY guard is now Cypher-aware: a read-only GRAPH.QUERY (MATCH/RETURN) is served by replicas; only write queries (CREATE/DELETE/ SET/MERGE tokens) are rejected. Previously the blanket W flag rejected all GRAPH.QUERY on replicas.
  • New e2e tests/replication_graph.rs: live-stream parity (nodes, properties, GRAPH.LIST) with a zero-reconnect stream-health assertion that would have caught the masked dead stream, plus snapshot backfill + post-snapshot live growth.

Fixed — PSYNC attach races closed (adversarial-review findings on R0/R0.5)

  • Registration-bounded catch-up: the master now registers the replica with the event loop BEFORE reading backlog catch-up bytes, and the event loop replies with the exact offset where live fan-out begins (RegisterReplica.registered reply channel). Catch-up sends exactly [snapshot_offset, registration_offset) — previously a write drained between the catch-up read and registration reached neither the RDB, the catch-up, nor the live stream: a silent, unlogged replica gap (worst for FT.* def mutations, whose fanout always crosses the SPSC queue).
  • Atomic snapshot capture: the FULLRESYNC snapshot offset is read in the same synchronous stretch as the RDB capture (no .await between), closing the inverse race where a write landed inside the RDB and above the advertised offset — double-applying non-idempotent commands (INCR) on the replica.
  • Backlog eviction during catch-up now aborts the sync loudly (replica retries a fresh full resync) instead of silently skipping the missing bytes.
  • A replica whose full resync carries NO index definitions (pre-R0.5 master or def-serialization failure) now logs a warning when the authoritative replace drops local indexes with no replacement.
  • New race-guard e2e replica_attach_races_live_ft_create (30 FT.CREATEs racing a mid-stream attach, exact index-list parity required).
  • Hygiene: read_moon_aux validates RDB version bytes like load_rdb; the FT.* fanout hook skips the serialize+SPSC round trip until a replica or backlog exists.

Added — vector/text index-plane replication sync (v0.7 R0.5)

  • Before this change a replica synchronized only the KV keyspace — FT index definitions never left the master (FT.CREATE/FT.DROPINDEX/FT.CONFIG are handled at the connection layer and never reached the replication fanout), and the replica's apply path ran generic dispatch only, skipping the master's auto-index parity hooks. A replica answered FT._LIST with nothing and FT.SEARCH with errors even while its hashes matched the master.
  • Snapshot leg: the FULLRESYNC RDB now carries the master's vector + text index definitions as moon-private RDB AUX fields (opcode 0xFA, keys moon-vector-defs / moon-text-defs), reusing the sidecar codecs (serialize_index_metas_v5, serialize_text_index_metas). Standard RDB loaders skip AUX fields, so the snapshot stays Redis-tool-compatible. On load the replica drops all local indexes (full resync = authoritative replace), installs the master's definitions, and backfills them by rescanning matching HASH keys — the same "restart semantics" rescan restart recovery performs.
  • Live leg: successful FT.CREATE / FT.DROPINDEX / FT.CONFIG SET on the master now fan out verbatim to the backlog + connected replicas via a new ShardMessage::ReplicateVerbatim (offset-accounted like any replicated write; durability remains the sidecar's job — no WAL/AOF leg). The replica applies them through the same ft_create/ft_dropindex/ft_config handlers, and runs the master's index-parity hooks after every applied KV write (HSET auto-index, DEL/UNLINK tombstone, HDEL vector-field tombstone, FLUSHDB/FLUSHALL content clear) so replica indexes track replica keyspace.
  • New black-box acceptance test (replica_syncs_vector_index_defs_and_contents) covering snapshot defs + backfill, live HSET indexing, live DEL tombstoning, live FT.CREATE streaming, and FLUSHALL clear-contents-keep-defs semantics.
  • Scope: single-shard master (matches R0); multi-shard FT.* replication rides the R2 broadcast redesign. Graph-plane replication is a separate v0.7 workstream.

Added — replica now applies the replication stream end-to-end (v0.7 R0)

  • Foundational fix: before this change a REPLICAOF replica completed the PSYNC2 handshake but applied nothing — run_handshake_and_stream discarded the FULLRESYNC RDB (logged "received … bytes" only) and stream_commands did buf.clear() after advancing the offset. A freshly attached replica reported DBSIZE 0. This went unnoticed because every replication_hardening.rs test is #[ignore]d and never runs in CI.
  • The replica now (1) loads the full-resync RDB snapshot into its local shard and (2) parses the live RESP command stream and applies each write, tracking SELECT. Both runtime variants (monoio, tokio). Apply runs synchronously on the shard thread via the thread-local ShardSlice (single-shard: every command is local, no SPSC self-hop), bypassing the connection-layer read-only guard as a replica must.
  • Also fixes a wire bug: the replica read the diskless full-resync RDB bulk as len + 2 bytes (assuming a trailing \r\n that diskless replication does not send), stealing the first two bytes of the command stream and desyncing it. Now reads exactly len. The replication offset advances by bytes consumed by complete frames, not the raw socket read count.
  • MOVE and cross-db COPY ... DB n are applied through the same two-db core helpers the master's handler-level intercept uses (generic dispatch cannot apply them), so they replicate correctly. A replicated command that still fails to apply is logged loudly rather than dropped silently.
  • New pure, unit-tested stream router (replication::apply) and a black-box streaming acceptance test (tests/replication_streaming.rs, covering snapshot + live SET/DEL/MOVE/COPY).
  • Scope: single-shard (--shards 1), logical db 0. A multi-shard replica refuses to start replication loudly (rather than diverging silently). Non-default-db writes are a known follow-up — the master's live fanout does not yet stream SELECT, so SELECT n; SET replicates into db 0. Per-shard WAIT/ACK and multi-shard PSYNC build on this in R1/R2.

Added — --repl-backlog-size (Redis repl-backlog-size parity)

  • New flag: per-shard replication backlog capacity in raw bytes (default 1 MiB, clamped to the 16 KiB Redis floor). Bounds how far a disconnected replica may fall behind and still partial-resync. Previously the capacity was hardcoded at three sites (handshake allocation ×2, SPSC lazy fallback) and replication_hardening::full_resync_outside_backlog failed at spawn with unexpected argument — the master process never started. ReplicationState.backlog_capacity is now the single source of truth, carried into ShardMessage::RegisterReplica for the fallback-init and reported truthfully by INFO replication repl_backlog_size.

Fixed — replication_hardening harness could never complete

  • Test teardown used SHUTDOWN NOSAVE + Child::wait(), but SHUTDOWN is not implemented on the production path (the dispatch arm is an error stub and no connection handler intercepts it) — every test hung forever at cleanup and a mid-test assert failure leaked live servers that wedged later tests' ports. Teardown is now a kill-on-drop guard (panic-safe). Implementing a real SHUTDOWN [NOSAVE|SAVE] is tracked separately.

Fixed — txn_kv_wiring integration test port-collision flake

  • start_txn_server picked a port from a throwaway bind(:0) probe, dropped it, then let run_sharded rebind it — a TOCTOU window where a parallel test (this binary or another) could steal the freed ephemeral port, after which a blind 200ms sleep handed the caller a dead/wrong port. Now: a process-global reservation set (reserve_unique_port) guarantees no two tests in the same binary are handed the same recycled port, and an active connect+PING readiness probe (await_server_ready) replaces the fixed sleep and lets the start retry on a fresh port when a bind is lost to another process. Verified flake-free across repeated fully-parallel runs (no --test-threads=1).

Fixed — --disk-offload without a durability backstop broke the default LRU cache recipe

  • GCP benchmark finding (2026-07-10): with --disk-offload enable (the default) but --appendonly no and no --save, the durable-spill eviction path added to fix the crash window above (evict_batch_durable_no_aof) never runs, because it needs a ShardManifest that is only threaded through the tick-driven memory-pressure cascade (shard::persistence_tick::handle_memory_pressure), itself gated on persistence_dir, which main.rs only constructs when appendonly == "yes" || save.is_some(). The inline write-path eviction gate has no manifest access, so durable spill is impossible from that call site — but the pre-fix manifest is None branch OOM'd unconditionally regardless of --maxmemory-policy, silently breaking the documented "Pure cache (no durability)" recipe (--appendonly no --maxmemory <bytes> --maxmemory-policy allkeys-lru): a classic LRU cache rejected writes at the cap instead of evicting.
  • Behavior fix (src/storage/eviction.rs, try_evict_if_needed_async_spill_with_total_budget's manifest is None branch): made the fail-close policy-aware, matching Redis semantics. When total_memory <= budget nothing happens; when policy == NoEviction it still OOMs (correct — noeviction always rejects); otherwise (any evicting policy — allkeys-*/volatile-*) it now reclaims the budget by dropping victims via evict_one_with_spill(.., spill: None) (the same plain-drop helper the disk-offload-disabled write path already uses — no new victim-selection logic), looping until the budget is satisfied or no eligible victim remains (then OOM). No durable spill is attempted here — that still only happens via the tick's evict_batch_durable_no_aof when a manifest is reachable. Three new tests in src/storage/eviction.rs: async_spill_no_aof_backstop_no_manifest_allkeys_lru_evicts_pure_cache (the regression test — fails with OOM on pre-fix code, passes after), async_spill_no_aof_backstop_no_manifest_noeviction_still_rejects (unchanged noeviction OOM behavior), and async_spill_no_aof_backstop_no_manifest_volatile_lru_scoped_to_ttl_keys (volatile-* only evicts keys with a TTL, never a persistent key). The --appendonly yes async-spill path and the manifest-reachable durable batch-spill path are untouched.
  • Startup warning: ServerConfig::disk_offload_spill_inert (new predicate, src/config.rs) plus ServerConfig::warn_disk_offload_without_durability (mirrors warn_deprecated_cold_tier_flags) emit a single tracing::warn! at startup when --disk-offload enable + no durability backstop is detected, called from main.rs right after the existing cold-tier-flags warning: disk-offload's cold-spill tiering is still inert in this combination (evicting policies drop eligible victims instead of tiering; noeviction — and any evicting policy with no eligible victim left, e.g. volatile-* once its TTL-bearing keys are gone — returns OOM). Documented in docs/guides/tuning.md's "Tiered memory offload" section.

Fixed — test flakiness: oom_bypass_closure readiness on loaded CI runners

  • tests/oom_bypass_closure.rs: the readiness path used a fixed 30s connect deadline with no CI allowance and no dead-child detection — test_case_e_cross_db_copy_oom lost that race on a loaded macOS CI runner (2026-07-10). Same remedy as the sigterm/bgsave harnesses (#264): wait_ready now fails fast via try_wait() when the server process has exited (surfacing moon.stderr.log/moon.stdout.log tails instead of burning the deadline), retries in bounded 1s connect windows, and extends the deadline 30s → 120s under CI.

Fixed — CI: fts_query_parse missing from fuzz matrix

  • .github/workflows/fuzz.yml: the fts_query_parse fuzz target was registered in fuzz/Cargo.toml but absent from both the PR and nightly CI matrices, so it never ran. Added to both. Also corrected CLAUDE.md's stale fuzz-target count (7 → 12) and target list.

Fixed — RDB stream consumer-group allocation-DoS gap (src/persistence/rdb.rs)

  • Four more untrusted-length counts in the native RDB TYPE_STREAM decoder were unvalidated, the same vulnerability class fixed for snapshot.rs/redis_rdb.rs above: read_entry and its zero-copy twin read_entry_zero_copy (src/persistence/rdb.rs) read a stream's group_count, pel_count, consumer_count, and pending_count straight off the wire without bounding them against the bytes actually remaining in the cursor, unlike every sibling count in the same functions (entry_count, field_count, hash/list/set/zset counts), which already went through rdb::validate_count. None of the four drives an eager with_capacity/zero-fill today (they grow a BTreeMap/HashMap incrementally), so this wasn't a single-allocation DoS, but it left a structural inconsistency that a future refactor could easily turn into one and skipped the fail-fast rejection every other count gets. Fixed by adding validate_count calls with the true structural minimum bytes per item (28 for a group: 4-byte empty name + 16-byte last-delivered-id + 4-byte pel_count + 4-byte consumer_count; 36 for a PEL entry: 16-byte StreamId + 4-byte empty consumer name + 16-byte delivery time/count; 16 for a consumer: 4-byte empty name + 8-byte seen_time + 4-byte pending_count; 16 for a pending id: StreamId) — chosen so no legitimate file is ever rejected. Red/green TDD: 4 crafted-blob tests (one per count) confirmed failing before the fix and passing after, plus a control test asserting a fully-populated 1-group/1-pel/1-consumer/1-pending stream is still accepted by both read_entry and read_entry_zero_copy.

Fixed — KV disk-offload: make eviction spill durable before dropping the hot value

  • Crash window (src/storage/eviction.rs): under memory pressure with --disk-offload enable, the async-spill eviction path (evict_one_async_spill, driven from the shard event loop's memory-pressure tick) queued the SpillRequest to the background SpillThread and immediately dropped the hot value from RAM — before the pwrite, fsync, or manifest commit had happened. Under --appendonly yes this is safe (a crash in that window is covered by AOF replay of the original write), but under --appendonly no there is no AOF to fall back on: the value existed nowhere (not in RAM, not yet on disk, no WAL/AOF record) for the entire SpillThread batching window (up to 256 entries or its 100ms tick, longer under backpressure). A kill -9 in that window was unrecoverable, silent data loss.
  • Fix: try_evict_if_needed_async_spill_with_total_budget now branches on --appendonly. With an AOF backstop the fast fire-and-forget path (evict_one_async_spill, queue-then-drop over the SpillThread channel) is unchanged (documented, intentional trade-off). Without one, it drives a new evict_batch_durable_no_aof: collects up to 256 victims (mirroring SpillThread's own FLUSH_ENTRY_CAP) without removing them from RAM, writes the whole batch as one durable, synchronous call to spill_thread::flush_buffer (pwrite + fsync + manifest commit — the same inline/oversized page routing the background thread itself uses, now made pub(crate) for reuse), and only then removes from RAM the keys whose file both pwrite-succeeded and manifest-committed, immediately populating ColdIndex so each stays read-through-able (previously only a later SpillCompletion — which never arrives on this path — did that). Batching is required, not optional: an earlier draft did one fsync per evicted key and stalled the single-threaded shard event loop under sustained pressure. try_evict_if_needed_async_spill_with_total_budget gained an Option<&mut ShardManifest> parameter to carry the manifest down; the memory-pressure cascade (handle_memory_pressure in src/shard/persistence_tick.rs) is the one call site that has a manifest and passes it through. The four other callers (inline per-connection write-path gate, sharded/monoio handlers, spsc_handler, the Lua scripting bridge) have no manifest reachable; under --appendonly no they now bail (retain the hot value, surface OOM) instead of risking the same crash window — the next 100ms memory-pressure tick reclaims the memory via the durable path instead. Trade-off: those inline call sites no longer evict opportunistically under --appendonly no (that unconditional fast-path eviction, regardless of durability, was the bug being fixed), so a write burst that crosses maxmemory now converges to budget only as fast as the 100ms tick reclaims it, evicting the minimum needed each pass rather than eagerly on every write — verified live (per-shard estimated_memory() accounting, not RSS) to converge correctly and stay converged.
  • Second bug, same file (evict_one_with_spill, the legacy fully-sync spill path used when no SpillThread exists): on a spill I/O error it logged a warning and still evicted the key from RAM — unconditional data loss on every spill failure, independent of any crash. Now returns without evicting on failure, matching the async path's fail-closed contract.
  • Regression coverage: sync_spill_failure_retains_hot_value_no_silent_drop, async_spill_no_aof_backstop_durably_spills_before_drop (locks the durable pwrite+fsync+manifest-commit-before-drop ordering and immediate ColdIndex availability), async_spill_no_aof_backstop_no_manifest_retains_hot_value, async_spill_no_aof_backstop_spill_failure_retains_hot_value — plus a db.remove() (DEL) check after the new durable-spill path to guard against reintroducing the DEL/cold-tombstone resurrection class (#257): an earlier draft of this fix inserted into ColdIndex before db.remove(), and Database::remove's own cold-tier cleanup silently wiped the insert back out — caught by this test suite, not by inspection.

Fixed — test flakiness: sigterm-shutdown readiness + BGSAVE file polling

  • tests/sigterm_shutdown.rs: wait_for_ready now distinguishes a crashed child (fails fast via try_wait() instead of burning the full deadline) from a genuinely slow-to-start one, and every readiness panic surfaces the captured moon.stdout.log / moon.stderr.log so a real failure stays diagnosable. The readiness deadline also doubles under CI (60s → 120s) — the observed flake was macOS CI runner startup latency under load, not a hang in the SIGTERM path under test.
  • tests/integration.rs: test_bgsave_creates_rdb_file and its two siblings that shut down and restart a server right after bgsave_with_retry (test_rdb_restore_on_startup, test_aof_priority_over_rdb) asserted dump.rdb existence off a fixed 200ms guess-sleep — BGSAVE queues the write on a spawn_blocking thread and replies before it lands on disk. Added a wait_for_file bounded-poll helper (50ms interval, 10s deadline) and used it at all three call sites in place of the point-in-time assumption. The MOON-magic-bytes assertion is unchanged.

Fixed — CI: nightly fuzz shard (graph_props_record) + stale deny.toml comment

  • Missing fuzz harness: fuzz/Cargo.toml and .github/workflows/fuzz.yml both declared a graph_props_record fuzz target, but fuzz/fuzz_targets/graph_props_record.rs didn't exist — that matrix shard failed every nightly run (audit finding, docs/PRODUCTION-CONTRACT.md FUZZ-01). The target function is real: src/graph/csr/props.rs's module doc even names it (fuzz target: graph_props_record) — the v5 node/edge property-record codec (decode_node_props, decode_node_prop, decode_node_embedding, node_record_len, decode_edge_weight, decode_edge_props, decode_edge_prop, edge_record_len) decodes bytes read straight off a .csr segment file and is documented to return None/empty on malformed input rather than panic. Added the missing harness so the shard actually fuzzes it, following the exact pattern of the sibling targets (wal_v3_record.rs, gossip_deser.rs).
  • Stale comment: deny.toml's header claimed a ci.yml "safety-audit" job runs cargo audit/cargo deny; no such job exists. The actual "Lint" job runs scripts/audit-unsafe.sh + scripts/audit-unwrap.sh (SAFETY- comment coverage / unwrap ratchet) — a different, unrelated check. Corrected the comment to describe reality and left a TODO(ci): marking the gap (wiring cargo deny check into CI is a separate decision, out of scope here).

Changed — Tiering: drop unused meta/undo placeholder files

  • transition_to_warm (HOT->WARM segment seal) no longer writes meta.mpf (VecMeta) or undo.mpf (VecUndo) — both were always-empty placeholders with zero readers anywhere in the codebase (the WARM load path WarmSearchSegment::from_files only ever opens codes.mpf/graph.mpf/ mvcc.mpf/optional vectors.mpf by name). Removes 2 extra File::create + fsync round-trips per WARM segment seal. write_meta_mpf and write_undo_mpf (src/vector/persistence/warm_segment.rs) are removed as they had no other callers.
  • src/vector/persistence/segment_io.rs: removed the dead sq_vectors.bin/f32_vectors.bin comment-only branch (never emitted since SQ8/f32 storage was dropped from ImmutableSegment in favor of TQ-ADC) and corrected the module doc header, which still listed both files as if written.
  • Back-compat: old WARM segments written by a prior build that still contain meta.mpf/undo.mpf continue to load — the loader opens files by name and never scans the segment directory, so the extra files are silently ignored (verified by a new regression test with stray legacy files on disk).

Fixed — RDB DoS hardening

  • Two untrusted-input allocation-DoS vectors in RDB/snapshot loading are now bounds-checked. shard_snapshot_load's SEGMENT_BLOCK_MARKER branch (src/persistence/snapshot.rs) read an untrusted entry_count: u32 and passed it straight to Vec::with_capacity with no gate; the Redis-RDB reader (src/persistence/redis_rdb.rs) did the same for read_redis_string's length-prefixed byte buffer and for the list/set/hash/zset collection decoders, where read_length's 8-byte (0x81) form can return up to u64::MAX. Both are on the server-startup / replica-full-sync path, so a single crafted or corrupt file could drive a multi-gigabyte allocation before a byte of the claimed data was read — aborting the process (release builds use panic = "abort") or OOM-killing it. Fixed by promoting rdb::validate_count to pub(crate) and calling it before the snapshot segment's Vec::with_capacity, and by adding a matching check_alloc_bound helper in redis_rdb.rs that bounds every untrusted length/count against the bytes actually remaining in the input before allocating. No wire-format or valid-file behavior changes — a legitimate length/count always fits within the remaining input.

Removed — Experimental DiskANN COLD vector tier

  • Deleted the experimental "COLD-ann" vector tier (src/storage/tiered/cold_tier.rs, the whole src/vector/diskann/ module — Vamana graph, product quantizer, co-located page format, io_uring beam search). It was gated off by default (MOON_VEC_COLD_TIER=1), incomplete by its own docs (no cold-segment deletion, ADC-only recall, no restart-time PQ-codebook reload), and had no restart recovery story. The M3-exit review decided delete-over-finish. The default COLD valve remains SegmentList.unloaded (UnloadedSegment, exact, near-zero RAM, reload-on-touch) plus WARM byte-budget LRU eviction (--vec-warm-mmap-budget) — neither is affected by this removal.
  • Removed the dead call chain: VectorIndex::try_cold_transitions[_all], VectorStore::register_cold_segments, the COLD 60s event-loop timer arm (tokio + monoio), cold_tier_experimental_enabled(), cold-segment manifest discovery in src/persistence/recovery.rs (RecoveryResult::cold_segments / cold_segments_loaded), and the snap.cold fan-out in FT.SEARCH / FT.INFO / SegmentHolder::resident_bytes/total_vectors.
  • --segment-cold-after, --segment-cold-min-qps, --vec-diskann-beam-width, --vec-diskann-cache-levels are kept as parseable, inert no-ops (rather than a hard CLI break) so existing moon.conf files/launch scripts do not fail to start; a startup tracing::warn! fires once if any is set away from its historical default (ServerConfig::warn_deprecated_cold_tier_flags, called from main.rs).
  • StorageTier::Cold/Archive (the shared on-disk tier-byte enum in src/persistence/manifest.rs, format §4.3) are left in place — they are part of the stable FileEntry.tier wire format and shared with the unrelated KV disk-offload subsystem's ColdIndex/sweep_orphans (which this change does not touch).

Fixed — Vector: WARM segments now survive a restart, and stop leaking superseded segment directories

  • Restart-recovery dead wiring fixed. VectorStore::register_warm_segments had zero non-test callers: Shard::restore_from_persistence's v3 recovery pass discovered WARM segments from the manifest (RecoveryResult.warm_segments) into a throwaway Shard-owned vector_store that event_loop.rs discards wholesale in favor of the live ShardSlice.vector_store — so the discovery result was thrown away with it, and WARM's RSS win evaporated on every restart. Fixed by staging the discovered segments on a new Shard::recovered_warm_segments field and draining them into the live store via register_warm_segments right after B3 recovery (RecoveryState::finish) completes in event_loop.rs.
  • Superseded segment-directory leak fixed. VectorIndex::try_warm_transitions_idle never told Stack B's durability manifest (vector/persistence/manifest.rs) that a segment had left the immutable tier, so the old idx-<hex>/segment-<old_id>/ directory it left behind was never garbage collected — permanent disk growth on every HOT→WARM/COLD transition, and (independently) the reason a restart used to silently reload the segment as a fully-materialized HOT copy of stale data instead of respecting the WARM tier it had just been moved to. Fixed by calling the existing persist_hook_after_install durability hook after every transition; its manifest diff (run_snapshot_job) now GCs the superseded directory the same way a normal compact/merge already does.
  • VectorIndex::persist_hook_after_install changed from &mut self to &self so try_warm_transitions_idle (which only holds &self) can call it; next_snapshot_seq changed from u64 to AtomicU64 to keep VectorIndex/VectorStore Sync (fetch_add, Relaxed — every actual writer is still the single owning shard thread, so this only needs to be a valid value, not a cross-thread synchronization point).
  • COLD (SegmentList.unloaded, the reload-on-touch stub tier) is unaffected in the case that already worked (WARM→COLD idle transitions — those segments were never in Stack B's segment_ids to begin with) and loses only an accidental, undocumented side effect in the other case (HOT→COLD-direct transitions used to reload as stale HOT on restart due to the same leak this fixes) — it still has no dedicated restart-recovery path of its own; a restart now correctly falls through to the existing dedup-rescan/AOF-replay self-heal instead (same fallback as any index with no durable manifest state), never silent data loss.
  • New TDD regression test (vector::persistence::recover_v2::tests:: warm_transition_leak_fix_and_restart_recovery) proves, on real force-compacted on-disk state: (a) the superseded segment directory is GC'd within a bounded wait, (b) a simulated restart does NOT reload it as a phantom HOT segment, (c) register_warm_segments reattaches it as WARM, and (d) post-restart search returns the identical top-1 global_id as the pre-transition baseline — no recall regression. Verified red (fails with the leak-fix call removed) before green.

Fixed — Vector: hand-review found register_warm_segments could permanently duplicate documents or attach a segment to the wrong index

  • CRITICAL — crash-before-GC restart could leave the same vectors live twice, forever. try_warm_transitions_idle commits Stack A's shard manifest durably (synchronous, fsync+commit) before persist_hook_after_install merely schedules the async Stack B snapshot job. A kill -9 landing in that window left the old segment still tracked by Stack B (reloaded as an ordinary HOT/immutable segment by recover_v2) and the warm copy discovered by Stack A (reattached WARM by the restart-recovery wiring above) — the same key_hashes live in two segments. search_mvcc's merge (src/vector/segment/holder.rs, all.sort_unstable(); all.truncate(k);) has no key_hash dedup, so this never self-healed: duplicate keys in FT.SEARCH results, inflated num_docs, and the next snapshot re-adopted the reloaded HOT copy into segment_ids forever. Fixed: register_warm_segments now checks a warm segment's own key_hash set (warm_search::peek_key_hashes, a cheap mvcc.mpf-only read — no codes/graph mmap, no CollectionMetadata dependency) against its owning index's current in-memory key_hash_to_key. If those keys are already covered by a live HOT segment, Stack B wins: the warm copy is left unattached and its on-disk directory is deleted (std::fs::remove_dir_all) instead of being registered, so Stack B stays the single source of truth.
  • HIGH — a warm segment could attach to the wrong index. The restart-recovery wiring above made a previously-dead code path live: register_warm_segments attached each segment to the first index for which WarmSearchSegment::from_files happened to succeed — from_files accepts any caller-supplied CollectionMetadata and never validates it against the on-disk codes/graph, so two indexes of the same dimension/quantization made the outcome a coin flip (wrong collection, wrong search results, no error). Fixed: ownership is now decided from key/keymap evidence, never from from_files success. Each candidate index's persisted Stack-B keymap (read straight off disk via a new read_manifest_and_keymap_consistent helper — bounded retry with backoff against the benign TOCTOU where a concurrent run_snapshot_job advances the epoch and GCs the old keymap file between the manifest read and the keymap read) is checked for key_hash overlap with the segment; the index with the most matches wins. No match, or a tie between two indexes, leaves the segment unregistered (files intact, logged via tracing::warn!) rather than guessing.
  • Two new TDD regression tests in vector::persistence::recover_v2::tests: warm_reattach_dedups_against_crash_before_gc_race stages the exact crash-before-GC state (rolls the on-disk manifest/segment/keymap back to pre-transition after waiting for the real background GC to land, so the rollback is deterministically the last write) and asserts zero duplicate global_ids in search results, warm.len() == 0, and the superseded warm directory is deleted; warm_reattach_picks_correct_index_among_same_dim_indexes builds two same-dimension indexes with disjoint keysets — where the old first-match behavior would have funneled both segments into one index regardless of HashMap iteration order — and proves each segment lands on its true owner via an end-to-end recall check (a cross-attached segment would still pass a bare warm.len() == 1 count but return the wrong top-1 result). Both verified red (fail at the exact assertion the fix satisfies) against the prior "first successful from_files" implementation before green.

Fixed — Vector: WARM restart recovery was still defeated on every NORMAL restart (not just a crash) by call-site ordering

  • CRITICAL — round 2 of hand-review found the previous fix's own wiring re-triggered its own duplication check. The full production B3 sequence is create_index -> keyspace dedup rescan (reconcile_key, once per matching HASH key) -> finish -> register_warm_segments. load_segments_and_keymap never populates key_hash_to_key/key_hash_to_global_id for WARM keys (see its segment_resident gate — nothing has attached the WARM segment to idx.segments yet at that point in create_index), so with register_warm_segments running LAST, the rescan saw every WARM key as unknown and re-encoded it into the mutable segment (full HNSW/TQ rebuild, not just metadata) on every boot. The duplication check added in the previous fix then saw those just-re-indexed keys as "already covered by a live HOT copy" and retired (permanently deleted) the WARM segment — the exact failure this feature exists to prevent, triggered by the fix's own ordering, on every normal restart. Fixed: VectorStore::register_warm_segments now runs immediately after the sidecar create_index loop, BEFORE the keyspace rescan (event_loop.rs) — ownership decisions only need every sidecar index to already exist, which they do by then.
  • On a clean attach, register_warm_segments now populates key_hash_to_key/key_hash_to_global_id/key_hash_to_vec_checksum for the segment's keys, sourced from the same owner-evidence persisted-keymap entries already read for the ownership decision (no extra disk read). This is what lets the rescan (which runs right after) see the keys as known/checksum-unchanged instead of re-encoding them, and lets FT.SEARCH resolve a WARM doc's real key bytes via key_hash_to_key (the exact map resolve_hybrid_doc_key consults) instead of falling through to a synthetic vec:<id>. A segment key_hash absent from the owner's persisted keymap (an earlier async-snapshot job never committed for it — the same class of race as the crash-before-GC finding, at per-key rather than per-segment granularity) is left out of the maps, so the rescan self-heals it into mutable, and is tombstoned in the warm copy for that one key_hash (WarmSearchSegment::seed_tombstones) so the stale warm copy and the freshly re-indexed mutable copy never coexist as live duplicates.
  • RecoveryState's deletion-probe baseline moved to a new explicit phase, snapshot_recovered_baseline. Previously create_index snapshotted the baseline immediately (before any WARM segment existed in-memory), which — now that register_warm_segments populates WARM keys into key_hash_to_key — would have permanently excluded every WARM key from finish()'s deletion probe: a WARM key whose underlying HASH was deleted while the server was down would never be recognized as "used to exist, now doesn't," and would silently survive as a stale search hit forever. event_loop.rs now calls snapshot_recovered_baseline once, after BOTH create_index (all indexes) AND register_warm_segments have settled, and before the keyspace rescan begins — the baseline is then complete for the deletion probe to catch a delete-while-down WARM key too (VectorIndex::mark_deleted_for_key_in_index already tombstones across every tier including WARM, so once the probe fires the fix is mechanical).
  • New TDD integration-level regression test, vector::persistence::recover_v2::tests::warm_tests::warm_recovery_full_production_sequence_matches_real_boot_order, drives the FULL real sequence in order (create_index -> register_warm_segments -> keyspace rescan, replaying every key except one to simulate a delete-while-down -> finish) and asserts: (a) the WARM segment attaches and is not retired, (b) every still-live WARM key is counted verified_unchanged (0 re_indexed, nothing lands in the mutable segment), (c) search resolves the ORIGINAL key bytes for a WARM doc via key_hash_to_key (not vec:<id>), (d) no duplicate key_hash across results, (e) the delete-while-down key is tombstoned by the deletion probe and absent from search. Verified red by hand (temporarily moving register_warm_segments back to run after the rescan/finish, matching the prior ordering): (a) fails first — the WARM segment is retired as a false-positive duplicate (warm.len() == 0) instead of attaching, which also makes (b)/(c) unreachable. The three existing WARM tests from the previous fix (warm_transition_leak_fix_and_restart_recovery, warm_reattach_dedups_against_crash_before_gc_race, warm_reattach_picks_correct_index_among_same_dim_indexes) and one pre-existing deletion-probe test (recover_finish_tombstones_keys_missing_from_rescan) were updated to call register_warm_segments/snapshot_recovered_baseline in the new production order so they keep testing real behavior instead of a now-stale sequence; two of them needed a full keyspace-rescan replay added (previously implicit/absent) to avoid the deletion probe wrongly tombstoning their now-baseline-visible WARM keys. crash_recovery_vector_durability.rs's six real-server-spawning scenarios (s1-s6) were re-run against a release-fast build as an additional non-WARM regression check on the reordered event_loop.rs sequence — all pass unchanged.
  • vector::persistence::recover_v2's test module, having grown past the 1500-line file size guideline with these additions, was split: recover_v2_tests.rs keeps the non-WARM B3 recovery tests + shared helpers, and a new nested recover_v2_warm_tests.rs (#[path]-included as mod warm_tests from within recover_v2_tests.rs) holds all WARM-specific tests.

Docs — Roadmap: native moon:// / moons:// connection URI scheme

  • Define Moon's native connection URI scheme in docs/roadmap/ROADMAP.md as a superset of Redis's redis:// / rediss:// (moons:// = TLS 1.3 variant, parity with rediss://). Adds the full cross-cutting spec (§8.5), the v0.6.1 spec-doc task (H-7), the v0.7.0 implementation workstream (R6), and a Protocol node in the features mindmap. redis:// / rediss:// remain fully supported; the native scheme adds a ?workspace=<tenant> selector and a fail-fast, no-opportunistic-downgrade TLS contract.

Docs — moon:// / moons:// connection URI scheme spec (H-7, PR #261)

  • New docs/protocol/moon-uri.md: the authoritative, doc-only spec for Moon's native moon:// / moons:// connection URI scheme — ABNF grammar, a field-by-field redis(s):// parity table, a query-parameter reference (parity + Moon-native ?workspace=), the design-for-failure section (no opportunistic downgrade, no auto-upgrade, fail-fast diagnostic text, unknown-scheme parse errors, bounded connect), server-participation semantics (--announce-url, INFO replication, cluster redirects, REPLICAOF), worked examples, and the conformance checklist the v0.7.0 Workstream R6 implementation + tests/uri_scheme.rs must satisfy. No Rust code changes — implementation is tracked separately as v0.7.0 R6.
  • Verified every cited flag/module against the current tree (--tls-port et al. in src/config.rs, TLS 1.3/rustls/aws-lc-rs + SIGHUP hot-reload in src/tls.rs, WS AUTH UUID-only argument in src/command/workspace.rs) rather than restating the roadmap's prose uncritically; flags the WS AUTH UUID-vs-name gap for the R6 implementer as an open decision instead of assuming it away.
  • mkdocs.yml: adds a "Protocol" nav section for the new page so mkdocs build --strict doesn't fail on an orphan doc.

Fixed — Cold-tier proactive TTL reclaim (H-2, R1: tmp/OFFLOAD-COMPRESSION-REVIEW.md)

  • TTL-expired disk-offload cold entries that are never re-read no longer leak forever. The on-read reclaim path (cold_read.rs) only reclaimed an expired ColdIndex entry when a caller issued a GET against it; a key that expired and was never touched again (the flagship offload use case — TTL'd sessions, caches) permanently leaked its index entry (RAM, full key bytes) and pinned its backing DataFile's refcount (disk), because nothing else in the system ever inspected a cold entry's TTL.
  • ColdLocation now carries ttl_ms: Option<u64>, a cached copy of the on-disk KvEntry::ttl_ms populated at spill time and re-derived fresh by rebuild_from_manifest on every restart — so the sweep can judge expiry from the in-RAM index alone, without a pread of the cold file. No on-disk format changed and ColdIndex has no serialized form of its own, so there is no index format to version and no old file that could fail to load.
  • New ColdIndex::sweep_expired runs in the same tick as the existing cold-tier orphan sweep (--cold-orphan-sweep-interval-secs, default 5 minutes), bounded to MAX_EXPIRED_SWEEP_BATCH (4096) entries per call so an expiry storm against a large cold index cannot stall the shard event loop; any remainder is picked up by the next tick. Batch-file colocation is preserved exactly like the orphan sweep (an expired key never drags down a co-located live key's file).
  • New INFO fields reclamation_cold_expired_reclaimed_total / reclamation_cold_expired_bytes_reclaimed_total (distinct from the existing orphan counters, which count any zero-ref file unlink regardless of cause).

Docs — H-4 doc reconciliation (ROADMAP §8.1)

  • docs/redis-compat.md: FT.AGGREGATE was listed as "Not implemented"; it is actually implemented and dispatched (src/command/vector_search/ft_aggregate.rs / src/text/aggregate.rs, wired through src/command/mod.rs and the single/sharded/monoio FT.* handlers on both read and write paths). Corrected to "Implemented", noting the APPLY-stage v1 stub, the FILTER-clause no-op, and that it has no phf COMMAND_META entry by design (so COMMAND INFO/COMMAND DOCS return nothing for it even though it works).
  • docs/comparison-valkey.md: "Hash field expiration — HEXPIRE/HTTL family Moon does not implement" was stale — Moon has shipped the full family (HEXPIRE, HPEXPIRE, HEXPIREAT, HPEXPIREAT, HTTL, HPTTL, HEXPIRETIME, HPEXPIRETIME, HPERSIST, HGETDEL, HGETEX) with full read+write dispatch and WAL replay since before this document's last refresh. Moved from "Valkey wins decisively" to "Effective parity" (Moon trails Valkey by 4-10% on these ops per docs/perf/2026-05-27-hash-ttl-3way-bench.md — a tracked perf gap, not a missing feature); also dropped it from the cluster-features risk bullet in §4, since it is unrelated to cluster mode.
  • SECURITY.md: supported-versions table said "0.1.x" (stale since well before v0.6.0). Updated to the current v0.6.x line, with an explicit pre-1.0 support policy (latest-minor-only; no LTS/backport promise until the v1.0.0 GA tag, per docs/roadmap/ROADMAP.md's stated 18-month LTS policy).

Docs — Production contract refresh + CI gate (ROADMAP H-5)

  • docs/PRODUCTION-CONTRACT.md refreshed from its stale v0.1.3-era draft to v0.6.0 reality and converted into a checked, evidence-linked ledger (56 GA-blocking rows across toolchain/CI, correctness hardening, durability, replication & HA, cluster mode, security, observability, compatibility, and release engineering). Every ✅ row cites a real file/CI job/script that was read and confirmed during this pass — no aspirational ticks. 33/56 GA-blocking rows are shipped today; 22 are honestly unticked (multi-shard replication, WAIT/keyspace-notifications/MONITOR, monoio cluster wiring, encryption at rest, and others already tracked in docs/roadmap/ROADMAP.md §8). The verification pass also surfaced two previously undocumented defects, left unticked rather than assumed fixed: deny.toml's header comment claims a ci.yml "safety-audit" job that does not exist (cargo audit/cargo deny are not CI-blocking), and fuzz/Cargo.toml + .github/workflows/fuzz.yml both reference a 12th fuzz target (graph_props_record) whose source file is missing from the tree, so that nightly fuzz matrix shard fails every run.
  • New scripts/check-production-contract.sh (grep-based, mirrors scripts/audit-unsafe.sh's style): reports unticked GA-blocking ledger rows on every tag, and hard-fails only on a v1.0* tag if any remain unticked. Wired into .github/workflows/release.yml's setup job alongside the existing RELEASES.md release-ledger gate, using the same env-var interpolation pattern (VERSION: ${{ steps.ver.outputs.version }}, never inline ${{ }} in a run: body).

Changed — Memory: vector tiering accounting spine (M1, tiering-v2 D3/D9)

  • --maxmemory now counts vector segment memory. The background eviction check compared only KV bytes against the per-shard budget, so a pure-vector workload could drive RSS to OOM while eviction reported "under budget". timers::run_eviction now gates on the shard AGGREGATE — Σ all dbs' KV + the shard's published vector bytes, computed once per 100ms tick — and evicts across dbs only until the aggregate is back under budget (adversarial review caught the initial per-db formulation, which both under-detected with KV spread across dbs and over-evicted sibling dbs). When the un-evictable vector term alone exceeds the budget, KV drains then errors OOM (shared-budget semantics; per-db quotas remain the tenant-isolation mechanism); the pressure cascade (which shrinks vectors via offload) fires earlier at --disk-offload-threshold, so with disk-offload enabled vectors shed first. Known limitation: the on-write eviction gate still checks KV-only, so under noeviction a vector-heavy shard over the vector-aware budget does not yet reject client writes — RSS is bounded by the pressure cascade and the RSS watchdog instead (write-gate consistency is a tracked M1 follow-up).
  • Elastic budget classification is vector-aware. A vector-heavy/KV-light shard was misclassified as an idle donor, lending headroom to siblings while its true footprint was over base — and the pressure cascade compared a vector-inclusive used-term against a budget inflated by that donation. The donor/hot snapshot now sums KV + vector per shard; a vector-heavy shard is classified hot and borrows instead of donating.
  • --vec-warm-mmap-budget is now an instance-total cap divided across shards (matching --maxmemory semantics). Each shard previously applied the full value — an N-shard instance silently allowed N× the configured WARM memory. Behavior change: multi-shard deployments relying on the old per-shard meaning should multiply their flag value by the shard count. A nonzero total floors at 1 byte/shard ("0" still disables enforcement).
  • IVF and DiskANN-cold segments report real resident_bytes(). Both tiers contributed a hardcoded 0 to the roll-up (untracked RAM: IVF centroids + posting lists; DiskANN PQ codes + codebook). They now feed the pressure trigger, MEMORY DOCTOR, and Prometheus like every other tier.
  • INFO reclamation_mmap_warm_bytes reports the live WARM counter. The field read a never-incremented local static (permanent 0) while the real MmapBudget counter was write-only. It now reads the live counter, and a new reclamation_mmap_budget_evictions_total field exposes cumulative byte-cap evictions.
  • D9 (DiskANN retention) quarantine: config docs now disambiguate COLD-stub (unloaded, exact reload-on-touch default valve) vs COLD-ann (cold, DiskANN serve-from-disk, inert behind MOON_VEC_COLD_TIER); the dead knobs are marked [reserved: M3/M5]. Keep-vs-delete is decided at the M3 exit-review on real per-index query-frequency telemetry.

Security — ACL now covers the early-intercepted command families (H-3) (PR #258)

  • CDC.READ bypassed ACL entirely on the monoio runtime (the default build). It was dispatched ~117 lines before the ACL gate, so any authenticated user — even a +get-only one — could run CDC.READ <any-wal-dir> <lsn> and read arbitrary WAL directories off the server disk. Moved the CDC.READ intercept to after the ACL gate (matching the tokio/sharded handler, which already ordered it correctly).
  • Category carve-outs silently didn't cover TXN/WS/MQ/TEMPORAL/CDC.READ (all runtimes). These connection-handler intercepts appeared in no get_category_commands() arm, so the common deny-list idiom (+@all -@dangerous, +@all -@transaction, +@all -@write) left them allowed. Added them to the relevant categories: ws/cdc.read → @admin + @dangerous; txn/temporal → @transaction; mq/txn → @write; all five → @all.
  • -@pubsub didn't block PUBLISH/SUBSCRIBE at the command level. The pub/sub intercepts consulted only the &pattern channel rule, so a -@pubsub carve-out was ineffective for a user with &*. Added a command-level ACL check to the PUBLISH and SUBSCRIBE/PSUBSCRIBE paths in the monoio and tokio-single handlers (the sharded handler already gated them post-ACL-gate).
  • CLIENT TRACKING ran before the ACL gate on monoio. Moved it to a post-ACL try_handle_client_tracking (it registers server-side invalidation state); CLIENT ID/SETNAME/GETNAME stay pre-ACL as connection-local metadata.
  • Tests: new unit test pins the category expansion for all five families; new integration tests prove end-to-end NOPERM for CDC.READ under -@dangerous and PUBLISH under -@pubsub.

Fixed — Vector: memory-aware WARM offload with a real, reloadable ceiling (PR #252)

  • Reloadable byte-cap eviction (A). MmapBudget::enforce_budget (the per-shard --vec-warm-mmap-budget cap, default 2gb) evicted a WARM segment by dropping its Arc outright with no COLD stub — so a byte-cap eviction silently removed the segment from search until process restart (recall loss), despite the doc claiming reload-on-touch. It now demotes each evicted segment to a reloadable UnloadedSegment stub in unloaded (every WARM segment is durably disk-backed via transition_to_warm, so the stub reloads byte-identically). This makes --vec-warm-mmap-budget a real, reloadable memory ceiling. The eviction tick now also takes the holder reload_lock so it serializes with the reload/install path instead of relying solely on shard-thread affinity.
  • WARM memory accounting fix. SegmentHolder::resident_bytes() hardcoded the WARM/COLD tiers to 0, so a shard whose HOT segments had aged into WARM reported ~0 vector memory — blinding both INFO/Prometheus and the memory-pressure trigger below. It now sums the WARM tier (the dominant term for a long-lived shard) plus COLD stubs. (IVF/DiskANN-cold still lack a resident accessor.)
  • Memory-triggered early offload (C). Vector segment memory was invisible to every pressure mechanism, so a vector-heavy shard (the primary disk-offload case) only ever offloaded on the wall-clock idle timer (--engine-offload-idle-secs, default 3600s), never on RAM pressure; and the pressure cascade's step 2 demoted HOT→WARM (no RSS win). Now the shard's vector resident bytes (HOT immutable + WARM) count toward the pressure trigger, and under pressure step 2 offloads idle vector segments straight to COLD (real RAM reclaim, reloadable) at an aggressive 60s idle floor. Known follow-up: HOT-immutable memory still has no standalone byte budget (only idle/pressure demotion bounds it).

Fixed — Disk-offload: crash recovery no longer resurrects deleted cold keys, nor drops the AOF after a spill (PR #257)

  • Deleted/flushed cold keys stayed deleted only until the next crash. A DEL/UNLINK/FLUSHALL/FLUSHDB of a spilled key tombstones the in-memory ColdIndex, but the manifest entry stays Active until the orphan sweep. A kill-9 inside that window lost the tombstone, and boot-time replay could not re-apply it: the replayed DEL ran against databases whose cold_index was None — a silent cold-plane no-op — so the key resurrected via cold read-through (measured 85–97/200 probes under --appendonly yes + --disk-offload). Two detach points fixed: (a) recover_shard_v3_pitr now attaches the Phase-3-rebuilt ColdIndex to the databases BEFORE Phase 4 WAL replay (was: stashed on RecoveryResult, merged only after recovery returned); (b) the per-shard/multi-part AOF replay now keeps cold wiring live during incr replay — main.rs re-attaches it after the pre-replay hot wipe, and replay_per_shard/replay_multi_part bridge it across rdb::load's wholesale database swap (which silently dropped it). New crash suite tests/crash_recovery_cold_del_resurrection.rs (DEL + FLUSHALL scenarios, kill-9 inside the pre-sweep window).
  • Any disk-offload spill silently discarded the AOF on the next restart. The "WAL replayed 0 commands → fall back to the AOF" gate counted vector + file-lifecycle records: after a spill wrote FileCreate records into the shard WAL, the gate saw a "non-empty" WAL, skipped appendonly.aof (the only complete KV history — --wal-kv-log is auto-off when the AOF is the authority), and every KV write since the last snapshot was lost. The gate now keys on KV Command records only.

[0.6.0] — 2026-07-08

Release highlights (full detail in the sections below; "PR #TBD" entries below all shipped together in this release's PR):

  • Multi-db isolation is now a first-class, enforced boundary: FT.*/graph/ full-text indexes are scoped to the db that created them (SELECT 1's indexes are invisible to db 0, verified across shards, restarts, and the recovery path); per-db memory quotas (--db-maxmemory <db>:<bytes>, CONFIG SET db-maxmemory) enforce noeviction/evicting policies per db slot with Redis-style deny-OOM semantics (shrink commands always pass, so a tenant can never wedge itself); named workspaces harden the prefix layer.
  • Memory accounting is now truthful under container growth: HSET/LPUSH/ SADD/ZADD growth into existing keys is charged to used_memory in O(1) (previously invisible to --maxmemory — an unbounded single-key hash could never trigger eviction).
  • Engines offload when idle: vector segments demote HOT→WARM (mmap)→COLD (unloaded stub, reload-on-search) on configurable idle/age thresholds (--engine-offload-idle-secs, --segment-warm-after), with DEL/HDEL tombstone correctness across all tiers — measured −26% process RSS on a 40K×768d corpus with search results identical after reload.
  • Single-node tuning preset: --profile standalone fills in the measured best flags for a shard-1 deployment (the p=1 busy-poll configuration that beats Redis on both GCE arches).
  • Command parity widened: SORT_RO / BITFIELD_RO / GEORADIUS_RO / GEORADIUSBYMEMBER_RO plus subcommand gaps; the compat matrix is regenerated. PFDEBUG/PFSELFTEST/FAILOVER/MODULE/SENTINEL are documented non-goals.

Operator callout — historical WS DROP key leak. Before this release, WS DROP's cleanup sweep was hardcoded to logical db 0: any workspace whose connection ever SELECTed a non-zero db before writing leaked those keys permanently on drop. v0.6.0 fixes the sweep (all dbs on the owning shard), but keys leaked by PAST drops are still resident. To detect/clean on an upgraded instance: SELECT <n> each non-zero db and SCAN 0 MATCH <ws-uuid>:* for workspace prefixes that no longer appear in WS LIST, then DEL the matches (or FLUSHDB if the db held nothing else).

Absorbed post-release-PR merges. The following sections merged after the v0.6.0 release PR but before the tag was cut (2026-07-10) and are part of the tagged v0.6.0 binary.

Performance — Vector: COLD-segment reload moved off the shard event loop (PR #251)

  • Under --disk-offload, an idle vector segment is demoted to the COLD tier (UnloadedSegment stub, everything in-memory dropped). The first FT.SEARCH to touch it must reload it — WarmSearchSegment::from_files: mmap + page-in (+ optional mlock) of a whole segment. That reload ran inline on the single-threaded shard event loop (promote_unloaded at the top of the yielding search capture), so it stalled EVERY other connection on the shard for the full reload (hundreds of ms–seconds on real NVMe/EBS with a cold page cache) — a multi-tenancy violation (audit finding 18).
  • A new process-global SegmentReloadPool (src/vector/reload_pool.rs, auto-on; MOON_VEC_RELOAD_WORKERS=0 disables) now performs the reload I/O on dedicated worker threads. The yielding capture path calls submit_unloaded_reloads (non-blocking): it submits each COLD stub off-loop and stashes the flume receivers in the SearchSnapshot; the handler awaits them (await_pending_reloads) before scanning, which parks only the triggering query's task (not the OS thread) and then splices the reloaded segments into that query's own snapshot. Result: the dense-KNN query that triggered the reload still sees full recall, while sibling connections on the shard keep running. Reloads are single-flight (concurrent first-touches of one segment share a job) and cached until the next capture installs them into WARM (reclaiming the stub's memory).
  • Design-for-failure: a reload that errors or whose worker panics/dies replies Err/disconnect to every waiter — the query answers with degraded recall (segment stays COLD, retried next touch) rather than hanging or crashing. When the pool is disabled, submit_unloaded_reloads falls back to the blocking promote_unloaded, preserving exact pre-#18 behavior. Non-yielding sync search paths are unchanged (still block); HYBRID/SPARSE/RANGE/non-default-field queries do not take the off-loop path (accepted follow-up).

Ops — v0.6.0 release-ledger closure + hygiene (PR #256)

  • RELEASES.md gains the missing v0.6.0 entry (shipped 2026-07-08, ledger lagged); release.yml now fails a version-tag push whose entry is absent from RELEASES.md — the ledger can never lag a tag again.
  • replication.state (root artifact from replication test runs) gitignored.

Docs — engineering roadmap suite under docs/roadmap/ (PR #254)

  • New planning docs (excluded from the published docs site via mkdocs.yml exclude_docs): product roadmap v0.7→v1.0 with features mindmap and task-level execution plans (ROADMAP.md), scale/HA architecture incl. multi-shard PSYNC wire-format draft and failover timeline (scale-ha-architecture.md), and the standalone horizontal-scale deep dive (standalone-horizontal-scale.md). Commercial/EE planning is maintained privately.
  • Docs only — no runtime code changes.

Changed — CI: give the macOS Test (tokio) step 30m (was 15m)

  • The macOS Test (tokio) step budget includes compilation. On a cold or rebased cache, ~11m compile + ~8m test suite exceeds 15m and trips a wall-clock timeout with zero actual test failures. Bumped to 30m to match the Windows Test step's headroom; removes recurring false-negative reds on rebased branches.

Fixed — Conn-plane: monoio central listener joins the SO_REUSEPORT group (PR #250)

  • Under the monoio runtime, each shard binds bind:port via SO_REUSEPORT, but the central listener bound the same port with a plain monoio::net::TcpListener::bind — unlike the tokio sibling, which uses create_reuseport_socket. Linux requires every socket sharing a REUSEPORT port to set the option, so depending on startup scheduling the central bind either lost the race (EADDRINUSE propagated out of run_sharded, silently killing the TLS listener + conn_rx fallback) or won it (forcing every shard's REUSEPORT bind to fail and collapsing all traffic onto the single central accept loop) — both silent and nondeterministic run-to-run. The monoio central listener now binds via create_reuseport_socket + from_std like the tokio path (plain-bind fallback on non-unix / unparseable addr). (audit finding 16)

Fixed — Durability: checkpoint Finalize backs off on repeated failure instead of flooding the WAL (PR #250)

  • When a checkpoint's Finalize step failed (wal.wait_durable, manifest.commit, graph_save, or the control-file write), the state machine stayed Finalizing and the next 1ms persistence tick re-ran finalize from step 1 — re-appending a WAL Checkpoint record every single millisecond with no backoff. Under a sustained failure (slow/degraded disk) this floods the WAL, and since last_checkpoint_lsn never advances, recycle_segments_before never fires — WAL disk usage grows fastest exactly during the disk-pressure incident it should be backing off from. CheckpointManager now arms an exponential backoff (50ms→5s cap) on each failed finalize; the tick checks finalize_ready(now) before doing any finalize I/O, so retries are bounded instead of per-tick, and a successful complete() clears it. (audit finding 13)

Fixed — Durability: legacy AOF replay stops at mid-stream corruption instead of resyncing (PR #250)

  • The default startup path replay_aof (single-file appendonly.aof, unframed RESP with no per-record length/CRC) treated ANY mid-stream parse error by discarding one byte, scanning forward to the next *, and resuming dispatch — but a * appears freely inside binary value blobs, so the resync could land mid-value and execute misaligned garbage as live SET/DEL/etc. against the recovering keyspace (silent data corruption, WARN-only). The per-shard shard_replay sibling already rejects exactly this. replay_aof now STOPS at the first genuine mid-stream corruption (clean tail-truncation is still handled gracefully), keeping only the valid prefix already applied and never dispatching past the corruption point; it logs a loud error with the byte offset and points at redis-check-aof. Operators who want the old best-effort behavior can opt in with MOON_AOF_BEST_EFFORT_RESYNC=1. (audit finding 12)

Fixed — Xshard: bound cross-shard reply wait so a wedged shard can't hang a client (PR #250)

  • Every cross-shard leg (MGET/MSET/DEL/UNLINK/EXISTS/BITOP/COPY/MSETNX remote legs, MULTI-on-owner, and the flush/scatter-gather loops) pushed a message via the backoff-retried spsc_send and then did reply_rx.recv().await with no timeout. spsc_send's retry budget only bounds getting the message INTO the target ring buffer; if the push succeeds but the target shard then stalls while executing (wedged-disk fsync during snapshot/AOF, uninterruptible D-state I/O, or a dead shard), the awaiting connection parked forever with no response and no way to cancel. Added a recv_reply_bounded helper (races the receiver against a runtime-agnostic 30s XSHARD_REPLY_TIMEOUT via race2) and applied it to all 11 reply-await sites in the coordinator — a genuinely wedged shard now surfaces the existing cross-shard-reply error instead of hanging. (audit finding 11)

Fixed — Hardening: bound TLS handshake + cluster-bus body reads against slow-loris (PR #250)

  • TLS handshake has no timeout (#17): after try_accept_connection consumed a maxclients slot, acceptor.accept(tcp_stream).await ran with no deadline on both the monoio and tokio TLS paths. A peer that completed the TCP handshake on the TLS port and then sent no (or a partial) ClientHello pinned the task, its fd, and its maxclients slot indefinitely at zero CPU cost — a classic slow-loris that could exhaust the connection limit. Wrapped both accepts in a 10s TLS_HANDSHAKE_TIMEOUT, releasing the slot via record_connection_closed() on expiry.
  • Cluster-bus gossip body read has no timeout (#10): handle_cluster_peer guarded only the 4-byte length-prefix read with a shutdown-cancel select; once a valid length was read, the body read_exact had no timeout and no shutdown arm. A client could send a valid length header and never send the body, blocking the spawned task forever (immune to graceful shutdown, socket never reclaimed). Both runtimes now read the body under a GOSSIP_BODY_READ_TIMEOUT (10s) + shutdown-cancel select.

Fixed — Txn: MULTI locality analysis honors SORT/GEORADIUS STORE destination (PR #250)

  • In a multi-shard deployment, MULTI; SORT src STORE dst; EXEC (or GEORADIUS ... STORE/STOREDIST dst) where src and dst hash to different shards was misclassified as single-shard and routed entirely to src's shard — writing dst into the WRONG shard's dataset, invisible to a later normally-routed GET dst. The STORE/STOREDIST destination is positional (the arg after the token) and so was not covered by the fixed command-metadata key specs analyze_txn_locality consulted. It now detects the STORE clause and forces a CrossShard classification, which the caller rejects with CROSSSLOT instead of silently misrouting. (audit finding 15)

Fixed — Vector: DEL tombstones secondary fields + FT.INFO num_docs accuracy (PR #250)

  • Multi-vector-field deletion resurrection (#20): DEL/UNLINK/HDEL on a document only tombstoned the default vector field. Secondary VECTOR fields live in a separate field_segments map that tombstone_key_in_index never iterated, so a subsequent FT.SEARCH idx '@field2:[...]' still returned the deleted document (under a synthetic vec:<id> key, since the shared key-hash→key map was cleared). Now every field's segments are tombstoned.
  • FT.INFO num_docs over/under-count (#28): the mutable segment's len() counts tombstoned entries in place until compaction, inflating num_docs by up to compact_threshold under DEL/HSET churn — switched to a new live_len() that excludes tombstones. Separately, DiskANN cold-tier segments (experimental, gated by MOON_VEC_COLD_TIER) were never summed at all even though cold docs ARE returned by FT.SEARCH; added a snap.cold counting loop. (audit findings 20, 28)

Fixed — Vector: GraphUnion merge recall gate no longer lets a total collapse through (PR #250)

  • The background/manual GraphUnion merge recall gate read recall < tolerance && recall > 0.0, so a merge whose verified recall was exactly 0.0 (total collapse) slipped past the guard and committed. Since verify_merge_recall returns 1.0 (not 0.0) for the too-few-vectors cases, 0.0 only ever means a real measured collapse — the && recall > 0.0 clause was removed so such a merge now aborts. (audit finding 21)

Fixed — Hardening: APPEND 512MB cap + reject FT.* inside MULTI on prod handlers (PR #250)

  • APPEND now enforces the 512MB max-string-size limit (matching SETRANGE/SETBIT); repeated APPENDs previously grew a value past the documented limit unbounded. (audit finding 22)
  • FT.* commands are now explicitly rejected inside MULTI/EXEC on the production sharded + monoio handlers (they aren't wired through the txn execution path), matching the existing handler_single guard — a clear "not supported inside MULTI/EXEC" error instead of an incidental one. (audit finding 25)

Changed — Vector storage: fence experimental DiskANN cold tier + drop dead vector WAL (PR #250)

  • The DiskANN cold tier (WARM→COLD segment demotion + on-disk Vamana beam search) was running by default (--disk-offload defaults on, --segment-cold-after defaults 86400) despite being incomplete — no cold-segment deletion (DEL/HDEL never removes cold docs), blocking per-hop uring I/O holding the segment mutex, FT.INFO num_docs under-counting cold segments, unfinished restart reload of PQ codebooks, and a library-code panic! in DiskAnnSegment::new. It is now fenced behind the experimental MOON_VEC_COLD_TIER=1 env gate (default off); the WARM→COLD transition is a no-op otherwise. Warm-tier mmap + LRU eviction (--vec-warm-mmap-budget) continues to handle out-of-RAM indexes. (audit findings 19, 24, 28)
  • DiskAnnSegment::new now returns io::Result instead of panicking on a file-open error, so a cold transition can never abort the shard.
  • Removed the dead vector WAL module (vector/persistence/wal_record.rs): it was never on the live recovery path (superseded by manifest + segment + keymap + dedup rescan) and carried a documented double-apply footgun.

Fixed — Cluster: gossip PING task leak + non-blocking accept send (PR #250)

  • The gossip PING ticker (both tokio + monoio) spawned an unbounded connect+write+read task every 100ms with no timeout; against a dead or partitioned peer the tasks piled up forever (task/FD/memory leak), and the "random peer" it documented was always the first non-self node. Now a single in-flight guard bounds outstanding probes to one, a probe timeout (node_timeout/2, ≥100ms) cancels a hung connect/read, and targets rotate across peers. (audit finding 7)
  • The monoio per-shard accept task used the blocking flume::Sender::send() on a bounded channel; because the accept task shares the shard's single monoio thread with the event-loop consumer, a full channel deadlocked the loop that drains it under connection-storm churn. Fixed by send_async().await (cooperative yield, keeps backpressure). (audit finding 8, Batch B)

Fixed — Security: bound client-controlled allocation counts (DoS class, PR #250)

  • Six wire-reachable command paths fed a client-supplied count straight into Vec::with_capacity / Vec::resize with no upper bound. On a real host the oversized request fails allocation and Rust's handle_alloc_error calls abort() — uncatchable, crashing every shard and connection. A single unauthenticated request was a full-process DoS. Fixed by clamping each count before allocating: FT.SEARCH HIGHLIGHT/SUMMARIZE FIELDS (bound by remaining tokens), HSCAN COUNT (count.min(total)), SRANDMEMBER / HRANDFIELD / ZRANDMEMBER negative COUNT (absolute 2²⁰ cap with a loud ERR COUNT is out of range beyond it — the first-cut relative len()*10 cap silently truncated legitimate requests on small collections, breaking Redis's exact-|COUNT|-with-duplicates contract; caught in review and fixed at all six sites including the pre-existing HRANDFIELD/ZRANDMEMBER ones), BITFIELD SET/INCRBY offset (reject past the 512MB limit, matching SETBIT; GET exempt), and ZMPOP COUNT (pop_count.min(card) — this last site was missed by the audit finders and caught by the allocation sweep). Each fix has a red/green test. From the production-hardening audit (Batch A).

Fixed — Review follow-ups on the hardening sweep (PR #250)

  • Gossip PONG fragmentation (monoio): the failure-detector's PING probe read the 4-byte length and the PONG body with single read() calls, so a valid frame arriving fragmented was silently dropped — leaving pong_recv_ms stale and pushing a healthy peer toward PFAIL. Both reads now use the bus listener's exact-read loop, and both runtimes' probes cap the peer-supplied PONG length at the bus's 64KB frame bound before sizing the buffer (the tokio probe previously allocated an unchecked vec![0u8; pong_len]).
  • AOF corruption offsets are file-absolute: replay's truncation/corruption logs reported offsets relative to the RESP section; with an RDB preamble the operator-facing repair guidance (redis-check-aof, truncation point) pointed at the wrong file location. Offsets now include the preamble length.
  • Cross-shard MSET no longer OKs an unconfirmed leg: a timed-out, closed, or errored remote MultiExecute ack was discarded and the client still got OK. All legs are drained (each was already dispatched, and the local leg still persists), but any failed leg now surfaces as an error instead of a false ack.
  • recv_reply_bounded extended to the handler plane: the coordinator's 30s bounded reply-await now also covers the 11 remaining direct cross-shard acks (WS.DROP cleanup, routed MQ commands, routed graph commands, TXN.COMMIT MQ materialization acks, and the txn-abort remote rollback ack, on both handler_monoio and handler_sharded) — a wedged owner shard can no longer park those connections forever. MQ/graph/TXN owner execution is synchronous on the shard thread (no blocking-wait commands route through these paths), so the bound cannot cut off a legitimate long wait.
  • Checkpoint finalize backoff arms from the failure time: the backoff was armed with a timestamp captured before the finalize I/O; when wait_durable/manifest.commit/graph_save/control-write took longer than the backoff window, the retry deadline was already in the past and the next 1ms tick retried immediately — exactly the flood the backoff exists to stop.

Added — CLI/moon.conf defaults for vector + graph tuning knobs (PR #248)

  • --vector-ef-runtime / --vector-rerank-mult / --vector-exact-beam set the server-wide starting values every NEW vector index is created with (FT.CREATE), so recall-sensitive fleets configure once instead of issuing FT.CONFIG per index. Per-index FT.CONFIG SET always overrides. Same ranges as FT.CONFIG (ef 10-4096 or 0=auto, mult 1-64), validated loudly at startup. All three work as moon.conf keys (vector-exact-beam yes|no).
  • --graph-result-cache-entries / --graph-result-cache-bytes size the per-graph Cypher result cache (previously hardcoded 256 entries / 4 MiB).
  • Startup wiring is first-write-wins process state, installed by both the binary entry and run_embedded; unit tests and library embedders that never install defaults keep the exact pre-flag behavior.

Added — Recall knobs: FT.CONFIG RERANK_MULT + EXACT_BEAM (PR #248)

  • FT.CONFIG SET <idx> RERANK_MULT <n> (1-64, default 4) deepens the HQ-1 exact-rerank stage: the top n·k beam candidates are re-scored with true f16 sidecar distances before top-k truncation, recovering true neighbors the quantized ADC ranking dropped below the default 4·k cut. Cost ~n·k·dim f16 decodes per segment.
  • FT.CONFIG SET <idx> EXACT_BEAM ON|OFF (default OFF) navigates the HNSW beam itself with exact f16 distances instead of quantized ADC estimates — beam candidate selection becomes exact, so recall is graph-limited (Qdrant-parity at equal ef) rather than quantization-limited. QPS cost grows with dimension. Segments without an exact-rerank sidecar (pre-HQ-1 disk reloads) silently keep the quantized beam.
  • Both knobs are per-index, apply to the next FT.SEARCH (all search paths: sync, yielding, worker-pool fan-out, FT.RECOMMEND, hybrid), broadcast to all shards, and persist across restarts via a new v4 index-meta sidecar format (v1-v3 sidecars load with defaults).

Added — Runtime-tunable EF_RUNTIME via FT.CONFIG (PR #248)

  • FT.CONFIG SET <idx> EF_RUNTIME <n> adjusts the HNSW search beam width at runtime without rebuilding the index (RediSearch parity). Range matches FT.CREATE (10-4096); 0 restores the auto heuristic. Applies to the next FT.SEARCH immediately and persists via the index meta sidecar. FT.CONFIG GET EF_RUNTIME added alongside.
  • Fixed — FT.CONFIG SET was local-shard only under monoio: at shards>1 the setting silently applied to 1/N index partitions (AUTOCOMPACT, COMPACTION_WEIGHT, MERGE_RECALL_TOLERANCE were equally affected). SET now broadcasts to all shards like FT.CREATE; GET stays local.

Fixed — TQ ADC estimator is now metric-faithful on raw L2 (PR #248)

  • TQ quantization ranked L2 queries by sphere_dist·‖a‖² — the unit-sphere distance between normalized directions scaled by the document norm. On unnormalized L2 data this makes every small-norm vector look near regardless of direction (gist-960: recall 0.002). All TQ scoring paths (HNSW beam + budgeted variant, mutable brute-force + FastScan pre-filter bound, flat scan, multibit brute-force, A2 decoded-L2) now reconstruct the true metric: ‖a−q‖² = (‖a‖−‖q‖)² + ‖a‖·‖q‖·d̂²_sphere. COSINE/IP scoring is unchanged. Unit tests pin recall on varying-norm data at both the mutable-scan and HNSW-beam levels (7+/10 and 6+/10 vs exact f32 ground truth; 4/10 and worse before the fix). GCE re-verification on gist-960 with explicit TQ4: recall@10 0.784/0.942/0.986 at ef 16/64/256 vs 0.002-0.003 pre-fix (~350x), within ~0.03 of SQ8.

Changed — FT.CREATE L2 indexes default to SQ8 quantization (PR #248)

  • FT.CREATE ... DISTANCE_METRIC L2 without an explicit QUANTIZATION now defaults to SQ8 instead of TQ4. TQ's norm-scaled ADC estimator assumes unit-sphere metrics (COSINE/IP) and collapses on unnormalized L2 data — recall 0.002 measured on gist-960-euclidean (1M × 960d). SQ8 is metric-faithful on raw L2. COSINE/IP defaults are unchanged (TQ4). An explicit QUANTIZATION TQ* + L2 is still honored but logs a warning.
  • Benchmarks: BENCHMARK.md §10.10 — ANN-benchmarks 4-way campaign (Moon vs RediSearch vs Qdrant vs turbovec on glove-200-angular 1.18M / gist-960 1M / glove-100K). Tuning guide gains bulk-load (--max-unflushed-immutable-segments 0), settle (VACUUM VECTOR), and runtime-EF_RUNTIME recipes.

Added — FastScan SIMD vector scan: NEON TBL kernel + live-path integration (PR #248)

  • NEON TBL FastScan kernel (src/vector/distance/fastscan.rs): register-resident 4-bit LUT accumulation via vqtbl1q_u8 for aarch64 — the primary platform previously fell back to the scalar kernel. 32 candidate distances per block at ~2 ns/candidate (128d), 17-21× faster than the per-candidate scalar ADC path.
  • AVX2 kernel overflow fix: nibble-pair distances were added as u8 before widening, silently wrapping for LUT pairs summing past 255 and corrupting rankings. All kernels now widen to u16 before adding and accumulate with saturating adds (scalar/NEON/AVX2 bit-identical across the full u8 LUT range; parity tests no longer mask LUTs to 0x7F).
  • Adaptive quantized LUT builder (build_quantized_lut): FAISS-style per-coordinate bias + global scale chosen so entries fit u8 AND the worst-case accumulated sum fits u16 (no kernel saturation for in-range data), with f32 reconstruction (acc/scale + bias). Replaces the legacy precompute_lut, which hardcoded the v1 1/sqrt(768) CENTROIDS table (encode/search codebook asymmetry — the same recall bug class fixed by codebook v2) and a fixed LUT_SCALE that could overflow both u8 entries and the u16 accumulator.
  • Mutable-segment FastScan pre-filter (src/vector/segment/mutable.rs): the MVCC brute-force scan now maintains a FAISS-interleaved shadow of the TQ4 codes (32-vector blocks, +padded_dim/2 B/vector) and screens candidates with the SIMD kernel; only candidates whose sound lower bound (approx − 0.5·padded_dim/scale) beats the current heap worst get the exact f32 ADC rescore. Results are bit-identical to the plain scan (the bound is a true lower bound; saturation only under-estimates). Measured: top-10 over 20K×128d 1.25 ms → 96 µs (~13×), 5K×768d 1.66 ms → 423 µs (3.9×) on Apple Silicon. Engages above 64 entries (per-query LUT build costs ~2-8 µs); SQ8/A2 collections and TQ-prod scoring keep the existing paths.
  • New criterion bench benches/fastscan_bench.rs (block kernels, LUT build, end-to-end mutable scan A/B); shadow-layout + FastScan-vs-plain equality tests (full scan, chunked scan, MVCC visibility + bitmap-filter paths).

Fixed — WS3 COLD/WARM tier adversarial-review fixes (round 2 follow-up, PR #TBD)

An adversarial review of the round-2 COLD tier (below) found one CRITICAL correctness bug and several hardening gaps. All fixed on the same branch:

  1. CRITICAL: DEL/HDEL resurrection through WARM/COLD. tombstone_key_in_index (src/vector/store.rs) only ever walked mutable + immutable — a HDEL against a key living in a WARM or COLD (unloaded) segment was silently dropped. Sequence: index idles overnight -> HDEL an old doc -> next-day query resurrects it as a phantom hit, and num_docs over-counts. Fixed by giving both tiers a live tombstone set:
  2. WarmSearchSegment gained the same interior-mutable tombstoned_keys/has_tombstones pattern ImmutableSegment already has (mark_deleted_by_key_hash, seed_tombstones, tombstoned_key_hashes, live_count) — search_filtered now filters tombstoned entries out before rerank/truncate, same ordering as the HOT path. Takes effect immediately, no reload required.
  3. UnloadedSegment gained a pending_tombstones set — a HDEL against a COLD key is queued in the stub (no reload forced just to record a delete) and replayed onto the freshly-reloaded WarmSearchSegment in reload() via seed_tombstones.
  4. tombstone_key_in_index now also calls mark_deleted_by_key_hash on every warm and unloaded entry.
  5. Bonus fix found while wiring this up: mvcc_raw_bytes() (the HOT->WARM transition's serialization of MVCC headers) only captures INSTALL-time deletes (mvcc[i].delete_lsn), never the STEADY-STATE tombstoned_keys interior set — so a HDEL landing shortly before a transition was ALSO silently lost, independent of the round-2 COLD work (this bug predates it, in the pre-existing age-based WARM path). try_warm_transitions_idle now seeds the new WarmSearchSegment with imm.tombstoned_key_hashes() right after opening it, closing this gap for both WARM and COLD destinations for free.
  6. FT.INFO's num_docs now sums warm.live_count() / stub.live_count() (total minus live tombstones) instead of the raw total_count(), at all 4 call sites (top-level + per-field, default + non-default field).
  7. New tests: tests/vector_idle_unload.rs's hdel_during_cold_does_not_resurrect and hdel_during_warm_does_not_resurrect — HDEL a doc, wait for the COLD/WARM transition (or vice versa), assert num_docs drops immediately and the doc never resurfaces in a post-transition FT.SEARCH, in both tier orderings.
  8. Reload thread placement — verified empirically, documented honestly. Traced all 3 promote_unloaded call sites (SegmentHolder::search_filtered, SegmentHolder::search_mvcc, and the FT.SEARCH yielding-path snapshot capture in command/vector_search/ft_search/dispatch.rs): all three run inside crate::shard::slice::with_shard(...), i.e. on the shard's own OS thread, before any .await boundary — PR #179's off-loop worker pool only carries the HNSW beam search itself off-loop, not the capture phase a COLD reload runs in. This means a first-touch reload is a SHARD-WIDE stall (every connection sharing that shard, not just the querying one), for the reload's duration. A real off-loop fix needs SegmentHolder reachable from the async continuation without an open VectorIndex borrow (e.g. Arc-wrapping it, a ~100-call-site change) — judged out-of-scope / too risky for this pass and tracked as a follow-up rather than attempted half-finished. docs/guides/tuning.md now documents this precisely instead of the earlier (incorrect) "one-query latency hit" framing, plus a same-instance measurement of the stall's actual duration: the first-touch FT.SEARCH reload took 79.59 ms round-trip, during which a concurrent PING on a second connection to the same shard peaked at 76.70 ms (vs sub-millisecond baseline) — confirming the block is shard-wide and roughly the full reload duration, not a per-connection cost.
  9. WARM segments could never reach COLD. try_warm_transitions_idle only ever scanned snapshot.immutable for idle/age eligibility — a segment that had already demoted to WARM (age-based) had no further path to COLD even if it then went idle, defeating the whole point of the COLD tier for that segment permanently. Fixed: the same function now also scans snapshot.warm (WarmSearchSegment::idle_secs, a new method mirroring ImmutableSegment::idle_secs) and demotes idle WARM segments straight to COLD stubs via the same UnloadedSegment::from_warm used for the immutable path — mechanical, since a WARM segment's .mpf files already exist on disk.
  10. Repo-rule compliance: no per-query SystemTime::now() syscalls. WarmSearchSegment::touch_last_access/from_files and ImmutableSegment's module-level now_micros() called SystemTime::now() directly on a path that runs once per query per segment. Both now read crate::storage::entry::current_time_ms() — the existing shard-tick-populated thread-local cached clock (TL_NOW_MS, updated once per ~1ms shard tick) that already backs the KV hot path, with its documented automatic syscall fallback on threads that never ticked a shard clock (tests, or an off-loop FT.SEARCH worker thread). Millisecond resolution (×1000 for unit parity) is sufficient since every caller only compares age_secs/idle_secs at second granularity.
  11. Tightened the RSS regression test. cold_unload_reduces_process_rss asserted a plain rss_after < rss_before, which would still pass for a regression that only frees (say) the HNSW graph while leaving TQ codes + sidecar resident. Now asserts rss_after < rss_before * 0.85 (a 15% floor — the real 40K×768d measurement showed −26.2%, so this has margin for the smaller CI-scale fixture and jemalloc arena noise without being toothless).
  12. SegmentList::with_unloaded / with_warm_and_unloaded clone-and-patch helpers. SegmentList now derives Clone (all fields are Arc/Vec<Arc<_>> — refcount bumps, never a data copy) so reconstruction sites can use Rust's functional-update syntax (SegmentList { changed_field, ..old.clone() }) instead of enumerating all 6 fields by hand. Applied at try_warm_transitions_idle's final swap (the site touched by this review's item 3 fix). This does NOT convert all ~17 SegmentList { ... } construction sites in the codebase — only the ones this pass's changes already touch, verified to compile. The remaining sites are left as full literals rather than converted speculatively without individual verification; they will fail to compile the moment a new field (e.g. WS5a's db_index) is added, which is itself the forcing function that gets them individually attended to. The pattern (derive Clone + functional update, plus the two named helpers) is now established for that future pass to reuse.
  13. Freshly compacted segments were unloaded to COLD before they were ever searchable. Two stacked bugs, found via cold_unload_reduces_process_rss timing out 600s waiting for graph_segments >= 1:
  14. ImmutableSegment seeds last_access_micros at construction (build start), but a multi-second HNSW build consumes the whole --engine-offload-idle-secs budget before the segment is published — the idle sweep then unloaded the brand-new segment in its very next tick (server log: "Cold (unload) transition: segment 1 (15000 vectors, idle 3s)" 4.4s after server start). Fixed with ImmutableSegment::mark_installed(), called at every point a built segment is swapped into the live SegmentList (inline + background compaction installs, both merge installs): idle time is now measured from visibility, not from the start of the build.
  15. mark_installed deliberately bypasses the cached clock (item 4 above) with a real SystemTime::now() syscall: the inline FT.COMPACT path builds ON the shard thread, so TL_NOW_MS hasn't advanced for the entire build and stamping the cached value would be build-duration stale — the first iteration of this fix did exactly that and still produced "idle 9s" on a just-installed segment. Install is a cold path; one syscall is fine (does not violate the per-query no-syscall rule).

Added — true COLD tier: idle segments now actually free memory (WS3 round 2, PR #TBD)

Round 1 (below) shipped an idle trigger but routed it into the pre-existing WARM tier, which copies .mpf payloads into owned buffers of the same size as the HOT segment it replaces -- measured flat-to-+1.1% RSS, not a reduction. Round 2 fixes the actual memory problem:

  • New tier, SegmentList.unloaded: Vec<Arc<UnloadedSegment>> (src/vector/persistence/unloaded_segment.rs, new file): a COLD segment stub holding only a handful of scalars (segment id, doc count, whether it had an exact-rerank sidecar, an mlock_codes flag) plus a cloned SegmentHandle keeping the on-disk directory alive. Nothing else -- no TQ/SQ8 codes, no HNSW graph, no f16 sidecar buffer.
  • Idle now routes straight to COLD, not WARM (VectorIndex::try_warm_transitions_idle, src/vector/store.rs): the two triggers now have different destinations. age_eligible (--segment-warm-after) still demotes to WARM exactly as before (an old-but-still-hot segment structurally simplifies to disk-backed storage, no memory-reduction promise). idle_eligible (--engine-offload-idle-secs, genuinely cold) now demotes straight to COLD -- this is the one that actually frees memory. If a segment satisfies both simultaneously, COLD wins. Both destinations write the identical .mpf files via the existing warm_tier::transition_to_warm (sidecar included); COLD additionally drops the materialized WarmSearchSegment immediately after capturing the stub, instead of keeping it resident.
  • Transparent, synchronous, single-flight reload on touch (SegmentHolder::promote_unloaded, src/vector/segment/holder/promote.rs -- split out as a child module to keep holder.rs under the 1500-line cap): every search path (search_filtered, search_mvcc, and the yielding worker-pool capture in command/vector_search/ft_search/dispatch.rs) calls this before scanning -- a KNN query must consider every segment for correctness, so a COLD segment cannot stay unloaded and still participate. Fast path (nothing unloaded) is a single is_empty() check, no lock, no allocation. When there IS something to reload, a blocking (never try_lock, never held across .await) reload_lock mutex makes it single-flight: N concurrent callers touching the same COLD segment all wait for the one reload (double-checked re-load after acquiring the lock) rather than racing ahead with a stale, missing-segment snapshot. Reload reuses WarmSearchSegment::from_files -- the same function the WARM tier and server-boot recovery already use -- so recall/exactness (sidecar included) is unaffected; a segment that fails to reload (corrupt/missing files) stays in unloaded and is retried on the next touch without poisoning the rest of the batch.
  • FT.INFO: new additive counters unloaded_segments / unloaded_segments_with_exact_rerank (the latter answered from the stub, no reload needed), and num_docs (both top-level and per-field) now also sums unloaded.iter().map(UnloadedSegment::total_count) -- the same regression class fixed for WARM in round 1, this time for COLD.
  • Lifecycle: VectorStore::drop_index / clear_all_contents (FLUSHALL/ FLUSHDB) now also tombstone every unloaded stub's SegmentHandle, or its on-disk directory would leak once the stub itself is dropped (WARM segments already got this treatment; COLD needed the same). GraphUnion merge scheduling (VectorIndex::needs_merge / begin_background_merge) only ever reads/filters snapshot.immutable (HOT segments) -- it never touches warm, cold (DiskAnn), or unloaded, so COLD segments are skipped by construction, not by added special-casing. Server restart: COLD segments have no restart-time restoration path, same as WARM already didn't -- the segment directories they point at are simply reloaded as fresh HOT/immutable segments by the existing boot-recovery scan (persistence/recover_v2.rs); "should be free" per the original ask.
  • RSS re-measurement (same methodology as round 1: real server, --shards 1, 40,000 x 768-dim vectors, SQ8, ps -o rss=) -- this time a real drop, not the round-1 flat-to-increase result:
Phase RSS (KB) Delta
Before unload (HOT) 400,544 --
After unload (COLD) 295,760 -104,784 KB (-26.2%)
After reload (touched, back to WARM) 395,376 +99,616 KB vs COLD

The reload number lands just under the original HOT figure because the segment reloads into the WARM (owned-buffer) representation, which is itself ~1.3% smaller than the original HOT ImmutableSegment layout in this run -- consistent with round 1's WARM-vs-HOT measurement, not a new effect. - Red/green: tests/vector_idle_unload.rs's hot_segment_idle_unloads_to_cold_and_preserves_recall (supersedes round 1's hot_segment_idle_unloads_and_preserves_recall, which polled warm_segments -- now stays 0 for a purely-idle transition) walks HOT -> COLD -> reload -> WARM, asserting FT.INFO counters at every stage (graph_segments, warm_segments, unloaded_segments, *_with_exact_rerank, num_docs) plus recall/exactness before and after. cold_unload_reduces_process_rss is the process-level RSS proxy (40,000 → scaled down to a CI-reasonable 3,000 × 256-dim fixture -- see the full 40K×768d numbers below for the production-scale claim) asserting a real ps RSS drop, not just a counter flip. unloaded_segment.rs's own unit tests cover stub-capture + reload round-trip and tombstone-triggers- directory-removal. - Correctness edges considered and their disposition: concurrent search during unload -- covered by single-flight reload_lock above; writes/ tombstones arriving for a COLD segment's key -- WarmSearchSegment (and therefore also the COLD stub, which reloads into one) has no per-key tombstone mechanism at all, a pre-existing gap discovered during this work (tombstone_key_in_index only ever touches mutable and immutable) -- COLD inherits exactly the same behavior as WARM already had, so this is a documented pre-existing limitation, not a regression; see "Known limitation" below. FLUSHALL/FLUSHDB/DROPINDEX -- handled (see Lifecycle above). GraphUnion merge -- skipped by construction (see Lifecycle above). Single-flight reload -- handled (see above).

  • ⚠ Known limitation (pre-existing, not introduced by WS3): no per-key tombstoning for WARM or COLD segments. tombstone_key_in_index (src/vector/store.rs) only marks deletions in the mutable and immutable (HOT) segments; WarmSearchSegment has no delete-bitmap or MVCC-visibility mechanism of its own. A key deleted (DEL/HDEL/UNLINK) after its vector's HOT segment has demoted to WARM or COLD stays returned by that segment's own local search until the segment is later merged away or the whole index is dropped/flushed. This predates WS3 (round 1's WARM tier already had it) and reloading a COLD stub does not make it worse or better -- it inherits WARM's exact behavior. Flagged here because implementing the new COLD tier required reading this code path closely enough to notice it; fixing it is out of scope for this workstream (adding per-segment tombstone tracking to WarmSearchSegment is its own, separately-scoped change) and is recorded as a follow-up.

Added — vector idle-unload with sidecar-preserving WARM transition (WS3, PR #TBD)

  • src/vector/segment/immutable.rs: ImmutableSegment gains a last_access_micros atomic (mirrors WarmSearchSegment's existing field), touched on every search/search_filtered, exposed via idle_secs().
  • src/config.rs: new --engine-offload-idle-secs <secs> (default 3600, 0 disables). A HOT segment now demotes to the mmap-backed WARM tier when EITHER --segment-warm-after (age) OR --engine-offload-idle-secs (idleness since last search) is reached — whichever fires first — so a segment that is old but still busy stays HOT instead of unloading mid-traffic. src/vector/store.rs adds try_warm_transitions_idle / try_warm_transitions_all_idle (the pre-existing age-only methods now delegate to these with idle disabled, so every existing caller is unaffected). src/shard/{persistence_tick, event_loop}.rs thread the new threshold through; the warm-check poll interval also adapts to the lower of the two thresholds for fast local testing.
  • Recall fix (was silently dropped, not new WS3 behavior): VectorIndex::try_warm_transitions always passed None for the HOT segment's f16 exact-rerank sidecar when writing vectors.mpf, and WarmSearchSegment had no sidecar field or rerank logic at all — every HOT->WARM transition silently downgraded to quantized ADC-only distances, contradicting the sidecar's whole purpose (HQ-1, PR #232). Fixed: src/vector/persistence/warm_search.rs's WarmSearchSegment now loads vectors.mpf (when present) into a RawF16Store, exposes raw_f16(), and runs the same rerank_exact HQ-1 logic ImmutableSegment does before truncating to k. VectorIndex::try_warm_transitions now passes imm.raw_f16()'s bytes through instead of None.
  • Correctness fix (also pre-existing, caught by the new integration test): warm_search.rs's parse_global_ids assumed a 24-byte MVCC entry with no key_hash field; the actual writer (ImmutableSegment::mvcc_raw_bytes) emits 32-byte entries (internal_id + global_id + key_hash + insert_lsn + delete_lsn). Every entry past the first was misaligned, and key_hash was dropped entirely — every WARM-tier FT.SEARCH hit resolved to a synthetic vec:<id> key instead of the real Redis key. Renamed to parse_mvcc_ids, fixed to the correct 32-byte stride, and key_hash now propagates onto SearchResult.key_hash via remap_to_global_ids.
  • FT.INFO: new additive (scatter-gathered across shards) counters warm_segments / warm_segments_with_exact_rerank, mirroring the existing graph_segments / segments_with_exact_rerank for the HOT tier.
  • Correctness fix (regression this WS3 work would otherwise have shipped, caught by RSS verification below): command/vector_search/ft_info.rs's num_docs computation (both the top-level and per-vector_fields entry) only summed mutable.len() + imm.live_count() over HOT segments — it never counted WARM segments. Before WS3 this was latent (age-based HOT->WARM demotion existed but nothing exercised it in a live num_docs check); with idle-unload now demoting segments routinely, a fully-idle single-segment index would report num_docs 0 even though every document was still present and searchable. Fixed by adding warm.iter().map(WarmSearchSegment::total_count) to both sums.
  • ⚠ Known limitation — RSS does NOT drop after HOT->WARM idle-unload. Verified with a real server (--shards 1, SQ8, 40,000 × 768-dim vectors, ps -o rss= sampled every 2s for 16s after the transition): RSS went from 395,168 KB to a flat 399,504–399,520 KB post-unload — a 1.1% increase, not a decrease, with zero decay over the observation window. Root cause is pre-existing (predates this WS3 change, confirmed by the standing doc comment on WarmSearchSegment::from_files in src/vector/persistence/warm_search.rs): the "mmap-backed" WARM tier opens each .mpf file's mmap only for the duration of from_files() and immediately copies every payload out into owned Vec<u8> / parsed-HnswGraph buffers of essentially the same size as the HOT segment's — the mmap is then dropped. So the two structures that dominate memory at scale (TQ codes + HNSW graph) are fully duplicated in heap memory on the WARM side, not lazily paged in from disk. The idle-based trigger added by this change is correct and verified end-to-end (FT.INFO counters flip, recall/exactness preserved, num_docs now stays correct — see tests/vector_idle_unload.rs), but it does not, by itself, achieve the memory-reduction goal this workstream set out to prove. A genuine RSS win requires reworking WarmSearchSegment to search directly over borrowed mmap'd bytes (HNSW traversal + TQ-ADC distance kernels operating on &[u8] instead of owned Vec), which is a substantially larger, higher-risk change (touches SIMD distance kernels, needs careful lifetime/alignment handling, and was explicitly out of scope for this pass per the "no new unsafe without approval" constraint). Tracked as an open follow-up; see docs/guides/tuning.md.
  • src/config.rs: corrected the --disk-offload-threshold doc comment, which claimed "parsed but not acted upon" — the memory-pressure cascade (persistence_tick::{should_run_pressure_cascade,handle_memory_pressure}) has actually acted on it since an earlier phase; the comment was stale.
  • src/storage/tiered/cold_read.rs: new _cached variants (cold_read_through_outcome_cached, read_cold_entry_at_cached) read the cold KV leaf page through PageCache when given one, instead of always preading — the existing (uncached) functions are now thin wrappers calling these with page_cache: None, so their behavior and every existing call site are unchanged. Not yet wired into Database::get()'s live cold-read path — that needs a PageCache handle threaded from the per-shard event loop into Database, a broader plumbing change deferred as a follow-up; the cache-capable read path itself is implemented and tested.
  • Docs: docs/guides/tuning.md gains "Tiered memory offload" and "Vector/FTS/graph idle-unload" sections documenting the new flags, FT.INFO counters, and the FTS/graph limitation (no equivalent idle-unload yet — FTS has no aggregate memory-accounting API, and the graph engine's mmap'd MmapCsrSegment has no LRU-eviction-driven unload comparable to vector's MmapBudget).
  • Red/green: tests/vector_idle_unload.rs (real server process, MOON_BIN / CARGO_BIN_EXE_moon) — idle-out a HOT segment with --segment-warm-after set far above the test's timeout so only --engine-offload-idle-secs can trigger the transition; asserts FT.INFO counters flip HOT->WARM and a post-transition FT.SEARCH returns the same key with the same near-zero exact-rerank distance (this is what caught both fixes above — it failed on the pre-fix code, first on distance, then on key identity). src/storage/tiered/cold_read.rs unit test proves a second cold read is served from PageCache even after the backing file is renamed away. src/vector/persistence/warm_search.rs unit tests cover the sidecar round-trip and the corrected MVCC entry parsing.

Fixed — WS6: self-inflicted write lockout on an over-quota key (HIGH, adversarial review 2026-07-08)

  • HIGH, release-blocking: making container growth visible to used_memory (see the accounting-gap fix below) made a second, previously-unreachable bug reachable: run_write_eviction_gate (src/server/conn/handler_monoio/mod.rs) and check_db_maxmemory_for_command (src/storage/db_quota.rs) applied the noeviction reject to every write command uniformly — so once a key's growth tripped the global --maxmemory gate or a db's --db-maxmemory quota, a pure-shrink command on that SAME key — HDEL, SREM, LPOP, ZREM, ... — was also rejected. A tenant that grew a key past the boundary had no self-recovery path short of FLUSHALL/restart, undermining the WS5b per-db-quota guarantee this same release ships.
  • Fixed with a static, provably shrink-only command classification, db_quota::is_shrink_only_command (src/storage/db_quota.rs), mirroring Redis's CMD_DENYOOM semantics: DEL/UNLINK, HDEL/HGETDEL, SREM/SPOP, LPOP/RPOP/LREM/LTRIM/LMPOP/ BLMPOP, ZREM/ZPOPMIN/ZPOPMAX/ZMPOP/BZPOPMIN/BZPOPMAX/ BZMPOP/ZREMRANGEBYSCORE/ZREMRANGEBYRANK/ZREMRANGEBYLEX, GETDEL, EXPIRE/PEXPIRE/EXPIREAT/PEXPIREAT/PERSIST, FLUSHDB/FLUSHALL now bypass the reject from both the global maxmemory gate and the per-db quota gate — eviction is still attempted first (an evicting policy may as well reclaim while the write lock is already held); only the reject is skipped. Deliberately conservative — commands that can grow a destination key (LMOVE/SMOVE/COPY/ RESTORE/any *STORE variant) or aren't statically classifiable (SET, even with a shorter value) are excluded ("when in doubt, exclude").
  • Applied at all three connection-handler write-path chokepoints: handler_monoio::run_write_eviction_gate, the inline eviction block in handler_sharded's sharded write path, and both inline blocks in handler_single (the pipelined-dispatch loop and the single-command write path).
  • RED/GREEN: test_hdel_self_recovery_past_maxmemory_boundary and test_hdel_self_recovery_past_db_maxmemory_boundary (tests/container_growth_memory_accounting.rs) — confirmed RED against the pre-fix handler/db_quota code (git stash): both failed with the exact OOM-on-HDEL symptom described above. Each test now: grows a hash past its cap (asserting the growing HSET IS rejected with the correct OOM flavor), asserts HDEL on that same over-cap key succeeds, drains the rest, then asserts a follow-up HSET succeeds once back under budget.

Fixed — WS6: container-growth memory accounting gap (HSET/LPUSH/SADD/ZADD... growth was invisible to --maxmemory)

  • Critical: used_memory was charged once at key-creation time for an EMPTY Hash/List/Set/ZSet container (entry_overhead() in src/storage/db.rs, called from get_or_create_hash/_list/_set/ _sorted_set), then a raw &mut reference into the map/deque/set/tree was handed to the command layer. Every subsequent mutation through that reference — HSET adding fields, LPUSH growing the deque, SADD/ ZADD growing the set/sorted-set, and their listpack/intset compact- encoding equivalents — never updated used_memory again, so growing a SINGLE key's contents was invisible to both the global --maxmemory gate and the per-db quotas (WS5b). Confirmed by adversarial review 2026-07-08 (flagged as a known limitation in docs/guides/isolation.md at the time).
  • Fixed via O(1) per-mutation delta accounting rather than a full recompute: RedisValue::estimate_memory() sums over every element for the expanded Hash/List/Set/SortedSetBPTree encodings, so recomputing it on every single-field HSET on a large hash would turn an O(1) command into O(n). New Database::charge_memory/credit_memory/adjust_memory (O(1), saturating) plus per-container byte-cost helpers next to entry_overhead (hash_field_cost, list_elem_cost, set_member_cost, zset_member_cost) that mirror RedisValue::estimate_memory's per-variant formulas exactly. The listpack/intset COMPACT encodings are the exception — their estimate_memory() is already O(1) (capacity-based), so those call sites snapshot before/after instead.
  • Every hash/list/set/zset mutation site instrumented: HSET/HMSET/ HINCRBY/HINCRBYFLOAT/HSETNX/HDEL/HGETDEL/HEXPIRE-family past-expiry delete (src/command/hash/hash_write.rs, src/storage/db.rs); LPUSH/RPUSH/LPOP/RPOP/LSET/LINSERT/ LREM/LTRIM/LMOVE/LPUSHX/RPUSHX/LMPOP (src/command/list/list_write.rs, src/storage/db.rs's list_push_front/_back/list_pop_front/_back); SADD (intset + HashSet paths)/SREM/SPOP/SMOVE (*STORE variants were already correct — they replace the whole entry via Database::set, which already accounted correctly); ZADD/ZREM/ZINCRBY/ZPOPMIN/ ZPOPMAX/ZMPOP/ZUNIONSTORE/ZINTERSTORE/ZRANGESTORE (src/command/sorted_set/sorted_set_write.rs). Whole-key removal paths (Database::remove/Database::set) needed no changes — they already recompute entry_overhead on the final (already-mutated) entry, an O(n) cost paid once per key lifetime, not per mutation, so they correctly credit a grown container's true size when it's fully removed or replaced.
  • The eviction loop's before/after estimated_memory() delta arithmetic (src/storage/eviction.rs) required no code changes — it already reads the (now-correct) live used_memory accumulator.
  • RED/GREEN: tests/container_growth_memory_accounting.rs (growing one HSET/LPUSH key past --maxmemory under noeviction now rejects further writes; draining a grown hash via HDEL correctly credits the freed bytes back). Unit tests per container family (hash/list/set/zset mod.rs) assert estimated_memory() rises/falls with insert/remove/ overwrite, including a same-length overwrite netting a zero delta.
  • Making this growth visible surfaced a second, previously-unreachable bug — a self-inflicted write lockout where even shrink commands like HDEL were rejected on an over-quota key. Now fixed; see the dedicated entry above.

Added — sharded MULTI/EXEC routes a single-owner-shard body to its owner (Phase B, PR #TBD)

  • src/shard/{dispatch,spsc_handler,coordinator}.rs, src/server/conn/{handler_sharded,handler_monoio}/write.rs: with keys routed per-key to their owning shard but MULTI/EXEC executing on the connection's (random, SO_REUSEPORT) accept shard, a hash-tagged {tag} transaction only worked from the ~1/N connections that happened to land on the owning shard — Phase A rejected the rest with CROSSSLOT. Phase B adds a boxed ShardMessage::TxnExecute (kept within the 64-byte cache-line cap like MqCommand) and a coordinator::execute_txn_on_owner hop: when analyze_txn_locality proves every key is owned by ONE shard that isn't the accept shard, the whole body is routed there, executed atomically on the owner's slice, and each write persisted to the OWNER's AOF/WAL via the same wal_append_and_fanout path as normal cross-shard writes. PUBLISH fan-out is deferred back to the originating connection (returned in the reply) so it keeps the normal scatter path; under appendfsync=always the originator issues one fsync_barrier to the owner before acking (H1-BARRIER parity), and owner-side append backpressure surfaces AOF_APPEND_LOST_ERR. Genuinely multi-shard bodies (CrossShard) stay rejected — a shared-nothing engine can't commit across shards atomically. Red/green: tests/sharded_multi_exec_routing.rs (hash-tagged txn succeeds from every connection + writes visible; routed writes survive kill-9 + restart via the owner's AOF; multi-shard span still CROSSSLOT). WATCH in sharded mode remains an unimplemented follow-up.

Fixed — sharded MULTI/EXEC writes are now persisted (were lost on restart) (PR #TBD)

  • src/server/conn/{shared,handler_sharded/write,handler_monoio/write}.rs: execute_transaction_sharded — the MULTI/EXEC executor used by the monoio handler at every shard count (including --shards 1) and by the tokio sharded handler at --shards ≥ 2 — appended nothing to the AOF. Every transactional write was silently lost on restart under appendonly=yes, while an identical write issued outside MULTI survived. The executor now returns the serialized AOF bytes for each successful write (mirroring the single-shard tokio execute_transaction), and a shared persist_txn_aof helper appends them to the owning shard's writer via the normal group-commit path, issuing one fsync_barrier under appendfsync=always before EXEC is acked. On a barrier failure EXEC returns AOF_FSYNC_ERR and suppresses any queued PUBLISH fan-out, rather than acking a durability it can't guarantee. try_handle_multi_exec is now async in both sharded handlers. Red/green: tests/sharded_multi_exec_durability.rs pins EXEC-committed write survives kill-9 + restart, exactly like a non-MULTI write (fails on the pre-fix path: the txn key returns nil after restart while the plain control key survives).

Security — sharded MULTI/EXEC no longer silently misplaces cross-shard writes (Phase A, PR #TBD)

  • src/server/conn/{shared,handler_sharded/write,handler_monoio/write}.rs: execute_transaction_sharded runs the whole queued body on the connection's OWN shard with no per-key routing, so at --shards ≥ 2 a key owned by another shard was silently written to / read from the wrong shard's table — EXEC reported success while the data diverged (silent lost updates; other connections routing to the true owner saw nothing). Hash tags did not help: they co-locate keys with each other but not with the (random, SO_REUSEPORT) execution shard, and migration is disabled during MULTI. A new analyze_txn_locality classifies the queued body via the command-metadata key specs + key_to_shard; EXEC now rejects with CROSSSLOT any transaction whose keys aren't all owned by the executing shard, instead of corrupting. Invariant restored: EXEC-success ⇒ every write is visible to other connections; CROSSSLOT ⇒ nothing was written. No effect at --shards 1. Red/green: tests/sharded_multi_exec_locality.rs (silent-divergence invariant, multi-shard span rejection, single-shard unaffected) + analyze_txn_locality unit tests. Phase B (above) now routes single-owner-shard bodies to their owner instead of rejecting them, so only genuinely cross-shard bodies still get CROSSSLOT. Note found in passing (unrelated, out of scope): a solo GET sent mid-MULTI on the monoio single-shard handler executes immediately instead of queueing.

Added — --profile standalone tuning preset (v0.6.0 WS4)

  • New --profile <name> flag (src/config.rs). standalone fills the proven single-instance p=1 recipe — --shards 1, --io-busy-poll-us 40 (implying --io-driver epoll) — for any of those flags the operator left unset. Fill-only precedence: an explicitly-passed flag (CLI or moon.conf) always wins over the preset. Startup logs exactly which flags the profile set; an unknown profile name is a startup error (exit 2), never a silent no-op.
  • ServerConfig::parse_from_with_matches + ServerConfig::apply_profile use clap::ArgMatches::value_source to distinguish explicit flags from defaults; FromArgMatches::from_arg_matches (non-consuming, clones internally) is used instead of from_arg_matches_mut, which removes consumed entries from ArgMatches and would erase this provenance.
  • Safety: --io-busy-poll-us busy-polls the shard thread and REGRESSES throughput on shared/unpinned cores (OrbStack default, laptops, noisy-neighbor cloud VMs) — standalone prints a prominent startup warning requiring pinned/dedicated cores, and docs/guides/tuning.md#profiles + docs/configuration.md document the same caveat. No raw flag default changed.
  • Tests: profile expansion, explicit-flag-wins precedence, no-profile no-op, and unknown-profile error (src/config.rs tests::test_profile_*).

Added — WS1 command parity: *_RO variants + ACL LOG/CLIENT LIST/OBJECT HELP audit (PR #TBD)

  • BITFIELD_RO, SORT_RO, GEORADIUS_RO, GEORADIUSBYMEMBER_RO: new read-only twins registered in the phf command registry with all three dispatch paths wired (mutable dispatch(), immutable dispatch_read(), and the is_dispatch_read_supported() fast-reject bucket list — a command missing that last one is unreachable over the wire despite compiling and passing unit tests). Each rejects its write-capable subcommand/option (SET/INCRBY/OVERFLOW for BITFIELD_RO; STORE for SORT_RO; STORE/STOREDIST for the GEORADIUS twins) via positional parsing (not a blind token scan, so a GET/BY pattern value that happens to read "STORE" is never misclassified).
  • CLIENT LIST/INFO: added the missing Redis fields (laddr, multi-mem, tot-net-in/tot-net-out, rbs/rbp/obl/oll/omem, events, cmd, redir, resp, lib-name, lib-ver) so key=value parsers no longer choke on absent keys; redir and the tracking flag stay at their "off" defaults; wiring them to the CLIENT TRACKING state shipped in PR #234 is a known follow-up.
  • Audit findings: ACL LOG (real entries + RESET) and OBJECT HELP were already fully implemented — added regression tests to lock in that coverage rather than re-implementing.
  • docs/redis-compat.md regenerated against src/command/metadata.rs (258 commands): fixed stale claims that WAIT and FUNCTION * are unimplemented (both are live), documented the four new _RO commands, and added an explicit non-goals section for PFDEBUG, PFSELFTEST, FAILOVER, MODULE *, SENTINEL * with one-line rationale each.
  • scripts/test-consistency.sh (+section 9b) and scripts/test-commands.sh (key-commands category) gained coverage for all four new commands.

Fixed — WS5b fix-first adversarial review: db-quota bypass on non-inline writes

  • Critical: --maxmemory 0 combined with --disk-offload disable (no disk-offload spill sender) silently bypassed the per-db quota gate for every non-inline write command (HSET/LPUSH/SADD/ZADD/INCR/APPEND/MSET/ SET-with-options/RESTORE/...) — plain SET/GET were unaffected because they take a separate, always-correctly-gated inline fast path. Root cause: handler_monoio's batch_eviction_active and its cross-shard-leg twin spsc_handler's evict_active only checked spill_sender.is_some() || maxmemory != 0, entirely skipping the write-eviction-gate call (and the db-quota check nested inside it) when neither was true. Fixed by adding db_quota::db_maxmemory_any_set() to both conditions, mirroring the Lua bridge's gate (which already had this term). RED/GREEN TDD via git apply -R; new integration test test_quota_rejects_non_inline_writes_without_spill_sender in tests/db_maxmemory_quota.rs (HSET + SET-with-EX). Manually re-verified RESTORE also rides the fixed gate (rejected before payload deserialization).
  • Along the way, found a second, independent, pre-existing bug: used_memory accounting for Hash (and, by the same code pattern, List/Set/ZSet) is only charged once at key-creation time for the empty container's overhead — subsequent field/element inserts into an EXISTING key never update used_memory. This defeats the global --maxmemory gate identically (verified: growing one hash key's fields never trips a small --maxmemory cap regardless of value size), so it predates and is independent of db-quota. At the time this was out of scope to fix here (touches every container-type mutation site, a much larger body of work) and was documented as a known limitation; now fixed — see "WS6: container-growth memory accounting gap" above.
  • --db-maxmemory CLI parsing now fails fast at startup (REFUSING TO START: + nonzero exit, matching this file's existing trusted-config validation convention) instead of silently warning and dropping a malformed/out-of-range entry — this is trusted operator config supplied at launch, not untrusted wire input, so a typo should refuse to start rather than silently leave a db unprotected. CONFIG SET db-maxmemory was already this strict; the CLI form now matches. New config::validate_db_maxmemory_cli(); new tests test_malformed_db_maxmemory_cli_refuses_to_start / test_out_of_range_db_maxmemory_cli_refuses_to_start.
  • Documented (not changed): WS DROP's all-dbs cascade-delete sweep is a synchronous, in-place O(total keys × --databases) scan on the owning shard's event-loop thread — it blocks every other connection pinned to that shard for the duration. Accepted trade-off for an admin-rare operation at the default --databases 16; flagged in docs/guides/isolation.md and inline code comments for large-keyspace, large---databases deployments.

Added — per-db resource quotas + workspace hardening sweep (WS5b, PR #TBD)

  • New per-db memory quota: --db-maxmemory <db>:<bytes> (repeatable CLI flag) and CONFIG SET/GET db-maxmemory <db> <bytes> (0 = unlimited, default). Enforcement mirrors global --maxmemory: noeviction rejects writes at quota with a MOONERR db maxmemory exceeded error; eviction policies shed keys from the offending db only. Zero-cost when unconfigured via a single relaxed-atomic pre-gate (DB_MAXMEMORY_ANY_SET), same pattern as the existing global-maxmemory gate. New src/storage/db_quota.rs.
  • Fixed a real bug found via this work: SELECT/SWAPDB are flagged is_write in moon's command-metadata table (for unrelated ACL/dispatch reasons), which meant a connection that filled a noeviction-quota'd db could not even SELECT away from it afterward — the eviction/quota gate ran against the pre-switch db index before SELECT updated connection state. Fixed for the new db-quota gate via check_db_maxmemory_for_command() (SELECT/SWAPDB exempt); the pre-existing identical quirk in the global --maxmemory gate is documented but deliberately left unfixed (out of scope — shared, widely-used eviction code).
  • Fixed a real, pre-existing bug: WS DROP's best-effort key cleanup only swept logical db 0 (hardcoded), silently leaking a workspace's keys forever if its connection ever SELECTed to a non-zero db before writing (WS AUTH and SELECT are orthogonal connection state). Now sweeps every db on the owning shard. Proven via RED/GREEN TDD (test_workspace_drop_cleans_keys_across_all_dbs).
  • New docs/guides/isolation.md — honest "isolation semantics" reference covering logical dbs, workspaces, and per-db quotas: what each guarantees, and every discovered limit (FLUSHDB is whole-db not workspace-scoped; MOVE/SWAPDB reconciled lazily by a background sweep, not synchronously; no disk-offload spill integration for db-quota eviction; FT.* indexes remain keyspace-global, not workspace-scoped — handoff note for the concurrent WS5a db-scoped-FT-index work).
  • New tests: 9 unit tests in src/config.rs (CLI/CONFIG parsing), 7 in src/storage/db_quota.rs (accounting + enforcement), 6 in src/command/config.rs (CONFIG GET/SET), 3 new workspace hardening tests in tests/workspace_integration.rs (WS DROP all-dbs regression, cross-workspace KEYS non-leakage, FLUSHDB-is-whole-db pin), and a new tests/db_maxmemory_quota.rs real-server integration test.

Fixed — WS5a round 4: write-path db-isolation leak in auto-index/auto-delete hooks (adversarial review)

A follow-up adversarial delta-review of round 3 (below) confirmed all of round 3's fixes, then found one more CRITICAL, data-integrity-severity gap: the auto-index/auto-unindex WRITE-path hooks were never db-scoped, despite the round-2 CHANGELOG entry (further below) explicitly claiming they were. Unlike the read-path leak fixed in round 3 (a foreign db could merely see another db's search results), this gap let a foreign db corrupt or erase another db's index contents — and it triggers at --shards 1 (no multi-shard fan-out required):

  • auto_index_hset (src/shard/spsc_handler.rs, private; reached via auto_index_hset_public/_public_txn from handler_single.rs, handler_monoio/mod.rs, handler_sharded/mod.rs, server/conn/shared.rs, and the sharded MULTI/EXEC replay path) called UNSCOPED vector_store.find_matching_index_names(key) / text_store.find_matching_index_names(key). An HSET issued from ANY db whose key happened to match another db's index PREFIX silently fed that foreign index. Fixed via find_matching_index_names_for_db, with db_index: u8 threaded through both public wrappers and every call site (conn.selected_db as u8 / sel_db as u8 / db_idx as u8, matching the idiom already established for FLUSHDB scoping at each of those sites).
  • auto_delete_vectors (same file) called UNSCOPED vector_store.mark_deleted_for_key, so a DEL/UNLINK issued from ANY db could tombstone another db's vector entry for a same-named key — the already-existing, already-unit-tested mark_deleted_for_key_for_db had zero production callers before this fix. Now threaded the same way.
  • auto_hdel_vectors (bonus fix, same bug class, not explicitly named by the review but caught while auditing the surrounding code): also called unscoped find_matching_index_names; now uses find_matching_index_names_for_db.
  • The inline DEL/UNLINK unindex path inside ShardMessage::Execute (spsc_handler.rs ~line 747) called unscoped mark_deleted_for_key directly (not through auto_delete_vectors) — fixed to mark_deleted_for_key_for_db(key, db_idx).
  • Restart-time reconciliation (RecoveryState::reconcile_key in src/vector/persistence/recover_v2.rs, called from event_loop.rs's per-db rescan loop) had the SAME unscoped find_matching_index_names call and delegated to the unscoped auto_index_hset_public. Left unfixed, this would have silently RE-INTRODUCED the exact write-path leak on every server restart even after the live-write fix above — closed by threading db_index: u8 through reconcile_key, sourced from event_loop.rs's existing for db_idx in 0..db_count loop (the caller already had the right value; it just wasn't being passed in).
  • Checked and cleared: active expiry (expire_cycle/expire_cycle_direct in src/server/expiration.rs, ticked per-db by run_active_expiry in src/shard/timers.rs) and hash-field lazy/active expiry do not call any vector/text unindex hook at all — Database::remove is a pure KV-layer operation with no vector_store/text_store hook. There is therefore no unscoped call to fix on the expiry path; this is a separate, pre-existing "TTL'd keys never get unindexed at all" gap (vectors/text entries for an expired key linger until an explicit DEL or HDEL), out of scope for a db-isolation fix and not new in this pass.
  • New integration tests in tests/vector_db_isolation.rs: write_path_auto_index_does_not_leak_across_db (FT.CREATE in db 0, HSET of a prefix-matching key from db 5, assert db 0's FT.INFO num_docs unchanged and FT.SEARCH does not return the foreign doc) and write_path_auto_delete_does_not_leak_across_db (HSET in db 0, DEL of the same key name from db 5, assert the db-0 document remains searchable). Both verified RED against the unfixed hooks (real cross-db corruption reproduced with the exact documented symptoms), GREEN after the fix. Note: FT.INFO num_docs is mutable_segment.len() (append-only; a tombstone doesn't decrement it), so the DEL-side test asserts on FT.SEARCH hit count instead — an early draft of that assertion only checked "not Nil", which a foreign-db tombstone's empty-but-OK [0] result also satisfies, and silently passed against the unfixed code; corrected to check for a real (> 0) hit count.

Fixed — WS5a round 3: multi-shard text-search db-isolation leak + hardening (adversarial review)

A fix-first adversarial review of round 2 (below) found the headline "index created in db N is invisible from every other db" promise still had a critical multi-shard gap: scatter_text_search (the plain-text/ BM25 FT.SEARCH scatter path), plus the remote legs of scatter_hybrid_search and scatter_text_aggregate, took no db_index at all — on --shards >= 2 a connection on any db got real results from another db's TEXT index. Round 2's own suite only spawned single-shard servers, so the bug was invisible to CI (single-shard scatter degenerates to the already-scoped local path).

  • scatter_text_search (src/shard/coordinator.rs) now threads db_index through its num_shards == 1 fast path AND its true multi-shard DFS fan-out (both the local DocFreq/TextSearch legs and the remote DocFreq/TextSearch SPSC legs), via new db_index: u8 fields on TextSearchPayload, InvertedSearchPayload, and a newly-boxed DocFreqPayload (boxing was required to keep ShardMessage within its 64-byte cache-line cap after adding the field to the previously-inline DocFreq variant). The dead scatter_text_search_filter was scoped the same way rather than left as a latent unscoped copy.
  • scatter_hybrid_search's remote leg and scatter_text_aggregate's remote leg — previously documented below as an open residual gap — are now also db_index-scoped (FtHybridPayload/TextAggregatePayload gain the field; execute_hybrid_search_local_raw_streams and execute_local_partial take it as a parameter). The "known gap" note in the round-2 entry below is superseded: both legs are closed.
  • VACUUM VECTOR (src/command/server_admin.rs::vacuum_vector) was an unscoped existence oracle — any db could merge/compact/probe another db's vector index by name. Now takes db_index and resolves ownership via get_index_for_db/get_index_mut_for_db before any mutation or existence check; the subsequent unscoped needs_merge/ immutable_segment_count/force_merge_index calls are safe unchanged (index names are globally unique per shard, so an ownership check by name is sufficient once performed).
  • --databases is now bounded to 256 (ServerConfig::MAX_DATABASES, validate_databases_bound(), checked unconditionally at both normal boot and --check-config). Vector/text index db-scoping uses a u8 tag; an unbounded --databases (e.g. 300) silently aliased SELECT 256 onto db 0's indexes. Deliberately NOT widened to u16 — 256 dbs is already far beyond Redis's default of 16.
  • New multi-shard coverage in tests/vector_db_isolation.rs (--shards 4): plain-text FT.SEARCH cross-db invisibility (the exact finding-1 scenario), KNN cross-db invisibility, and FT._LIST scoping, all under real multi-shard fan-out. Verified RED (real cross-db hits leaked) against the unfixed scatter_text_search, GREEN after the fix.
  • Stale comments in handler_monoio/ft.rs / handler_sharded/ft.rs / coordinator.rs's test module claiming execute_text_search_local is the multi-shard path were corrected — that function is dead code outside its own unit tests; the real multi-shard path is scatter_text_search → run_text_query_on_index.

Fixed — WS5a: db-scoped vector/text index isolation (headline bug closed)

Round 1 shipped only the data model (see below) — FT.CREATE still tagged every index db_index: 0, so no read-path scoping had any observable effect outside db 0. Round 2 closes the headline promise: an index created in db N is invisible from every other db, across all three dispatch paths (handler_single, handler_sharded/tokio, handler_monoio), including the multi-shard scatter/broadcast legs.

  • IndexMeta/TextIndex gain a db_index: u8 tag (vector sidecar format bumped to v4, text sidecar to v2; legacy sidecars default to db 0, fully backward compatible).
  • FT.CREATE now tags the connection's actual SELECTed db instead of a hardcoded 0. Naming decision: index names stay globally unique per shard (not a composite (db, name) key) — each index is bound to exactly one db; FT.CREATE of a name that already exists in ANY db (including the same db) errors "Index already exists", unchanged from pre-WS5a behavior. This avoids migrating the HashMap<Bytes, _> key type across ~150+ call sites for a benefit (same name reusable per db) nobody asked for; an operator wanting per-db name reuse renames (e.g. idx_db0, idx_db1).
  • Full read-path scoping: FT.SEARCH, FT.INFO, FT._LIST, FT.DROPINDEX, FT.COMPACT, FT.CONFIG (GET/SET), FT.AGGREGATE, FT.CACHESEARCH, FT.RECOMMEND, FT.NAVIGATE, FT.INVALIDATE_RANGE, and hybrid search all resolve indexes scoped to the caller's current db across every dispatch path — ~45 call sites migrated from get_index/get_index_mut to get_index_for_db/get_index_mut_for_db (or the equivalent scoped helper), no new hot-path allocations (same O(n) filter shape as the unscoped originals).
  • Write-path scoping (claim corrected by round 4 below): this entry originally claimed auto_index_hset and the DEL/HDEL auto-unindex hooks only touched indexes bound to the write's db via mark_deleted_for_key_for_db. That claim was FALSE — the round-4 adversarial review found auto_index_hset, auto_delete_vectors, and auto_hdel_vectors all called their UNSCOPED counterparts (find_matching_index_names/mark_deleted_for_key), and mark_deleted_for_key_for_db had zero production callers at the time this was written. See the "WS5a round 4" entry above for the actual fix.
  • Cross-shard scatter/broadcast fully threaded: ShardMessage::VectorCommand and VectorSearchPayload now carry db_index across the SPSC boundary; broadcast_vector_command, scatter_ft_info, scatter_invalidate_range, scatter_vector_search/_remote all honor it. scatter_hybrid_search and scatter_text_aggregate's multi-shard DFS legs honor it on the num_shards == 1 fast path; the true multi-shard remote leg of hybrid search and FT.AGGREGATE's execute_local_partial, and the plain-text/BM25 scatter_text_search scatter path itself, were left fully unscoped — a critical gap an adversarial review caught (single- shard-only test coverage hid it). Closed in the WS5a round 3 entry above — do not treat this note as still-open.
  • Bonus fix found in the same code path: scatter_text_aggregate's single-shard fast path always hardcoded s.databases.first() (db 0) for @field value materialization regardless of the caller's SELECTed db — fixed alongside the db_index threading (was a pre-existing bug, not a WS5a regression).
  • Known gaps, explicitly NOT implemented this pass (see .planning/v0.6.0-release/WS5A-NOTES.md for the full design rationale):
  • SWAPDB does not retag index ownership. SWAPDB a b swaps the KV keyspace but leaves every index's db_index tag untouched — an index created in db a stays visible only from db a even after its data moves to db b. Target design (swap db_index on every index owned by either swapped db) is documented but not coded.
  • MOVE/COPY of an indexed hash does not re-index the key into the target db's matching index — operator must re-HSET in the target db.
  • Graph engine (src/graph/store.rs) is a single per-shard GraphStore, structurally global across all 16 logical dbs — same shape vector/text stores had before this fix, but NOT scoped in this pass (see "Known Limitations" below). Timeboxed analysis confirmed no per-db concept exists anywhere in the graph engine (no db_index, no per-db FLUSHDB differentiation); scoping it would require the same class of work done here for vector/text, sized as its own follow-up.
  • Legacy (pre-v0.6.0) persisted sidecars continue to load with db_index defaulting to 0 (unit-tested, unchanged from round 1).
  • New integration suite tests/vector_db_isolation.rs (real spawned server, real TCP client): cross-db invisibility both directions, FT._LIST db scoping, FT.CREATE cross-db name-collision error, and a true process-restart round-trip proving a db-1-scoped index reloads still bound to db 1.

Fixed — WS5a: FLUSHDB no longer clears vector/text index contents across dbs (foundation)

  • IndexMeta/TextIndex gain a db_index: u8 tag (vector sidecar format bumped to v4, text sidecar to v2; legacy sidecars default to db 0, fully backward compatible).
  • FLUSHDB now scopes to the connection's selected db — VectorStore::clear_all_contents_for_db / TextStore::clear_all_contents_for_db, wired through all 3 dispatch paths (handler_single, handler_sharded, handler_monoio) plus the cross-shard MULTI/EXEC replay path. FLUSHALL is unchanged (still clears every db). Previously FLUSHDB in ANY db cleared ALL index contents keyspace-wide — direct fix, covered by a new integration test (tests/vector_flush_hdel_tombstone.rs).
  • Db-scoped lookup/listing/delete primitives (get_index_for_db, index_names_for_db, find_matching_index_names_for_db, mark_deleted_for_key_for_db, drop_index_for_db) added alongside the existing unscoped methods on both VectorStore and TextStore, unit tested.

Docs — tuning guide: vector bulk load & compaction (PR #TBD)

  • docs/guides/tuning.md: new "Vector bulk load and compaction" section — documents immediate-search-then-background-HNSW-build behavior, the COMPACT_THRESHOLD time-to-serve vs recall lever, the MOON_VEC_COMPACT_WORKERS pool knob, the ~10K-vector parallel-build threshold, per-shard bulk-load parallelism, and the ingest-rate/recall trade-offs. Operator guidance for the parallel-build + insert-path-compaction work shipped in #237.
  • Removes a stray <<<<<<< HEAD conflict marker accidentally left in the [Unreleased] heading by #237's squash merge.

Added — parallel HNSW build + insert-path compaction trigger: time-to-index-green 11× (PR #237)

  • src/vector/hnsw/parallel_build.rs (new): concurrent HNSW construction into one shared graph — per-node parking_lot::Mutex adjacency (copy-under-lock, distance math unlocked; single lock held at a time, deadlock-free by construction), entry point under a RwLock, levels pre-generated from the same seeded LCG as the sequential builder, sequential 1K warmup, then dynamic fan-out over an atomic cursor. Finalize adds a connectivity repair pass (concurrent back-link pruning orphaned ~0.2% of nodes = permanent recall loss; BFS + force-link makes every node reachable, unit-asserted) and the exact same BFS-reorder → HnswGraph path as the sequential builder — drop-in for search, persistence, and GraphUnion merge. compact() routes builds ≥ 10K vectors here; smaller segments keep the bitwise-deterministic single-threaded builder. Measured (50K × 384d, 6-core Linux VM): FT.COMPACT wall 30.2s → 2.71s; a 24K segment build 14.1s → 2.65s (~88% scaling efficiency).
  • Affinity-mask trap fixed (3 sites): shard threads are core-pinned and every thread spawned from one inherits the SINGLE-core mask on Linux, so std::thread::available_parallelism() returned 1 — the "parallel" build ran at exactly sequential speed and the background-compactor pool sized itself to one worker. New shard::numa::system_parallelism() (sysfs online-CPU count, affinity-independent, cached) now sizes both; parallel build workers and pool workers explicitly re-pin round-robin across the machine (pin_worker_to_core), pool workers from the last core downward so small builds stop time-slicing shard 0's core.
  • src/shard/spsc_handler.rs: HSET auto-index hook now calls try_compact() after appending a vector — a pure bulk load (no FT.SEARCH traffic) previously left everything in the brute-force mutable tier until the autovacuum backstop's 30s tick, so the whole HNSW build landed on the first explicit FT.COMPACT. Builds now start and install DURING ingest (in-flight guard unchanged; respects FT.CONFIG AUTOCOMPACT OFF). Red/green: new test_insert_path_triggers_background_compact_without_search.
  • Net effect on the §10.8 losing metric (bulk load 50K × 384d → HNSW-tier serving, VM): ~30s of post-ingest compaction becomes ~0 (already compacted by measure time) — Moon's time-to-green now beats Qdrant's optimizer on the same box (GCE re-validation pending). Multi-segment recall@10 0.9986 vs 0.9992 single-segment (−0.0006, within the ≥0.99 gate); multi-shard verified (--shards 4: per-shard triggers, 9 segments, FT.COMPACT 1.57s).

Docs — tuning guide: durability + pub/sub enhancements (PR #TBD)

  • docs/guides/tuning.md: rewrote the "Persistence: what durability costs" section to match the write-path campaign — everysec is now a win at pipeline depth (~1.32× Redis), always is disk-fsync-bound (parity floor) and safe to pipeline via group commit (P16 0.12×→0.91×), with a "which policy?" decision line. All framed as automatic behavior (no knobs).
  • New "Pub/sub fan-out" section: coalesced delivery (5.09M msg/s, zero drops) and the intentional slow-subscriber drop policy (256-msg queue cap).
  • Quick-recipe table: refreshed the durable-store row and added a pub/sub row. Cross-links to BENCHMARK.md §7.3.

Docs — durability write-path benchmark results (PR #TBD)

  • BENCHMARK.md §7.3: new section recording the 2026-07-08 durability write-path campaign (PRs #238–#242) — the vs-Redis matrix (always P16 0.12×→0.91×, everysec P16 1.32× win, everysec P1 0.80×→0.99× parity, pub/sub fan-out 438→5.09M msg/s), per-PR root-cause table (strace diagnostics), durability-invariant note, and A/B reproduction steps. Executive summary + header updated.
  • docs/production-guide.md, docs/benchmarks.md, docs/configuration.md: AOF appendfsync tuning guidance updated with the measured ratios and a plain-language explanation of group commit / coalesced writes / park-free writer poll (all automatic; durability unchanged).

Fixed — Windows build: accept_backoff used unix-only libc errnos (PR #TBD)

  • is_resource_exhaustion (accept-loop backoff, PR #230) referenced libc::EMFILE/ENFILE/ENOBUFS/ENOMEM unguarded; libc is not linked on Windows, breaking the main-push Windows check (PRs never caught it — Windows CI is skipped on PRs). Now cfg-split: unix keeps the errno match, Windows matches WSAEMFILE (10024) / WSAENOBUFS (10055) / ErrorKind::OutOfMemory. Follow-up: the module's unit test (resource_exhaustion_classification) also referenced libc errnos — now cfg-split the same way (unix errnos vs WSA codes).

Changed — AOF writer coalesces each group-commit batch into one write (PR #TBD)

  • The two monoio AOF writer paths (TopLevel via commit_group_commit_batch, PerShard framed loop) wrote each record with its own write(2) on the raw unbuffered File (the PerShard path even used a header+body pair). strace during always-P16 SET: 127,812 write calls / 1.25 s of an 8 s window vs 2,144 fdatasyncs / 0.2 s — the writer thread was write-syscall-bound, not fsync-bound, while Redis batches ~120 records per write via aof_buf. Each batch's records now coalesce into one contiguous buffer (reusable in the PerShard loop, capped at 1 MB high-water) and are written with ONE write_all before the single per-batch fsync. Single-message batches keep the zero-copy direct write. Failure semantics unchanged: a failed batch write acks every waiter WriteFailed and engages the torn-stream latch.
  • tokio writer loops already amortize via BufWriter — untouched.
  • wal_group_commit sink tests updated to the coalesced contract (one write, channel-order bytes, fsync after it).

Changed — AOF writer polls park-free under everysec/no (PR #TBD)

  • The two std-thread AOF writer loops (monoio TopLevel + PerShard) no longer park in recv_timeout under appendfsync everysec/no. A parked receiver makes every producer try_send pay a futex WAKE on the shard thread — measured at ~150k futex calls (63% of shard-thread syscall time) during an 8s non-pipelined SET run, the mechanism behind Moon's everysec P1 SET deficit vs Redis (whose AOF append is a plain memcpy). The writer now polls try_recv with adaptive sleeps (wait/16, clamped 500µs–50ms) so producer sends stay pure userspace atomics; post-fix the same run shows 96 futex calls. appendfsync always keeps the parked recv (its callers await the per-batch fsync ack, so receive latency is client-visible RTT).
  • EverySec durability bound unchanged: at the 50ms fast floor the deadline is re-checked within ~3ms slack; idle cost is ≤20 writer wakes/s at the 1s escalated wait (vs 1/s parked — still far below the pre-wave-5 fixed 50ms cadence).

Changed — Pub/sub subscriber delivery coalesces message bursts (PR #TBD)

  • All three connection handlers' subscriber delivery arms (handler_monoio, handler_sharded's run_subscriber_step, handler_single) now drain the subscriber queue (try_recv, capped at 64 KB) and deliver the burst with ONE write_all instead of one write syscall per message — the same batched write pattern the command path already uses. Raises the fan-out delivery ceiling and shortens the window in which a slow subscriber's bounded queue (256) overflows and drops messages. The single-message case keeps the zero-copy path (is_empty fast path, no allocation; handler_sharded reuses the connection's write_buf).
  • New wire-level integration test tests/pubsub_burst_delivery.rs: a pipelined publish burst must arrive complete, frame-boundary-intact, and in order at 1 and 4 shards (both runtimes).

Changed — appendfsync always: local writes now group-commit per pipeline batch (PR #TBD)

  • All three connection handlers (handler_monoio, handler_sharded, handler_single): plain local writes (plus MOVE/COPY and single-shard GRAPH. WAL records) no longer await one fsync ack per command* under appendfsync always. Appends are enqueued fire-and-forget (send_append_group) and the whole pipelined batch is confirmed by ONE fsync_barrier before response serialization — the same contract cross-shard writes and coordinator local legs already used (PR #213). The writer processes its channel in order, so an acked barrier proves every prior append durable; the fsync-before-ack H1 guarantee is unchanged (a failed barrier converts every joined response to an error, never a silent +OK).
  • Why: a 16-deep pipeline paid 16 serialized fsync round-trips per connection while Redis fsyncs once per event-loop iteration — measured 8× SET-throughput deficit at P16 (Moon 5.2k vs Redis 44k ops/s, GCE c3-standard-8, tmp/MOON-VS-REDIS-DURABILITY.md). Writer-side group commit existed but could only batch across connections; the per-command await defeated it within a connection.
  • everysec/no policies are unchanged (send_append_group degrades to the same bounded-backpressure enqueue).
  • flush_with_aof_ack's discriminating H1 ordering test updated to the batch protocol (one Append + one AppendSync barrier; ack still gates the first response).

Changed — WAL v3 fsync moved off the shard event loop (PR #TBD)

  • src/persistence/wal_v3/sync_agent.rs (new): per-shard WalSyncAgent thread receives fd-dup'd sync requests over a bounded flume channel and publishes a monotonic durable-LSN watermark after each fdatasync. WalWriterV3 gains request_sync() (non-blocking initiation; queue-full falls back to inline fsync — a durability request is never dropped) and wait_durable(lsn, timeout) (bounded blocking wait, used only by the two checkpoint ordering invariants and shutdown). An fsync error poisons the agent permanently and fails subsequent syncs loudly; the checkpoint then refuses to advance redo_lsn.
  • Why: flush_sync() ran fdatasync on the shard event-loop thread. Measured on GCE pd (tmp/WALV3-OFFLOOP-FSYNC.md): the everysec 1s timer froze every connection on the shard for 10–16 ms once per second when the WAL held real bytes, and appendfsync always paid −20% RPS / ~2× tail on top of the AOF cost.
  • Call sites: everysec timer (timers::sync_wal_v3) and the always-mode per-drain-batch sync (both runtimes) now use request_sync(); the checkpoint log-before-data page gate and WAL-before-manifest finalize use wait_durable (5s bound); shutdown paths keep the inline flush_sync(). everysec semantics are now "sync initiated every 1s, durable typically ms later" — the same window the AOF everysec writers provide.
  • Loom model for the watermark/poison state machine in tests/loom_wal_sync_agent.rs.

Changed — consolidated dependency bumps wave 2 (PR #TBD, supersedes dependabot #223–227)

  • Patch-level bumps rolled into one Cargo.lock update (no Cargo.toml edits — all within existing ranges): bumpalo 3.20.2→3.20.3, uuid 1.23.1→1.23.4, smallvec 1.15.1→1.15.2, unicode-segmentation 1.13.2→1.13.3, and the crypto-tls group rustls 0.23.40→0.23.41 + aws-lc-rs 1.17.0→1.17.1 (pulls aws-lc-sys 0.41.0→0.42.0).
  • Perf side-effect verified by same-instance p=1 KV A/B (deps vs main) on GCE c3-standard-4 (x86) + c4a-standard-4 (ARM), shard=1, spin={0,40}, clients={1,8}, best-of-5, steal=0%. All single-op cells land 0.974–1.008 (new_vs_base) on both arches — no regression from smallvec on the dispatch/protocol hot path, nor from the aws-lc-sys 0.42 codegen. See tmp/DEPS-WAVE2-AB.md.
  • CI-only GitHub Actions bumps folded in (no runtime effect): actions/setup-node 4→6 (release.yml), actions/setup-python 5→6 (docs.yml), taiki-e/install-action 2.81.10→2.82.9 (fuzz.yml) — supersedes dependabot #220/#221/#222.

Changed — BREAKING: WAL v3 is now the only WAL; XactCommit unified to the db-aware record (pre-1.0 format freeze)

Moon is pre-release: this drops on-disk format compatibility that no shipped release depended on. There is no migration path — operators upgrading a node that still has WAL v2 files (shard-N.wal) or pre-freeze XactCommit(0x51) records on disk must let those age out via the AOF-authoritative recovery path (below), or replay them on a pre-freeze build first and re-persist.

  • src/persistence/wal_v3/record.rs: deleted the pre-1.0 XactCommit (v1, tag 0x51) record type, its encoder, and its replay arm. Renamed XactCommitV2 -> XactCommit, keeping discriminant 0x53 (0x51 is retired, not reused — a stray 0x51 record in a v3 stream now decodes as an unrecognized tag rather than being silently reinterpreted). All producers (the TXN.COMMIT WAL record in handler_sharded/handler_monoio) already wrote only the v2 (now unified) record, so this is a pure simplification of the read side.
  • src/persistence/wal.rs deleted — WAL v2 (WalWriter, replay_wal, wal_path) is removed in its entirety. WAL v3 (WalWriterV3, src/persistence/wal_v3/) is the only WAL, and is now also created in the old v2 niche (--appendonly yes with --disk-offload disable), rooted at <persistence_dir>/shard-N/wal-v3/ instead of the old flat <persistence_dir>/shard-N.wal file — matching the disk-offload layout. This closes a latent gap: --appendfsync always previously only fsynced WAL v3 (a no-op in non-offload mode, since only v2 existed there); it now fsyncs the WAL unconditionally.
  • replay_wal_auto rejects a v2-format WAL file (RRDWAL magic + version byte 2) with a loud WalError::UnsupportedVersion instead of delegating to the deleted v2 replay path.
  • Recovery: the shard-N.wal (v2) fallback rung is removed from both persistence::recovery::recover_shard_v3_pitr (disk-offload path) and Shard::restore_from_persistence_v2 (legacy path) — appendonly.aof remains the recovery authority and is unaffected. shared_databases:: replay_graph_wal (graph-feature boot-time WAL scan) is ported from the v2 flat file to the v3 segment directory (shard-N/wal-v3/*.wal), matching replay_workspace_wal/replay_temporal_wal. Additionally, both legacy-dir recovery paths gain a last-resort WAL v3 fallback: when NO appendonly.aof exists, the legacy-mode shard-N/wal-v3/ directory is replayed (with a loud partial-coverage warning — WAL v3's KV coverage is intentionally partial post-#211, so it never shadows a present AOF), and a leftover legacy shard-N.wal (v2) file on disk now triggers a tracing::error! naming the file instead of being silently ignored.
  • wal_append_and_fanout / wal_fanout_has_work (src/shard/ spsc_handler.rs) drop the v2 wal_writer parameter and the "v3 supersedes v2" branch — a single Option<WalWriterV3> parameter, renamed wal_writer. finalize_snapshot_success no longer takes a WAL writer: v2's per-snapshot truncate_after_snapshot(epoch) had no v3 analogue — v3 retention is LSN-driven (WalWriterV3::recycle_aggressive / recycle_segments_before, run from autovacuum Pass C and the checkpoint protocol) and already ran independently of legacy snapshot epochs, so nothing was lost by not porting it over.
  • Deferred (unrelated to this change, found in passing): shared_databases:: replay_temporal_wal scans shard-N/ directly for *.wal files instead of shard-N/wal-v3/ like its siblings — looks like a pre-existing directory mismatch, left as-is pending its own investigation.

Fixed — CLIENT TRACKING invalidation actually works on sharded servers (PR #TBD)

  • src/tracking/, all three handler stacks: RESP3 client-side-caching invalidation was non-functional outside the single-shard handler via five stacked defects: (1) sharded handlers registered the per-connection invalidation channel but never drained it — pushes went into a channel nobody read; (2) invalidation was gated on the WRITER's own tracking state (backwards — a non-tracking writer never invalidated anyone); (3) only the first key of a multi-key write was invalidated (DEL k1 k2 left k2 stale); (4) the tracking table was per-shard, so a reader and writer accepted on different shards never saw each other; (5) FLUSHALL/FLUSHDB never invalidated. The table is now process-global (OnceLock, gated by a relaxed ACTIVE_TRACKERS atomic so the hot path pays one load when tracking is unused), keys are extracted via the command-metadata key specs, and all write paths (local, cross-shard remote leg, multi-key coordinator, monoio inline fast-path via a tracking-aware gate, FLUSH) hook invalidation. Red/green: tests/client_tracking_invalidation.rs (4 tests, raw RESP3 client) fails 0/4 on the previous binary.

Changed — pub/sub publish fan-out moved outside the registry lock (P1, PR #TBD)

  • src/pubsub/mod.rs (publish_shared): PUBLISH previously held the per-shard registry write lock for the whole O(N) subscriber fan-out — and the SPSC drain held it across the entire drain cycle. High fan-out (10K cache clients on one invalidation channel) stalled every concurrent (UN)SUBSCRIBE and the shard's cross-shard message drain. The new 3-phase path snapshots matching subscribers under a brief read lock, serializes + try_sends lock-free, and takes the write lock only to remove slow subscribers actually hit. Handler PUBLISH sites (both runtimes), the SPSC PubSubPublish/PubSubPublishBatch arms, and the MQ trigger-notification loop all switched. Documented at-most-once trade (matches Redis).

Fixed — PUBLISH inside MULTI executed immediately instead of queueing (C2, PR #TBD)

  • All three handler stacks: the PUBLISH arm ran before the in_multi queue check, so MULTI; SET k v; PUBLISH ch m; EXEC fanned the message out at queue time — subscribers received it before the transaction's writes applied (even on DISCARD). PUBLISH now queues like any other command; both transaction executors intercept it and the handler fans it out after the transaction body, patching the reply placeholder with the real receiver count (remote shards awaited via targeted PubSubPublish). Discovered but NOT fixed (pre-existing, tracked as follow-up): sharded MULTI/EXEC executes the queued body on the connection's local shard regardless of key ownership — at shards=4 most txn-written keys are silently invisible to other connections; the new test documents this and runs on --shards 1.

Added — pub/sub ↔ KV ordering guarantee documented + locked in (C1/C2, PR #TBD)

  • src/pubsub/mod.rs module docs now state the contract: same-connection write→publish ordering IS guaranteed (a received message implies every preceding write on the publisher's connection is visible), cross-connection ordering is NOT, delivery is at-most-once. tests/pubsub_kv_ordering.rs pins both the pipelined SET;PUBLISH case and MULTI{SET,PUBLISH}EXEC across a 4-shard server. (P2 of the review — event-loop/scheduler coupling benchmark — deferred to a Linux box; macOS numbers are not decision-grade.)

Security — channel ACL enforced on the transactional PUBLISH path (C2 follow-up, PR #TBD)

  • src/server/conn/{shared,handler_single,handler_sharded/mod,handler_monoio/mod}.rs: the C2 ordering fix queues PUBLISH inside MULTI and fans it out after the transaction body, but the fan-out skipped the per-channel ACL check that the immediate path enforces — so a client denied a channel could wrap PUBLISH in MULTI/EXEC to reach it. A shared publish_channel_acl_deny helper now gates every fan-out site (all three handlers); a denied channel is patched into the EXEC reply slot as NOPERM and never delivered. The check runs at fan-out time because Moon has no queue-time EXECABORT machinery. Also closes a pre-existing gap where the single-handler immediate PUBLISH never checked channel ACLs at all (the sharded/monoio immediate paths already did). Red/green: tests/pubsub_multi_channel_acl.rs (2 tests, 1- and 4-shard) — a restricted user's MULTI; PUBLISH denied m; EXEC returned :0 on the old binary, NOPERM on the fixed one.

Fixed — CLIENT TRACKING invalidation + cleanup edge cases (review follow-up, PR #TBD)

  • EXEC now invalidates tracked keys: SET/DEL/MSET applied inside MULTI/EXEC bypassed the post-write invalidation hook, leaving RESP3 client-side caches stale until the next non-txn write. All three handlers now run invalidate_after_write for each successful queued write (self-gated on tracking_active(), so no cost when tracking is unused).
  • monoio disconnect releases tracking registration: a client that disconnected without CLIENT TRACKING OFF left ACTIVE_TRACKERS nonzero and kept tracking_active() hot for the rest of the process (the single/sharded handlers already cleaned up on close). The monoio disconnect path now calls untrack_all, gated on tracking_active() to keep the common no-tracking close lock-free.

Fixed — TopLevel-monoio AOF writer: EverySec fsync deferred indefinitely when idle (PR #TBD)

  • src/persistence/aof/writer_task.rs: the TopLevel monoio AOF writer blocked on an untimed rx.recv(), with its EverySec deadline check living only inside the batch-commit path. A batch written under appendfsync everysec gets no per-batch fsync, so if the client stopped writing right after a burst, the buffered bytes only became durable when the NEXT message happened to arrive — the 1s fsync bound was deferred indefinitely while idle. Exposure is host-crash-only (the per-batch flush() already reaches the kernel page cache, so a plain process kill loses nothing), which is exactly the window EverySec exists to bound. The loop now mirrors the audited PerShard writers: bounded recv_timeout on the wave-5 IdleWait ladder (50ms → 250ms → 1s, pinned at the floor while a batch awaits its fsync) plus an end-of-loop proactive fsync that runs on message AND timeout iterations. This supersedes wave 5's "TopLevel monoio needs none of this" note — it was the one writer loop left without the ≤ ~1s idle durability bound.
  • tests/crash_matrix_per_shard_aof.rs harness hardening (found while validating the above): (1) redis_set asserted only redis-cli's exit status, which is 0 even for server ERROR replies — a tripped diskfull guard (host <5% free) turned every SET into a silent no-op and surfaced as a bogus "200 keys missing after recovery"; the helper now pins the reply to +OK. (2) The three server-spawning tests split-brain when run in parallel: unique_port() hands out OS-sequential ephemeral ports, the other tests offset +1/+2, and SO_REUSEPORT lets two tests bind the SAME port without an error — one test's redis-cli traffic lands on another test's server (rotating total-loss false alarms). A shared mutex now serializes them.

Fixed — RSS/CPU remediation wave 5 (PR #TBD)

  • Item A — mmap the exact-rerank f16 sidecar on segment reload (src/vector/segment/raw_f16_store.rs, new RawF16Store enum): a segment reloaded from disk used to fs::read the entire raw_f16.bin sidecar into a second heap Vec<u16>, doubling resident vector memory for reload-heavy deployments (warm starts, segment promotion). Reload now memory-maps the file (memmap2, already a workspace dependency) and hands out a zero-copy &[u16] view backed by the kernel page cache — RSS only grows for pages the rerank path actually touches. Freshly-built segments (compaction/merge) are unaffected — they keep their owned buffer. Rerank parity (Owned vs Mapped, byte-identical sidecar + identical search() output) is pinned by test_reload_raw_f16_sidecar_uses_mmap_and_matches_owned_rerank.
  • Item A follow-up — text posting-list capacity reclaim (src/text/posting.rs): PostingList::term_freqs/positions grow to the peak document count ever seen for a term and, per the existing remove_doc contract, the postings HashMap entry is kept forever even once a term has zero live documents. The buffers now shrink_to_fit() once the last document leaves a posting, releasing peak capacity for terms that go idle without changing the "entry survives" contract.
  • Item B — AOF writer idle wake made adaptive (src/persistence/aof/writer_task.rs, new IdleWait state machine): the 3 steady-state writer loops that need a bounded channel poll to service the EverySec proactive-fsync deadline (TopLevel tokio, PerShard tokio, PerShard monoio — TopLevel monoio blocks on an untimed rx.recv() and needed no change) used to poll at a FIXED cadence forever (50ms monoio / 200ms tokio), waking an idle server's AOF writer thread 5-20 times a second doing nothing. The wait now escalates 50ms → 250ms → 1s once a poll times out with nothing queued, and resets to the floor the instant any message arrives — a real write always wakes the loop immediately regardless of the current timeout, since the poll races a message against the deadline. Escalation is refused (pinned at the floor) whenever a write is buffered under FsyncPolicy::EverySec without an immediate fsync, or last_fsync was manually back-dated (the F6 post-fold drain trick) — the ~1.2s EverySec bound is provably unchanged. FsyncPolicy::Always/No have no such deadline and escalate freely once idle.
  • Item C1 — WAL v3 write buffer shrinks after an oversized flush (src/persistence/wal_v3/segment.rs): a single large record (e.g. a FullPageImage) grew the 8KB write buffer to fit it, and clear() alone never released that capacity — the peak allocation was pinned for the writer's lifetime. flush_write/rotate_segment now shrink_to the 8KB default once capacity exceeds 4x that, a no-op for the common small-record case.
  • Item C2 — SearchScratch visited-set: already bitset-based (SKIP) (src/vector/hnsw/search.rs): the per-query search hot path already uses a word-based BitVec (u64 words, test_and_set/clear_all memset), thread-cached and reused across queries — no change needed. The other Vec<bool> visited sets found in the vector module are all build-time/ compaction/merge-oracle code, not the per-query path; search_sq.rs in particular carries an explicit comment warning that a prior BitVec conversion there caused correctness issues, so it was left untouched.
  • Item C3 — SmallVec the per-tick elastic-budget shard snapshot (src/shard/shared_databases.rs): recompute_elastic_budget (called from every shard's 100ms eviction tick) collect()ed a fresh Vec<usize> snapshot of all shards' published memory on every call. Switched to SmallVec<[usize; 16]> — stack-only for the common <=16 shard case, unchanged single heap allocation beyond that.
  • Item C4 — Lua script-cache byte estimate exposed via INFO/MEMORY DOCTOR (src/scripting/cache.rs, ScriptCache::resident_bytes()): the per-shard Lua cache was invisible to observability — its growth folded silently into "allocator overhead." Added a byte-estimate accounting method, published per-shard via the existing C5/M4 ShardStoreMemory tick pattern (new lua atomic), and surfaced in both the Prometheus moon_memory_bytes{kind="lua_scripts"} gauge and MEMORY DOCTOR's text report. The cache itself remains intentionally unbounded (Redis parity — SCRIPT FLUSH is the only eviction path); this is observability only.
  • Item C5 — removed dead parse_single_frame_zc RESP parser (src/protocol/parse.rs): a full ~150-line RESP2/RESP3 parser superseded by the current validate_frame + parse_frame_zerocopy pipeline, with zero external callers (only self-recursion) — silently masked by the file's #![allow(dead_code)]. Removed along with its exclusively-private helper read_decimal_zc.
  • Item C6 — jemalloc decay policy audited, docs added (SKIP code change) (CLAUDE.md): the baked-in _rjem_malloc_conf static and the --memory-arenas-cap re-spawn override already carry byte-identical dirty_decay_ms:1000,muzzy_decay_ms:5000,background_thread:true tuning — no drift to reconcile. Added the missing operator-facing _RJEM_MALLOC_CONF documentation (docs-only, no code changed).
  • Item C7 — tokio 1ms shard tick idle cost audited (SKIP) (src/shard/event_loop.rs, src/shard/spsc_handler.rs): the 1ms periodic_interval tick's SPSC drain is already a non-blocking, zero-allocation try_pop() loop, and every downstream side effect is already gated behind a cheap conditional. The one unconditional cost (cached_clock.update(), a single clock_gettime) is the documented "Timestamp caching" design. Unlike item B's AOF writer poll, this 1ms cadence IS the low-latency WAL-flush contract (CLAUDE.md), not incidental idle waste — escalating it would widen that bound. No code change; a real fix would be event-driven WAL triggering, an architectural change out of scope here.
  • Item C8 (sigterm readiness deadline) landed early via the Windows-CI PR (#229) — see the CI section below.

CI — fix Windows main-push test failures (PR #TBD)

  • test_poll_real_process_smoke is now gated to Linux/macOS: get_rss_bytes() returns the documented 0 fallback on every other platform, so asserting rss_bytes() > 0 can never pass on Windows.
  • test_scatter_text_aggregate_single_shard_skips_spsc now runs on a fresh OS thread and calls init_shard itself (via a new shared slice::test_support::make_init fixture). As a plain #[tokio::test] it only passed when libtest scheduled it onto a harness thread where an earlier test had left a ShardSlice behind — a latent order-dependent flake on all platforms that failed deterministically on Windows CI.
  • All 8 find_moon_binary test helpers now fall back to env!("CARGO_BIN_EXE_moon") (after the MOON_BIN override) instead of probing target/{release,debug}/moon: the old probing found nothing on Windows (missing .exe), ignored CARGO_TARGET_DIR, and preferred a possibly-stale release binary over the one cargo just built for the run.
  • New MOON_DISK_FREE_MIN_PCT env override for --disk-free-min-pct (clap env feature; CLI flag still wins). Windows CI exports it as 0: windows-latest runners sit below the 5% free-disk default, so every server-spawning suite failed with MOONERR diskfull instead of OK.
  • shardslice_shape.rs::split_off_test_module computed line offsets with a 1-byte \n assumption; on CRLF checkouts (Windows runners set core.autocrlf=true) the split point drifted one byte per line and panicked slicing inside a multibyte comment char. Now uses split_inclusive('\n') for byte-exact offsets on both line endings.
  • mem_watchdog integration cases A/B (guard-engages assertions) are gated to Linux/macOS for the same get_rss_bytes() 0-fallback reason as the lib smoke test: with RSS reported as 0 the memfull guard is structurally inert on Windows. Cases C/D (guard-disabled directions) still run there.
  • sigterm_shutdown.rs readiness deadline widened 15s → 60s behind a shared READY_TIMEOUT constant (poll-based loop notices readiness within 100ms either way; the fixed 15s was a documented host-load flake, observed on the macOS CI runner).

Fixed — connection-plane durability & lifecycle wave (PR #230)

  • D-2 — FLUSHDB/FLUSHALL now clear every shard, not just the local one (src/shard/coordinator.rs coordinate_flush_broadcast, both sharded handlers): FLUSHDB/FLUSHALL are keyless, so key-based routing executed them only on the issuing connection's shard — on --shards 4 a FLUSHALL left ~3/4 of the keyspace intact (red/green: 49/64 keys survived, now 0). The originating shard now broadcasts the flush to every peer shard as a MultiExecute leg, which reuses the existing remote dispatch + per-shard AOF/WAL persistence + index-flush path. Any failed leg returns a loud MOONERR FLUSH partial error (local shard already flushed; client should retry). Known limits: the flush is not atomic across shards (same relaxed semantics as SWAPDB), and a FLUSH issued inside MULTI/EXEC still executes local-only (follow-up); FLUSHALL still clears only the selected db (documented v0.1.5 single-active-DB behavior — all-dbs semantics is a separate follow-up needing replay parity).
  • D-1 — cross-store TXN crash replay restores into the correct db (src/transaction/mod.rs, wal_v3/record.rs, wal_v3/replay.rs, txn handlers): CrossStoreTxn had no db field and WAL replay hardcoded databases[0], so a TXN.BEGIN…COMMIT issued under SELECT n≠0 replayed its KV ops into db 0 after a crash — silent recovery corruption. Commits now write a new XactCommitV2 record (0x53) whose header carries the BEGIN-time db index; replay targets that db (out-of-range falls back to db 0 with a loud warning). v1 records still replay into db 0 (faithful to the builds that wrote them).
  • R-6 — connection migration fails open (conn_accept.rs, handler_sharded/mod.rs): the source shard dropped its socket before the SPSC hand-off to the target was confirmed, so a full ring lost the client — precisely under the overload that triggers migration. Both runtimes now keep the original stream alive until the push succeeds (bounded retry, 8 × 100µs) and on give-up resume serving the connection on the source shard with migration disabled, using the state recovered from the undelivered message. Also fixes the tokio path's connected_clients leak on the old loss path.
  • Observability (admin/metrics_setup.rs): new moon_xshard_backpressure_drops_total{target_shard} counter + warn on every R-1 give-up, and moon_shard_connected_clients{shard} gauge (maintained at registry register/deregister) to surface SO_REUSEPORT imbalance and affinity funnels without parsing CLIENT LIST.
  • P-1 — lazy per-connection FunctionRegistry (both handlers, conn/core.rs): the Functions API registry + LuaEvictionCtx (6 Arc/Rc clones) were built eagerly for every connection; now built on first FUNCTION/FCALL/FCALL_RO, cutting per-connection setup cost at high connection counts.
  • S-5 — --tcp-backlog (config.rs, conn_accept.rs, main.rs): the per-socket listen backlog was hardcoded 1024; now configurable (default unchanged) for connection-storm tuning alongside ulimit -n.

Fixed — connection-plane robustness hardening (Track B) (PR #230)

  • R-3 — CLIENT KILL force-closes idle connections (src/client_registry.rs, handlers, conn_accept.rs): CLIENT KILL only set a cooperative kill_flag checked once per batch, so a connection parked in read().await (idle, the default timeout 0) was never torn down until it next sent bytes. Now the registry stores each connection's socket fd and kill_clients also shutdown(2)s it, so the parked read returns Ok(0)/Err immediately (the existing disconnect path). The raw-fd close is race-free: kill_clients holds the registry read lock, and a connection's RegistryGuard deregisters (needs the write lock) strictly before its stream drops — so a visible entry always has a live fd. No-op on non-unix (CLIENT KILL stays cooperative there). Self-kill (the filter matching the executing connection) stays cooperative — flag only, no fd shutdown — so the +count reply is still delivered before the handler closes, matching Redis's reply-first self-kill semantics.
  • R-2 — inline-command length cap (src/protocol/inline.rs, frame.rs): the RESP-less inline path had no maximum line length, so a client that never sends \r\n (raw non-RESP bytes) grew the connection read buffer without bound — a per-connection memory-exhaustion vector. Added ParseConfig::max_inline_size (default 64 KB, mirroring Redis's PROTO_INLINE_MAX_SIZE); parse_inline now rejects an over-cap line (complete or incomplete) with Protocol error: too big inline request instead of buffering forever.
  • R-1 — bounded cross-shard spsc_send (src/shard/coordinator.rs): the single dispatch primitive behind every scatter-gather command (MGET/MSET/ DEL/EXISTS/SCAN/KEYS/DBSIZE, vector scatter, FT.INFO, graph traverse) retried a full SPSC ring in an unbounded loop { try_push; yield/sleep } — on tokio yield_now() reschedules with no backoff, busy-spinning a full core, and a wedged target could block graceful shutdown forever. Replaced with a bounded retry (CROSS_SHARD_PUSH_MAX_RETRIES × CROSS_SHARD_PUSH_BACKOFF, the same budget as push_with_backpressure) returning PushOutcome. On give-up the message is dropped; reply-carrying callers observe the closed reply channel and synthesize a per-shard error via the path they already handle. PushOutcome is #[must_use] so no call site can drop a message silently by accident, and the two side-effect-bearing senders fail loud: a dropped MqTxnMaterialize leg turns TXN.COMMIT's reply into an explicit MOONERR TXN.COMMIT partial error (the commit is already WAL-durable; only foreign MQ materialization was lost) and a dropped GraphRollback logs at error!. A brief bounded spin (≤64 iterations) before the first timed backoff preserves the old ~10µs-class latency when the target ring is only transiently full (the 100µs timer sleep may round up to ~1ms).
  • R-4 — accept-loop backoff on fd exhaustion (src/server/accept_backoff.rs, listener.rs, shard/event_loop.rs): all six socket-accept loops (tokio and monoio; sharded, non-sharded, and TLS) retried accept() errors with zero backoff, so EMFILE/ENFILE under a connection storm pinned a shard core at 100% and flooded the log. Added AcceptBackoff: capped exponential backoff (1 ms → 100 ms) on resource-exhaustion errors only, plus rate-limited logging; benign per-connection errors (e.g. ECONNABORTED) log without sleeping. The cap is 100 ms (not 1 s) because the sleep runs inside select! arms where the shutdown branch cannot preempt it — this bounds shutdown latency during a storm. The io_uring multishot-accept resubmit (opt-in MOON_URING=1 bridge) is documented as a scoped follow-up (synchronous CQE handler, no async context to sleep in).

Changed — graph engine deep-review fixes + optimization waves (PR #TBD)

  • Freeze-boundary correctness (P0): CSR v5 segments now persist node properties, embeddings, and edge weights/properties across freeze; Cypher reads see the frozen tier via MergedNodeView (NodeScan/eval/ IndexScan); cross-tier delta edges; GRAPH.NEIGHBORS existence + direction fixes; GRAPH.HYBRID/VSEARCH operate across both tiers. New red/green graph_freeze_boundary suite (11 tests) pins the lifecycle.
  • Point-query indexes (P1): lazy per-segment property indexes (built once at first use from freeze-time data) + PhysicalOp::IndexScan in the planner/executor — MATCH (n:L {p: v}) narrows instead of label-scanning.
  • Query-path cost (P2): plan-cache auto-parameterization (literal variants share one cached plan; cache hits skip parse+compile entirely) with LRU eviction; slot-indexed executor rows (no per-row HashMap allocation, no per-insert String clone); write-path/WAL allocation cleanups (ADDNODE double WAL-encode deleted, itoa/ryu statement temporaries, FxHash for slotmap-keyed sets, Dijkstra three-maps-to-one merge, allocation-free neighbor iteration in Cypher Expand).
  • Traversal engine (P3): CSR row-space BFS fast path on fully-frozen graphs — dense-bitmap visited set, TRUE parallel frontier expansion (thread::scope over Send+Sync CsrStorage), direction-optimizing Beamer push/pull over the incoming index; TraversalGuard wall-clock budget (30s default) now enforced per hop in Cypher variable-length Expand and ShortestPath (ExecErrorKind::Timeout); GRAPH.HYBRID's HnswPreFilter strategy is real — a lazy per-segment HNSW bridge (src/graph/hnsw_bridge.rs, raw-f32 cosine over v5 embeddings, >= 4096 vectors) replaces the silent brute-force stub, with exact-scoring fallback whenever the approximate beam under-fills.
  • Pre-merge review fixes: (1) restart NodeKey aliasing (P0) — a fresh post-recovery MemGraph's deterministic SlotMap could mint keys bit-identical to loaded CSR segments' external_ids, silently shadowing frozen nodes; recovery now seeds the mutable tier via MemGraph::with_id_offset (watermark past the largest persisted id) and WAL replay dedup-skips AddNodes already resident in a loaded segment (red/green graph_restart_id_aliasing suite). (2) Plan-cache raw-hash fast path — exact-repeat queries hit the cache without the parameterize() lexer pass, and the SlotTable is cached alongside the plan instead of being rebuilt per execution; cache backing moved to FxHashMap. (3) CsrStorage::resident_bytes now counts the v5 property blobs, lazily-built per-segment property indexes, and the HNSW bridge (heap-owned; mmap-backed sections count 0 per the RawF16Store precedent), keeping the elastic memory budget honest.

Changed — graph engine wave 2: durability, Cypher coverage, hardening (PR #TBD)

  • Durability (P0, found by GCP production-hardening soak): graph data now survives kill -9 under the DEFAULT disk-offload (WAL v3) configuration. Two stacked pre-existing bugs erased graphs on crash: (A) v3 recovery collected graph WAL commands into a throwaway replay engine — graph replay was a complete no-op in v3 mode (a dedicated replay_graph_wal_v3 boot pass now applies them); (B) the checkpoint advanced the WAL replay floor and recycled segments without ever snapshotting the graph store — save_graph_store ran only on graceful shutdown and never persisted the mutable tier (checkpoint finalize now freezes each graph's write buffer, persists segments + a snapshot_lsn replay floor, and ABORTS the checkpoint if the graph snapshot fails, keeping the old floor). The prior crash suite ran exclusively with --disk-offload disable; new G4/G5 scenarios cover the default config with and without a checkpoint crossing.
  • Durability: stable external node/edge ids across WAL replay (handles handed to clients before a crash resolve to the same rows after recovery); Cypher SET property/label writes now emit WAL records (GRAPH.SETPROP/GRAPH.SETLABEL — previously a documented durability gap: SET mutations silently vanished on kill -9); new kill-9 crash-recovery suite (tests/crash_recovery_graph_durability.rs: mutable tier, frozen 64K-edge tier rebuilt through replay-time freeze, double-crash idempotence).
  • Cypher coverage: aggregations count/sum/avg/min/max/collect with implicit grouping, DISTINCT, and count-over-zero-rows semantics; OPTIONAL MATCH null-pads unmatched expansions (previously compiled silently to inner MATCH; unsupported shapes now reject loudly); WITH rebinds the pipeline mid-query (aggregate + WHERE-as-HAVING, ORDER BY/SKIP/LIMIT between WITH and RETURN; previously every clause after WITH ran on an empty row stream). WITH * rejects loudly.
  • Copy-up writes: SET/DELETE/MERGE on frozen rows copy the row up into the write buffer instead of silently missing the frozen tier.
  • Query performance: IndexScan range predicates (WHERE n.p > x prunes via per-segment B-tree-ish numeric index, superset semantics + residual filter); write-side plan cache (repeated write shapes skip parse+compile); row-BFS fast path engages on multi-segment fully-frozen graphs (~4× vs reader fallback, criterion-pinned); Value::String holds Bytes for a zero-copy reply path; HYB-02/HYB-04 use the HNSW bridge with off-thread bridge builds.
  • Query performance — mutable-tier property index (task #31): MATCH (a:N {id: X}) on the write buffer used to degrade to a full O(live_node_count) linear scan even though PhysicalOp::IndexScan was already the plan (profiling on the 5K-node/15K-edge Cypher point-query bench put this scan at 82.5% of shard CPU — the mutable tier never freezes below edge_threshold, default 64,000). A new incrementally maintained MutablePropertyIndex (src/graph/index.rs, mirrors the frozen tier's SegmentPropertyIndexes numeric-BTree/string-hash split, keyed by NodeKey) turns the mutable-tail probe into an O(log N + |result|) index seed; index_scan_keys's residual label/MVCC/property checks are unchanged (SUPERSET contract preserved, no planner changes). MemGraph::set_node_property/remove_node_property/undelete_node are now the single source of truth for node-property mutation, replacing 5 previously hand-rolled call sites (Cypher SET, MERGE ON CREATE/MATCH SET, TXN.ABORT undo, WAL replay) that could silently let the index drift from live state. Closes a TXN.ABORT UndeleteNode gap found during design review: undoing a DELETE inside a transaction now re-indexes the restored node instead of leaving it live but permanently unreachable by index probes.
  • Query performance — Cypher result cache (task #32): read-only GRAPH.QUERY/GRAPH.RO_QUERY now cache the fully-encoded RESP reply bytes for repeated identical queries (same raw Cypher text + args), keyed by (query_hash, args_hash) and served via the existing Frame::PreSerialized passthrough — no new protocol variant needed, and both RESP2 and RESP3 encodings are cached in separate slots so a protocol-version switch is a clean miss rather than a stale re-encode. Invalidation is a per-graph monotonic write_gen: u64 bumped by NamedGraph::touch() at every real mutation site (plain GRAPH.ADDNODE/ GRAPH.ADDEDGE, the Cypher write-plan mutation loop, TS.* temporal invalidation, and TXN.ABORT rollback of graph undo-ops) — deliberately NOT at freeze_and_compact, which reshapes storage without changing query-visible content. The write_gen is captured before executing a read and the encoded reply is only cached if it is still unchanged after — a miss is always safe (SUPERSET semantics), so a benign race just skips caching rather than serving stale data. Decayed queries and errored/timed-out results are never cached. ResultCache is a bounded per-graph LRU (256 entries / 4MiB default, src/graph/cypher/ result_cache.rs) whose resident bytes are folded into GraphStore::resident_bytes(); dropping a graph (GRAPH.DELETE) drops its cache with it. The cache is only consulted on the connection-local dispatch path where the negotiated protocol version is reliably known (dispatch_graph_read, threaded as Option<u8>); the cross-shard GraphCommand hop and graph_query_or_write's internal read-execute calls pass None and bypass the cache entirely rather than risk an unverified protocol version.
  • Query performance — FTS reuse for Cypher text predicates (task #33): first-class CONTAINS/STARTS WITH/ENDS WITH Cypher operators (pure syntax sugar over the existing =~ byte-level checks — src/graph/ cypher/lexer.rs, ast.rs, parser/expr.rs, executor/eval.rs), plus a new per-frozen-segment SegmentTextIndex (src/graph/text_index.rs, lazy OnceLock + resident_bytes accounting, same pattern as SegmentPropertyIndexes/hnsw_bridge) that accelerates CONTAINS/ STARTS WITH/ENDS WITH/=~ conjuncts in WHERE (planner.rs's new extract_text_conjuncts, mirroring W2-3's extract_range_conjuncts). Reuses crate::text::posting::PostingStore/crate::text::bm25:: {FieldStats, bm25_score}/crate::text::term_dict::TermDictionary verbatim (already ID-space-agnostic, zero-fork). Correctness note: the index prunes on PROPERTY PRESENCE only, not token identity — a tokenized/stemmed/case-folded posting lookup cannot safely decide substring/prefix/suffix containment (e.g. CONTAINS 'rust' must also match "trusted", which tokenizes to a different term entirely), so candidate_rows returns every row whose target property is a String/Bytes at all, and the existing residual Filter remains the sole decider (SUPERSET contract, same as prop_eq/prop_range). The MUTABLE tier gets no acceleration (falls back to an exact scan of the write buffer, pre-approved scope decision — mirrors the pre-task-#31 numeric story); a GraphUnion-merged segment does not remap posting row ids and simply rebuilds its text index lazily on first use post-merge. SegmentTextIndex::bm25_score_for is a tested-but-unwired internal BM25 relevance-scoring hook (no Cypher ORDER BY grammar was specified for it) — proves the bm25_score reuse end-to-end; surface syntax is a documented follow-up.
  • Operability: traversal timeout is configurable — --graph-timeout-ms server default plus per-query GRAPH.QUERY ... TIMEOUT <ms> (RedisGraph parity, 0 = unlimited).
  • Dead code: cross-shard traverse scaffolding deleted (532 lines, no sender existed); the single-shard-per-graph sharding model is now documented in src/graph/mod.rs.

  • Cargo: ringbuf 0.4.8 → 0.5.0 and metrics-exporter-prometheus 0.16.2 → 0.18.3 (semver-major; both compile and test green with no code changes — verified by dependabot's per-bump CI runs and locally), bytes 1.11.1 → 1.12.0, rand 0.10.1 → 0.10.2, lz4_flex 0.13.0 → 0.13.1, cmov 0.5.3 → 0.5.4, cudarc 0.19.4 → 0.19.8 (gpu feature, not covered by CI — patch bump), and openssl 0.10.78 → 0.10.81 in sdk/rust (security bump).

  • GitHub Actions: actions/checkout v4 → v7 (including one straggler in docker-publish.yml dependabot missed), actions/deploy-pages v4 → v5.
  • Console: vite 7.3.2 → 7.3.5 (dev dependency).
  • Dependabot: dtolnay/rust-toolchain is now ignored — its pin IS the MSRV (1.94) gate, so auto-bumping it to latest stable (rejected PR #195) silently defeats the MSRV check.
  • Supersedes dependabot PRs #191, #194, #196–#201, #208–#210 (their Lint failures were the CHANGELOG-entry gate, which this consolidated PR satisfies once).

Fixed — proactive RSS watchdog pauses writes before kernel OOM (PR #TBD)

  • New mem_monitor guard (src/shard/mem_monitor.rs), the memory analogue of the existing diskfull guard (MA12): pauses writes once process RSS crosses --mem-full-pct (default 95, 0 = disabled) of the DETECTED system/cgroup memory limit, not the configured --maxmemory (which can be an unconfigured 0/unlimited). Mirrors disk_monitor's hysteresis state machine, inverted for direction (high RSS is bad): pauses at rss% >= mem_full_pct, resumes only at rss% <= mem_full_pct - 5. MOONERR memfull: writes paused until memory pressure recovers joins the existing MOONERR diskfull / MOONERR busy message-selection chain in both dispatch paths (order: diskfull, memfull, segment-backlog busy); the hot-path cost is one additional AtomicBool::load(Relaxed) inside the already write-gated branch. Polled on shard 0's existing 5s disk-monitor timer tick (both event-loop poll sites), plus one immediate poll at startup (init_global) so the AOF-replay/segment-load recovery peak is visible before the first tick. New INFO fields reclamation_mem_rss_bytes / reclamation_mem_watchdog_active (OR'd into the existing reclamation_write_stall_active). Same accepted trade-off as the diskfull guard: DEL/UNLINK/EXPIRE/FLUSHALL are write-flagged and blocked too while paused — no allowlist.

Fixed — vector keymap CoW amplification + bounded snapshot queue (PR #TBD)

  • VectorIndex's three key_hash -> V maps (key_hash_to_key, key_hash_to_global_id, key_hash_to_vec_checksum) were each a single Arc<HashMap<u64, V>>. Under ft-search-off-eventloop, a live FT.SEARCH snapshot holds an extra Arc clone of the whole map while the event loop keeps processing writes; the next write's Arc::make_mut then sees refcount > 1 and full-clones the entire map synchronously on the shard event loop (~47MB @1M vectors). Replaced with BucketedKeyMap<V> (src/vector/keymap.rs): 256 fixed buckets, each independently Arc<HashMap<u64, V>>, bucket selected by (key_hash >> 56). A concurrent snapshot now pins at most 1/256th of the map per write instead of the whole thing. Same treatment applied to SnapshotJob's mirrored fields and SearchSnapshot.key_hash_to_key. On-disk keymap.bin format is unchanged (read-compatible).
  • SnapshotPool (src/vector/persistence/manifest.rs) used an unbounded flume queue with a single worker; if compact/merge cadence outpaced the worker, queued jobs pinned every submitter's Arcs indefinitely. Replaced with a coalescing slot keyed by index directory: a new submit for an index with a job still pending (not yet dequeued) replaces it in place, dropping the stale job's Arcs immediately; an in-flight (already-dequeued) job is unaffected. Submit never blocks the caller.

Fixed — CI: de-flake OOM case E + Windows test-compile gate (PR #TBD)

  • test_case_e_cross_db_copy_oom (was red on main's post-merge CI for PR #217/#218/#219, 0/300 OOM on both platforms while passing locally): the single-shot statistical assert sat inside the slack GAP-1's elastic budget + its 100ms-stale usage snapshots can grant. Rewritten as bounded escalation: up to 8 rounds of COPYs into fresh destination keys until the destination db's per-shard usage (~2.4MB) provably exceeds the ABSOLUTE budget ceiling (compute_elastic_budget ≤ maxmemory = 2MB) — timing-independent on any runner speed, and the RED floor stays exact (gate stashed = 0 OOM through all rounds; re-verified red/green).
  • tests/quickwins_red_api.rs broke the Windows CI leg at COMPILE time (also red on main): qw1_accepted_socket_has_nodelay calls apply_client_socket_opts, which is #[cfg(unix)] (takes AsFd). The test (and its imports) are now unix-gated to match.
  • write_stall_segment_backlog: two tests raced on the process-global RECL_SEGMENT_STALL_ACTIVE atomic (test harness runs same-binary tests on parallel threads; observed on CI macOS as store(1) reading back 0). All mutation of the global now lives in one merged test.

Fixed — FT.COMPACT left mutable residue behind when draining a background build (PR #TBD)

  • Pre-existing (main, not introduced by this branch): when an explicit FT.COMPACT/force_compact found a background auto-compaction already in flight, it drained that build, installed it, persisted, and returned — WITHOUT compacting the documents inserted while the background build was running (the clone_suffix(frozen_len) mutable tail). Those docs stayed mutable-only (no durable segment) until some future compact fired, breaking force_compact's documented full-drain contract (frozen == mutable on reply) and — combined with the B3 phantom-keymap hole below — silently losing them on a kill -9 (observed stuck durable state: segments live=1961, keymap=2000, converging never). Load-dependent: only reproduces when insert cadence keeps a background build in flight at FT.COMPACT time (this is what made the crash suite's S1/S2 flake by machine load, masquerading as a branch regression). force_compact now falls through to the inline drain loop after installing the in-flight build (a no-op when the tail is empty).

Fixed — B3 vector recovery: phantom keymap entries silently dropped docs (PR #TBD)

  • Pre-existing silent data loss on kill -9 (found while validating this branch against the crash_recovery_vector_durability suite; reproduces on main): the durable keymap-<epoch>.bin covers EVERY indexed key (mutable + immutable) at snapshot-submit time, while the paired manifest.json lists only installed immutable segments. A crash landing after one snapshot commits but before the next (covering a just-frozen segment) leaves the on-disk keymap a strict superset of the on-disk segments. The B3 dedup rescan then "verified" those uncovered keys as unchanged (keymap checksum matches the AOF blob — exactly the crash-window signature) and never re-indexed them: their documents vanished from search with num_docs under-reporting (observed: 1950/2000) while recovery logged verified unchanged for all keys. Fix: recover_v2::load_segments_and_keymap now drops any keymap entry whose key_hash is not LIVE in a loaded segment (ImmutableSegment::live_key_hashes), so the rescan re-indexes those keys from the AOF; a loud log line reports the dropped-phantom count. New deterministic unit regression (recover_reindexes_keymap_entries_not_backed_by_any_segment, red/green verified) plus integration scenario S6 (kill inside the async-snapshot window; asserts only the crash contract: num_docs == N, every key answers its own exact-match query). S1/S2's re_indexed == 0 fast-path determinism is restored by a new full-durability pre-kill gate (wait_for_durable_docs: manifest'd segment live-counts AND keymap entries both equal N).

Fixed — zombie/CPU hardening wave 1 (PR #TBD)

  • Busy-poll spin idle-disengages (--io-busy-poll-us / vendored monoio legacy driver): an idle server no longer burns the spin budget on every park forever (measured ~2.4s CPU per 3s wall on an idle 4-shard server at 200µs). The driver tracks the last real event per thread and skips the spin window after MOON_SPIN_IDLE_DISENGAGE_US (default 10ms; 0 = old always-spin) of quiet; the next event re-arms it via the normal wake path, so steady traffic — the GCE p=1 win path — never disengages.
  • Shard-thread panics abort the whole process instead of leaving an N-1-shard server silently answering with a dead shard's keys gone; the shutdown join loop logs instead of swallowing panics. New DEBUG PANIC subcommand (Redis parity) as the crash-handling test aid.
  • SIGTERM shutdown regression matrix: shards × held-connections × busy-poll × 16-writer AOF write-storm (7 cases, both runtimes, 3 platforms). The documented SIGTERM+SO_REUSEPORT bench hang did NOT reproduce at HEAD in any composition; the matrix stays as the guard and the detached monoio accept-task change is deferred until a reproducer exists.
  • wait_ready test harness hardening: readiness probes retry with a fresh connection on mid-startup connection resets (CI flake: per-shard SO_REUSEPORT listeners accept-then-reset during init).

Fixed — OOM eviction bypass closure (PR #TBD)

  • Cross-shard SPSC write legs now enforce --maxmemory. Execute, MultiExecute, PipelineBatch and their *Slotted variants (src/shard/spsc_handler.rs) executed writes against the TARGET shard's Database directly, bypassing the connection handlers' eviction gate entirely — a scatter-gather write (e.g. a pipeline of individually-routed SETs hashing to a remote shard) could grow that shard's memory past maxmemory without limit. A new spsc_eviction_gate helper mirrors run_write_eviction_gate (handler_monoio/mod.rs) exactly — same elastic per-shard budget (GAP-1), same spill-vs-plain branching — threaded through drain_spsc_shared/handle_shard_message_shared and gated by a once-per-drain-cycle evict_active snapshot (perf parity with batch_eviction_active: zero extra lock acquires when maxmemory is unset). Perf follow-up (same PR): that snapshot, and the Lua bridge's gate below, each used to answer "is maxmemory nonzero?" with either a runtime_config.read() per drain cycle or a generation/Cell snapshot. Both now read a single process-global AtomicU64 (storage::eviction::{publish_maxmemory, maxmemory_is_set}), published at every production write site of RuntimeConfig.maxmemory (CONFIG SET maxmemory, server startup) — one Relaxed load, no lock, no per-script Cell bookkeeping.
  • Lua redis.call/redis.pcall writes now enforce --maxmemory. EVAL/EVALSHA carry no WRITE command flag, so the dispatch-level OOM check never saw them, and the bridge (src/scripting/bridge.rs) never ran the check before executing a write inside a script — a tight redis.call('SET', ...) loop could write arbitrarily far past the cap. A new LuaEvictionCtx, captured by value into the redis.call/redis.pcall closures at VM-setup time (setup_lua_vm, once per shard — zero new unsafe, zero per-call allocation), runs the same gate before a WRITE redis.call executes; redis.call raises the OOM error as a Lua runtime error (script aborts, matching Redis semantics), redis.pcall returns it as an {err = ...} table. Read-only scripts are unaffected.
  • FCALL-internal redis.call/redis.pcall writes now enforce --maxmemory too. FUNCTION LOAD (src/scripting/functions.rs) creates its own per-library sandboxed Lua VM, separate from the shard's shared EVAL/EVALSHA VM — that per-library VM registered redis.call/ redis.pcall with LuaEvictionCtx::disabled() unconditionally, so a FUNCTION whose body wrote in a loop could grow memory past the cap independent of the EVAL/EVALSHA fix above. FunctionRegistry::new now takes a LuaEvictionCtx, stored on the registry and cloned into every library's VM at create_library time; both production construction sites (handler_monoio/mod.rs, handler_sharded/mod.rs) build a real ctx from the same shard handles setup_lua_vm already uses for the shared VM (ConnectionContext::{shard_databases, runtime_config, shard_id, spill_sender, spill_file_id, disk_offload_dir}); only test-only callers pass LuaEvictionCtx::disabled().
  • Cross-shard MOVE/COPY ... DB n two-database intercept now runs on EVERY ShardMessage SPSC arm, not just the plain Execute arm. The intercept (special-cased ahead of the generic single-db write path because both commands need two &mut Database borrows at once) previously existed only in Execute. Every other arm — including PipelineBatchSlotted, the one real client traffic (handler_monoio/handler_sharded's pipelined remote-write path) actually uses — fell through to the generic single-db key_extra::copy, which parses and silently ignores the DB clause: a remote COPY src dst DB n performed a same-db copy instead (data corruption, confirmed pre-fix via manual probe: 4/20 remote COPYs landed in the wrong db), and a remote MOVE returned a loud-but-wrong "MOVE requires handler-level dispatch (cross-db operation)" error instead of moving the key. The intercept is now extracted into a shared helper (src/shard/spsc_two_db::try_two_db_intercept) that Execute, MultiExecute, PipelineBatch, ExecuteSlotted, MultiExecuteSlotted and PipelineBatchSlotted all call verbatim; cross-db COPY still runs spsc_eviction_gate against the destination db before copy_core (COPY duplicates the value, so it needs the gate; same-shard MOVE is net-zero and stays ungated, same rationale as DEL).
  • Found, NOT fixed (out of scope): cross-shard COPY's destination key is only reliably retrievable when it hash-tag-colocates with the source key. Moon shards purely by key hash, with no db-awareness. A COPY src dst DB n executes on whichever shard owns src (that's the routing key for the whole command) and writes dst into that SAME shard's db space — it does not (and cannot, without a real cross-shard two-phase protocol) relocate the write to whichever shard dst's OWN hash would naturally pick. A later plain GET dst routes by hash(dst), so it only reliably finds the value when src and dst share a hash tag (or --shards 1). MOVE is exempt (the key name is unchanged, so hash(key) is identical before and after). This is a pre-existing, inherent property of the per-key-hash sharding model applied to COPY with a renamed destination — the plain Execute arm this fix mirrors has the identical property — not a regression from extending the intercept to every arm. Fixing it is a materially larger change (routing the destination write to dst's own natural shard) than "apply the existing intercept to more arms" and is out of scope here.
  • Found, NOT fixed (out of scope): Moon's eviction gate is per-DATABASE, not per-shard-aggregate-across-databases. try_evict_if_needed_budget (and every SPSC write arm, including the new cross-db COPY gate above) measures db.estimated_memory() for the ONE database being written to, not the sum across all 16 logical dbs a shard holds — consistent throughout Moon's write path, but a deviation from real Redis's instance-wide maxmemory. A cross-db COPY into a mostly-empty destination db can undercount total shard pressure. Uniform, pre-existing behavior; fixing it is a cross-cutting change to the eviction model, not a COPY/MOVE-specific one — out of scope here.
  • New wire-level regression suite tests/spsc_two_db.rs (3 cases, both runtimes) exercises the one arm real client traffic uses (PipelineBatchSlotted): pipelined cross-shard COPY ... DB n (RED before the fix — MOVE's loud wrong error and COPY's silent same-db copy — both confirmed via a git-stash re-run against the pre-fix source), pipelined cross-shard MOVE, and a same-db COPY control (routes through the coordinator's multi-key scatter path, unaffected by this fix, kept as a negative control). The COPY case uses {i}-hash-tagged src/dst key pairs so the destination key is independently GET-able post-fix — see the "destination hash-tag" caveat above; an untagged version of the same test only succeeds for the ~1-in-4 keys where hash(dst) collides with hash(src) by chance.
  • New wire-level regression suite tests/oom_bypass_closure.rs (8 cases, both runtimes): direct-SET control (A), cross-shard pipeline (B), Lua EVAL loop (C), read-only EVAL under pressure not blocked (D), cross-shard COPY under pressure (E), CONFIG SET maxmemory atomic publish/un-publish (F), FCALL-internal write loop (G), FCALL control with no maxmemory (H). B, C, E and G RED before their respective fixes. D locks the write-only scope of the Lua gate. Case F (added with the atomic-snapshot perf follow-up) starts the server with --maxmemory 0 --disk-offload disable (both required so the atomic starts genuinely unset — omitting --maxmemory trips the auto-guardrail's nonzero default, and disk-offload's spill-sender presence ORs into the same gate) and drives a cross-shard pipeline before and after CONFIG SET maxmemory/CONFIG SET maxmemory 0; RED before the CONFIG SET call site published the atomic (775/3000 OOM — the local-only share) GREEN after (3000/3000). Case G (FCALL) RED before the FunctionRegistry fix — a 2000-iteration redis.call('SET', ...) loop inside an FCALL'd function completed with 'DONE' instead of an OOM error against a 2MB-capped noeviction server; GREEN after. Case H locks the write-only/OOM-only scope of the FCALL gate, mirroring case D for EVAL. Case E was revised for the cross-shard SPSC intercept fix above: it now pre-loads db 1 (the COPY destination) with the same ballast as db 0 before the OOM phase, since Moon's per-database eviction gate (see the "found, not fixed" note above) needs the destination db under its OWN pressure to trip — the original version passed only as a side effect of the pre-fix same-db-copy bug (0/300 OOM with the destination-db gate disabled and routing intact, confirmed via a targeted stash of just that gate block; 76/300 fully fixed).

Added — int8 symmetric ADC for SQ8 vector search (task #13)

  • New per-candidate integer dot-product path for SQ8 asymmetric distance computation (ADC), replacing the f32-widening sq8_stats kernel on the hot per-candidate pass with exact-integer int8 dot products: NEON widen-multiply (vmull_s8 + the c^0x80 offset trick, since baseline ARMv8-A has no mixed u8×i8 widening multiply), AVX2 (cvtepu8/cvtepi8_epi16 → mullo_epi16 → madd_epi16, saturation-free), and AVX512 VNNI (_mm512_dpbusd_epi32, feature-gated behind simd-avx512). Query is quantized to symmetric int8 once per search (turbo_quant::sq8::sq8_quantize_query_scalar); sum_c/sumsq_c are exact integers straight from the u8 codes. Combines through the unchanged sq8_l2_from_stats — the on-disk SQ8 format and the exact-rerank sidecar (HQ-1) are untouched. Wired into both hnsw/search.rs beam-search closures and all three SQ8 call sites in segment/mutable.rs brute-force search. Local (dev-Mac, NEON) directional speedup vs the existing f32 SIMD path: ~1.5x at dim 128, ~2.0x at dim 384, ~2.15x at dim 768 (benches/sq8_adc_bench.rs, int8_dispatch variant) — absolute/cross-arch numbers are GCE-validation scope (task #14). Recall A/B-gated (turbo_quant::sq8 test module): R@10 delta ≤ 0.01 vs the f32 ADC path and ≥ 0.98 top-10 overlap, both L2 and Cosine, at dim 128/384. MOON_SQ8_INT8_ADC=0 forces the f32 kernels (bench/diagnostic escape hatch, not yet added to CLAUDE.md's env list — flagged for a follow-up doc pass).

Fixed — merge recall gate self-exclusion bias (manual merges could never pass)

  • verify_merge_recall compared HNSW results including the query point (every sampled query is a database point, distance 0 → always rank 1) against a ground truth that excluded it, structurally capping measurable recall at (k−1)/k = 0.90 — exactly the manual FT.COMPACT/VACUUM VECTOR merge gate, so any manual merge of an index with ≥ 50 vectors was rejected with merge recall 0.9000 < tolerance 0.9000 regardless of actual graph quality (the 0.70 background gate masked this). The merged-graph search now requests k+1 and drops the self-point before comparison.

Added — vector-index durability: kill-9 crash-recovery test suite (B4)

  • New tests/crash_recovery_vector_durability.rs: five end-to-end scenarios against real server processes (SIGKILL + restart): unchanged-key dedup fast path, update/delete reconcile, orphan sweep on boot, collection_id pin surviving a post-recovery compact + GraphUnion merge, and a no-persistence regression guard. All waits are bounded condition polls; servers run with --disk-free-min-pct 0 so the suite is immune to the dev volume hovering at the 5% diskfull guard.

Added — vector-index durability: startup recovery from disk (B1-B3)

  • Vector indexes now persist their segments across restarts instead of paying a full re-index (TQ encode + HNSW build) of every matching hash key on every restart. Each index gets an idx-<hex(name)>/ directory holding an atomically-written manifest.json (collection_id / segment ids / id-allocator floors), a checksummed keymap-<epoch>.bin (key_hash → global_id + vector checksum + original key), and its immutable HNSW segments (staged write: staging-<id> → fsync → atomic rename), written in the background after each compact/merge install (B1/B2).
  • Startup now loads that state instead of discarding it: segments are read back and reattached (pinning the recreated index's collection_id to the persisted value — required for the HNSW QJL rotation seed to match), and the keyspace rescan is now a dedup rescan: a key whose vector bytes checksum-match the last snapshot is rebuilt as metadata-only (no HNSW/TQ re-encode); only genuinely changed or unknown keys are fully re-indexed; keys removed from the keyspace are tombstoned (B3).
  • Crash-safety contract: any corrupt/missing artifact (segment, checksum mismatch, unreadable headers) degrades to a rescan-rebuild of exactly the affected keys, never to wrong search results; a manifest may understate what's durable (costs a rescan) but never overstates. Startup also sweeps orphaned segment/staging/keymap files and stale idx-* directories left by an interrupted drop.
  • Multi-field indexes: only the default vector field's segments/checksum are persisted; additional named vector fields always re-encode on restart (documented gap, safe/conservative).
  • Removed the vestigial WAL-replay vector recovery path (src/vector/persistence/recovery.rs, VectorStore::attach_recovered/ pending_segments): it only ever populated a VectorStore that was discarded before the shard event loop started (recovery authority is now exclusively the manifest/segment/keymap layout above).

Fixed — FLUSHALL/FLUSHDB ghost vectors + HDEL stale-vector gap

  • FLUSHALL/FLUSHDB never touched the FT indexes (persistence-review R3): flushed hashes stayed searchable as ghost vectors/documents until restart. Both commands now clear every vector + text index's CONTENTS (segments, key-hash maps, postings, TAG/NUMERIC indexes) while KEEPING the FT.CREATE definitions — matching restart semantics, where definitions come from the sidecar and contents are re-derived from the (now empty) keyspace. Recovered-but-unclaimed pending_segments are discarded too, so a post-flush FT.CREATE cannot resurrect pre-flush contents.
  • HDEL key <vector-field> left the vector searchable (R4): only whole-key DEL/UNLINK tombstoned. A successful HDEL that removes an index's vector field now tombstones that key in exactly the affected indexes (per-index mark_deleted_for_key_in_index; sibling indexes keyed on other fields keep their entries). Known follow-ups documented in auto_hdel_vectors: multi-vector-field indexes tombstone the whole doc, and TEXT/TAG/NUMERIC field removal is not yet re-indexed.
  • Both hooks wired with wire-parity on ALL dispatch paths (single-shard, monoio conn-local, tokio sharded, SPSC execute, shared batch, MULTI/EXEC). New integration suite tests/vector_flush_hdel_tombstone.rs (red→green).

Observability — exact-rerank sidecar coverage is visible and loud (R5)

  • A GraphUnion merge with even one sidecar-less source segment silently dropped the exact-rerank f16 sidecar for the ENTIRE merged segment (all-or-nothing propagation — a partial sidecar would mix exact and ADC distances within one segment). The drop now logs a tracing::warn with source counts, and FT.INFO exposes additive graph_segments / segments_with_exact_rerank counters (merged across shards) so ADC-only segments are observable in the steady state.

Performance — AE-1: saturation-gated per-segment adaptive ef

  • With G graph segments, every segment searched at the FULL resolved ef — ~G× the CPU of one ef-wide beam (the root cause of the pooled regression on SMT-constrained boxes in PR #214's GCE run). Compact/GraphUnion builds now run a one-time self-probe against the exact f16-sidecar ground truth (leave-self-out R@10 across an ef ladder, estimate_suggested_ef): a segment whose curve is FULLY SATURATED from the minimum rung (flat at ≈1.0) is certified trivially easy and searches at min-ef (24); every other segment keeps the full resolved ef. Applied identically to the sync, serial-yielding, and worker-pool paths (pooled == serial identity holds); never active when the user pins EF_RUNTIME; in-memory only (reloaded segments fall back to the full beam).
  • Two weaker designs were tried and REJECTED by recall A/B (20k×384d SQ8, 5 segments): a blanket ef/√G split (EF-SPLIT) silently cost gaussian R@10 0.9915 → 0.9295 (its original "recall preserved" validation had exercised a stale binary), and letting the self-probe pick a mid-ladder knee cost 0.9915 → 0.939 — self-sampled queries are optimistically biased on unstructured data, so the probe's ONLY trustworthy verdict is "saturated at min-ef".
  • Recall/latency gate (same seeds/harness end-to-end): clustered 5-seg R@10 1.0000 at p50 0.79 → 0.26 ms (~3×), clustered 1-seg 1.0000 → 0.9960 (within the probe's ε=0.005 contract) at 1.78 → 1.46 ms; gaussian 5-seg 0.9915 at ~0.79 ms and 1-seg 0.7710 — byte-identical to the pre-split baseline (probe self-rejects, full beam).

Performance — FT.SEARCH per-query fixed-cost cleanup (QP-2/QP-3)

  • Thread-cached SearchScratch (QP-3): the wire path built a fresh scratch per query — heaps, visited bitmap, and the 32–65KB ADC LUT reallocated every FT.SEARCH. Captures now take a thread-local recycled scratch (exact padded-dim match, mirroring the worker pool's cache) and the yielding search returns it on completion.
  • Amortized committed-treemap capture (QP-2): every search cloned the MVCC committed RoaringTreemap; TransactionManager::committed_snapshot() now hands out a cached Arc refreshed only on the first capture after a commit/prune — read-heavy captures are one refcount bump, write-heavy is never worse than before. All 7 capture sites switched.
  • ryu score formatting: RESP __vec_score replies formatted f32 via Display (grisu, 1.4% of a matched-recall query); now ryu.
  • VM A/B (same box/config as the f16 kernel entry): ef64 8,468 → 8,810 QPS (+4.0%; +13% cumulative with the f16 kernels).

Performance — SIMD f16 exact-rerank kernels (NEON + F16C)

  • The HQ-1 exact-rerank pass decoded its f16 sidecar with scalar software f16_to_f32 — 17% of a matched-recall (ef 64) query. New fused decode+distance kernels in the distance dispatch table (f16_l2, f16_dot_normsq): aarch64 uses a baseline-NEON integer-rescale decode (no ARMv8.2 FP16 intrinsics needed; exact for all finite f16, Inf/NaN bit-selected), x86_64 uses F16C vcvtph2ps + FMA (runtime-detected with AVX2, scalar fallback). Subnormal/Inf/NaN semantics match the scalar reference exactly (pinned by tests). VM A/B (20k×384d clustered SQ8, single conn, ef 64): 7,795 → 8,468 QPS (+8.6%).
  • Negative result, kept for the record: skipping the rerank pass entirely for SQ8 was A/B'd and REFUTED — it costs real recall (R@10 −0.010 clustered, −0.017 gaussian 5-segment). The pass stays unconditional; the SIMD kernels make it cheap instead.

Performance — FT.SEARCH intra-query worker pool + bounded bulk compaction

  • --ft-search-workers N worker pool (src/vector/search_pool.rs): the per-segment HNSW searches of ONE KNN query fan out across a pool of searcher threads while the shard task scans the mutable segment; replies are awaited with flume::recv_async, so the shard event loop is never blocked. Default 0 (off, opt-in) — each segment pays the full resolved ef, so a pooled N-segment query does ~N× the CPU work for its latency win: a 4.9× QPS win on physical-core-rich boxes but a measured regression on SMT-constrained ones (same posture as --io-busy-poll-us). Good opt-in size: physical cores − shards, cap 8. Results are identical pooled or serial (enforced by pooled-vs-serial identity + 8-thread concurrency stress tests); worker panics are contained per-job (empty segment result + warn, never a hang). Runtime-agnostic (monoio + tokio).
  • Bounded bulk compaction: with a pool active and COMPACT_THRESHOLD > 0, one compaction build freezes at most max(threshold, len/8) entries (MutableSegment::freeze_prefix), so FT.COMPACT after a bulk load yields several independently searchable segments instead of one giant graph — bounded build memory and pool-parallel search. Pool-less deployments keep single-segment builds (multi-segment serial search would be strictly slower). Red/green: test_force_compact_bulk_bounded_segments, test_bg_compact_bulk_bounded_segments.
  • Measured (20k×384d clustered SQ8, single connection, post-FT.COMPACT, R@10 = 1.0 in all rows): macOS 10-core p50 2.06 ms → 0.46 ms, 473 → 2,321 QPS (4.9×); GCE c3-standard-8 (4 physical cores) regresses at default ef (1,732 → 1,165 QPS) — hence the opt-in default. The GCE probe matrix also showed the per-index EF_RUNTIME knob alone closes most of the vs-RediSearch clustered gap (ef 64: 4,493 QPS serial at R@10 = 1.0). Full GCE vs-RediSearch tables in PR #214.

Fixed — update-churn no longer mass-deletes vectors at compact/merge install (soak-diagnostic find)

  • Compact install: dead window entries no longer kill their own update. snap_and_reconcile treated ANY tombstoned entry in the frozen window as "key deleted" and applied a key_hash-wide tombstone to the new immutable — but since the VEC-1 update path tombstones the old copy in place, a key updated before the freeze had a dead old copy AND a live new copy in the same window, and the install deleted the new copy out of the segment. Every key updated-then-compacted silently vanished from FT.SEARCH (32% of live keys lost / live-set recall 0.985 → 0.685 in a 1-minute churn soak; regression vs v0.5.1). A dead window entry now only proves deletion when the key has no live window sibling.
  • Merge install: source tombstones are origin-gated. The merge replay applied each source segment's lifetime interior tombstone set key_hash-wide to the merged output, so a key whose old copy was tombstoned-by-update in one source killed its current copy merged in from a sibling segment. The replay now only tombstones entries whose global_id originated in the tombstone's own source; DEL/UNLINK tombstones land in every source's set and still apply everywhere.
  • VectorStore::insert_vector allocates monotonic LSNs (same allocator as the wire path) instead of mutable.len()+1, which restarted after every compaction and made merge dedup keep a stale copy over the current one.
  • Red/green: test_bg_compact_update_before_freeze_survives_install, test_bg_merge_update_across_segments_survives; end-to-end churn repro (24k mixed ops): LOST 1054 → 0, live-set recall 0.685 → 0.98.
  • Deleted keys no longer resurface in FT.SEARCH — the vector auto-delete hook (mark_deleted_for_key) existed only on the cross-shard SPSC Execute arm and the tokio sharded handler. The monoio conn-local path (the default runtime's only path at --shards 1), handler_single, the MULTI/EXEC batch paths, and the SPSC pipeline arms never tombstoned vectors on DEL/UNLINK, so deleted keys kept matching KNN searches forever (20% of results after one minute of mixed churn; live-set recall collapsed 0.985 → 0.735 in the Bundle-5 soak diagnostic). All paths now share one auto_delete_vectors parity helper; wire-level red/green coverage in tests/vector_del_unindex.rs.

Added — long-run vector reliability harness

  • scripts/vector-validate.py — on-target validation driver: recall/QPS comparison between two moon binaries (SQ8 + TQ4, ground-truth brute force), churn soak with live-set recall / RSS / resurrection sampling, and kill -9 durability (settled-write survival + double-crash restart) under appendonly yes.
  • scripts/gcloud-vector-soak.sh — GCE orchestration (c4a ARM + c3 x86): ships the working branch via git bundle, builds baseline + branch binaries, runs the validation driver, fetches JSON results, tears down on exit.
  • scripts/bench-vector-vs-redisearch.py — Moon vs RediSearch head-to-head driver (same FT.* wire dialect for both engines): insert throughput, KNN-10 QPS/p50/p99, and R@10 vs numpy ground truth on random-Gaussian and clustered-mixture 384d datasets, with an EF_RUNTIME sweep so QPS is compared at matched recall.

Fixed — vector search correctness bundle (deep-review VEC-1/XC-SHARD-1/XC-3/VEC-4/VEC-7)

  • HSET update no longer duplicates a vector — re-indexing an existing key tombstones the old copy first (O(1) fast path in the mutable segment via the key→global-id map, scan fallback across mutable + immutable segments), so KNN totals and results no longer count both the stale and the fresh vector.
  • FT.INFO is now cluster-wide at --shards N — previously answered from the local shard only (~1/N of num_docs). Now scatter-gathers to every shard and merges additively (top-level counters + per-field stats), on both the sharded and monoio handlers.
  • Filtered FT.SEARCH on the cooperative-yield path honors FilterStrategy — the yielding search always graph-filtered; the post-filter strategy (selective filters) now searches unfiltered with 3×k oversampling and bitmap post-filter, matching the non-yielding path's recall behavior.
  • FT.CREATE ... MERGE_MODE KEEP_RAW is rejected fail-loud — it silently behaved as graph-union; now errors until the raw-vector sidecar is implemented (the separate KEEP_RAW ON flag is unchanged).
  • MEMORY DOCTOR / Prometheus KV memory no longer report 0 under unlimited maxmemory — the per-shard KV memory publish on the 100ms eviction tick had been gated on maxmemory > 0 since the GAP-1 elastic-budget work, permanently zeroing the DashTable + entries line for the default config. The publish (an O(1) accumulator read per DB) now runs unconditionally; elastic-budget recompute stays gated on a finite cap.

Added

  • FT.CONFIG SET/GET <index> MERGE_RECALL_TOLERANCE <0.0..=1.0> — per-index recall gate for unattended (background/vacuum) GraphUnion merges; default 0.70 unchanged.

Added — exact rerank stage (deep-review HQ-1)

  • FT.SEARCH distances on compacted segments are now (near-)exact. Immutable segments carry an f16 sidecar of the original vectors (built at compaction, BFS-ordered); the top 4·k beam candidates are re-scored with true metric distances (L2: squared L2; Cosine/InnerProduct: normalized-pair squared L2) before top-k truncation, replacing pure quantized ADC estimates — the recall lever vs engines that keep full-precision vectors. SQ8, previously ZERO-refinement, benefits most; TQ4 estimates improve from percent-level error to f16 tolerance (~1e-3).
  • The sidecar persists (raw_f16.bin per segment dir, missing file = no sidecar, fully backward/forward compatible), survives GraphUnion merges (all-or-nothing propagation), and is MEMORY DOCTOR-accounted. Memory cost: +2·dim bytes per vector in both the mutable segment and compacted segments (384d ≈ +768 B/vector); an opt-out knob is a follow-up.

Performance — SIMD SQ8 ADC kernels (deep-review HQ-2)

  • SQ8 asymmetric distance is now SIMD-dispatched (NEON / AVX2+FMA / AVX-512F, scalar fallback) via an algebraic decomposition: per-query constants (Σq, Σq²) computed once, per-candidate work reduced to three fused widen-u8→f32 FMA sums (Σq·c, Σc, Σc²) combined in O(1). Wired into HNSW beam search and both mutable-segment brute-force paths. Criterion (aarch64 NEON): 2.7×/1.8×/1.6× faster per candidate at 128/384/768d. Note: the pre-AVX2 x86 scalar fallback is ~1.65× slower than the old naive loop (the decomposition only pays with SIMD); all supported targets (aarch64 NEON, x86-64 AVX2+) are wins.

Performance — vector search hot-path quick wins

  • Copy-on-write Arc key map (no full HashMap clone per snapshot), hoisted brute-force query prep, striped search metrics counters, stable aarch64 prfm node prefetch, dead quantize removed from the insert path, SQ8 code_len fix in compaction, and raw-buffer work skipped for SQ8 segments.

Performance — coordinator local legs ride group commit under appendfsync=always (PR #TBD)

  • The cross-shard coordinator's LOCAL-leg persist (co-located MSET/MSETNX and scattered-MSET local slices) awaited one fsync ack per command, each bounded by --aof-fsync-timeout-ms (default 2000ms) — a pipeline of coordinated writes stacked these serially into the 2000–3000ms always far-tail measured on Linux/GCE (c2d-standard-16, pd-ssd; see the v3-4 bench notes — OrbStack/macOS fsync is near-free and does not reproduce it). Local legs now enqueue fire-and-forget (bounded backpressure, same contract as the remote SPSC legs) and the connection handler confirms them with one fsync_barrier on the local shard per pipeline batch, before responses are serialized — so +OK still implies confirmed durability, but a batch of N coordinated writes costs 1 awaited fsync instead of N. everysec/no behavior is unchanged. On barrier failure every affected response is replaced with MOONERR AOF fsync (never a false +OK).
  • The cross-shard coordinator's in-process legs for BITOP (dest write), COPY (dst write + TTL restore), and multi-key DEL/UNLINK (co-located fast path AND the scattered local slice) executed in memory but never reached the owning shard's AOF — deleted keys resurrected from their seed writes on restart, and BITOP/COPY results on the connection's own shard silently vanished (carried v3-4 follow-up; remote legs were always durable via MultiExecute). All four now persist through the same persist_local_leg group-commit path as MSET/MSETNX: synthesized over only locally-owned keys, skipped when nothing was written (DEL of missing keys), and confirmed by the batch-end fsync barrier under appendfsync=always.

Fixed — cold-tier (disk offload) correctness & reliability (PR #TBD)

  • DEL/UNLINK/FLUSHALL now reach the cold tier — deleting a key whose value had been offloaded to disk left its ColdIndex entry alive, so the next GET resurrected the deleted value from the .mpf heap file (DEL even returned 0 for cold-only keys). Database::remove/clear now drop the cold-index entry (and queue file unlinks) alongside the hot entry, and DEL/UNLINK count cold-only keys correctly.
  • Expired cold reads reclaim their index entry — a cold read that found an expired entry returned nothing but left the index entry + file refcount behind forever (nothing else ever reclaims them). The read path now distinguishes Expired from Miss and removes the index entry on expiry; transient I/O errors still leave the entry alone.
  • Spill files survive a crash at the directory level — spill writes fsynced the file but never the directory, so after a power loss the (dir-fsynced) manifest could reference a heap file whose directory entry vanished. Both the batch (tmp+rename) and single-file spill paths now fsync data/ after publishing the file.
  • Crash-orphaned heap files are swept at startup — a crash between the spill write and the manifest commit left heap-*.mpf/.tmp files on disk that no manifest references (invisible to the cold index, leaked disk forever). Recovery now unlinks unregistered heap files once the manifest has opened successfully.

Added

  • Spill-thread liveness metrics in INFO persistence — spill_batches_flushed, spill_completions_dropped, and spill_last_heartbeat_ms (0 = never ran) expose a silently-dead spill thread, whose only prior symptom was an unbounded eviction backlog.

[0.5.1] — 2026-07-04

Fixed

  • AOF+WAL double-write eliminated for the common config (PR #211) — new --wal-kv-log auto|on|off (default auto). At --appendonly yes every SPSC-executed KV write was logged to BOTH the per-shard AOF and the per-shard WAL (measured 2.7× file-byte / 4.1× device write amplification at --shards 4), yet startup recovery wipes WAL-replayed state and replays the AOF — the WAL copy was pure disk wear. auto skips WAL KV records while the AOF is the recovery authority and no CDC subscriber is attached (logging re-engages dynamically when one attaches); on restores the pre-0.6 always-log behavior (needed for PITR / full CDC history alongside AOF); off never logs KV records. FPI/checkpoint/ feature records are unaffected.
  • everysec/no AOF appends are no longer silently dropped under backpressure (PR #211) — a full writer channel used to warn! + drop the record while the client still received +OK (client-acked write loss + AOF/memory divergence on replay). Durable handler paths now await enqueue under --aof-fsync-timeout-ms and surface ChannelFull/WriteFailed as an error frame; the synchronous SPSC-drain and monoio inline-SET paths apply a bounded blocking send (ONE 5ms budget shared across a whole pipeline/MULTI batch) and, on loss, replace the success frame with MOONERR AOF backpressure instead of acking — every drop is error!-logged and counted. --wal-kv-log also rejects unknown values at parse time. Watch aof_backpressure_dropped in INFO.

  • Docker image build — the multi-stage Dockerfile now copies the vendored vendor/monoio source into the cargo-chef planner and cook stages. The v0.5.0 release introduced a [patch.crates-io] monoio = { path = "vendor/monoio" } path dependency, but the chef stages only copied Cargo.toml/Cargo.lock/src, so cargo chef cook panicked (failed to load source for dependency monoio … vendor/monoio/Cargo.toml: No such file or directory) and the Docker Image release job failed. Adds a standalone docker-publish.yml workflow to (re)publish only the container image for a tag whose other assets are already live.

[0.5.0] — 2026-07-04

Two milestones since v0.4.1:

  • v3-4 KV Write Correctness & Data-Integrity Parity — the primary KV engine now honors Redis's data-integrity contracts (no wrong-type data loss, no integer-overflow panics/wraps, past-expire deletes, atomic MSETNX).
  • p=1 & multi-shard throughput — Moon now beats Redis on non-pipelined (p=1) GET/SET on both GCE arches (ARM 1.19–1.21×, x86 1.65–1.66×), and multi-shard wins from 8 concurrent connections up (s4 c8 1.57–2.0×, c64 2.50×), via a poll-mode park (--io-busy-poll-us, backed by a vendored monoio fork) plus a cross-shard reply-path convoy fix. All perf gains are bench-only knobs / opt-in flags — the default runtime behavior is unchanged unless you set --io-busy-poll-us.

Added

  • MSETNX — atomic multi-key string write (set all pairs iff none of the keys exist; returns 1/0). Single-shard atomic (two-phase check-then-set, no await between phases). The cross-shard coordinator rejects a key span with CROSSSLOT (by design — no two-phase commit) and runs co-located {hash-tag} keys atomically on the owning shard. New --shards 4 regression suite tests/msetnx_cross_shard_reject.rs.
  • --io-busy-poll-us <µs> — poll-mode park for the monoio legacy (epoll/kqueue) driver: the shard thread busy-polls readiness for the given budget before blocking, deleting the per-op scheduler sleep+wake. This is the lever that flips non-pipelined (p=1) GET/SET to a win vs Redis on both GCE arches. Defaults to 0 (off); implies --io-driver epoll. Costs up to budget-µs CPU per idle park — only a win on pinned, disjoint cores (a regression on shared-core hosts). Env equivalent MOON_EPOLL_SPIN_US.
  • --io-driver <auto|epoll> for the monoio runtime — force the epoll/kqueue LegacyDriver instead of io_uring (some platforms, e.g. GCE ARM Axion, run KV faster on epoll).
  • Per-use-case tuning guide at docs/guides/tuning.md (quick recipes, shard-count guidance, busy-poll trade-offs, persistence cost, platform notes), linked from the docs-site Operations nav.

Performance

  • Multi-shard now wins from 8 concurrent connections up, even without pipelining. Fixed the C2 reply-spin convoy: the cross-shard reply busy-poll was synchronous on the shard thread and, at ≥2 connections per shard, a spinning connection starved its sibling and the shard's SPSC drain — the root cause of the --shards 4 c8-P1 0.45× collapse. A solo-conn gate now spins only when a connection is alone on its shard. Result (s4, P1, --io-busy-poll-us 40, vs Redis): c8 GET/SET ARM 1.57/1.71×, x86 1.9/2.0×; c64 2.50× both arches. (c1-P1 remains the structural single-hop ceiling, 0.91–0.96×.)
  • monoio cross-shard replies unified on the zero-allocation ResponseSlotPool (L3b), removing the per-op flume-oneshot allocation and chunked park from the hot reply path (+11–14% at c8).
  • Non-pipelined GET reply framed from a borrow — dropped a per-GET Vec copy.
  • Stripped four per-op lock/atomic costs from the p=1 dispatch hot path.
  • Tuned the default monoio io_uring driver (COOP_TASKRUN | SINGLE_ISSUER | DEFER_TASKRUN) on Linux.

Changed

  • Corrected default-config documentation: --shards default is 1 (not "0/auto") and --appendonly default is yes (not "no") — both config pages were wrong.

Fixed

  • Cross-shard reply use-after-free on panic-unwind (memory safety). The cross-shard reply path sent a raw *const ResponseSlot into a pool living on the connection task's stack; a panic unwinding through the reply dispatch/drain would drop that pool while a target shard still held the pointer, making the target's later slot.fill() a use-after-free (heap corruption). ResponseSlotPtr now carries an Arc<ResponseSlot>, so the slot outlives whichever of {connection pool, in-flight message} drops last — structurally closing the window on both dispatch and drain. This also retroactively fixes the same latent hazard on the tokio runtime (shipped since v0.4.0) and removes four unsafe blocks (net unsafe reduction). Follow-up: a shutdown-aware bound on the reply await (now safe to add) to close a partial-shutdown liveness hang.
  • MOON_NO_URING=1 now actually switches the monoio driver. It was a silent no-op for the monoio FusionDriver (io_uring was selected regardless); it now forces the epoll/kqueue LegacyDriver, matching its documented contract and the new --io-driver epoll.

  • KV data-integrity (P0): wrong-type GETDEL / GETSET / SET … GET no longer silently delete or overwrite a key holding a non-string — they return WRONGTYPE and preserve the value (check-before-mutate).

  • Integer overflow (P0): DECRBY key -9223372036854775808 no longer panics; every expiry-setting command (SET EX/PX/EXAT, SETEX, PSETEX, EXPIRE/EXPIREAT/PEXPIRE) now rejects a time whose absolute expiry falls outside the i64 millisecond domain with an "invalid expire time" error — matching Redis (when > LLONG_MAX/1000). This closes a latent wrap where an accepted-but-huge TTL (e.g. EXPIRE k 15000000000000000) surfaced as a negative PTTL/PEXPIRETIME on a live key, and makes an extreme-negative EXPIRE/EXPIREAT (< i64::MIN/1000) error instead of deleting.
  • Expire-in-the-past semantics (P0): EXPIRE/PEXPIRE/EXPIREAT with a non-positive or already-past time now deletes the key and returns 1 (Redis parity); verified to propagate through command-based WAL replay so replicas stay consistent.
  • Cross-shard coordinator local-leg durability (P0, pre-existing): a co-located MSET/MSETNX whose keys hash to the connection's own shard now appends to that shard's AOF. The coordinator's local leg previously executed the write in memory but never persisted it (the remote MultiExecute leg always did), so with --shards >1 --appendonly yes an own-shard co-located MSET/MSETNX could be lost on crash. It now persists via the same append path every local single-key write uses — the whole command for a co-located owner, and a synthesized MSET over only the local keys for a scattered MSET's local slice — returning an AOF error instead of a false +OK on append failure. Crash-recovery verified on both monoio and tokio (tests/coordinator_local_leg_durability.rs). Single-shard (the default) was never affected. The same local-leg gap remains for coordinator BITOP / COPY / DEL / UNLINK (tracked follow-up).

Dependencies

  • Vendored monoio 0.2.4 (vendor/monoio, wired via [patch.crates-io]; upstream git sha f7827ddd, recorded in vendor/monoio/.cargo_vcs_info.json). A fork of upstream monoio 0.2.4 with a bounded "moon patch" (marked // moon patch, confined to 4 files — lib.rs, driver/mod.rs, driver/uring/mod.rs, driver/legacy/mod.rs — verified against pristine upstream) that adds the env/CLI-gated poll-mode park behind --io-busy-poll-us and a per-thread spin-hook handshake. Adds zero new unsafe; the dependency list is byte-identical to upstream (no added transitive deps); behavior is byte-identical to upstream when the MOON_* spin knobs are unset. Introduced to enable the p=1 poll-mode park.

[0.4.1] — 2026-06-23

Measure-only validation release — no server behavior change. Bundles the closed milestone Cross-Shard Read Absolute Validation (v2-2): bare-metal GCloud proof that the v0.4.0 cross-shard read C2 fast-path performs as claimed, plus a reusable absolute-latency benchmark harness and a data-backed decision to not pursue further cross-shard-read optimization.

Added

  • scripts/gcloud-xshard-absolute.sh — reusable dual-vendor (Intel c3 + AMD c2d), dual-runtime (monoio + tokio) absolute cross-shard latency harness: best-of-5 RPS floor, fail-closed validity gates (load / steal / clean / s1-LOCAL-control), per-run provision + teardown, and an --self-test that exercises the gates with no cloud cost.

Validated

  • The shipped C2 cross-shard read fast-path, measured absolute on bare-metal across four instruments (Intel + AMD × monoio + tokio): C2 recovers 38–49% of the SPSC-migration regression, leaving a vendor/runtime-converging ~10µs (±1µs) structural residual — one irreducible cross-thread reply hop. monoio is the high-confidence set (tokio is rep-bimodal; best-of-5 floor clean). Disposition: close-the-line — coalescing cannot help the singleton path (the concurrent guard is already flat) and RCU's per-shard RSS cost is unjustified by a 10µs gap. Mechanism asserted byte-identical (tests/xshard_mechanism_unchanged.rs); git diff --stat src/ empty.

Risk-accepted (shipped, disclosed)

  • The shardslice cross-shard-read fast-path removal waiver (PR #175, owner Tin Dang, expires 2026-08-01) rides into this release. This release validates it: the regression is quantified and the C2 recovery confirmed on bare-metal, so the waiver is now retirable and its follow-up task (cross-shard-read-acceleration) is closed as "no further work warranted."

[0.4.0] — 2026-06-22

Bundles 5 closed milestones (first ADD-recorded cut; see RELEASES.md for attribution): Shared-Nothing Integrity & Cross-Shard Latency (3 carried) · Multi-Core Throughput Hardening (12 carried) · Throughput Polish (3 carried) · FTS Hardening (8 carried) · Graph Correctness & Cypher Filtering. Headlines: working FTS query combinators (OR / TEXT+TAG / NUMERIC) with true total-matched counts and the O(M²) high-DF cliff gone; correct Cypher inline-property narrowing, incoming-edge traversal after compaction, and >32 graph labels; WAL group-commit, cross-shard read fast-path, and off-event-loop FT.SEARCH throughput wins. Per-change detail below.

Risk-accepted (shipped, disclosed) — cross-shard read fast-path removed (PR #175)

The shared-nothing migration (PR #175) deleted the direct-RwLock cross-shard read fast-path; cross-shard reads now route through the owner shard's SPSC ring, so cross-shard-read cells regress versus the pre-migration lock-read. This is a signed RISK-ACCEPTED waiver (owner: Tin Dang) carried into this release, with a follow-up task cross-shard-read-acceleration (lock-free cross-shard read acceleration) tracked to land before 2026-08-01. No correctness impact — consistency is 197/197 across 1/4/12 shards on both runtimes; the cost is latency on the cross-shard read path only.

Docs — Verified cross-arch GCloud benchmark confirms v3-1 FTS + v3-2 graph fixes (PR #194)

Re-ran the 4-feature concurrent-vs-competitor benchmark (KV / Vector / Graph / FTS vs Redis / RediSearch / FalkorDB) on GCloud x86 c3-standard-8 (Sapphire Rapids) + ARM t2a-standard-8 (Neoverse-N1) against build 8238515, with the graph harness corrected to filter on the node id property value (not the internal GRAPH.ADDNODE handle that produced a false cypher_match_rows=0 in the prior re-baseline). Proves in-harness that v3-2's inline filter narrows Cypher point-queries (cypher_match_rows 14991 → 4, identical to FalkorDB; cypher_1hop ~28–30× faster than its old full-scan self) and that v3-1's FTS hardening makes every query combinator return RediSearch-exact total counts (OR 2072, TAG 5064, TEXT+TAG 253, NUMERIC 10119) while indexing ~40× faster (376 → 15052 docs/s) with the O(M²) high-DF cliff gone (419 ms → 33 ms). KV and Vector unchanged (no regression). Folded as layered BENCHMARK.md §2.9 / §10.6 / §11.5 / §12.3 over the 2026-06-16 record; full report + corrected harness + raw logs in docs/reviews/2026-06-17/.

Docs — Wider cross-arch benchmark: shard scaling (1/4/12) × structures × pipeline/datasize × production (PR #194)

Ran the repo-canonical scripts/bench-compare.sh (all-commands × pipeline 1–128 × datasize 8B–64KB × connections) at --shards 1/4/12 plus scripts/bench-production.sh (10 scenarios at shards 1 and 12) on GCloud x86 c3-standard-8 + ARM t2a-standard-8, Moon 8238515 vs Redis 7.0.15. Quantifies the core scaling claim honestly: Moon's advantage is pipeline depth, not shard count (break-even p=16; GET p=128 ~2.8×, SET p=64 ~1.85×), every core data structure tracks GET/SET at p=1 (~0.79×, TCP-RTT bound), and adding shards hurts uniform single-key non-pipelined throughput (12 shards → 0.46–0.51× Redis vs ~0.79× at 1 shard). Folded as BENCHMARK.md §4.5 (structure coverage) + §6.4 (cross-arch shard scaling); full detail + logs in docs/reviews/2026-06-17/WIDER-BENCH.md. The run also patched a latent unfairness in the bench scripts (they start Redis AOF-off but Moon AOF-on, flooding ~1M WARN lines).

Follow-up root cause for the shards=12, pipeline=1 worst case: a controlled single-VM A/B (docs/reviews/2026-06-22/) isolates it as a three-factor interaction — cross-shard dispatch (2 cross-thread wakes/op) × p=1 (no batch amortization) × random keys (defeat the ≥62.5% AffinityTracker migration, so the connection never goes local). The collapse is confined to the s12 × random cell (GET 0.47× / SET 0.41×); s1 random is 0.95×/1.15× and s12 single self-heals to 0.74×/0.73×. WIDER-BENCH.md §4a + XSHARD-P1-ROOTCAUSE.md.

Fixed — bench-compare.sh / bench-production.sh start moon AOF-off to match Redis (fair, no flood) (PR #194)

Both benchmark scripts started Redis in-memory (--save "" --appendonly no) but moon with persistence defaults (AOF on). Under SET load moon's AOF channel saturated, flooding AOF append dropped … channel full WARN logs (~1M lines, which also pinned the port so later shard configs failed to bind) and unfairly handicapping moon against a no-AOF Redis. Both scripts now start moon with --appendonly no --disk-offload disable (and redirect its stdout/stderr), so the comparison is AOF-off on both sides and the report stays clean.

Fixed — Graph node labels with id ≥ 32 are stored, matched, and persisted (no silent truncation) (PR #193)

CSR node labels were packed into a 32-bit NodeMeta::label_bitmap, so any label id ≥ 32 was silently dropped at from_frozen (if label < 32) — MATCH (a:Label40) matched nothing and a graph using more than 32 distinct labels lost the rest. Labels 0–31 keep the u32 bitmap fast path; ids ≥ 32 now live in a sparse, per-row label_overflow store, and LabelIndex is built from both sources, so every u16 label id is matchable. The on-disk CSR format is bumped to version 4 with a new version-gated trailing section ([entry_count][ (row, label_count, label_id…) ], ascending by row, covered by the existing CRC); the GraphSegmentHeader gains a label_overflow_offset field reusing existing padding (on-disk header size unchanged). Both the heap (from_bytes) and mmap (from_mmap_file) loaders parse the section into an owned map via one shared bounds-checked parser (returns CsrError on truncation, never panics; transitively fuzzed by csr_from_bytes), and compact_segments re-keys overflow labels to each surviving node's new merged row. Backward-compatible: version ≤ 3 segments load with an empty overflow store and behave exactly as before; NodeMeta's 48-byte layout is unchanged.

Fixed — Graph Incoming / Both traversals return incoming edges after compaction (PR #193)

A compacted graph stores edges only in immutable CSR segments, which keep just OUTGOING adjacency. SegmentMergeReader therefore continued past the CSR for Direction::Incoming (returning nothing) and silently dropped the incoming half for Direction::Both — so MATCH (a)<-[]-(b) / undirected traversals over compacted data missed every predecessor. The reader now serves predecessors from a derived reverse index (IncomingIndex) built once per segment from the existing forward arrays (row_offsets + col_indices) and cached in a non-persisted OnceLock — no segment format/version change, so on-disk heap and mmap segments answer incoming queries unchanged. Incoming edges reuse the same edge_meta / edge_created_ms / validity / node-visibility by edge index, so the edge-type filter, tombstones, decay stamp, deleted-node exclusion, snapshot rules, and NodeKey dedup are identical to the outgoing path. Direction::Outgoing is unchanged. MemGraph already handled all three directions; this fixes the immutable-CSR path only.

Fixed — Cypher MATCH narrows on inline node-property predicates instead of full-scanning the label (PR #193)

compile_match built each pattern node's NodeScan/Expand but silently dropped the node's inline properties (v {k:e, …}), so MATCH (a:Person {id:1})-[]->(b) ignored {id:1} and returned every edge of the Person label (≈|E| rows) instead of node 1's out-edges. The planner now emits a Filter — the equality conjunction v.k = e AND … built by the existing properties_to_filter — immediately after the op that binds each inline-propertied node: after its NodeScan for the first node, after the Expand that produces any subsequent node. It reuses the existing Filter/Expr evaluation, so parameters ({id:$p}), cross-type compares, and missing-property ⇒ excluded all follow the same semantics as WHERE; an empty {} pattern emits no Filter, leaving plain label/all scans byte-identical. No new PhysicalOp, no index, no executor change.

Fixed — FT.SEARCH routing: prose "knn" searches as text, standalone SPARSE reaches the vector engine, no panic on the BM25 AND path (PR #192)

is_text_query uppercased the query and matched the bare substring KNN, so a text search whose terms merely contained the word knn (e.g. FT.SEARCH idx "knn tutorial") was misclassified as a vector query and fell to the KNN parser, returning ERR invalid KNN query syntax instead of text results. It now keys on the canonical [KNN vector-query bracket — prose containing "knn" searches as text while *=>[KNN …] still routes to the vector engine. A standalone SPARSE @field $param clause paired with a text-looking query string was silently dropped onto the text path (is_text_query only sees args[1]); a has_sparse_clause guard now defers any SPARSE-carrying query to the vector engine at every text-route gate. The three .expect("posting exists") on TextStore::search_field's BM25 AND + scoring path are replaced with defensive control flow, so a vanished posting yields empty results instead of panicking the server. Output is byte-identical for all existing queries.

Fixed — FT.SEARCH integer reply is the true total-matched, not the page size (PR #192)

The first element of an FT.SEARCH reply (the match count) was capped at the LIMIT/top_k page size: FT.SEARCH idx "term" LIMIT 0 5 over 100 matches reported 5, not 100, because the count was read from the already-truncated result page. On multi-shard indexes the coordinator compounded it by counting the merged-and-truncated returned docs rather than each shard's true matched count. The reply now reports the true number of matched, key-resolvable documents (RediSearch semantics): the evaluator surfaces the count before truncation, and the multi-shard merge sums each shard's local matched count (keys partition to exactly one shard, so the sum is exact; errored shards contribute 0). Verified identical on 1- and 4-shard servers. Returned document pages (count, order, scores, keys) are unchanged.

Performance — FT.SEARCH upsert / bulk re-index no longer O(V) per document (PR #192)

Re-indexing a document (an HSET upsert, or a bulk re-index pass) called PostingStore::remove_doc, which scanned EVERY term's posting list to clear one document — O(total-vocabulary) per doc — so per-upsert cost grew with the corpus (the indexing-rate cliff: ~376 vs ~18,052 docs/s as vocabulary V grew). PostingStore now keeps a reverse doc_id → term_ids index, populated once per (doc, term) edge as terms are added, so remove_doc visits only the terms the document actually contributed — O(terms-in-doc), independent of V. Search output is byte-identical (the rank-aligned posting contract is unchanged): matched doc sets, BM25 scores, doc_freq, num_docs, and avgdl are unaffected. A missing or already-cleared reverse entry is skipped defensively — never an unwrap/panic.

Fixed — FT.SEARCH OR and multi-clause queries return correct result sets (PR #190)

FT.SEARCH query combinators were silently broken: OR (alpha | beta) collapsed to an intersection (returning only docs matching both terms), and a multi-clause query such as @body:foo @tag:{bar} returned zero results instead of the intersection. A code trace showed the root cause was a missing parser layer — four runtime handlers each re-parsed the query inline with a flat, AND-only shape that could not represent an OR/grouped AST, and the cross-shard scatter carried that same flat shape so OR was unfixable at the leaf. This lands a recursive-descent query parser producing a frozen QueryNode AST and one centralized evaluator (eval_set folds the AST to a RoaringBitmap — AND = intersection, OR = union — over the shared doc-id space; eval_query adds best-effort BM25 with deterministic score-desc / doc-id-asc ordering). All four FT.SEARCH text branches plus the cross-shard Phase-2 scatter now route raw query bytes through parse_query → eval_query → build_text_response; the cross-shard payload carries the opaque query bytes and each shard re-parses, so OR/grouping/TEXT+TAG/TEXT+NUMERIC work in every shard configuration. Malformed queries return a coded Frame::Error (one of five frozen codes) and never panic the server. Verified end-to-end over the wire on 1- and 4-shard servers (alpha | beta → union, @body:foo @tag:{bar} → intersection, 1-shard vs 4-shard result sets identical). HYBRID/vector-KNN/ SPARSE/SESSION/RANGE dispatch is untouched.

Performance — high-DF FT.SEARCH term queries no longer O(M²) (PR #190)

The per-document term-frequency lookup on the BM25 path was an O(N) .position(|id| id == doc_id) linear scan over a posting list, so a query on a high-document-frequency term degraded to O(M²) (a ~5%-of-corpus term took ~419 ms on a 100K-doc index). PostingList now keeps term_freqs/positions in sorted-doc-id (rank) order aligned with the doc-id RoaringBitmap, and tf()/positions_for() resolve via RoaringBitmap::rank (sub-linear). The rank alignment also fixed a latent BM25 corruption where re-indexing an updated low-id document pushed its term frequency out of position, silently misaligning scores.

Performance — monoio FT.SEARCH yield is now cost-free; brute-force knee raised to 1024 (PR #189)

PR #179's monoio cooperative yield reaped the io_uring completion queue by parking on sleep(ZERO) — correct, but ~1746 µs per yield, which is what forced

179 to keep the brute-force chunk small and defer ~22% of transient throughput.

The yield now reaps the CQ by parking on a pre-armed UnixStream::pair read (~0.317 µs — effectively free), so the brute-force chunk knee walks back up from 256 to 1024 elements per yield without reintroducing the co-located latency

179 fixed. The knee was A/B confirmed per-architecture on GCloud: K=512 breached

the 5% latency budget on x86 (+6–8%) but held on aarch64, so 1024 was chosen as the value safe on both. #179's co-located p99 relief is fully preserved — this reclaims only the throughput #179 had to trade away. tokio is unaffected (it already used yield_now()). Companion benchmark/review docs land alongside: BENCHMARK.md §2.8 (GCloud cross-arch re-measurement) and a 4-feature deep review + concurrent-vs-competitor benchmark under docs/reviews/2026-06-16/.

Performance — FT.SEARCH no longer stalls the shard event loop (PR #179)

A heavy FT.SEARCH (large brute-force mutable segment or deep HNSW traversal) used to run fully synchronously on its shard's event loop, blocking the 1ms tick and every co-located command on that shard until the search returned. The per-shard local search slice now captures an owned, point-in-time snapshot of the index at entry and walks the same logical steps cooperatively, yielding to the event loop between bounded chunks so the tick fires and co-located PING/GET are serviced while the search is in flight. Results are byte-identical to the synchronous path (the mutable segment is append-only, so a chunked scan over the captured length matches an atomic scan), MVCC snapshot isolation is preserved, and there is no new cross-thread lock and no steady-state RSS growth. The cooperative yield is runtime-specific: tokio uses yield_now(); monoio uses a zero-duration timer park, because a bare self-wake never lets monoio's io_uring loop reap the completion queue (it would be a silent no-op). Measured co-located PING p99 during a heavy search drops from ~48 ms (1 client) / ~300 ms (3 clients) to 6.6 ms / 27 ms — roughly 7–11× relief — on both runtimes. The yield costs throughput only on a large uncompacted brute-force scan (operator-tunable via MOON_FT_YIELD_CHUNK); light and HNSW searches whose work is under one chunk never yield and pay nothing. The cross-shard scatter-gather, merge/rerank, and result semantics are unchanged.

Performance — WAL group commit under appendfsync=always (PR #178)

Concurrent writes pending at the same shard now coalesce into a single fsync instead of one sync per write, collapsing a large share of the ~11× appendfsync=always throughput penalty that appears when multiple clients write in parallel. The batching is opportunistic — a writer drains every AppendSync already queued on its shard channel and issues one barrier fsync for the whole group — and is wired into all four AOF writer loops ({TopLevel, PerShard} × {monoio sync, tokio async}). Durability is unchanged: control records route to a separate deferred queue so a batch can never straddle a non-data message, and the exactly-once-under-crash invariant (crash-matrix + SIGKILL integration tests) holds on both runtimes. The absolute throughput gain is disk-dependent — structurally unmeasurable on the near-free virtio fsync of the dev VM, so the coalescing mechanism (K AppendSync → 1 fsync) is pinned by a deterministic batching seam test and the wall-clock magnitude is deferred to real-disk hardware.

Removed (BREAKING) — --cross-shard-fast-path flag and its dead telemetry

The orphaned --cross-shard-fast-path CLI flag (with its CrossShardFastPath config enum and cross_shard_fast_path_enabled()) is deleted, along with the dead moon_cross_shard_lock_contention_total metric, the moon_dispatch_cross_read_fastpath_latency_us histogram, the record_dispatch_cross_read_fastpath_* recorders, and the hardcoded-0 cross_read_fast_dispatches stat. These were Phase-0 scaffolding for the pre-shared-nothing RwLock read path (gone since PR #175) and had zero production callers. Breaking: moon --cross-shard-fast-path … now exits non-zero with a clap unknown-argument error — remove it from any launch script. The internal scripts/bench-cross-shard-fastpath.sh is also deleted.

Performance — Lock-free cross-shard read recovery (idle-gated reply spin) (PR #177)

Cross-shard single-key reads recover a meaningful share of the latency the shared-nothing migration (PR #175) gave up, without re-introducing a cross-thread lock or growing per-key memory. When a shard is near-idle the requesting connection briefly busy-polls its cross-shard reply instead of parking — skipping one of the round-trip's two cross-thread wakes. The poll is gated on a thread-local in-flight counter (Cell<u32>, no atomic/lock), so a busy shard never spins and high-concurrency throughput is unaffected. Same-run quiesced-VM measurement (monoio, 4 shards): single-client cross-shard GET +18.0% (and cross-shard SET +21%, sharing the reply path); c100 GET +8.4% with no starvation. Consistency 197/197 across 1/4/12 shards; dual-runtime (monoio + tokio). Cross-connection read coalescing — the second half of the original design — is deferred to a dedicated follow-up.

Changed — Shared-nothing thread-local shard storage (PR #175)

Shard storage moved from RwLock/Mutex-wrapped shared databases to thread-local ShardSlices owned exclusively by each shard's event loop, on both runtimes (monoio and tokio). The lock wrappers are deleted; all cross-shard access routes through the owner shard's SPSC ring. Routed multi-shard workloads are parity or better (up to +12% on pipelined GET at 4 shards); the cross-shard read fast-path (direct RwLock read of another shard's data) is definitionally gone and those cells regress — accepted with a follow-up task for lock-free cross-shard read acceleration. BGREWRITEAOF on every layout now uses the C4 cooperative-fold protocol: the AOF writer requests an atomic snapshot from the owning shard and drains pre-snapshot appends with an exact bound, so concurrent writes are neither lost nor double-applied. This also fixes BGREWRITEAOF being silently inoperative on the top-level AOF layout (--shards 1), where the rewrite previously errored internally while the client was told it had started.

Added — HYBRID FT.SEARCH FILTER push-down on both branches (PR #174)

FT.SEARCH ... HYBRID now accepts an optional FILTER clause that is applied to both the BM25 and the dense-KNN streams before RRF fusion, closing a bypass where a filtered hybrid query leaked foreign-field hits through the unfiltered dense branch. The filter is a recursive, arity-counted grammar — TAG @field value (exact, or prefix when value ends in *), NUMERIC @field min max (inclusive range), and AND/OR combinators — bounded at depth 4 / 16 leaves and parsed without panicking on malformed input. Filtering is per-shard on the scatter/gather path, so multi-shard results stay correct. Absent a FILTER clause the wire format and behavior are byte-identical to before (fully backward compatible). Exposed through the Rust SDK as TextClient::hybrid_search(..., filter: Option<&HybridFilter>).

Fixed — Cross-shard commands now route to the owning shard (PR #173)

On multi-shard servers, SO_REUSEPORT spreads client connections across shard accept loops — and several command families operated on the connection's shard instead of the shard that owns the data. Whether these commands worked depended on which shard the kernel happened to accept the connection on. All four groups are fixed; the consistency suite now passes 197/197 at 1, 4, and 12 shards:

  • BITOP routed by its sub-operation literal (AND/OR/…): sources were read and the destination written on an arbitrary shard. A new coordinator gathers sources per owning shard and writes the result on the destination's owner (Redis semantics: NOT arity, zero-padding, all-missing sources delete the destination and return 0).
  • COPY wrote the destination into the source's shard, making it unreadable at its own. Cross-shard string copies preserve value + TTL and honor REPLACE/NX; cross-shard non-string COPY returns an explicit error instead of corrupting state.
  • GRAPH.*, TEMPORAL.INVALIDATE, and TXN.ABORT graph rollback ran on the connection's per-shard graph store — graphs randomly "did not exist" from ~3/4 of connections at 4 shards. Graph commands now hop to the graph-name-owning shard; transactional rollback partitions undo work by owner and awaits acknowledgements.
  • WS.* and MQ.*: the workspace registry was per-shard (WS AUTH from another connection answered "workspace not found"; WS LIST diverged per connection) and durable queues lived on the creating connection's shard (MQ PUSH/POP/ACK/DLQLEN from elsewhere answered "queue is not durable"). The workspace registry is now global with a single WAL stream, and every MQ operation targets the queue key's owning shard — including dead-letter routing, trigger registration, and MQ.PUBLISH materialization at TXN.COMMIT.

Changed — Event-driven cross-shard wake (the ~1ms monoio floor is gone)

  • The monoio shard event loop is now event-driven. Its single await point was the 1ms periodic tick, so every cross-shard hop queued up to 1ms on the target shard plus up to 1ms on the origin's reply sweep. The loop now races the tick against the shard's SPSC Notify (hand-rolled allocation-free race2 future — monoio::select! remains banned), and connection tasks await cross-shard replies directly. Cadence work (WAL flush, cached clock, snapshot/auto-save sub-timers) stays pinned to the timer arm. Measured on a 4-shard Linux server: single-client cross-shard SET p99 4.07 ms → 0.071 ms (57×), pipelined SET +368%, non-pipelined SET +404%; single-shard throughput +63–80%. (PR #172)
  • Drain-cap self-re-notify: a >256-message cross-shard burst no longer strands its tail until the next tick — a capped drain re-arms the wake immediately (both runtimes).
  • New INFO Stats counters: spsc_notify_wakes and spsc_drain_renotify make the event-driven path observable from a black-box client.

Changed — Hot-path lock quick wins

  • Per-command global lock acquisitions removed from the command dispatch path (replication offset reads, shared-databases lookups, socket-option setup, accept-path state), with the replication backlog append switched to a bulk-copy. (PR #172)

Fixed

  • uring_handler in-flight send queue used Vec API on a VecDeque (.push → .push_back) — the Linux+tokio io_uring bridge path failed to compile on its own. (PR #172)

[0.3.0] — 2026-06-11

Hot-shard elasticity + vector engine maturity: background HNSW compaction (no more shard freezes), working SQ8 quantization, elastic per-shard memory budgets, HOTKEYS detection, and a +33%/+25% pipelined SET/GET perf pass.

Added — Docs: Design Advantages page with architecture diagrams

  • New docs page (design-advantages.md, MkDocs nav → Concepts) telling the full architecture story with 6 whiteboard diagrams: platform hero, memory engine, persistence, vector engine, graph engine, converged GraphRAG workflow, plus an "AI agent era" section mapping agent memory types to Moon surfaces. Originally authored for the retired Mintlify site (PR #145), converted to MkDocs Material and refreshed: quantization guidance now says SQ8 (TQ8 never existed), and the segment-lifecycle and disk offload sections reflect background compaction (PR #165) and elastic memory budgets (PR #170). README "Why Moon" embeds the hero diagram.

Changed — Background FT compaction + segment merge (no more shard freezes)

  • HNSW index builds no longer block the shard event loop. try_compact (auto-compaction on search) now begins the build on a background worker pool and installs the finished immutable segment on a later poll; the triggering search continues against the still-present brute-force mutable segment. Inline builds previously froze the shard for the full build (0.42s @2k → 24.9s @50k vectors, measured); the event-loop stall is now ~4.4ms and PING latency stays flat during compaction.
  • Background immutable-segment merge: many small immutable segments are merged into one on the same worker pool (graph-union, no lossy decode/re-encode), bounding per-search segment fan-out. Real-MiniLM recall preserved across the merge (0.9447 → 0.9440 R@10, 18 segments → 1).
  • Deletes that race a background build are reconciled at install time (delete-window replay) — no resurrection of removed vectors.
  • FT.COMPACT (explicit user intent) stays synchronous: it drains any in-flight background build, then compacts inline. Tradeoff: searches run brute-force on the mutable segment until the background build installs.

Added — Elastic per-shard memory budgets (hot-shard headroom borrowing)

  • Hot shards now borrow idle siblings' unused maxmemory headroom instead of being capped at a static maxmemory / num_shards. Each shard publishes its used memory on its existing 100ms eviction tick and recomputes an elastic budget from the published snapshot: base + Σ(under-budget surplus) / hot_shard_count, clamped to total maxmemory. No new locks or channels — two relaxed atomics per shard, read once per write-path eviction check.
  • Snapshot invariant: hot shards' budgets plus under-budget shards' usage never exceed N × base at recompute time; transient overshoot is bounded by one tick of donor write growth and converges via write-path eviction.
  • Live 4-shard validation (32MB cap, allkeys-lru, all writes to one shard via a {tag}): hot shard retained ~28k × 1KB keys (~30MB) vs the ~7.2k (8MB) static ceiling — 3.9× more of the configured memory actually usable under skewed load. Uniform load is unchanged (same global ceiling, budgets stay at base).
  • Disk-offload pressure cascade and timer-based eviction enforce the same elastic budget; budget 0 (not yet published, single shard, or maxmemory=0) falls back to the static per-shard cap everywhere.

Added — Hot-key detection: HOTKEYS command + per-shard SpaceSaving sketch

  • New HOTKEYS [COUNT n] command (Moon extension, @server ACL, read-only): returns the top sampled keys as [key, sampled_count] pairs, count descending. Counts are 1-in-64 samples of keyed commands — multiply by 64 for an approximate command rate. In multi-shard mode the connection handler merges per-shard sketches (DBSIZE-style scatter-gather), so clients always see the global view.
  • Per-shard SpaceSaving top-K sketch (K=128, ~5 KB per database, L1-resident, single-pass scan): fed by all three execution paths — write dispatch, the shared read path, and the inline GET/SET fast path. The tick counter is one relaxed fetch_add per command; the O(K) sketch update runs only on sampled ticks and try_locks (a contended sample is dropped — the hot path never blocks). Kill switch: MOON_NO_HOTKEYS=1.

Fixed — OBJECT was unreachable over the wire (monoio handler)

  • OBJECT ENCODING/FREQ/IDLETIME/REFCOUNT returned ERR unknown command on a live server even though dispatch implemented them: the monoio handler routes every non-write command through the read dispatcher, which had no OBJECT arm. Added a read-only OBJECT handler (get_if_alive, never mutates expiry state) and enabled it on the cross-shard fast path.
  • OBJECT <subcmd> <key> routed by the subcommand name in multi-shard mode (extract_primary_key hashes args[0], which is FREQ/ENCODING, not the key) — the command landed on an arbitrary shard and answered ERR no such key. Key extraction now uses the real key (arg 2).

Changed — Hot-path performance: deep-review optimization pass (PR #168)

  • +33% pipelined SET, +25% pipelined GET (single shard) from a deep performance review covering protocol → dispatch → shard → storage → persistence → vector. 11 verified findings implemented:
  • Database::get: 3 → 2 DashTable probes on a hit, 2 → 1 on a miss (single-probe live/expired/absent classification).
  • WAL v2 flush: 3 → 1 write(2) syscalls per 1ms tick via a reusable staging buffer; on-disk format unchanged.
  • WAL v3 record append no longer heap-allocates per non-FPI record (Cow::Borrowed); checkpoint state machine no longer clones per tick.
  • SPSC drain (drain_spsc_shared) reuses thread-local batch buffers instead of allocating two Vecs per call.
  • Frame::Double serialization is zero-alloc (stack f64 Display formatter, byte-identical wire output) in RESP2 and RESP3.
  • SQ8 vector search no longer heap-allocates per query (HNSW + both brute-force paths); IVF rotation/LUT buffers hoisted out of the per-segment loop in search_filtered (parity with search_mvcc).
  • io_uring drain_completions reuses persistent CQE/event buffers (allocation-free drain loop at steady state).
  • Multi-shard (2/4/8/16/auto): no regressions; +9–32% pipelined SET at 2 and 8 shards. The unpipelined cross-shard latency floor (~1.7ms p50, waker-relay + dispatch tick) is unchanged and tracked for follow-up.

Fixed — SQ8 vector quantization now works across the full FT lifecycle

  • FT.CREATE ... QUANTIZATION SQ8 was a declared-but-unimplemented quantizer — vectors fell through to the TQ4 encoder against an empty codebook, producing degenerate codes and random search (recall 0.014, exact-match returned the wrong key) on real embeddings. Implemented a real per-vector affine scalar-8-bit quantizer (dim u8 codes + (min, scale) trailer) and wired it through the entire segment lifecycle: append, brute-force search, compact, immutable HNSW search, multi-segment search fan-out, segment merge, and persistence reload. Recall on real all-MiniLM-L6-v2 384d embeddings: 0.014 → 0.897, exact-match correct.
  • append_transactional() corrupted SQ8 vectors — it wrote the TQ slot layout (padded/2 + 4) into dim + 8 SQ8 slots, corrupting the code stride for every transactionally-inserted or WAL-recovered SQ8 vector. Both append paths now share one encode_sq8_slot() helper.
  • SQ8 ignored the index metric — an InnerProduct index ranked non-normalized inputs by magnitude. SQ8 now normalizes for both unit-sphere metrics (Cosine + InnerProduct), matching the rest of the engine; L2 keeps raw vectors for true Euclidean ranking.
  • Docs: corrected the ≤384d quantization guidance to recommend SQ8 (the real 8-bit option) instead of the nonexistent "TQ8".

Changed — Rust SDK released as moondb 0.2.0 on crates.io

  • moondb crate 0.1.1 → 0.2.0 — publishes the v0.2-era client API that has lived in-tree since the hybrid_search sparse upgrade (f4fcd5a). Breaking: text().hybrid_search() now takes sparse_field: Option<&str> and weights: [f64; 3] (was two-way [f64; 2]) and speaks the PARAMS-based wire format with @-prefixed field refs. Added: Client::connect_with_timeout for bulk-write workloads. Released as 0.2.0 — not 0.1.2 — because cargo treats 0.1.x versions as compatible, and the signature change would break published 0.1.x consumers (lunaris-retrieve / lunaris-storage-moon 0.2.1 pin "0.1.1" with the two-way call). Aligns the SDK version with the Moon v0.2.0 server release.

[0.2.0] — 2026-06-06

The v0.2 enterprise beachhead. Built additively on per-shard WAL v3 + the dual-root manifest; no changes to the KV hot path, MVCC, page format, or transaction layer.

Changed — default persistence directory is the platform user-data dir (was: current directory)

  • --dir now defaults to the platform user-data directory, created on first run: Linux $XDG_DATA_HOME/moon (or ~/.local/share/moon), macOS ~/Library/Application Support/moon, Windows %LOCALAPPDATA%\moon. Installed binaries no longer litter whatever directory they happen to be started from.
  • Back-compat guard: if the startup directory already contains moon persistence data (appendonlydir, shard-0, dump.rdb, replication.state — the pre-v0.2.0 default layout), moon keeps using it and logs a warning, so upgrades never silently boot with an empty keyspace away from their data.
  • Explicit --dir <path> / conf dir (including --dir .) opt out of auto-resolution entirely; environments with no HOME/LOCALAPPDATA fall back to the current directory with a warning. Docker (--dir /data) and the systemd package (dir /var/lib/moon) already pass explicit paths and are unaffected.

Changed — default shard count is now 1 (was: auto-detect)

  • --shards defaults to 1 instead of auto-detecting the CPU count. Single-shard gives the best throughput for non-pipelined workloads (cross-shard SPSC dispatch dominates local lookups) and a deterministic persistence layout across hosts. --shards 0 remains the explicit auto-detect opt-in; nothing changes for deployments that pass --shards explicitly.
  • Upgrade note: deployments that previously relied on the auto-detect default with appendonly yes will refuse to start after upgrading (ERR shard count changed (manifest=N, config=1)) — this is the intended data-loss guard. Start with --shards <N> matching the manifest, or see the new shard-count-change runbook (also linked from the error message itself, now as a full GitHub URL so installed binaries point somewhere reachable).

Fixed — release pipeline verification + prerelease safety

  • release.yml sign job now publishes Fulcio certificates (*.crt) alongside signatures: keyless cosign verify-blob requires the certificate — the previous .sig-only output was unverifiable by users.
  • Docker job no longer moves ghcr.io/pilotspace/moon:latest on prerelease tags (e.g. v0.2.0-rc.1); only stable releases repoint :latest. The versioned tag is always pushed.

Changed — v0.2.0 release prep

  • Crate version bumped 0.1.12 → 0.2.0; INFO server moon_version now reports 0.2.0 (previously released binaries would have self-reported the stale crate version regardless of the git tag).
  • release.yml: the homebrew-tap bump job is gated behind the HOMEBREW_TAP_ENABLED repository variable. Homebrew distribution is deferred — v0.2.0 ships via curl install.sh/install.ps1, .deb/.rpm, and Docker. Without the gate the job would hard-fail on every stable tag (missing HOMEBREW_TAP_TOKEN secret). packaging/bump-homebrew.sh and the formula template stay in-tree for later enablement.

Added — Temporal-decay traversal scoring (agent-memory recency)

User-facing recency bias for graph traversal: paths through recently created edges win over stale ones. The decay engine (scorers, Dijkstra composite cost) existed but was unreachable — shortestPath() hardcoded lambda = 0. This wires it end to end.

  • GRAPH.QUERY ... --decay <λ> [--time-weight <w>] — per-edge cost becomes |weight| + λ·w·age_seconds for shortestPath(); λ is 1/seconds, strictly validated. Decay off (no flag) keeps exact distance-only behavior — the age term contributes zero to every edge cost, so path choice is identical to pre-decay Moon. Applies to the read-only and GRAPH.PROFILE paths via ExecutionContext (same pattern as VALID_AT); write queries (CREATE/SET/DELETE/MERGE) reject the flag instead of silently ignoring it.
  • FT.NAVIGATE ... DECAY <λ> — graph-expanded hits pay λ × age_seconds of their discovery edge on top of the hop penalty (a re-rank of the already-explored expansion, not a steer of the expansion itself); KNN direct hits unaffected.
  • Edges stamp created_ms at insert from the shard-cached clock (zero syscall on the insert path). 0 = unknown is decay-neutral — pre-upgrade edges never look maximally old. Distinct from the user-owned bi-temporal valid_from/valid_to.
  • CSR segment format v3 — per-edge created_ms array (parallel to col_indices) survives freeze → disk → mmap → compaction, so decay sees true edge age after segments rotate. v1/v2 files keep loading (empty stamps = neutral); both parsers (heap + mmap zero-copy) are version-gated, plus a new csr_from_bytes fuzz target.
  • Fixed (latent durability bug): compact_segments stamped merged segments version: 1 while the serializer always writes 48-byte v2+ NodeMeta records — a vacuumed segment written to disk misparsed on reload (panic in debug, silent node_meta corruption in release). Merged segments are now v3 and carry per-edge stamps through dedup.
  • Docs: guides/temporal.mdx decay section + commands.mdx; script coverage in test-commands.sh (6 DECAY cases) and test-consistency.sh (1/4/12-shard path-flip parity).
  • Known gap: WAL-replayed not-yet-frozen edges re-stamp to replay time on restart (newest edges look new — bias direction preserved); CSR-resident edges keep exact age.

Added — Ship moon as an installable application on macOS, Linux, and Windows

Five-PR milestone making moon installable on all three platforms (consolidated entry; the sibling PRs carry the skip-changelog label).

  • Windows native port — cross-platform positioned-I/O helper (src/util/file_ext.rs; pwrite/pread on unix, seek_write/seek_read loops on Windows), portable RawSocketFd alias with #[cfg(unix)]-gated connection-migration fd ops, compile_error! guard for jemalloc on Windows MSVC (mimalloc fallback), and Windows-only startup warnings (VirtualLock working-set, no memory guardrail). Build: --no-default-features --features runtime-tokio,graph,text-index.
  • Graceful shutdown on SIGTERM — ctrlc termination feature with a single centralized handler in main.rs; systemctl stop / launchctl stop now flush AOF before exit (previously only SIGINT was caught).
  • Version strings — INFO server reports redis_version:7.4.0 (client feature-gating) plus a new moon_version: field from CARGO_PKG_VERSION; HELLO/LOLWUT report the real moon version (previously hardcoded 0.1.0).
  • moon.conf config file — redis-style key value config (moon /etc/moon/moon.conf or --config), CLI flags override the file; commented default shipped as packaging/moon.conf.example and installed to /etc/moon/moon.conf (deb/rpm, config|noreplace).
  • Install channels — install.sh (Linux/macOS) and install.ps1 (Windows) one-liner installers with SHA256 verification; Homebrew tap automation (packaging/homebrew/moon.rb.tmpl + bump-homebrew.sh publishing to pilotspace/homebrew-moon on release).
  • Release pipeline overhaul — 8-target matrix (linux gnu/musl x86_64/aarch64, macOS arm64/Intel, Windows x86_64 MSVC), all release binaries now include graph,text-index,console (previously missing from releases), prepare-console pnpm build job, checksums + cosign over all artifacts including .deb/.rpm, --prerelease for -rc tags, and a workflow_dispatch dry-run mode.
  • Packaging fixes — nfpm license corrected to Apache-2.0, moon.service ExecStart path fixed (/usr/bin/moon), CI gains main-push-only check-windows and check-console jobs plus a macOS Intel cross-build check.
  • Console build fix (PR #152) — graph components aligned with the pinned @cosmos.gl/graph 2.6.4 API (setConfigPartial → setConfig, no async ready hook); pnpm run build type-checks again, unblocking the Console Integration workflow and the release prepare-console job.
  • RESP3 negotiation fix (PR #153) — the monoio (default) handler now syncs the wire codec when HELLO 3 negotiates RESP3; previously the codec stayed on RESP2 and silently flattened map/set/push frames, breaking RESP3 clients on connect (redis-py 8 defaults to RESP3).
  • Windows runtime fixes (PR #154) — first fully green check-windows run: fsync helpers reworked for Windows semantics (no directory handles, writable handle for FlushFileBuffers); unsigned-Instant underflows fixed (autovacuum scheduling, rate-limiter cutoffs, mvcc test anchors); the tokio sharded handler no longer initiates connection migration on non-unix platforms (previously aborted clients mid-session in multishard workloads); GraphManifest::save now propagates parent-dir fsync errors.
  • Known limitations (v0.2.0) — Windows binaries are not Authenticode-signed (SmartScreen warning); connection migration is unix-only (connections stay on the originating shard on Windows); Windows service wrapper + MSI deferred.

Fixed — CodeRabbit PR #136 durability follow-ups + decomposition + test isolation (PR #144)

Closes the 8 CodeRabbit findings left open after PR #136, plus two PR #144-review Majors, oversized-file decomposition, and a parallel-test flake. No production hot-path behaviour change.

  • Disk-offload spill (data-loss fixes): block instead of dropping spill completions; salvage inline-batch spill failures per-entry rather than wholesale; preserve spill context across tokio connection migration; recover the cold tier under appendonly=no.
  • Per-shard AOF rewrite robustness: clear the rewrite flag when fan-out fails partway; roll per-shard rewrite writers back to the committed generation on abort (barrier-before-resume + panic-safe ShardDoneGuard); ack drained AppendSync only after the boundary fsync (issue #140 ordering).
  • Async-spill eviction: corrected a stale doc comment, removed the dead remove-first eviction path, and added a regression test locking the fail-safe send-before-remove ordering (a full spill channel keeps the victim resident — no data loss).
  • Platform hygiene: gate the migrated-connection spawn fns behind cfg(all(..., unix)) to match their RawFd usage.
  • File decomposition (1500-line cap): split aof_manifest.rs (3058 → mod/shard_replay/shard_rewrite) and aof.rs (4379 → mod/pool/writer_task/ rewrite); pure code relocation, verified line-exact on both runtimes.
  • Test isolation: fixed the VECTOR_INDEXES counter flake — the process- global metrics counter is now guarded by an RwLock (delta-reader tests take write(), mutator tests take read()), making the index-count delta assertions deterministic under the parallel test harness.

Persistence — Per-shard AOF migration complete (PR #129)

Closes the P0 multi-shard AOF data-loss bug (~50% loss on SIGKILL with --shards >= 2 + --appendonly yes) by shipping the full per-shard AOF architecture (Option B of tmp/rfc-per-shard-aof-v02.md).

  • H2 closed — src/main.rs no longer skips multi-shard AOF replay. Each shard owns its own writer task; recovery walks every shard's segment manifest independently. Shard replay is parallel (recovery time does not grow linearly with shard count).
  • H1 closed — new AppendSync { bytes, ack } rendezvous variant ensures +OK is on the wire only after fsync ack under appendfsync=always. try_send (everysec/no) paths unchanged.
  • --unsafe-multishard-aof deprecated — was the v0.1.13 escape hatch acknowledging the ~50% loss risk. The flag is now a no-op that prints [DEPRECATED] at startup; will be removed in v0.2.0-rc.
  • CRASH-01-LITE matrix — 200/200 SIGKILL recoveries across --shards 1/2/4/8 × appendfsync always/everysec. Gated in .github/workflows/integration-tests.yml; run locally with cargo test --release crash_01_lite on the moon-dev OrbStack VM.
  • TopLevel-manifest safety guard — Moon refuses to start (exit 2) when it finds an existing v0.1.13-style TopLevel manifest with --shards >= 2, to prevent silent data loss from replaying a non-routed log. Migration via docs/runbooks/multi-shard-aof-rewrite.md Option A.
  • INFO persistence — new field aof_backpressure_dropped:<N> exposes per-shard writer drop counts; non-zero indicates the AOF writer is falling behind write throughput.
  • Per-shard layout on disk: appendonlydir/shard-{N}/moon.aof.{seq}.base.rdb
  • moon.aof.{seq}.incr.aof mirrors the per-shard WAL v3 design.

Architectural follow-ups parked for v0.2.0:

  • Rule 3 (LSN ordering invariant) under per-shard topology — issue #131
  • Always-fsync handler-layer integration audit — issue #132
  • OrderedAcrossShards merge-replay correctness on large transcripts — issue #133

Per-shard BGREWRITEAOF (step 6 of the migration RFC) is not yet in this PR; rewriting a per-shard AOF returns ERR BGREWRITEAOF is not yet supported under per-shard AOF layout. Tracked for v0.2.0.

Headline capabilities landed in alpha:

  • Point-in-Time Recovery (PITR) — --recovery-target-lsn / --recovery-target-time restore to any LSN or wall-clock boundary inside the WAL retention window.
  • Change Data Capture (CDC) — CDC.READ polling command with Debezium-compatible JSON envelopes, resumable cursors, segment-rotation safety.
  • Hash-field TTL — full Valkey 9.0 / 9.1 surface (HEXPIRE / HPEXPIRE / HEXPIREAT / HPEXPIREAT / HEXPIRETIME / HPEXPIRETIME / HTTL / HPTTL / HPERSIST / HGETDEL / HGETEX) with O(1) HGET + HLEN fast path. Three-way benchmark vs Redis 8.0.2 / Valkey 9.1.0 ships in docs/perf/2026-05-27-hash-ttl-3way-bench.md.
  • Tier 2 Lane A — SWAPDB, MOVE, COPY ... DB n, CLUSTER REPLICAS / SLAVES, CLUSTER COUNT-FAILURE-REPORTS. All WAL-durable with cross-shard atomic semantics.
  • Storage format v1 commitment — RDB v2 + WAL v3 + multi-part AOF manifest grouped under a single --storage-format v1 umbrella with ≥18-month LTS forward-read guarantees.
  • Embedded sharded server — server::embedded::run_embedded(config, cancel) exposes the full sharded handler (with TXN.*) to in-process embedders.

What is not yet in alpha: PITR live-snapshot LSN wiring (P3c), CDC.SUBSCRIBE push channel (C3b), and the multi-shard master PSYNC deferred from v0.1.10. Tracked in .planning/rfcs/v02-enterprise-architecture.md.

Fixed — Multishard idle-RAM blowup + tokio-Linux serving hang + graph parser DoS (PR #136)

Closes the "multishard RAM zombie": a fresh multishard instance with no maxmemory could commit multiple GB of RSS while idle, and the tokio runtime could hang its accept loop on Linux. Three independent root causes, all fixed and verified in the OrbStack VM under a cgroup memory cap (idle RSS 3791 MB → 29 MB for a 4-shard no-maxmemory instance; ~45 MB under load):

  • PageCache eager pre-allocation. Each shard committed num_frames × PAGE zeroed bytes at construction, sized to 25% of the whole-instance maxmemory (or a large default when unset). Now the page buffers allocate lazily and the frame budget is divided by shard count (per_shard_pagecache_budget); a startup WARN reports the resolved per-shard budget.
  • tokio listener bind-race. The central accept listener bound the port without SO_REUSEPORT while per-shard listeners bound with it, producing a bind-order race that could leave the port served by a shard that never received connections. The central listener now also binds SO_REUSEPORT; --per-shard-accept defaults to false under tokio.
  • io_uring-under-tokio default-off. The tokio→io_uring bridge floods errors under load. It is now opt-in via MOON_URING=1; tokio shards run plain epoll/kqueue by default. MOON_NO_URING=1 still force-disables io_uring everywhere. The monoio runtime is unaffected (always io_uring unless MOON_NO_URING).

Two durability defects found during review were fixed in-PR with red/green TDD:

  • Manifest compact-reopen failure no longer silently loses commits. manifest.rs compact() reopened self.file after rename(tmp, path); if that reopen failed, later commits silently wrote to an orphaned inode. A needs_reopen flag now reattaches to self.path before any subsequent commit (or fails loudly).
  • tokio per-shard AOF writer now latches after a torn write. The tokio per-shard writer lacked the write_error latch present on the single-file and monoio writers; a torn write (header OK, data fails) corrupted the frame stream. It now latches on any write failure and acks WriteFailed.

  • Cypher parser stack-overflow DoS bounded. The precedence-climbing expression grammar pushes ~9 stack frames per source nesting level but check_depth() counts only one level per recursion, so the previous limit of 64 overflowed a 2 MiB async-worker stack (SIGABRT) before the guard could fire. DEFAULT_MAX_NESTING_DEPTH is now 32 (~288 worst-case debug frames, ~50% margin); deep nesting returns NestingDepthExceeded gracefully.

Low-severity durability edge cases filed as follow-ups: #137 (apply_spill_completions failed-commit cold-key window), #138 (do_rewrite_per_shard panic wedges --experimental-per-shard-rewrite),

139 (multi-DB SELECT >0 cold recovery restores db0 only).

Fixed — maxmemory is now a whole-instance cap across shards (G2)

Behavior change for multishard deployments. Previously each shard enforced eviction against the full maxmemory, so an N-shard server tolerated ~N× the configured cap before evicting (a 4-shard server at --maxmemory 100mb retained ~307 MB / 768K keys vs a 1-shard server's 124 MB / 322K — the "RAM keeps growing in multishard mode" report).

maxmemory is now a true whole-instance cap. Each shard enforces eviction against maxmemory / num_shards, so aggregate RSS converges on the configured value regardless of shard count.

  • CONFIG GET maxmemory / INFO are unchanged — they report the whole-instance value verbatim (Redis-compatible). Division happens only at enforcement.
  • Operators running explicit --maxmemory N --shards M (M>1) now get an effective ceiling of N (not N×M). A startup log line states the resolved per-shard budget so the change is visible: maxmemory <N> bytes is a whole-instance cap; each of <M> shards enforces eviction against a per-shard budget of <N/M> bytes.
  • Single-shard servers are byte-for-byte unaffected (num_shards == 1 ⇒ no division).

Docs — Hash-field TTL three-way benchmark suite (PR #127)

  • scripts/bench-hash-ttl.sh (2-way harness) + scripts/bench-hash-ttl-3way.sh (3-way harness with Valkey 9.1.0). Both use redis-benchmark against concurrent Moon / Redis / Valkey servers on distinct ports with trap-based cleanup, FLUSHALL + re-seed between scenarios, and median-of-3 RPS reporting.
  • docs/perf/2026-05-26-hash-ttl-bench.md — pre-fix 2-way baseline that surfaced the two HashWithTtl perf issues fixed in PR #126.
  • docs/perf/2026-05-27-hash-ttl-3way-bench.md — headline Moon vs Redis 8.0.2 vs Valkey 9.1.0 comparison across 26 scenarios. Plain HGET p=16 ties both competitors (1.00–1.01×). HEXPIRE-family Moon vs Valkey: 0.90–0.99× across the surface; HGETEX hits parity at 0.99×. Redis 8.x has no HEXPIRE-family — Moon is the only Redis-compatible alternative aside from Valkey.

Performance — HashWithTtl HGET + HLEN O(1) fast path (PR #126)

Resolves the two HashWithTtl perf issues surfaced by the 2026-05-26 bench:

  • HGET on HashWithTtl was 39.9% slower than plain Hash at p=16; now +4% faster (within VM measurement noise — effective parity).
  • HLEN on HashWithTtl was 80.5× slower than plain Hash at 1000 fields; now 1.00–1.03× parity (the O(N) live-count scan is fully eliminated when no field has expired).

Two changes shipped together (variant layout forces both at once):

  • HashWithTtl.ttls BTreeMap → HashMap. Per-field TTL probe becomes O(1) HashMap lookup instead of O(log N) BTreeMap descent. Active-expire iteration doesn't require ordered keys.
  • HashWithTtl.min_expiry_ms: u64 cached minimum. Tracks the smallest expiry across all per-field TTLs. Invariant: min_expiry_ms = min(ttls.values()). When cached_now_ms < min_expiry_ms, no field can be expired, so HGET skips the ttls probe and HLEN returns fields.len() directly. The hot path is a single u64 < u64 compare. Invariant maintenance is amortized O(1) (one min(min, ts_ms) per HEXPIRE; conditional recompute on HPERSIST / overwrite / active reap when the removed TTL equalled the min).

No on-disk format change. min_expiry_ms is recomputed at load time from the existing v2 RDB per-field TTL trailer. 6 new invariant tests cover HEXPIRE / HPERSIST / HSET-overwrite / active-reap / persistence-decode paths.

Docs — Hash-field TTL audit follow-up (PR #123)

  • docs/commands.mdx — Hashes section bumped from "(14)" → "(25)" and now lists all 11 new Valkey 9.0/9.1 commands plus a paragraph on the per-field return convention and the active-expiry downgrade behaviour.
  • docs/STORAGE-FORMAT-V1.md §3.2 — added the per-field hash TTL trailer format: every TYPE_HASH body is followed by [ttl_count u32][field_len varint | field_bytes | ttl_ms u64]* in v2 RDB files; v1 readers stop after the hash body.
  • src/command/metadata.rs — HEXPIRETIME / HPEXPIRETIME / HTTL / HPTTL PHF flags flipped from R (READONLY) to RF (READONLY|FAST) — per-field TTL lookup is O(1) thanks to the PR #126 fast path. Cosmetic; no behavioural impact.

Added — HGETDEL / HGETEX atomic compound hash commands (phase 199, issue #110)

Two new Valkey 9.1 atomic compound hash commands:

  • HGETDEL key FIELDS numfields field [field ...] — returns the values of the specified fields and deletes them from the hash atomically. Returns a RESP Array with one BulkString(value) per found field and Null for missing fields. If the hash becomes empty after all deletes the key is removed entirely (auto-cleanup).

  • HGETEX key [EX s | PX ms | EXAT unix-s | PXAT unix-ms | PERSIST] FIELDS numfields field [field ...] — returns the values of the specified fields and optionally updates (or removes) their per-field TTLs atomically. TTL modes:

  • EX s — set relative expiry in seconds from now.
  • PX ms — set relative expiry in milliseconds from now.
  • EXAT unix-s — set absolute expiry as unix seconds.
  • PXAT unix-ms — set absolute expiry as unix milliseconds.
  • PERSIST — remove any existing per-field TTL.
  • (no mode) — pure read; no TTL change (fast path, zero DB mutation). TTL changes apply only to live (non-expired) fields. Missing / expired fields return Null and leave TTLs untouched.

Atomicity: per-shard single-threaded execution gives atomicity for free across the entire field list — no client can observe a partial state between reads and deletes/TTL-updates within a single call.

Implementation: - Two new Database primitives in src/storage/db.rs: - hash_get_and_delete_field — atomically reads and removes a single field; handles Hash, HashListpack, and HashWithTtl (also removes TTL sidecar). Downgrades HashWithTtl → Hash when the last TTL is removed. - cleanup_empty_hash — removes the key when its hash has become empty; called once after a HGETDEL/HGETEX field loop. - HGETDEL handler uses parse_key_and_fields (phase-198 shared parser) plus a single cleanup_empty_hash call after the field loop. - HGETEX parser (parse_hgetex_args) scans the optional mode token(s) before FIELDS with mutual-exclusion enforcement; uses saturating i128 arithmetic + u64 clamp for safe overflow handling (mirrors phase 196). - Both commands dispatch as writes (WF flags, &mut Database); neither is added to is_dispatch_read_supported. - 17 new unit tests: 8 for HGETDEL, 9 for HGETEX.

Added — Hash-field TTL read + persist (phase 198, issue #109)

Five new Valkey 9.0 hash-field TTL commands:

  • HEXPIRETIME key FIELDS numfields field [field ...] — absolute expiry per field as a unix timestamp in seconds.
  • HPEXPIRETIME key FIELDS numfields field [field ...] — absolute expiry per field as a unix timestamp in milliseconds.
  • HTTL key FIELDS numfields field [field ...] — remaining TTL per field in seconds; already-expired-but-not-reaped fields return 0.
  • HPTTL key FIELDS numfields field [field ...] — remaining TTL per field in milliseconds; same 0 edge-case for expired-but-not-reaped.
  • HPERSIST key FIELDS numfields field [field ...] — removes the per-field TTL; downgrades HashWithTtl back to plain Hash when the last TTL is removed (handled by the phase-195 hash_persist_field primitive).

Per-field return codes (Valkey 9.0): - -2 — field does not exist (or key is missing — not a WRONGTYPE error) - -1 — field exists but has no TTL - ≥0 — absolute unix time or remaining duration (HEXPIRETIME/HPEXPIRETIME) - 1 — TTL successfully removed (HPERSIST only)

WRONGTYPE is returned immediately (before field iteration) when the key holds a non-hash value.

Implementation: - FieldState tri-state enum and hash_field_state helper (pre-landed in phase 197) provide the zero-allocation field-state read used by all five commands. - parse_key_and_fields shared parser (hash_write.rs) extracts key FIELDS numfields field [field ...]; reuses SmallVec<[&[u8]; 4]> to avoid heap allocation for the common ≤4-field case. - Four read handlers (HEXPIRETIME, HPEXPIRETIME, HTTL, HPTTL) take &Database; HPERSIST takes &mut Database. - All five commands routed in both dispatch() (mutable path) and dispatch_read_inner() / is_dispatch_read_supported() (shared-read path) for the four read commands. - 14 unit tests cover all return-code variants, the already-expired edge case, WRONGTYPE, missing key, numfields=0, and encoding downgrade.

Added — Hash-field active expiration (phase 197, issue #108)

All 9 hash read commands (HGET, HMGET, HGETALL, HEXISTS, HLEN, HKEYS, HVALS, HSCAN, HRANDFIELD) now respect per-field TTLs set by the HEXPIRE family. Expired fields are invisible to callers without requiring the active-expiry tick to have run first (lazy expiry).

Lazy expiry (read path, &Database): HashRef gains a third variant WithTtl { fields, ttls, now_ms } that filters expired fields on every field-level operation. The get_hash_ref_if_alive accessor now returns this variant for HashWithTtl entries instead of falling through to WRONGTYPE. All 9 mutable read commands now call get_hash_ref_if_alive instead of the unfiltered get_hash, so expired fields are never returned.

Active expiry (tick path, &mut Database): the per-shard expire_cycle gains a second sweep via Database::hashes_with_field_expiry() and Database::reap_expired_fields_one_hash(). The reaper removes expired fields from the fields and ttls maps, downgrades the hash back to plain Hash when the last TTL sidecar entry is drained, and signals KeyDeleted when all fields expire (the key is then removed by the caller).

The maybe_has_expiring_keys fast-path flag is now cleared only when both the whole-key sweep and the hash-field sweep return empty, preventing premature flag-clearing that would have silenced future field reaping.

Complexity change: HLEN is now O(N) for HashWithTtl hashes (counts only live fields). Plain Hash and HashListpack remain O(1) / O(N) respectively — no regression.

Added — HEXPIRE-family write commands (phase 196, issue #107)

  • HEXPIRE key seconds [NX|XX|GT|LT] FIELDS numfields field [field ...]
  • HPEXPIRE key milliseconds [NX|XX|GT|LT] FIELDS numfields field [...]
  • HEXPIREAT key unix-seconds [NX|XX|GT|LT] FIELDS numfields field [...]
  • HPEXPIREAT key unix-ms [NX|XX|GT|LT] FIELDS numfields field [...]

Per-field return codes match Valkey 9.0: 0 (no such field), 1 (TTL set or updated), 2 (field expired during this call and was deleted), -2 (NX / XX / GT / LT condition not met). Wrong-type key returns WRONGTYPE; missing key returns one 0 per requested field. PHF + dispatch routes the four commands to hash::{hexpire,hpexpire,hexpireat,hpexpireat}; AOF replay already handled by the phase-200 intercepts.

Side fix — HashWithTtl-aware hash-write path: HSET, HMSET, HSETNX, HDEL, HINCRBY, and HINCRBYFLOAT previously returned WRONGTYPE once HEXPIRE had promoted a hash to the HashWithTtl encoding. Extended Database::get_or_create_hash / get_or_create_hash_listpack to handle HashWithTtl (returning the inner fields map / Ok(None) respectively). Added Database::hash_clear_field_ttls (called by HSET/HMSET on overwrite, no-op on plain Hash) and Database::hash_delete_field (used by HDEL — removes from both the fields map and the per-field TTL sidecar, downgrades back to plain Hash when the last TTL is dropped). HSETNX correctly leaves existing TTLs intact; HINCRBY / HINCRBYFLOAT preserve TTL on the incremented field.

Fixed — Test infra: txn_kv_wiring flake diagnosis

  • test_txn_commit_wal_crash_recovery previously masked moon-server crashes as "Connection refused (os error 61) after 60s" because the spawned moon binary's stdout / stderr were piped to Stdio::null(). Hardened the child-process harness:
  • ChildGuard::spawn now redirects child stdout/stderr to moon-phase-{1,2}.log inside the per-test temp dir.
  • ChildGuard::poll_exit checks try_wait() inside the connection retry loop — if the child exited, the test fails immediately with the exit status instead of timing out for a useless 60 s.
  • ChildGuard::dump_log is called on every failure path (wait_for_server timeout, connect_redis_with_retry timeout, child crash), so the CI log records the actual server output.
  • wait_for_server deadline widened from 5 s → 15 s, matching the realistic CI-runner spawn + WAL boot envelope.
  • No semantic change for the happy path. Future flakes become diagnosable instead of silently retrying for 60 s.

Docs — Storage Format v1 commitment (phase 192, PR #115)

  • docs/STORAGE-FORMAT-V1.md — public on-disk format contract for v0.2.x. Documents the WAL v3 / RDB v2 / AOF multi-part sub-formats as a single "storage format v1" umbrella with explicit forward-read, reverse-read, crash-recovery, and migration guarantees through ≥18 months of LTS. Adds cross-reference doc-comments to src/persistence/aof_manifest.rs, snapshot.rs, and wal_v3/segment.rs pointing readers at the canonical on-disk markers. Reserves a --storage-format <v1> CLI flag for the follow-up code PR closing issue #103's second checkbox.

Added — Hash-Field TTL primitive (phase 195, PR #116)

  • RedisValue::HashWithTtl { fields, ttls } storage variant + borrowed RedisValueRef::HashWithTtl view. Per-field TTL sidecar (BTreeMap<Bytes, u64>) carries absolute unix-ms expiry alongside an unchanged HashMap<Bytes, Bytes> field map. OBJECT ENCODING reports hashtable.
  • Database::hash_set_field_ttl(key, field, abs_ms, cond) — Valkey 9.0 parity result codes (0 missing / 1 set / 2 deleted-on-set / -2 cond not met). NX / XX / GT / LT semantics matching Valkey (non-volatile is +∞ for GT/LT). Auto-promotes Hash / HashListpack → HashWithTtl on first per-field TTL; past-expiry short-circuits to in-place delete with code 2.
  • Database::hash_persist_field(key, field) — clears one field's TTL; downgrades back to plain Hash when the last TTL is removed.
  • Database::hash_get_field_ttl_ms + Database::hash_field_state — read helpers returning the tri-state FieldState::{Missing, NoTtl, Ttl(u64)} consumed by phase-198 HTTL / HEXPIRETIME / HPTTL / HPEXPIRETIME read commands.
  • Additive HashWithTtl match arms in command::key::should_async_drop, server_admin::estimate_serialized_length, eviction::evict_one_async_spill, tiered::kv_spill, tiered::kv_serde, persistence::rdb::write_entry, persistence::redis_rdb, and persistence::aof::generate_rewrite_commands. Persistence-side TTL payload lands in PR #117 (phase 200); arms here are TTL-stripping placeholders so HEXPIRE handlers (phase 196) cannot be merged ahead of the persistence wiring.

Added — Hash-Field TTL persistence wiring (phase 200, PR #117)

  • RDB v2 — bumped RDB_VERSION 1 → 2. New per-hash trailer [ttl_count u32][field, ttl_ms u64]* follows every TYPE_HASH body (count=0 for non-TTL hashes). Reader accepts both v1 and v2 files; count_entries_per_db / skip_entry / read_entry_zero_copy / read_entry all plumb a has_hash_ttl_trailer flag so the format is parsed correctly on both code paths (file load + in-memory bytes).
  • Shard RDB V3 — SHARD_RDB_VERSION 2 → 3 with the same trailer plumbed through rdb::read_entry. V1 (legacy) / V2 (PITR LSN+ts) / V3 (TTL trailer) all load via per-version preamble + min-file-size branching.
  • Tiered KV serde — per-hash trailer on disk-offload blobs. Plain Hash + HashListpack serializers now append ttl_count = 0; HashWithTtl serializer writes the real trailer. Deserializer treats a truncated trailer as zero TTLs for graceful migration from pre-trailer in-process spill blobs.
  • BGREWRITEAOF — RedisValueRef::HashWithTtl arm emits HSET key f1 v1 f2 v2 ... followed by per-field HPEXPIREAT key abs_ms FIELDS 1 field, one per TTL'd field. Per- field framing keeps the replay shim simple (single-field parse).
  • Replay shim — CommandReplayEngine::replay_command intercepts HEXPIRE / HPEXPIRE / HEXPIREAT / HPEXPIREAT / HPERSIST (case-insensitive) before command::dispatch and routes directly to Database::hash_set_field_ttl / hash_persist_field. This bypasses the phase-196 command handlers (which do not yet exist) so crash-restart restores per-field TTLs from any AOF stream emitted by either user-typed HEXPIRE or BGREWRITEAOF.
  • Redis-compat RDB — emits tracing::warn! when dropping per-field TTLs on Redis-compat export. Redis 7.4 hash-field-TTL opcode emission deferred to a future cross-vendor compat phase.

Added — Tier 2 Lane A (PR #100)

  • T2.1 c381b31 — SWAPDB cross-shard atomic swap via ShardMessage::SwapDb; WAL-durable; BGREWRITEAOF concurrency guard; restart-replay test.
  • T2.2 4958dc9 — MOVE key db with with_two_dbs_locked (lower-index-first lock ordering); WAL-durable; intercept in all four handler paths.
  • T2.3 bbc6117 — COPY ... DB n cross-database; reuses with_two_dbs_locked; WAL-durable.
  • T2.4 f538589 — CLUSTER REPLICAS / CLUSTER SLAVES; shared format_node_line(node, self_node_id) helper extracted from CLUSTER NODES.
  • T2.5 ebd240a — CLUSTER COUNT-FAILURE-REPORTS; counts non-stale pfail_reports; exposes DEFAULT_NODE_TIMEOUT_MS as pub(crate).

Fixed

  • PERF 608e2d1 — collapse duplicate is_write PHF gate on MOVE/COPY hot path; restores s=1 SET p=1 throughput (−9.5 % → +0.9 % vs. merge base).
  • CR (this PR) — SWAPDB now runs after the ACL gate in handler_monoio, closing a runtime-specific authorization bypass.
  • CR (this PR) — with_two_dbs_locked and ShardDatabases::swap_dbs now hard-assert non-equal indices in release builds, preventing same-index self-deadlock.
  • CR (this PR) — SWAPDB strict arity (exactly two args) across all three handlers; rejects SWAPDB 0 1 extra with the canonical wrong-arity error.
  • CR (this PR) — DashTable::Segment::insert_or_update_at now sets has_non_home_keys = true whenever the chosen free slot is in a non-home group; fixes a latent miss where find() could not locate a fallback-placed key on subsequent lookups.
  • CR (this PR) — Local MOVE / COPY AOF append is gated on Frame::Integer(1) (success) rather than !Error, matching the handler_single behavior and suppressing no-op :0 log entries.
  • CR (this PR) — MOVE, COPY ... DB n, and SWAPDB are rejected with ERR_TXN_CROSS_SHARD while an active_cross_txn is in flight; previously the intercepts bypassed undo/intents bookkeeping and escaped TXN.ABORT rollback.
  • CR (this PR) — spsc_handler MOVE/COPY arms call refresh_now_from_cache on both source and destination DBs before move_core/copy_core; fixes expired-key visibility skew on the local-write path.

Refactor

  • e429b2b — src/storage/dashtable/segment.rs (1587 LOC) split into segment/{mod,find,insert,ops}.rs; mechanical refactor, zero semantic change, brings all files under the 1500-LOC limit ahead of future hot-path additions.

Added — Point-in-Time Recovery (PITR)

  • P0 ac3aa92 — WalWriterV3::new() now scans existing .wal segments on open and resumes next_lsn from max_observed_lsn + 1. Fixes a latent durability bug where LSN reset to 1 on every restart and blocked both PITR and CDC.
  • P1 e1e9bda — FileEntry extended with last_modified_lsn (offset 48..56, struct size 48 → 56). Manifest format_version bumped to v2; backward-compat reader synthesizes last_modified_lsn = created_lsn for v1 entries. New CLI flags --recovery-target-lsn and --recovery-target-time.
  • P2 25ece4b — Snapshot header bumped to SHARD_RDB_VERSION = 2, embedding last_lsn and created_at_unix_ms. v1 snapshots load with last_lsn = 0 and are conservatively skipped by PITR. Adds read_snapshot_metadata peek API.
  • P3a 048a883 — replay_wal_v3_dir_until(stop_at_lsn) and resolve_target_time_to_lsn() (scans TemporalUpsert / GraphTemporal records for system_from anchors).
  • P3b a496413 — recover_shard_v3_pitr() honors target_lsn: skips snapshots whose last_lsn is unknown or past the target, then stops replay at the cutoff without advancing wal_flush_lsn.

See docs/guides/pitr.md for operator usage.

Added — Change Data Capture (CDC)

  • C1 b271e21 — WalTailReader with resumable TailCursor { segment_seq, byte_offset, last_lsn }. Re-stats segment metadata on each call for torn-write safety, auto-advances on segment rotation.
  • C2 e97b80d — New src/cdc/ module with typed CdcEvent enum, decode_wal_record() translator, and hand-rolled Debezium-compatible JSON envelope serializer (encode_debezium). Non-UTF-8 keys fall back to {"_b64":"..."}.
  • C3 v1 bf4230b — CDC.READ <wal_dir> <from_lsn> [LIMIT N] polling command. Returns RESP array [next_lsn, env1, env2, ...]; default LIMIT 256, hard ceiling 10 000. Idle response is [from_lsn] (length 1) — stable no-new-data signal.

See docs/guides/cdc.md for consumer integration.

Added — Embedded sharded server

  • PR #95 — server::embedded::run_embedded(config, cancel) exposes the full sharded handler (with TXN.* cross-store transactions) to in-process embedders such as helios moon-daemon. The existing run_with_shutdown drives handler_single, which deliberately does not implement TXN. Embedded mode skips TLS, console, cluster bus, admin port, and multi-part AOF manifest replay; it does include per-shard RDB + WAL recovery, graph/temporal/workspace/MQ WAL replay, SO_REUSEPORT, NUMA pinning, and cancel-driven graceful shutdown.

Fixed — PR #96 test deflake + tokio AOF replay

  • main.rs AOF recovery: gated the multi-part AOF manifest replay block to #[cfg(feature = "runtime-monoio")]. Under tokio, the legacy single-file appendonly.aof is loaded via the v2 recovery chain; the multi-part loader no longer creates an empty manifest at first boot that wiped v2-loaded state on the next restart (every tokio SET was lost on restart).
  • tests/txn_kv_wiring.rs: made test_txn_commit_wal_crash_recovery runtime-agnostic by polling for either the monoio multi-part .base.rdb artifact or the tokio single-file appendonly.aof. Added connect_redis_with_retry helper to bound and retry the post-bind RESP handshake (was racing the shard accept loop, surfacing as EAGAIN on Linux and ECONNRESET on macOS CI).

Fixed — PR #95 review hardening

  • main.rs malloc_conf symbol: replaced the union-based unsafe pun with a #[repr(transparent)] Sync wrapper around a c"..." literal.
  • command/server_admin.rs get_vsz_bytes: replaced four unsafe libc::{open,read,close,sysconf} blocks with safe /proc/self/status parsing on the cold MEMORY DOCTOR path.
  • main.rs arena scan now uses env::args_os() so non-UTF-8 argv no longer panics before clap reports the error.
  • server/embedded.rs shutdown sequence: cancel → join shard threads → drop the outer aof_tx → join the AOF thread, so the writer never exits while shards are still queuing appends. Thread join panics are now propagated through the function result.
  • storage/db.rs annotated two .expect() calls in Database::set's insert_or_update closures for the hot-path unwrap ratchet.

Deferred to v0.2 follow-ups

  • P3c — wire SnapshotState::set_last_lsn(wal_flush_lsn) into the live persistence tick so freshly-written snapshots embed their LSN (currently PITR falls back to full WAL replay when only v1-shaped snapshots exist).
  • C3b — push-based CDC.SUBSCRIBE over RESP3 Push frames, per-shard subscriber registry hooked into wal_append_and_fanout, slow-consumer disconnect policy. Envelope format unchanged.
  • Integration suites (tests/pitr_integration.rs, tests/cdc_integration.rs), scripts/test-pitr-cdc.sh end-to-end smoke, and the benchmark gates (PITR restart ±10%, CDC ≥100K events/s/shard, write p99 ±5%).

[0.1.12] — 2026-05-12

Performance & memory observability release. 50 commits since v0.1.11, no public API breaks, no on-disk format change. Validated on OrbStack moon-dev (2026-05-12) and locally green for both runtime-monoio and runtime-tokio,jemalloc.

Performance — DashTable hot-path (Phase 189, PERF-07 + PERF-09)

  • Pre-sized DashTable. DashTable::with_capacity() plus the new --initial-keyspace-hint <N> flag size the segment array up front so steady-state operation hits zero split_segment calls. Pre-size invariant test confirms zero splits at 1 M keys. The 27 % CPU spent in split_segment during SET p=16 (PERF-07) is fully eliminated.
  • Database::set rewrite. New DashTable::insert_or_update / Segment::insert_or_update_at single-probe helpers replace the previous find + remove + insert triple-probe pattern.
  • Segment::find fallback elimination + force-inlined SIMD. The cold "key spilled to non-home group" fallback path is removed once has_non_home_keys is invariant-tracked on insert (the insert_or_update_at change above already maintains the flag); the SIMD probe helpers are #[inline(always)]. PERF-09 attributed 12.65 % of Segment::find self-time to the fallback; remaining cost is the irreducible per-hit memcmp confirm (threshold amended to <3 %). 1 M-key correctness gate validates zero false positives/negatives.

Performance — Memory observability (Phase 190)

  • moon_memory_bytes{kind=…} Prometheus gauge. Seven subsystem labels — dashtable, hnsw, csr, wal, sealed_replication_backlog, allocator_overhead, and the rolled-up total. Updated every scrape via a single hook so the sum reconciles to RSS within the CI tolerance window.
  • MEMORY DOCTOR full schema. Multi-line RESP response covering every subsystem, the rolled-up total, and a derived allocator_overhead pseudo-kind (RSS − Σ subsystems). Adds operator triage signal beyond the legacy single-line summary.
  • resident_bytes() trait implemented across Database, DashTable, VectorStore (HNSW + IVF), GraphStore (CSR + SlotMap), WalWriter, ReplicationBacklog (sealed-segment side), and AllocatorOverhead. Zero-allocation, on-demand poll.
  • Memory steady-state CI job. scripts/bench-memory-steady-state.sh
  • baseline fixture; gate widened to ±10 % on RSS / Σ ratio after a Linux-CI tolerance pass.

Changed — Allocator UX (Phase 191)

  • jemalloc narenas:8 cap with --memory-arenas-cap <N> CLI override. Caps the per-CPU arena explosion that inflates VSZ on high-core hosts; mostly a cosmetic fix on Linux containers but produces a meaningfully tighter top/ps reading for operators.
  • Tri-state allocator selection. New mimalloc-alt cargo feature alongside the existing jemalloc / mimalloc (fallback) paths; mutually exclusive at compile time. A/B benchmark script scripts/bench-allocator-ab.sh ships with the release.
  • docs/OPERATOR-GUIDE.md — Memory Accounting section. Documents the VSZ-vs-RSS distinction, MEMORY DOCTOR field-by-field, and the --memory-arenas-cap / mimalloc-alt tuning knobs.

Added — Dispatch Observability (Phase 177)

  • moon_dispatch_path_total{path=...} Prometheus counter: four-way classification of every command by shard-routing decision — local_inline (SIMD fast path), local (standard local branch), cross_read_fast (RwLock shared-read bypass of SPSC), cross_spsc (deferred cross-shard write via PipelineBatchSlotted). Ratio cross_spsc / Σ is the ground-truth signal for dispatch-layer optimization work. Zero-allocation hot-path overhead (&'static str labels, #[inline] with early-return on !METRICS_INITIALIZED). Verified on macOS + Linux: counter sums close exactly to driven traffic, no overcount.

Changed

  • text-index is now a default feature. BM25 full-text search (FT.SEARCH BM25 mode), FT.AGGREGATE, and three-way RRF hybrid fusion are included in all standard builds. No longer requires --features text-index. To exclude it (e.g. minimal embedded builds): --no-default-features --features runtime-monoio,jemalloc,graph.

Added — SDK Validation

  • Python SDK sdk/python/examples/validate.py: End-to-end live validator for all SDK sub-clients: ping, strings, counter, hash, list, set, zset, vector index lifecycle, graph engine, session search, semantic cache, text search (BM25 + aggregate + hybrid), and server info. Result against Moon with text-index: 114 PASS / 0 FAIL / 0 SKIP. Gracefully skips text sections when server built without text-index.
  • Rust SDK sdk/rust/examples/validate.rs: Re-validated against text-index build — 85 PASS / 0 FAIL.

Fixed — Python SDK

  • moondb.graph._parse_neighbors: server returns alternating [edge_map, node_map, ...] as flat key-value arrays (b'id', int, b'src', int, b'dst', int, …). Previous parser expected positional [node_id, label, props] — caused int() on b'id' crash. Now correctly identifies node entries by labels key and parses them from the flat kv format.

Fixed — CI Hygiene

  • tests/pipeline_auto_index.rs: tighten outer cfg from runtime-tokio to all(runtime-tokio, text-index) so the file compiles to zero tests when text-index is disabled. Previously the file compiled but the FT.SEARCH text fast path was #[cfg]-ed out, causing @name:corpus queries to fall through to the KNN-only parser and panic with "invalid KNN query syntax".
  • 4 FT unwraps: add inline #[allow(clippy::unwrap_used)] with invariant justifications in vector_search/ft_text_search.rs (3 sites inside apply_post_processing where do_summarize / do_highlight implies the Option is Some) and handler_monoio/ft.rs:165 (is_text was derived from query_bytes.as_ref().map_or(false, _)). Restores the audit-unwrap baseline to 0.

Compatibility

  • Wire protocol: unchanged. Drop-in replacement for v0.1.11.
  • Persistence on-disk format: unchanged.
  • Default feature set: text-index is now on by default. Minimal embedded builds need an explicit --no-default-features --features runtime-monoio,jemalloc,graph.

[0.1.11] — 2026-04-27

Hot-path perf release — eliminates two atomic-CAS hot paths in the write dispatch loop discovered via ARM perf annotate on c4a-16 (GCloud Axion). Empirically validated on the same hardware: 8-shard SET p=64 c=200 throughput 1.84M → 3.87M RPS (+110%) when run with --disk-offload disable, or +15% under default flags. No public API change.

Sprint 3.5a and 3.5b from .planning/rfcs/v02-enterprise-architecture.md.

Performance — Sprint 3.5a: Lock-free is_replica mirror

try_enforce_readonly was taking RwLock::try_read() on Arc<RwLock<ReplicationState>> for every command before dispatch — an atomic CAS on the per-command hot path. ARM annotate showed mov w8, #0xfffd; cmp w11, w9 consuming 84% of self-time inside the function (10% of total CPU on 8-shard SET p=64).

Fix: replace the per-command lock probe with a single AtomicBool::load(Acquire).

  • ReplicationState::is_replica_mirror: Arc<AtomicBool> — lock-free mirror of role == Replica { .. }, kept in sync via the new ReplicationState::set_role(&mut self, role) method (single owner of the invariant).
  • ConnectionContext::is_replica_mirror: Option<Arc<AtomicBool>> — snapshotted from ReplicationState once at connection setup; per-command try_enforce_readonly is now just an atomic load with no lock acquisition.
  • All 6 production rs_guard.role = ... sites in handler_single, handler_monoio/dispatch, and handler_sharded/dispatch migrated to set_role(). Test fixtures in replication/handshake.rs migrated too so the invariant holds in test code.
  • Round-5 verification on commit 32f48c4 (c4a-16, 8-shard SET p=64 c=200): try_enforce_readonly is now 0% of profile (down from 10%). Sprint 3 acceptance criterion <1% met.

Performance — Sprint 3.5b: WAL no-op bypass

wal_append_and_fanout was acquiring a parking_lot::Mutex (replication backlog) and a std::sync::RwLock (replication state) on every write, even when no replica was connected and no WAL writer existed. ARM annotate showed caslb/casab ARM CAS-byte atomics dominating self-time (~21% of total CPU on 8-shard SET p=64 with --appendonly no and zero replicas).

Fix: hoist a single early-return at the top of the function:

if wal_writer.is_none() && wal_v3_writer.is_none() && replica_txs.is_empty() {
    return;
}

The criterion is fully derivable from existing inputs — no new shared state. Skips both the backlog Mutex::lock and the repl_state RwLock::read on the cold path.

  • Round-5 verification with --disk-offload disable (commit 32f48c4): wal_append_and_fanout is now 0.05% of profile (down from 21%). Sprint 3 acceptance criterion <2% met.
  • Operator note: the bypass only fires when (a) --appendonly no, (b) --disk-offload disable, AND (c) no replicas are connected. Default builds with disk-offload on (the production default) keep wal_v3_writer = Some(_) and the function continues to do real WAL v3 work — that path is unaffected.

Throughput Impact

Workload (8-shard SET p=64 c=200, c4a-16, frame pointers ON) v0.1.10 v0.1.11 Δ
Default flags 1.84M RPS 2.11M RPS +15%
--disk-offload disable ~2.95M RPS (projected) 3.87M RPS sustained +31%
try_enforce_readonly self-time 10.0% 0% -10pp
wal_append_and_fanout self-time 21.2% 0.05% (with disk-offload disable) -21.1pp

Production builds without frame pointers should clear 5.5–6.5M RPS at the same flag set. Round-4 baseline data: memory/benchmark_perf_round4_2026_04_27.md. Round-5 verification data: memory/benchmark_perf_round5_2026_04_27.md.

Tests Added

  • replication::state::tests::test_set_role_updates_is_replica_mirror
  • replication::state::tests::test_is_replica_mirror_default_false
  • shard::spsc_handler::wal_append_tests::test_wal_append_bypass_when_no_writers_no_replicas
  • shard::spsc_handler::wal_append_tests::test_wal_append_writes_backlog_when_replicas_present

All 2450 lib tests passing locally (cargo test --no-default-features --features runtime-tokio,jemalloc --lib); clippy clean (cargo clippy --no-default-features --features runtime-tokio,jemalloc -- -D warnings).

Drive-by Fixes

  • src/shard/mod.rs: pre-existing test compile failure where two drain_spsc_shared call sites passed &mut None for repl_backlog instead of a SharedBacklog (Arc<Mutex<Option<...>>>). Fixed by constructing Arc::new(parking_lot::Mutex::new(None)) in the test fixture. This was blocking cargo test --lib on main independent of this release.
  • src/shard/dispatch.rs: pre-existing clippy doc_lazy_continuation error on the TextSearchPayload doc comment, blocking cargo clippy -- -D warnings on main.

Compatibility

  • Wire protocol: unchanged. Drop-in replacement for v0.1.10.
  • Public API: only additive (ReplicationState::set_role, ReplicationState::is_replica_mirror field). No breaking changes.
  • Persistence on-disk format: unchanged.
  • Replication wire format: unchanged.

[0.1.10] — 2026-04-23

Stable replication marker. Single-shard PSYNC2 wired end-to-end and production-ready for --shards 1 master with any --shards N replica topology. Multi-shard master PSYNC is scheduled for v0.2 (see .planning/rfcs/multi-shard-replication-design.md).

  • Replication (081c43b): single-shard master PSYNC2 end-to-end wired, REPLCONF validated, master_link_status reports the actual handshake state instead of the legacy up stub.
  • Performance: batch-level eviction gate; try_handle_* paths #[inline]-ed; DashTable carries through the v0.1.10 pre-size groundwork (capacity hint + headroom).
  • Docs: BENCHMARK.md §2.7 updated with the 2026-04-22 GCloud re-measurement; v0.1.x replication scope documented under docs/guides/clustering.mdx#replication.

[0.1.9] — 2026-04-19

Lunaris Retriever Gap Closure. Every v0.1.8 client-side fallback in the Lunaris SDK is now closed so HybridRRFRetriever (dense path), GraphFirstRetriever, and PathReasoningRetriever run Moon-native.

  • Phase 167 CYP-01/02: Cypher CREATE / MERGE writes participate in CrossStoreTxn via record_graph(); TXN.ABORT rolls them back.
  • Phase 168 CYP-03/06: coalesce() built-in + single-hop edge-var binding in variable-length EXPAND.
  • Phase 169 CYP-04/05: shortestPath() parser + Dijkstra executor bridge with path-variable binding.
  • Phase 170 HYB-01/02/04: FT.SEARCH HYBRID dense stream honours as_of_lsn.
  • Phase 171 SCAT-01/02/03: ShardMessage::VectorSearch + FtHybridPayload carry as_of_lsn for multi-shard AS_OF correctness.
  • Phase 172 PIPE-01/02/03: pipeline-aware HSET auto-indexing regression guard (3-test suite).

Audit status: PASSED_WITH_DOCUMENTED_DEFERRALS. 15 / 20 requirements fully satisfied; HYB-03 BM25 MVCC deferred and closed in v0.1.10 follow-up (G-1); Phase 173 hygiene HYG-02 handler split RFC'd.

Stats: 6 phases shipped, 17 plans, 27 files changed, +2924 / −376 LOC.

[0.1.8] — 2026-04-18

Added — Cross-Store ACID Transactions (Phases 157, 161-163)

  • TXN.BEGIN: Start a cross-store transaction — buffers KV, vector, and graph writes as intents.
  • TXN.COMMIT: Commit all changes atomically with WAL record (XactCommit 0x34) for crash recovery.
  • TXN.ABORT: Roll back all changes via undo-log replay with before-images.
  • KvWriteIntents: Sparse MVCC side-table for uncommitted KV writes during transactions.
  • DeferredHnswInserts: Vector index inserts deferred until commit, avoiding partial graph states.
  • UndoLog: SmallVec-based KV rollback with before-image recording for all mutation types.
  • WAL transaction records: XactBegin (0x33), XactCommit (0x34), XactAbort (0x37) with crash recovery replay.
  • Mutual exclusion: TXN and MULTI/EXEC cannot be mixed (enforced at handler level).

Added — Bi-Temporal MVCC (Phase 158)

  • TEMPORAL.SNAPSHOT_AT: Record wall-clock → WAL LSN binding for point-in-time queries.
  • TEMPORAL.INVALIDATE: Set valid_to on a graph entity (NODE or EDGE) for temporal visibility control.
  • FT.SEARCH AS_OF: Query vector indexes at a historical timestamp via LSN resolution.
  • GRAPH.QUERY VALID_AT: Execute Cypher queries against graph state valid at a specific timestamp.
  • TemporalRegistry: BTreeMap-backed wall-clock → LSN mappings with O(log n) range lookups.
  • TemporalKvIndex: Sparse versioned KV index with lazy initialization.
  • CSR segment format v2: Bi-temporal NodeMeta with valid_from/valid_to fields.
  • WAL temporal records: TemporalUpsert (0x35) and GraphTemporal (0x36) for crash recovery.

Added — Workspace Partitioning (Phase 159)

  • WS CREATE: Create a workspace with UUID v7 (time-ordered, 74-bit random). Name max 64 bytes.
  • WS DROP: Delete a workspace and its registry entry.
  • WS AUTH: Bind a connection to a workspace — all subsequent commands transparently prefixed.
  • WS INFO: Return workspace metadata (name, creation timestamp).
  • WS LIST: Enumerate all registered workspaces.
  • Transparent key rewriting: workspace_rewrite_args() injects {ws_hex}: hash tag prefix on key arguments and strips it from responses.
  • WorkspaceRegistry: Per-shard metadata with creation timestamps and WAL persistence.

Added — Durable Message Queues (Phase 160)

  • MQ CREATE: Create a durable queue with MAXDELIVERY (default 3) and DEBOUNCE options.
  • MQ PUSH: Enqueue messages with field/value pairs (returns stream ID).
  • MQ POP: Claim messages with optional COUNT (defaults to 1). Increments delivery counter.
  • MQ ACK: Acknowledge messages by stream ID.
  • MQ DLQLEN: Return dead-letter queue depth.
  • MQ TRIGGER: Register debounced trigger callbacks with configurable debounce interval.
  • MQ PUBLISH: Transactional enqueue within a TXN block — applied on TXN COMMIT.
  • Dead-letter queue: Automatic DLQ at {queue_key}::mq:dlq after exceeding MAXDELIVERY attempts.
  • TriggerRegistry: Debounced callback execution via pub/sub publish.
  • WAL recovery: replay_mq_wal() with cursor rollback for durable queue state restoration.

Added — Handler Parity

  • All TXN/TEMPORAL/MQ/WS commands wired into handler_monoio.rs, handler_sharded.rs, uring_handler.rs, and handler_single.rs.
  • 12 KV transaction integration tests, 14 MQ integration tests, workspace cross-shard dispatch tests.
  • TEMPORAL/MQ/WS/TXN entries added to test-commands.sh and test-consistency.sh.

[0.1.7] — 2026-04-17

Added — BM25 Full-Text Search Engine (Phases 149-156)

  • BM25 inverted index: Full-text search with multi-field boosting and per-field term frequency tracking.
  • TEXT field type: Unicode tokenization with stemming and normalization in FT.CREATE.
  • TAG field type: Categorical tag filtering with multi-value support (@field:{val1|val2}).
  • NUMERIC field type: Range filtering (@field:[min max]) in FT.CREATE and FT.SEARCH.
  • FT.AGGREGATE: Aggregation pipeline with GROUPBY/REDUCE, scatter-gather across shards, HLL COUNT_DISTINCT.
  • Three-way RRF hybrid fusion: Combines BM25 + dense vector + sparse vector results via Reciprocal Rank Fusion.
  • Typo tolerance: FST Levenshtein fuzzy matching (%%term%%) and prefix search (term*).
  • HIGHLIGHT/SUMMARIZE: Post-processors for formatting search results with matched term highlighting.
  • Multi-shard DFS global IDF: Distributed frequency statistics for accurate BM25 scoring regardless of shard count.
  • FT.DROPINDEX DD: Atomic index + document deletion flag — deletes all hash keys matching index prefixes.
  • Python SDK text module: client.text.text_search(), client.text.aggregate(), client.text.hybrid_search() with typed pipeline DSL.
  • LangChain/LlamaIndex hybrid adapters: Framework integrations updated for three-way hybrid search.

Statistics

  • 8 phases (149-156), 27 plans, 26 requirements, 122 commits.

[0.1.6] — 2026-04-15

Added — AI-Native Data Primitives

  • Multi-field vector indexes (FT.CREATE): multiple VECTOR fields per index, per-field segment storage, field-targeted @field_name syntax in FT.SEARCH KNN clause.
  • Sparse vector module (src/vector/sparse/): inverted index with SparseStore, enabling BM25-style sparse retrieval alongside dense HNSW.
  • Hybrid dense+sparse search: SPARSE clause in FT.SEARCH with Reciprocal Rank Fusion (RRF) for combining dense and sparse results.
  • Text index (src/vector/text_index.rs): Unicode tokenization pipeline with stemming and normalization, feature-gated under text-index.
  • Boolean and geo filter expressions: BoolEq and GeoRadius filter variants with evaluation logic in FilterExpr.
  • FT.RECOMMEND: centroid-based recommendation over vector indexes — computes centroid of seed vectors and returns nearest neighbors.
  • FT.NAVIGATE: multi-hop knowledge graph navigation from vector search results, bridging vector and graph queries.
  • FT.EXPAND: GraphRAG expansion command — traverses graph edges from vector search results to discover related entities, with configurable depth.
  • FT.CACHESEARCH: semantic cache-or-search command — returns cached results on similarity hit, falls back to full search on miss.
  • FT.CONFIG SET/GET: runtime configuration for per-index knobs (e.g., AUTOCOMPACT toggle).
  • SESSION clause: session-scoped filtering in FT.SEARCH for multi-tenant and agent memory isolation.
  • RANGE threshold post-filter: distance threshold filtering in FT.SEARCH results.
  • LIMIT pagination: LIMIT offset count support in FT.SEARCH with multi-segment merge.

Added — Production Infrastructure

  • moondb Python SDK (python/moondb/): high-level client with vector, graph, session, cache, and framework integrations (LangChain, LlamaIndex).
  • 5 quickstart examples: RAG, semantic cache, GraphRAG, AI agent tools, memory engine — all using real MiniLM embeddings.
  • Production deployment guide (docs/production-guide.md): configuration reference, TLS setup, monitoring, ACL, tuning.
  • Dockerfile improvements: OCI labels, admin port exposure, production defaults.
  • docker-compose.yml: production configuration with resource limits, health checks, ulimits.
  • Prometheus metrics: counters and histograms for v0.1.6 commands (FT.RECOMMEND, FT.NAVIGATE, FT.EXPAND, FT.CACHESEARCH, FT.CONFIG).
  • OpenTelemetry stubs: otel feature flag with tracing infrastructure for future OTLP export.
  • Benchmark script (scripts/bench-v0.1.6.sh): automated benchmarks for new vector search features.

Added — Graph-Vector Integration

  • EXPAND GRAPH clause in FT.SEARCH: inline graph expansion during vector search with configurable depth.
  • graph_expand.rs bridge module: connects vector search results to graph traversal engine.
  • key_to_node mapping on NamedGraph: enables graph expansion from vector search key hashes.

Fixed — Critical Production Bugs

  • Deadlock in cross-shard HSET auto-index: parking_lot::Mutex re-entry in PipelineBatch and PipelineBatchSlotted handlers — shard_databases.vector_store(shard_id) attempted to re-lock a non-reentrant mutex already held by the caller. Fixed by using the passed-in reference.
  • Auto-index HSET inside MULTI/EXEC: vector auto-indexing now works correctly within transactions and pipeline batch paths.
  • FT.RECOMMEND filter bug: filter expressions were not applied correctly to recommendation results.
  • FT.CONFIG/RECOMMEND/NAVIGATE/EXPAND routing: commands now dispatched correctly in all connection handlers (monoio, tokio, single-threaded).
  • Session deduplication: fixed duplicate session entries in search results.
  • Inline filter parsing: corrected parsing of filter expressions in inline command mode.
  • FT.COMPACT, FT.CONFIG, FT.CACHESEARCH metadata: registered in phf command metadata table for ACL and COMMAND DOCS.

Fixed — CI/Release Pipeline

  • nfpm download URL: pinned v2.46.1 with correct Linux_x86_64 asset name (old nfpm_linux_amd64.tar.gz returns 404).
  • SBOM filenames: cargo-cyclonedx v0.5+ writes .json not .cdx.json.
  • Cosign keyless signing: added id-token: write permission for Sigstore OIDC (was falling back to interactive device flow).
  • Package job race condition: deb/rpm upload now waits for GitHub release to be created first.
  • upload-artifact: upgraded v4 → v7 (Node.js 20 deprecation).
  • Test compilation: added tokio dev-dependency with process feature for blocking_list_timeout.rs under default (monoio) features.

Changed

  • Graph enabled by default: graph feature now included in default feature set.
  • Vector index persistence: v2 format with backward-compatible v1 migration.
  • Memory engine example: rewritten as 142-line script with real MiniLM embeddings (was 487-line complex agent loop).
  • Clippy 1.94 compliance: all warnings resolved for Rust 1.94 MSRV.

Validation

  • 2,613+ unit tests pass (release mode, default features).
  • 2,139 library tests pass under runtime-tokio feature set.
  • 184 unsafe blocks, all with SAFETY comments (audit pass).
  • 0 unannotated unwraps on hot paths (ratchet pass).
  • Zero clippy warnings (default + runtime-tokio,jemalloc feature sets).
  • cargo fmt --check clean.
  • 8 fuzz targets in CI.
  • Full release pipeline validated: 6 binary targets, Docker image, deb/rpm packages, SBOMs, cosign signatures.

[0.1.5] — 2026-04-12

Added — Moon Console (Interactive Data Client)

  • HTTP/WebSocket gateway (src/admin/): REST endpoints (/api/v1/info, /api/v1/command, /api/v1/keys, /api/v1/key/*, /api/v1/memory/treemap, /api/v1/hnsw/trace), WebSocket-to-RESP3 bridge at /ws/console, SSE metrics stream at /sse/metrics (1 Hz), CORS allowlist, per-IP token-bucket rate limit, HMAC-SHA256 Bearer auth, HTTP/2 support, static file serving via rust-embed.
  • React 19 console (console/): 7-view SPA (Dashboard, Browser, Console, Vector Explorer, Graph Explorer, Memory, Help) served at /ui/. 50.9 KB gzipped initial bundle (6× under 300 KB target) via Vite 8 + manual chunk splitting (Three.js/Monaco/Recharts lazy).
  • Real-time Dashboard: 7 widgets (QPS, latency P50/P99, memory, clients, ops by type, keyspace) driven by SSE stream.
  • KV Data Browser: namespace tree, virtual-scrolled key list (TanStack Virtual), type-specific editors for Strings/Hashes/Lists/Sets/Sorted Sets/Streams, TTL display + edit, bulk delete with toasts.
  • Query Console: Monaco editor with RESP + Cypher Monarch syntax, 233-command auto-complete, multi-tab, history, Cmd+Enter (current line) / Cmd+Shift+Enter (whole buffer), line-by-line execution for paste safety.
  • Vector 3D Explorer: UMAP projection in a web worker, HNSW layer overlay, KNN search with distance rings, lasso selection, Three.js r183 + React Three Fiber.
  • Graph 3D Explorer: force-directed layout (d3-force-3d worker), Cypher editor, node/edge property inspector, hybrid query integration.
  • Memory view: keyspace treemap (server-aggregated /api/v1/memory/treemap), slowlog table, command stats.
  • Built-in Help guide: 427-line Getting Started tutorial with seed examples.
  • Core admin commands (src/command/server_admin.rs): FLUSHALL, FLUSHDB, DBSIZE, DEBUG OBJECT/SLEEP/JMAP, MEMORY USAGE — closing pre-existing dispatch gaps.
  • Multi-shard SCAN fan-out (src/admin/scan_fanout.rs): composite cursor {shard_id}:{cursor} so Browser sees unified keyspace.
  • Frontend test infrastructure: 56 Vitest unit tests + 9 Playwright E2E specs. New scripts/test-integration.sh harness and .github/workflows/console-integration.yml.
  • Admin-port hardening (src/admin/{auth,cors,rate_limit,middleware}.rs): Bearer auth, CORS allowlist, per-IP rate limit.

Fixed

  • WebSocket request ID echo: errors now echo back the client's id, preventing client-side promise timeouts on malformed input.
  • Console type badges: execCommand response was returning the full {result, type} envelope; now unwraps .result correctly.
  • Multi-line paste in Console: Cmd+Enter now executes the current line only (redis-cli paste behavior). Cmd+Shift+Enter executes the whole buffer line-by-line.

Validation

  • 101+ Rust unit/integration tests pass on both runtime-tokio and runtime-monoio.
  • 56 Vitest tests + 9 Playwright specs pass.
  • Zero clippy warnings (default + runtime-tokio,jemalloc feature sets).
  • cargo fmt --check clean.
  • 8 fuzz targets in CI.

[0.1.4] — 2026-04-11

Added — Graph Engine Integration (v0.1.4, 2026-04-11)

  • Property graph engine (src/graph/, feature-gated under graph): segment-aligned CSR storage with SlotMap generational indices, ArcSwap lock-free reads, Roaring validity bitmaps, and Rabbit Order compaction for cache locality. 8,500+ LOC, 319 tests.
  • 12 GRAPH.* commands: CREATE, ADDNODE, ADDEDGE, NEIGHBORS, QUERY, RO_QUERY, EXPLAIN, VSEARCH, HYBRID, INFO, LIST, DELETE — all with RESP3 Map responses and ACL annotations.
  • Cypher subset parser: hand-rolled recursive descent with logos lexer, 12 clauses (MATCH/WHERE/RETURN/CREATE/DELETE/SET/MERGE/WITH/UNWIND/CALL/ORDER/LIMIT), parameterized queries ($param), nesting depth limit (64), plan caching.
  • Hybrid graph+vector queries: graph-filtered vector search, vector-to-graph expansion, vector-guided walk with automatic strategy selection.
  • Traversal engine: BFS/DFS/Dijkstra with bounded frontiers (100K cap), temporal decay + distance scoring, segment merge reader across mutable + immutable segments.
  • Graph indexes: per-label/type Roaring bitmaps, boomphf minimal perfect hash (~3 bits/key), property B-tree for range queries.
  • Cross-shard traversal: scatter-gather via SPSC mesh, graph hash tags for shard co-location, snapshot-LSN forwarding, configurable depth limit.
  • Graph MVCC: extends existing TransactionManager with graph write intents, snapshot-isolated multi-hop traversal, bounded epoch hold (30s).
  • Graph WAL durability: RESP-encoded graph commands in per-shard WAL, two-pass replay (nodes before edges), CRC32-validated CSR segment persistence.
  • Cost-based planner: GraphStats with incremental degree tracking, graph-first vs vector-first strategy selection, P99 hub detection.
  • Criterion benchmarks: CSR 1-hop 1.02ns, edge insert 64.8ns, 2-hop BFS 4.99µs, CSR freeze 5.12ms, SIMD cosine 384d 33.9ns.
  • Fair comparison benchmark (tests/graph_bench_compare.rs): Moon 2.4x FalkorDB on Cypher MATCH, 19x on native 1-hop, 23x on population.
  • New dependencies: slotmap 1.x (generational indices), boomphf 0.6 (MPH), logos 0.14 (Cypher lexer, optional).

Added — High-Impact Redis Command Parity (2026-04-10)

  • COPY command — atomic key duplication with DESTINATION, REPLACE options (Redis 6.2+).
  • Bit operations — GETBIT, SETBIT, BITCOUNT (byte/bit range modes), BITOP (AND/OR/XOR/NOT), BITPOS (byte/bit range modes) with read-only dispatch variants.
  • SORT command — full BY/GET/LIMIT/ALPHA/ASC/DESC/STORE support for lists, sets, and sorted sets.
  • Geospatial commands — GEOADD (NX/XX/CH), GEOPOS, GEODIST (M/KM/FT/MI), GEOHASH (11-char base32), GEOSEARCH (FROMLONLAT/FROMMEMBER, BYRADIUS/BYBOX, WITHCOORD/WITHDIST/WITHHASH), GEOSEARCHSTORE.
  • CONFIG REWRITE — atomic write of runtime config to <dir>/moon.conf (tmpfile + rename). CONFIG RESETSTAT stub.
  • CLIENT PAUSE/UNPAUSE — delays command processing with WRITE-only mode support. CLIENT INFO, CLIENT LIST (stub), CLIENT NO-EVICT/NO-TOUCH accepted.
  • MEMORY USAGE/DOCTOR/HELP — key memory estimation via estimate_memory().
  • Lazyfree threshold — configurable via CONFIG SET lazyfree-threshold N (default 64).
  • GETBIT/SETBIT metadata — added to PHF command registry.
  • GEOADD/GEOSEARCHSTORE — added to AOF write commands test list.
  • EXPIREAT/PEXPIREAT — absolute Unix timestamp expiry (seconds/milliseconds).
  • EXPIRETIME/PEXPIRETIME — read back absolute expiry timestamp.
  • FLUSHDB/FLUSHALL — clear all keys in current database.
  • TIME — server clock as [seconds, microseconds].
  • RANDOMKEY — return a random key from the database.
  • TOUCH — refresh LRU/LFU access time without reading value.
  • SHUTDOWN — dispatch entry (graceful stop via signal handler).
  • BITFIELD — GET/SET/INCRBY with type specifiers (u8/i16/u32/...), OVERFLOW WRAP/SAT/FAIL.
  • LCS — Longest Common Substring with LEN option.
  • XSETID — set stream last-delivered ID without adding entries.
  • GEORADIUS/GEORADIUSBYMEMBER — deprecated wrappers translating to GEOSEARCH.
  • OBJECT FREQ/IDLETIME/REFCOUNT — LFU counter, idle seconds, reference count introspection.
  • LOLWUT — Easter egg returning Moon version.

Added — Client Connection Security Hardening (2026-04-10)

  • --maxclients (P0): Connection limit with atomic CAS rejection (default 10000, 0=unlimited). Returns -ERR max number of clients reached when exceeded.
  • --timeout (P0): Client idle timeout in seconds (default 0=disabled). Disconnects idle clients via tokio::time::timeout / monoio::select!.
  • --tcp-keepalive (P0): TCP keepalive interval (default 300s, 0=disabled). Sets SO_KEEPALIVE + TCP_KEEPIDLE on accepted sockets via socket2.
  • AUTH rate limiting (P0): Per-IP exponential backoff on AUTH failures (100ms base, 10s cap, 60s auto-reset). New module src/auth_ratelimit.rs.
  • CLIENT LIST / INFO / KILL (P1): Global client registry with Drop-guard deregister. Redis-compatible output format. Kill by ID/ADDR/USER. New module src/client_registry.rs.
  • CLIENT PAUSE / UNPAUSE (P1): Server-wide pause with ALL/WRITE modes and auto-expiry. New module src/client_pause.rs.
  • CLIENT NO-EVICT / NO-TOUCH (P1): Accepted stubs for Redis compatibility.
  • ACL GENPASS (P1): Cryptographically secure random password generation (1-4096 bits, hex output).
  • CONFIG GET/SET support for maxclients, timeout, tcp-keepalive (runtime-mutable).
  • Monoio connection tracking: Added missing record_connection_opened / record_connection_closed for accurate connected_clients metric.

Fixed — Deep Review Findings (2026-04-11)

  • DoS protection: execute_profile and execute_mut Cypher paths now enforce MAX_HOPS_LIMIT=20 and MAX_RESULT_ROWS=100K (were unbounded).
  • WAL correctness: Cypher DELETE passes actual LSN to remove_node/remove_edge (was hardcoded to 0).
  • GRAPH.DROP metadata: added missing phf dispatch table entry.
  • SAFETY comments: added to all 7 unsafe SIMD/mmap functions.
  • BFS 30% faster: scratch buffer reuse in SegmentMergeReader, zero-alloc CsrStorage callback, MergedNeighbor derives Copy.
  • ParallelBfs: uses plain HashSet on sequential path (was DashSet with 64 shards overhead).
  • Recovery hardening: CSR manifest path traversal validation, WAL embedding dimension cap (65536), LSN saturating_add.
  • CI optimized: consolidated 26 jobs → 4 per PR, concurrency groups cancel superseded runs, fixed org runner group for public repos.

Fixed — Wave 0-4 Gap Closure (2026-04-09)

  • ZREVRANGEBYSCORE/ZREVRANGEBYLEX correctness bug: Fixed double-swap of min/max bounds in zrange_by_score and zrange_by_lex that caused empty results for finite score ranges (e.g., ZREVRANGEBYSCORE key 3 1). Added finite-range test to test-commands.sh.
  • INFO command enriched: Clients section now reports connected_clients, Memory section reports used_memory/used_memory_human/used_memory_rss (from /proc/self/status), Replication section wired to actual ReplicationState (role, connected_slaves, master_replid, master_repl_offset).
  • Tracing spans: Added #[instrument] to connection handlers (single, monoio), replication master (tokio, monoio), HNSW compaction, and AOF rewrite — 6 new spans.
  • Replication lag metric wired: moon_replication_lag_bytes Prometheus gauge now updated from get_replication_info().
  • CI supply chain security: cargo deny check + cargo audit added to CI pipeline (deny.toml was previously unenforced).
  • Release pipeline: aarch64-unknown-linux-gnu build added via cross for primary production target.
  • Crash matrix expanded: BGSAVE and BGREWRITEAOF crash cells added (6/7 coverage).
  • Compatibility tests expanded: Stream (XADD/XLEN/XRANGE/XTRIM), Lua scripting (EVAL/EVALSHA/SCRIPT), and ACL (WHOAMI/LIST) tests added to redis_compat.rs.

Added — Production Contract (Phase 87, 2026-04-08)

  • docs/PRODUCTION-CONTRACT.md — Moon's v1.0 promises: per-command-class SLOs (provisional until Phase 97), supported platform matrix (Linux aarch64 primary, Linux x86_64 secondary contingent on PERF-04, macOS dev-only via OrbStack), durability mode semantics per appendfsync × failure-class, availability & replication guarantees, security guarantees, explicit out-of-scope list, and a machine-checkable GA Exit Criteria checklist that every v0.1.3 phase ticks off. This is the contract every downstream hardening phase (88–100) tests against.
  • docs/runbooks/ — stub directory for operator runbooks authored in Phase 99 (REL-05).

Changed — Toolchain Upgrade (Phase 88, 2026-04-08)

  • MSRV bumped from Rust 1.85 to 1.94.0. rust-toolchain.toml committed so fresh clones auto-install the pinned version; CI workflows (ci.yml, codeql.yml, release.yml) and OrbStack moon-dev VM provisioning in CLAUDE.md updated. No language/runtime behavior change; downstream phases benefit from new clippy lints and std/compiler improvements. Contributors must run rustup update on next pull.

Added — Production Readiness Phases 92-105 (2026-04-09)

  • Observability: Prometheus /metrics on --admin-port, SLOWLOG GET/LEN/RESET/HELP, HEALTHZ + READYZ commands, /healthz + /readyz HTTP endpoints, INFO extended with Server/Clients/Memory/Stats/CPU sections, --check-config flag, per-command latency histograms + connection metrics wired into dispatch
  • Durability proof: Crash-injection test matrix, torn-write WAL v3 tests (CRC32C validated), Jepsen-lite linearizability harness, backup/restore workflow test
  • Replication hardening: PSYNC partial resync, full resync, network partition, kill-restart, replica promotion tests
  • Client compatibility: CI matrix (redis-py, go-redis, jedis, ioredis, node-redis, redis-rs, hiredis), 24 Redis compat tests, vector client smoke script, docs/redis-compat.md
  • Performance gates: Criterion regression CI with baseline caching, RSS-per-key memory gate script
  • Security hardening: deny.toml (cargo-deny), SECURITY.md, docs/THREAT-MODEL.md, docs/security/lua-sandbox.md, TLS cipher suite freeze
  • Release engineering: docs/versioning.md, 6 operator runbooks, CHANGELOG CI gate, user docs (getting-started, configuration, monitoring), release pipeline SHA256 checksums + SBOM + cosign

[0.1.3] — 2026-04-10

Production-readiness foundation: dispatch hot-path recovery, vector-search 4× QPS + correctness fixes, and the tiered disk-offload landing with 100 % crash recovery across 7 persistence configurations. Bundles three work streams originally tracked as separate Unreleased blocks (Apr 7–8).

Dispatch Hot-Path Recovery (2026-04-08)

Pipelined SET +37%, pipelined GET +68% at p=16 after PR #43 regression recovery.

Three targeted perf fixes landed after flamegraph-driven analysis of pipelined SET on aarch64 (OrbStack moon-dev, 1 shard, default config, redis-benchmark -c 50 -n 3M -P 16 -r 100000 -d 64):

Metric Broken baseline After T0a+T0b+T0c Δ
SET p=1 (ratio Redis) 0.99x 1.12x +13pp
SET p=16 1.42M/s 1.94M/s +37%
SET p=32 2.06M/s 2.26M/s +10%
GET p=16 2.40M/s 4.04M/s +68%
GET p=128 vs Redis 1.87x 1.91x +4pp

Perf fixes

  • T0a — Thread-local cached clock (4041b0d). Entry::new_* constructors were calling SystemTime::now() / clock_gettime on every write, showing up at 10.14% of CPU in the perf profile. Added a thread-local Cell<u32> / Cell<u64> refreshed once per shard tick (~1 ms) from CachedClock::update(). current_secs / current_time_ms now read the Cell and fall back to the syscall only on tests / cold init. __kernel_clock_gettime dropped from 10.14% → 0% of CPU.

  • T0b — Hot command dispatch bypasses phf SipHasher (4b0eec3). The command metadata registry is a phf::Map keyed by &'static str using SipHasher — cryptographic overkill for a 173-entry ASCII table. Combined phf::Map::get

  • SipHasher::write + hash_one was ~6% of CPU. Added a direct match path in command::metadata::lookup: pack the first ≤8 bytes of the command name as a u64 with ASCII letters uppercased, match against 24 hand-picked hot commands (GET/SET/DEL/TTL/MGET/MSET/INCR/DECR/HSET/HGET/HDEL/HLEN/LPOP/ RPOP/LLEN/PING/LPUSH/RPUSH/EXPIRE/EXISTS/INCRBY/DECRBY/SELECT/HGETALL). Hot-path resolves through a pre-resolved LazyLock<[&'static CommandMeta; 24]> — single array index, no hashing. Cold commands fall through to phf unchanged. Correctness asserted by hot_path_matches_phf_map test: every hot entry must return the same &'static pointer as a direct phf probe, in both upper and lowercase.

  • T0c — ACL unrestricted-user short-circuit (4603511). Every command executed check_command_permission + check_key_permission even for the default on nopass ~* &* +@all user, burning 2.11% of CPU on lowercasing, extract_command_keys, and glob matching. Added a cached unrestricted: bool field to AclUser, true iff the user is enabled, has AllAllowed commands, only ~* read/write key patterns, and only * channel patterns. The three check_*_permission methods early-return None on unrestricted before any allocation or iteration. The cache is recomputed once at the end of apply_rule (the single mutation entry point used by ACL SETUSER / LOAD / reset). Correctness covered by three new tests (default_user_is_unrestricted, restrictions_clear_unrestricted_flag, unrestricted_user_passes_all_checks).

Correctness fix (PR #43 review)

  • Inline monoio fast-path restricted to GET (613c164). The previous inline dispatch in try_inline_dispatch handled both GET and SET directly against the DashTable, bypassing replica READONLY enforcement, ACL checks, maxmemory eviction, client-side tracking invalidation, keyspace notifications, replication propagation, and blocking-waiter wakeups. Under any of those configurations the inlined SET would silently diverge from the normal path — accepted writes on replicas, ACL-denied clients writing, maxmemory overshoot, stale client-side caches. Fix: inline only handles *2\r\n$3\r\nGET now; SET and everything else fall through to the full dispatcher where all side-effects run.

Cold-tier lock hygiene (PR #43 review)

  • Release shard read guard before cold-tier disk read (ff51135). The cold-tier fallback in server::conn::blocking previously called get_cold_value() — which does a synchronous std::fs::read() — while still holding the per-shard read guard, blocking all concurrent operations on that shard during disk I/O. Split the path: Database::cold_lookup_location returns the (ColdLocation, PathBuf) under the lock, the guard is dropped, and cold_read::read_cold_entry_at performs the disk read unlocked.

Additional PR #43 fixes

  • read_overflow_chain now bounded at 1000 iterations (cycle guard against corrupted next_page links)
  • recovery.rs FPI replay replaces .unwrap() on try_into() with explicit byte-array construction (coding-guidelines compliance)
  • bench-production.sh: fixed unsupported -t zrangebyscore (→ zpopmin), MSET rps parser for "MSET (10 keys):" output, heredoc $(date) expansion, and Redis RSS probe (pgrep//proc instead of missing lsof)
  • bench-cold-tier.sh: removed stray & backgrounding FT.CREATE
  • test-recovery-all-cases.sh: NoPersistence case now PASSes at 0 keys
  • benches/resp_parsing.rs, benches/get_hotpath.rs: wrap Vec<Frame> in FrameVec via .into() after frame.rs type change

All 1872 unit tests pass under --no-default-features --features runtime-tokio,jemalloc. Follow-up work (T1 dispatch_raw zero-alloc entry point, Tier 2 storage/DashTable optimization, residual ACL SipHash elimination) captured as todo in .planning/todos/pending/.


Vector Search 4× QPS + Correctness (2026-04-07)

4x search QPS, 4.1x lower latency, 2.56x faster than Qdrant on real MiniLM data.

Performance (perf-profiled on GCloud c3-standard-8, Intel Xeon 8481C)

  • 8-wide ILP unrolled dist_bfs_budgeted subcent path (the real hot loop, 90% of search time per perf profile). Loads 4 code bytes + 1 sign byte per iteration, 8 independent f32 accumulators. Confirmed via objdump: parallel vaddss into xmm3-xmm8 (vs serial single-xmm0 chain before).
  • 4-way unrolled dist_bfs non-subcent path with unsafe pointer arithmetic
  • Pre-allocated ADC LUT in SearchScratch (eliminates 32-65KB heap alloc per query)
  • Hoisted IVF q_rotated and lut_buf allocation out of per-segment loop

Correctness fixes

  • FT.COMPACT silent no-op: split try_compact (threshold-gated) from force_compact (unconditional). Previously FT.COMPACT returned OK without compacting when compact_threshold >= mutable_len, leaving all vectors in brute-force O(n) mutable segment.
  • key_hash_to_key mapping restored (lost in earlier refactor). FT.SEARCH now returns original Redis keys (doc:N) instead of vec:<internal_id>. Carried through SearchResult.key_hash and populated by remap_to_global_ids.
  • FT.INFO num_docs now sums mutable + immutable segments (was 0 after compact)
  • Vector index recovery metadata loads without --disk-offload flag (was gated behind server_config.disk_offload_enabled())

Real MiniLM benchmarks (10K vectors, 384d, x86 Xeon 8481C)

Metric Mar 31 (M4 Pro) Apr 7 (Xeon 8481C) Δ
Recall@10 0.9250 0.9670 +4.5%
QPS 1,126 1,296 +15%
p50 0.878 ms 0.783 ms -11%
Moon Qdrant 1.12 FP32 Ratio
QPS (10K MiniLM) 1,296 507 2.56x
p50 0.783 ms 1.79 ms 2.29x lower
Recall@10 0.967 ~0.95 +1.7%

Infrastructure (for future segment merge work)

  • ImmutableSegment::decode_vector / iter_live_decoded
  • MutableSegment::iter_live

Attempted and reverted

Segment merge on FT.COMPACT via TQ4 decode → re-encode. Dropped recall from 0.73 → 0.0005 due to accumulated quantization error across 14 segments. Proper fix requires retaining f32/f16 vectors alongside TQ codes in immutable segments.

Known limitation

TQ4 quantization at 384d with random Gaussian inputs hits ~0.73 recall floor (curse of dimensionality — all points nearly equidistant). Real semantic embeddings (clustered) achieve 0.92-0.97 recall with the same code.


Disk Offload & x86_64 Performance (2026-04-06)

Tiered storage, crash recovery, and 2× Redis on x86_64 (Intel Xeon, io_uring).

Added

Disk Offload (Tiered Storage)

  • --disk-offload enable — evicted keys under maxmemory are spilled to NVMe instead of being deleted
  • Async SpillThread: background pwrite via dedicated std::thread per shard (no event loop blocking)
  • Cold read-through: GET transparently reads spilled keys from NVMe DataFiles
  • ColdIndex: in-memory key→file mapping, updated immediately on eviction for consistent reads
  • SpillThread channel capacity: 4096 bounded flume channel for burst absorption
  • --disk-offload-dir, --disk-offload-threshold configuration flags

Crash Recovery

  • V3 recovery falls back to appendonly.aof when WAL v3 has 0 commands
  • V2 recovery falls back to appendonly.aof when shard WAL has 0 commands
  • Automatic --dir creation before AOF writer starts (fixes silent write failure)
  • Cold index rebuilt from manifest during v3 recovery
  • Verified: 100% recovery (5000/5000 keys) across 7 persistence configurations after SIGKILL

Inline GET Optimization

  • read_db + get_if_alive replaces write_db + triple-lookup get() — single DashTable probe
  • Removed unnecessary write lock for timestamp refresh before inline dispatch
  • Multi-shard inline dispatch: local keys bypass Frame construction via key_to_shard() check
  • Cold storage fallback in get_readonly and inline GET dispatch paths

Changed

  • Connection handler eviction uses try_evict_if_needed_async_spill when disk offload enabled
  • spawn_monoio_connection passes spill sender, file ID counter, and offload dir to handlers
  • Event loop syncs next_file_id between Rc<Cell<u64>> (handlers) and local variable (timer tick)
  • Inline dispatch try_inline_dispatch takes now_ms and num_shards parameters

Fixed

  • Data loss under maxmemory: evicted keys were silently deleted instead of spilled to disk (6 bugs)
  • Crash recovery = 0 keys: appendonly.aof never tried as fallback source
  • AOF writer silent failure: --dir directory not created before AOF writer task started
  • Cold read miss: get_if_alive (read path) didn't check cold storage; get_readonly returned NULL for spilled keys
  • ColdIndex never initialized: cold_index and cold_shard_dir were None on all databases at startup

Performance (GCP c3-standard-8, Intel Xeon 8481C, CPU-pinned)

Metric Before After
c=1 p=1 GET vs Redis 0.35x (47K) 1.0x (47K) — parity
c=10 p=64 GET 2.29M 4.71M (2.06x Redis)
c=50 p=64 GET 2.36M 4.81M (2.04x Redis)
Disk offload GET overhead N/A <1% vs no-persist
Recovery (SIGKILL) 0/5000 5000/5000 (100%)

[0.1.2] - 2026-03-29

Multi-shard scaling milestone. Eliminated negative scaling, achieving 5M GET/s and 2.5M SET/s at 4 shards — both exceeding Redis 8.6.1.

Added

Shared-Read Direct Access (Phase 49)

  • Arc<ShardDatabases> with parking_lot::RwLock<Database> replaces Rc<RefCell<Vec<Database>>>
  • Cross-shard read commands (GET, HGET, SCARD, ZRANGE, etc.) bypass SPSC channels entirely via read_db() + dispatch_read() — reduces cross-shard read latency from ~88μs to ~56ns
  • Local read path uses shared read_db() lock instead of exclusive write_db() — eliminates RwLock contention between shards

Connection Affinity (Phase 50)

  • AffinityTracker samples first 16 commands per connection to detect dominant shard
  • Lazy FD migration: if ≥60% of keys target a non-local shard, migrates the TCP connection's file descriptor to the target shard via ShardMessage::MigrateConnection
  • MigratedConnectionState preserves selected_db, client_name, protocol_version across migration
  • Graceful fallback: if migration fails, connection stays on current shard with shared-read

Pre-Allocated Response Slots (Phase 51)

  • ResponseSlotPool with lock-free AtomicU8 state machine for zero-allocation cross-shard write dispatch (Tokio path)
  • Eliminates per-dispatch channel::oneshot() heap allocation (~80-120ns savings per cross-shard write)

SO_REUSEPORT Per-Shard Accept (Phase 52)

  • Each shard opens its own TCP listener with SO_REUSEPORT on Linux via socket2 crate
  • Kernel distributes connections across shard listeners using consistent 4-tuple hashing
  • macOS/non-Linux: falls back to single-listener + MPSC round-robin (no behavior change)

jemalloc Production Tuning (Phase 53)

  • malloc_conf static: percpu_arena:percpu, background_thread:true, metadata_thp:auto, dirty_decay_ms:5000, muzzy_decay_ms:30000, abort_conf:true
  • Closes ~50% of allocation speed gap with mimalloc while retaining jemalloc's superior fragmentation behavior

New Commands (Phase 55)

  • GETRANGE — return substring of stored string value
  • SETRANGE — overwrite part of stored string at offset with zero-fill
  • SUBSTR — alias for GETRANGE (Redis 1.x compatibility)

Changed

  • Custom AtomicU8 oneshot channel replaced with flume::bounded(1) for cross-thread safety on monoio's !Send executor
  • pending_wakers relay pattern: event loop locally wakes connection tasks after SPSC processing, bridging monoio's cross-thread waker limitation
  • write_db() uses try_write() spin loop instead of blocking write() — prevents OS thread freeze on monoio when cross-shard readers hold locks
  • Benchmark scripts: scripts/bench-scaling.sh for multi-shard test matrix, scripts/bench-production.sh updated

Fixed

  • ResponseSlot UnsafeCell<Option<Waker>> data race on ARM64 — replaced with AtomicWaker
  • Local read path took exclusive write lock (write_db()) even for GET — split into read_db() + dispatch_read()
  • Monoio local write path silently dropped responses (responses.push(response) missing after read/write split) — all write commands (SET, INCR, LPUSH, etc.) hung on monoio
  • Pipeline ordering guard: !remote_groups.contains_key(&target) prevents stale reads when batch has pending writes for same shard

Performance

Metric Before (v0.1.0) After (v0.1.2) Change
Multi-shard GET p=16 688K (0.38x Redis) 1,923K (1.17x Redis) 2.8x
Multi-shard GET p=64 N/A 5,002K (1.60x Redis) New
Multi-shard SET p=16 N/A 1,515K (1.32x Redis) New
Multi-shard SET p=64 N/A 2,500K (1.55x Redis) New
Monoio 1s p=128 GET 5,407K 5,005K (1.25x Redis) Maintained
Negative scaling -25% at 12 shards Zero at 1-8 shards Eliminated
Command coverage p=1 Parity Monoio beats Redis 8/10 Improved

[0.1.1] - 2026-03-28

Structural stability milestone. Codebase refactoring for maintainability — no feature changes, no performance changes.

Changed

Error and State Foundations (Phase 44)

  • Unified MoonError type hierarchy with structured #[source] on I/O variants carrying PathBuf context
  • ConnectionContext struct for connection state (selected_db, authenticated, client_name, protocol_version)
  • Criterion benchmark baseline (GET dispatch 69.1ns) to guard against regressions

Command Metadata Registry (Phase 45)

  • phf static perfect hash map for O(1) command lookup (112 commands)
  • CommandMeta struct: name, arity, flags (read/write/fast/admin), key positions, ACL categories
  • is_write() classification via const bitflags — replaces duplicated match arms across codebase

Persistence Hardening (Phase 46)

  • Eliminated server-crashing unwrap() calls in WAL, AOF, and RDB persistence code
  • Corruption recovery: WAL uses per-block CRC32 log+skip, AOF seeks to next RESP * marker, RDB breaks on mid-stream corruption
  • WalWriter methods remain std::io::Result (must-panic on flush = data loss prevention)

AOF Replay Decoupling (Phase 47)

  • CommandReplayEngine trait breaks circular dependency between persistence and command dispatch
  • StorageEngine trait boundary for persistence replay and Lua scripting
  • execute_command() at command level (not individual get/set methods)

God-File Decomposition (Phase 48)

  • connection.rs (5,102 lines) → 6 sub-modules in conn/: handler_sharded.rs, handler_monoio.rs, handler_single.rs, shared.rs, blocking.rs, conn_state.rs
  • shard/mod.rs (2,004 lines) → 6 sub-modules: event_loop.rs, spsc_handler.rs, persistence_tick.rs, conn_accept.rs, timers.rs, uring_handler.rs
  • Module facade pattern with pub(crate) re-exports preserving all external import paths
  • No single file exceeds 800 lines

Added

  • Docker: optimized multi-stage build (113MB → 41MB)
  • Mintlify documentation site
  • Claude Code GitHub workflow for PR reviews
  • scripts/bench-resources.sh for memory/CPU efficiency benchmarking

0.1.0 - 2026-03-27

Initial release. A Redis-compatible in-memory data store written in Rust, achieving 1.84-1.99x Redis throughput at 8 shards and 27-35% less memory for 1KB+ values.

Added

Core Data Types (Phases 1-5)

  • RESP2 protocol parser and serializer with inline command support
  • TCP server with concurrent connections, graceful shutdown, and redis-cli compatibility
  • String commands: GET, SET, MGET, MSET, INCR/DECR, APPEND, GETEX, GETDEL (17 commands)
  • Hash commands: HSET, HGET, HGETALL, HINCRBY, HSCAN (14 commands)
  • List commands: LPUSH, RPUSH, LPOP, RPOP, LRANGE, LPOS (12 commands)
  • Set commands: SADD, SREM, SINTER, SUNION, SDIFF, SPOP (15 commands)
  • Sorted Set commands: ZADD, ZRANGE, ZRANGEBYSCORE, ZINCRBY, ZPOPMIN (18 commands)
  • Key management: DEL, EXISTS, EXPIRE, TTL, SCAN, KEYS, RENAME (13 commands)
  • Lazy + active key expiration with probabilistic sampling
  • RDB persistence with point-in-time snapshots
  • AOF persistence with configurable fsync (always/everysec/no)
  • Pub/Sub messaging: SUBSCRIBE, PUBLISH, PSUBSCRIBE (4 commands)
  • Transactions: MULTI/EXEC/DISCARD with WATCH optimistic locking
  • LRU/LFU/random eviction policies with configurable maxmemory

Performance Architecture (Phases 6-15)

  • SIMD-accelerated RESP parsing via memchr CRLF scanning and atoi
  • CompactValue 16-byte SSO struct with embedded TTL delta
  • DashTable segmented hash table with Swiss Table SIMD probing
  • Thread-per-core shared-nothing architecture with per-shard event loops
  • io_uring networking layer with multishot accept/recv and registered buffers
  • Per-shard memory management with jemalloc and bumpalo arenas
  • Forkless compartmentalized persistence (no COW memory spike)
  • B+ tree sorted sets replacing BTreeMap for cache-friendly access
  • Per-connection arena allocation with bumpalo

Protocol & Data Types (Phases 16-18)

  • RESP3 protocol: Map, Set, Double, Boolean, VerbatimString, Push frames
  • HELLO command for protocol negotiation
  • Client-side caching invalidation via Push frames
  • Blocking commands: BLPOP, BRPOP, BLMOVE, BZPOPMIN, BZPOPMAX
  • Streams data type: XADD, XREAD, XRANGE, XGROUP, XREADGROUP, XACK, XPENDING, XCLAIM, XAUTOCLAIM

Clustering & Replication (Phases 19-20, 26)

  • PSYNC2-compatible replication with per-shard WAL streaming
  • Partial resync support with replication backlog
  • Cluster mode with 16,384 hash slots and gossip protocol
  • MOVED/ASK redirections and live slot migration
  • Majority consensus failover election with automatic promotion

Scripting & Security (Phases 21-22, 43)

  • Lua 5.4 scripting via mlua: EVAL, EVALSHA, SCRIPT LOAD/EXISTS/FLUSH
  • Sandboxed Lua VM with Redis API bindings (redis.call, redis.pcall)
  • ACL system: per-user command/key/channel permissions
  • ACL SETUSER, GETUSER, DELUSER, LIST, WHOAMI, LOG, SAVE, LOAD
  • TLS 1.3 via rustls + aws-lc-rs with dual-port support
  • mTLS client authentication
  • Protected mode (reject non-loopback when no password set)

Optimization (Phases 24-42)

  • WAL v2 format: checksums, header, block framing, corruption isolation
  • CompactKey SSO: 23-byte inline keys, eliminating heap allocation
  • Response buffer pooling and adaptive pipeline batching
  • Dual runtime: Tokio (all platforms) + Monoio (Linux io_uring / macOS kqueue)
  • Full Monoio migration: channels, TCP, codec, spawn, persistence, replication
  • Direct GET serialization bypassing Frame allocation
  • Zero-copy argument slicing from parse buffer
  • Lock-free oneshot channels (12% CPU reduction vs tokio::oneshot)
  • CachedClock timestamp caching (4% throughput gain)
  • HeapString for values (eliminates Arc overhead)
  • Inline dispatch for single-shard commands

Performance

Benchmark Result
Peak GET throughput 3.79M ops/sec (4 shards, p=64)
Peak SET with AOF 2.78M ops/sec (AOF everysec, p=64)
vs Redis (pipeline=64) 3.17x SET, 2.50x GET
vs Redis (8 shards, p=16) 1.84-1.99x
vs Redis with AOF 2.75x (per-shard WAL vs global)
Memory (1KB+ values) 27-35% less than Redis
Memory (empty server) Identical 7.0 MB baseline
p50 latency (8 shards) 0.031ms (Redis: 0.26ms)
Data consistency 132/132 tests pass

Technical Details

  • Language: Rust (stable, edition 2024)
  • Lines of code: ~54,000 across 96 files
  • Dependencies: tokio, monoio, jemalloc, rustls, mlua, bumpalo, bytes, clap
  • Supported platforms: Linux (io_uring via Monoio), macOS (kqueue via Monoio or Tokio)
  • Build time: ~50s release build