Skip to content
Merged
Show file tree
Hide file tree
Changes from 2 commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
42 changes: 42 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,48 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

## [0.7.0] — 2026-07-15

**Replication GA for multi-shard masters.** Moon now supports real Redis-compatible
asynchronous replication with `WAIT`/`ACK` acknowledgement semantics: a multi-shard
master (`--shards N`) streams a merged, exactly-once command feed to a single-shard
streaming replica across all data planes — KV, vector/text index, graph, workspace (WS),
message-queue (MQ), and temporal. Every write plane is crash-durable and survives
`kill -9` on either side, validated by a 24h continuous-load kill-9 soak (alternating
master/replica restarts every 12 min) that asserts zero loss of any `WAIT`-acknowledged
write. Also folds in the full v0.6.1 hardening scope (WAL v3 storage-kernel M1–M4:
cross-plane crash matrix, unified per-shard WAL-recycle floor, atomic durable writes,
FTS term-dict durability) and a supply-chain CI gate (`cargo audit` + `cargo deny`).

Soak evidence (release gate REPL-SOAK-01): `SOAK-PASS duration=86400s cycles=114
acked=82044 inflight=7 master_kills=57 replica_kills=57` — 82,044 WAIT-acked writes
preserved across 114 alternating kill-9 cycles, zero acked-write loss (2026-07-15,
run dir `moon-soak/runs/20260714-141946`, RC `e2d87893`).
Comment on lines +22 to +25

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift

The soak is documented as a tag gate, but the supplied release workflow does not enforce it for v0.7.0.

  • CHANGELOG.md#L22-L25: either wire REPL-SOAK-01 into an enforcing release step or describe it as completed/manual evidence.
  • docs/PRODUCTION-CONTRACT.md#L129-L129: remove the claim that the row gates the v0.7.0 tag unless the workflow is changed to block that tag.
📍 Affects 2 files
  • CHANGELOG.md#L22-L25 (this comment)
  • docs/PRODUCTION-CONTRACT.md#L129-L129
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@CHANGELOG.md` around lines 22 - 25, Update CHANGELOG.md lines 22-25 to
describe REPL-SOAK-01 as completed/manual evidence, unless the release workflow
is changed to enforce it for v0.7.0. Update docs/PRODUCTION-CONTRACT.md line 129
to remove the claim that this evidence gates the v0.7.0 tag unless such
enforcement is added.


Known limitation: the streaming replica is single-shard only (`--shards 1`) — the
multi-shard work in this release is master-side (merged N-shard PSYNC feed). Multi-shard
replicas are roadmapped for v0.8/v0.9.

Replica TTL semantics (disclosure): relative-expire commands (`EXPIRE`, `SETEX`, `PEXPIRE`,
`GETEX` with a relative TTL) currently replicate verbatim rather than being rewritten to
absolute `PEXPIREAT` on the master, and replicas run their own active-expiry cycle
regardless of role. In practice keys expire correctly on both sides under normal clock
sync, but a master/replica clock skew can shift a relative-TTL key's expiry moment between
the two by up to that skew. Absolute-expiry rewrite + role-gated passive expiry land in
v0.7.1 (task #71b). Applications needing exact cross-node expiry parity should set
absolute deadlines with `PEXPIREAT` until then.

### Documentation — replication configuration & tuning

Documented the v0.7 replication surface: corrected the stale `guides/clustering.md`
replica walkthrough (it still showed the pre-GA topology — single-shard leader,
multi-shard replica — the reverse of the shipped shape) and its outdated
"WS/MQ not replicated yet" note (all six planes replicate as of Wave B). Added the
`WAIT`-durability ladder, replica promotion (`REPLICAOF NO ONE`), `INFO replication`
monitoring, and the replica-TTL caveat. New **Replication** section in
`configuration.md` and a **Replication durability** section in `guides/tuning.md`
(RPO/latency trade-offs, read-scaling, lag alerting).

### Fixed — `segment_plane_scan` missed v0.6.0's nested-Command plane framing, risking WS/MQ/temporal data loss on upgrade (task #69)

`WalWriterV3::recycle_aggressive`/`recycle_segments_before` gate deletion of
Expand Down
2 changes: 1 addition & 1 deletion Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion Cargo.toml
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
[package]
name = "moon"
version = "0.6.0"
version = "0.7.0"
edition = "2024"
rust-version = "1.94"
description = "A high-performance Redis-compatible server written in Rust"
Expand Down
5 changes: 5 additions & 0 deletions RELEASES.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,10 @@
# Releases

## v0.7.0 — 2026-07-15
milestones: v0.7.0 "Replication GA for multi-shard masters" (soak-gated tag; v0.6.1 hardening folded in)
waivers: none — the v0.6.0 `shardslice-migration` waiver (was: expires 2026-08-01) is **retired**; lock-free cross-shard read work is now tracked as an open GA gap (`XSHARD-READ-01`, ROADMAP R4), not a time-boxed waiver.
Comment on lines +3 to +5

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Synchronize the waiver status with the production-contract ledger.

This release entry says the shardslice-migration waiver is retired, but docs/PRODUCTION-CONTRACT.md Line 132 still says it expires on August 1, 2026 and remains unstarted. Update one ledger so both documents consistently represent whether XSHARD-READ-01 is a retired waiver or an active open obligation.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@RELEASES.md` around lines 3 - 5, Synchronize the production-contract ledger
with the v0.7.0 release entry: update the XSHARD-READ-01/shardslice-migration
status in the ledger so it no longer shows an active August 1, 2026 waiver and
instead matches the release’s retired-waiver, open-GA-gap state.

evidence: Real async replication with WAIT/ACK across all six data planes (KV, vector/text, graph, WS, MQ, temporal); multi-shard master → single-shard streaming replica; every plane crash-durable under kill -9. Plus full v0.6.1 hardening (WAL v3 storage-kernel M1–M4) and a supply-chain CI gate (cargo audit + cargo deny). Release gate — 24h replication kill-9 soak (REPL-SOAK-01), zero acked-write loss: `SOAK-PASS duration=86400s cycles=114 acked=82044 inflight=7 master_kills=57 replica_kills=57` (2026-07-15; run dir moon-dev:~/moon-soak/runs/20260714-141946; RC main e2d87893; 114 alternating master/replica kill -9 cycles, 82,044 WAIT-acked writes preserved). Known limitation: streaming replica is single-shard only. Replica relative-TTL semantics: verbatim replication (no PEXPIREAT rewrite) + role-agnostic active expiry — absolute-rewrite lands in v0.7.1 (task #71b).

## v0.6.0 — 2026-07-08
milestones: v0.6.0-release (PR #249, six workstreams)
waivers: shardslice-migration (expires 2026-08-01; retirement planned with v0.7 L4 cross-shard-read work)
Expand Down
4 changes: 2 additions & 2 deletions docs/PRODUCTION-CONTRACT.md
Original file line number Diff line number Diff line change
Expand Up @@ -89,7 +89,7 @@ per the CI gate · Blocking `—` = tracked but never blocks a tag (see Out of S
| ✅ | CHANGELOG-01 | CHANGELOG.md CI gate on every PR (skip-changelog label escape hatch) | `ci.yml` step `CHANGELOG check` | GA |
| ✅ | REL-LEDGER-01 | Release tag requires a matching `RELEASES.md` entry | `release.yml` step `Require RELEASES.md entry for this tag` | GA |
| ✅ | CONTRACT-01 | This document is a checked ledger with a CI-wired gate | `scripts/check-production-contract.sh`, `release.yml` | GA |
| | SUPPLY-01 | `cargo audit` + `cargo deny` CI-blocking | `deny.toml` exists but no CI job runs the tools. Dedicated `supply-chain.yml` workflow in flight (task #63, 2026-07-14) — tick when merged and green. | GA |
| | SUPPLY-01 | `cargo audit` + `cargo deny` CI-blocking | Shipped (task #63, PR #326): `.github/workflows/supply-chain.yml` runs `cargo audit` + `cargo deny` on push/PR with `deny.toml` policy. Folded into v0.7.0. | GA |

### B. Correctness Hardening

Expand Down Expand Up @@ -126,7 +126,7 @@ per the CI gate · Blocking `—` = tracked but never blocks a tag (see Out of S
| ✅ | REPL-MULTISHARD-01 | Multi-shard master replication (a `--shards N>1` master can be replicated at all) | R2 (task #20): `ShardMessage::PrepareReplicaSync` per-shard atomic snapshot legs + merged Redis-format RDB + per-record SELECT framing on the merged wire. monoio only; replicas run `--shards 1`; partial resync degrades to full at N>1. `tests/replication_multishard.rs` (2/4/8-shard resync, interleaved multi-db parity, graph, partial→full). | GA |
| ✅ | WAIT-01 | `WAIT` reflects real replica ACK state | R1 (task #19, PR #282): replica 1s `REPLCONF ACK` ticker on the split PSYNC socket; master `ack_read_loop` + `drain_ack_offsets` record into `ReplicaInfo.ack_offsets`; connection-layer `try_handle_wait` blocks until ACK ≥ target or timeout. `wait_returns_acked_replica_count` e2e; exact on multi-shard masters too (summed snapshot offset). | GA |
| ✅ | REPL-PLANES-01 | Every write plane replicates, not just KV: eviction/expiry DELs, Lua effects, graph, vector/text index defs+contents, WS.*, MQ.*, TEMPORAL.* | Wave A (PR #285): eviction/expiry DELs + Lua effects to both planes (EVAL was previously durable in neither). Wave B (PR #294 + task #34): WS/MQ deterministic records + replica apply + PSYNC registry blob. Graph plane (task #25): live GRAPH.* streaming + snapshot backfill. Suites: `tests/replication_planes.rs`, `replication_graph.rs`, `replication_mq.rs`, `replication_readonly_ws_mq.rs`. Unified poison-record policy for replica apply (task #48). | GA |
| | REPL-SOAK-01 | 24h replication soak: kill -9 either side under WAIT-confirmed load, zero acked-write loss | Harness in flight (task #61: `scripts/soak-replication-24h.sh`); the run (task #62) gates the v0.7.0 tag. | GA |
| | REPL-SOAK-01 | 24h replication soak: kill -9 either side under WAIT-confirmed load, zero acked-write loss | **PASSED 2026-07-15** — `SOAK-PASS duration=86400s cycles=114 acked=82044 inflight=7 master_kills=57 replica_kills=57`; 82,044 WAIT-acked writes preserved across 114 alternating kill-9 cycles, zero acked-loss. Run dir `moon-soak/runs/20260714-141946`, RC `e2d87893`. Gates the v0.7.0 tag (task #65). | GA |
| ⬜ | KEYSPACE-NOTIF-01 | `notify-keyspace-events` keyspace notifications | No implementation found in `src/`. ROADMAP v0.7.0 workstream R5 — deferred to v0.7.1 (one-headline rule). | GA |
| ⬜ | MONITOR-01 | `MONITOR` command | No implementation found in `src/command/`. ROADMAP v0.7.0 workstream R5. | GA |
| ⬜ | XSHARD-READ-01 | Lock-free cross-shard read path (retire the shardslice waiver) | Waiver **expires 2026-08-01** per `RELEASES.md` v0.6.0 entry and ROADMAP §5; L4 redesign (`tmp/MULTISHARD-REDESIGN.md`) unstarted. ROADMAP v0.7.0 workstream R4. | GA |
Expand Down
19 changes: 19 additions & 0 deletions docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,25 @@ All options are available as command-line flags. Run `moon --help` for the full
| `--cluster-enabled` | `false` | Enable cluster mode |
| `--cluster-node-timeout` | `15000` | Node timeout in ms |

## Replication

Replication (v0.7 GA) is initiated at runtime with the `REPLICAOF <host> <port>`
command — there is **no startup flag**. The relevant startup flags shape the
topology and durability of the pair:

| Flag | On | Effect for replication |
|------|-----|------------------------|
| `--shards N` | master | Multi-core writer; the master merges all shards into one exactly-once replication feed |
| `--shards 1` | replica | **Required** — replicas are single-shard; scale reads by adding replicas |
| `--appendonly yes` | both | Persist the AOF so a restarted node recovers before re-syncing |
| `--appendfsync always` | master | RPO 0 on the master; pair with `WAIT N` for zero-RPO cross-node durability |

Replicas are read-only (`slave_read_only:1`; writes return `-READONLY`). `WAIT
numreplicas timeout` reports real replica ACKs. Full setup, `WAIT` durability,
promotion (`REPLICAOF NO ONE`), and the replica TTL caveat: see the
**[clustering & replication guide](guides/clustering.md#replication)** and the
[tuning guide](guides/tuning.md#replication-durability).

## ACL

| Flag | Default | Description |
Expand Down
70 changes: 58 additions & 12 deletions docs/guides/clustering.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,42 +30,88 @@ Moon implements PSYNC2-compatible replication with per-shard WAL streaming and p
- `WAIT N timeout` reflects real replica ACKs (1s `REPLCONF ACK` cadence).
- `master_link_status` in `INFO replication` reflects the handshake state — use it to detect a failed REPLICAOF.
- `CLIENT LIST TYPE replica` has no predicate yet; returns all clients.
- WS.\*/MQ.\* planes are **not replicated yet** (the master logs one warning
when a replica is attached); vector/text/graph planes replicate fully.
- **All six data planes replicate** (v0.7 GA): KV, vector/text index, graph,
workspace (WS.\*), message-queue (MQ.\*), and temporal. A full resync ships
each plane's snapshot; the live stream carries every plane's effect records.

### Set up a replica

### Start the leader (must be --shards 1 in v0.1.x)
**Topology (v0.7):** the **master** runs any `--shards N` (multi-core writer); each
**replica** runs `--shards 1`. Scale reads by adding replicas, not replica shards.
Replication is initiated at runtime with the `REPLICAOF` command — there is no
startup flag; operators script it after the replica is up (e.g. via an init hook).

#### 1. Start the master (any shard count)

```bash
./target/release/moon --port 6379 --shards 1
./target/release/moon --port 6379 --shards 4 --appendonly yes --appendfsync always
```

### Start the replica (any shard count)
#### 2. Start the replica (must be --shards 1)

```bash
./target/release/moon --port 6380 --shards 4
./target/release/moon --port 6380 --shards 1 --appendonly yes
```
Comment thread
coderabbitai[bot] marked this conversation as resolved.

### Connect the replica to the leader
#### 3. Attach the replica to the master

```bash
redis-cli -p 6380 REPLICAOF 127.0.0.1 6379
```

### Verify link status
The replica performs a full resync (one merged Redis-format RDB across all planes),
then applies the live stream. Replicas are **read-only**: writes return
`-READONLY` (`INFO replication` reports `slave_read_only:1`).
Comment on lines +68 to +70

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Do not imply standard Redis replica interoperability.

The supplied architecture contract describes the full-resync payload as RRDSHARD and says conversion to standard Redis RDB for heterogeneous Redis replicas is out of scope. Describe the Moon-specific payload and limit the compatibility claim to the PSYNC2 control protocol.

Also applies to: 100-103

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/guides/clustering.md` around lines 62 - 64, Update the clustering
guide’s replica synchronization description to identify the full-resync payload
as Moon-specific RRDSHARD, not a standard Redis-format RDB. Remove any
implication of interoperability with heterogeneous Redis replicas, and limit
compatibility claims to the PSYNC2 control protocol while preserving the
read-only and live-stream behavior.


#### 4. Verify the link

```bash
redis-cli -p 6380 INFO replication | grep master_link_status
# Expect: master_link_status:up
# Expect: master_link_status:up (anything else = handshake not complete)
```

### Acknowledged writes with WAIT

`WAIT numreplicas timeout` blocks until at least `numreplicas` replicas have ACKed
every write issued on the connection (replicas send `REPLCONF ACK` on a ~1 s
cadence). Combine it with `appendfsync always` on both sides for a **zero-RPO,
cross-node durable write**:

```bash
redis-cli -p 6379 SET k v
redis-cli -p 6379 WAIT 1 1000 # returns the number of replicas that ACKed within 1000 ms
```
Comment thread
coderabbitai[bot] marked this conversation as resolved.

On the master, `INFO replication` lists each replica's `offset` and `lag`; on the
replica it reports `slave_repl_offset`. See the [tuning guide](tuning.md#replication-durability)
for the durability/latency trade-offs.

### Promote a replica (failover)

```bash
redis-cli -p 6380 REPLICAOF NO ONE # replica stops applying and becomes a writable master
```

Repoint surviving replicas at the new master with `REPLICAOF <new-host> <new-port>`.
There is no automatic replication failover outside cluster mode (see below).

### Replication features

- **All six data planes** — KV, vector/text index, graph, WS, MQ, temporal (v0.7 GA)
- **PSYNC2 protocol** — compatible with Redis replication clients
- **Per-shard WAL streaming** — each shard streams its own WAL independently (once connected)
- **Partial resync** — reconnecting replicas resume from where they left off via the replication backlog
- **Lazy backlog** — the replication backlog is only allocated when the first replica handshake begins (REPLCONF), saving ~12 MB baseline memory
- **Per-shard WAL streaming** — each master shard streams its own WAL; a multi-shard master merges them into one exactly-once feed
- **Partial resync** — single-shard masters resume reconnecting replicas from the backlog window; multi-shard masters answer every reconnect with a full resync
- **Lazy backlog** — allocated only when the first replica handshake begins (REPLCONF), saving ~12 MB baseline memory
- **Validated** — 24 h continuous-load kill-9 soak (alternating master/replica restarts), zero loss of any WAIT-acknowledged write

!!! warning "Replica TTL semantics (v0.7)"
Relative-expire commands (`EXPIRE`, `SETEX`, `PEXPIRE`, `GETEX` with a relative
TTL) replicate verbatim rather than being rewritten to absolute `PEXPIREAT`, and
replicas run their own active-expiry cycle. Under normal clock sync keys expire
correctly on both sides, but master/replica clock skew can shift a relative-TTL
key's expiry moment by up to that skew. For exact cross-node expiry parity, set
absolute deadlines with `PEXPIREAT`. Absolute-rewrite + role-gated expiry land in
v0.7.1.
Comment thread
coderabbitai[bot] marked this conversation as resolved.

## Cluster mode

Expand Down
28 changes: 28 additions & 0 deletions docs/guides/tuning.md
Original file line number Diff line number Diff line change
Expand Up @@ -157,6 +157,34 @@ data loss on a host crash; then use `always` and size expectations to your disk'
rate (or provision faster storage). Both preserve their guarantee under SIGKILL —
validated by the crash-recovery matrix (100% of acked writes recovered).

## Replication durability

AOF durability protects a **single node** against a crash. Replication (v0.7 GA)
adds a **second node** so a write survives losing the master's disk entirely. The
two combine on a latency/RPO ladder — pick the rung your workload needs:

| Goal | Master | Client | RPO | Cost |
|------|--------|--------|-----|------|
| Fast, replica for read-scaling/DR | `--appendfsync everysec` | fire-and-forget | ≤1 s on master crash; replica lag on failover | lowest latency |
| Durable on the master | `--appendfsync always` | — | 0 on master crash (disk-bound) | fsync per write |
| **Zero-RPO across nodes** | `--appendfsync always` | `WAIT 1 <timeout>` after the write | 0 even if the master's disk is lost | fsync + one replica round-trip |

`WAIT numreplicas timeout` blocks until `numreplicas` replicas have ACKed the
write (replicas ACK on a ~1 s cadence, so a `WAIT` timeout below ~1 s may return
before an idle replica reports in — size the timeout accordingly, or keep the
write stream busy). It reports the count that ACKed; it never rolls the write back.

**Read scaling:** replicas are `--shards 1` and read-only. Add *more replicas* for
read throughput and DR breadth — you cannot add shards to a replica. Route reads
to replicas and writes to the master at the client/proxy layer.

**Monitoring:** on the master, `INFO replication` lists each replica's `offset`
and `lag` (bytes behind); alert on sustained lag growth. On a replica,
`master_link_status:up` confirms a live stream — treat `down` as an outage even if
TCP is connected (it means the PSYNC handshake has not completed).

Full setup and failover: [clustering & replication guide](clustering.md#replication).

## Pub/sub fan-out

Moon's subscriber delivery path **coalesces automatically** — when a publish burst queues
Expand Down
Loading