From 6cd3c6e078ed3ad6be4c5b44fe3966a6b11bdd6d Mon Sep 17 00:00:00 2001 From: Brylie Christopher Oxley Date: Sun, 30 Aug 2026 15:48:25 +0300 Subject: [PATCH 1/3] docs: propose workspace sharding architecture --- docs/specifications/README.md | 1 + docs/specifications/workspace-sharding.md | 220 ++++++++++++++++++++++ 2 files changed, 221 insertions(+) create mode 100644 docs/specifications/workspace-sharding.md diff --git a/docs/specifications/README.md b/docs/specifications/README.md index 967b9f4..80006df 100644 --- a/docs/specifications/README.md +++ b/docs/specifications/README.md @@ -3,6 +3,7 @@ Canonical specification for each implemented subsystem, translating [`prd.md`](../prd.md) into concrete engineering decisions. Split by feature rather than kept as one document, so a change to one subsystem doesn't require editing (or risk going stale in) an unrelated one. - [`architecture.md`](./architecture.md) — process model (one Node process, `/ws` + `/mcp`) and SvelteKit app structure. +- [`workspace-sharding.md`](./workspace-sharding.md) — proposed Workspace catalog, content-shard, and committed-catalog-write/SSE design for #112; pending #31 measurement before approval. - [`data-model.md`](./data-model.md) — the Record/Document/Collection types and their Yjs mapping. - [`collaboration.md`](./collaboration.md) — presence and holds, built on Yjs Awareness. - [`mcp-tools.md`](./mcp-tools.md) — the MCP tool surface and permission scoping. diff --git a/docs/specifications/workspace-sharding.md b/docs/specifications/workspace-sharding.md new file mode 100644 index 0000000..7111dc5 --- /dev/null +++ b/docs/specifications/workspace-sharding.md @@ -0,0 +1,220 @@ +# Workspace catalog and CRDT sharding + +**Status:** Proposed — requires the capacity baseline from #31 before approval. + +**Depends on:** [`prd.md`](../prd.md), [`architecture.md`](./architecture.md), +[`persistence.md`](./persistence.md), [`collaboration.md`](./collaboration.md), +and [`mcp-tools.md`](./mcp-tools.md) + +**Tracked by:** #112 (design), #113 (implementation), #114 (migration and +isolation), and #6 (multi-space experience) + +## 1. Decision to validate + +Compendium uses CRDTs where several clients concurrently edit rich content, +not as the universal transport for every reactive interface element. + +- A server-owned **workspace catalog** is durable SQLite state. It owns Spaces, + document and Collection titles, hierarchy, membership, access scopes, and the + current catalog revision. +- A **Document shard** is one Y.Doc containing one Document's ordered blocks + and rich text. +- A **Collection shard** is one Y.Doc containing one Collection's schema and + rows. It is the initial unit; #31 determines whether a later row partition is + needed and under which measured conditions. +- The browser uses a small server-sent-events feed for catalog changes. SSE + delivers invalidations; it never becomes a second source of truth or a + replacement for collaborative content sync. +- Yjs WebSocket and Awareness connections are scoped to the explicitly opened + Document or Collection shard. + +This is intentionally a proposed decision, not an implementation commitment. +#31 must establish the Phase-0 baseline and identify whether a one-Y.Doc-per +Document/Collection design meets the expected operating envelope. The approval +record for #112 must state which assumptions survived that measurement. + +## 2. Terms and boundaries + +| Term | Meaning | Authority | +| ------------------- | --------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------- | +| Deployment instance | One locally configured Compendium process and database target. It separates dogfooding, development, and test environments. | Server configuration; see #111. | +| Workspace | The top-level durable knowledge environment within a deployment instance. | Server-authorized request context. | +| Space | An organizational and permission boundary within one Workspace. A person switches Spaces to work on different projects. | Workspace catalog. | +| Catalog | The SQLite metadata projection needed to route, list, search, and authorize content without opening every content shard. | Catalog service and committed database transactions. | +| Content shard | The Y.Doc for one Document or Collection. | Authorized Yjs and service-layer mutations. | + +A deployment instance is not a Workspace, and a Space is not a tenant. This +keeps local operational isolation (#111), single-user multi-space organization +(#6), and eventual multi-user authentication (#3) independently evolvable. + +## 3. Ownership model + +### 3.1 Workspace catalog + +The catalog is the authoritative location for the metadata required to navigate +and authorize a Workspace without subscribing to its content shards: + +- Workspace and Space identities, titles, ordering, and membership. +- Document and Collection identities, titles, owning Space, hierarchy/order, + and shard locator. +- Space-level access grants and the derived scopes assigned to access tokens. +- Catalog revision and committed event records. + +Document titles and hierarchy are catalog fields. They are not duplicated into a +Document Y.Doc and derived later, because the sidebar needs them before opening +the Document. Record IDs remain stable global identities; a title is never an +identity or authorization key. + +The existing `snapshots`, `access_tokens`, and `audit_log` storage must become +explicitly scoped by Workspace and, where applicable, Space and shard. A +catalog migration must preserve existing IDs and create one default Space for +the current single-space workspace. + +### 3.2 Document and Collection shards + +A Document shard owns its block records, their order, Y.Text content, and the +rich-content collaboration state for that Document. A Collection shard owns its +schema, rows, and Collection-view source data. A content shard does not own its +title, parent hierarchy, or Space membership. + +Relations, page links, and collection-view references remain ID-backed. A +cross-Space reference is valid only when the caller's permission scope permits +both sides; rendering and search must not use a catalog lookup to reveal an +inaccessible target's title. + +### 3.3 Ephemeral coordination + +Awareness belongs to the active content shard: cursors, active editing state, +and visual agent placeholders must not broadcast to clients outside that shard. + +The server retains a workspace-scoped, ephemeral hold coordinator keyed by +`workspaceId` and record ID. It aggregates the active shard Awareness state +needed to decide an MCP `hold_records` request, while projects only the result +into shards whose authorized clients need to render it. This preserves a +cross-document agent batch without leaking cursor or presence state to an +unrelated Space. + +## 4. Committed catalog writes and SSE + +Catalog mutations use one service-layer transaction: + +```text +authorize caller and target scope + → update catalog rows + → increment workspace catalog revision + → append catalog outbox event and audit entry + → commit SQLite transaction + → publish the committed event to active SSE subscribers +``` + +SSE is driven by the committed catalog write, never by the periodic persistence +of a Y.Doc snapshot. A snapshot callback would be delayed, might batch unrelated +content changes, and cannot safely identify which navigation metadata changed. + +The SSE endpoint emits compact invalidations, for example: + +```text +id: 184 +event: catalog-changed +data: {"workspaceId":"…","revision":184,"documents":["…"],"spaces":["…"]} +``` + +The client treats this event as a hint to fetch authoritative catalog state. It +does not mutate its catalog cache solely from event payload. On reconnect, the +client supplies `Last-Event-ID` (or a known revision); if retained events no +longer cover that revision, the server sends a resync signal and the client +reloads the authorized catalog snapshot. + +An SSE subscriber is authorized for a Workspace and a set of Spaces before it +is registered. Filtering happens before an event is emitted, not after it +reaches the browser. A permission revocation closes or resynchronizes affected +subscriptions immediately. + +For the initial single-process deployment, an in-process publisher may drain +committed outbox rows immediately. Multiple application processes require a +shared outbox consumer/event backplane and shard ownership; independent mutable +in-memory replicas are invalid. + +## 5. Trusted routing + +Every boundary resolves a server-authorized context before reading or mutating +state: + +| Boundary | Required context | Result | +| ------------------------- | ---------------------------------------------------------------- | -------------------------------------------- | +| Page load and catalog API | deployment instance, Workspace, authorized Spaces | Catalog snapshot or a scoped subset. | +| Yjs WebSocket upgrade | deployment instance, Workspace, content shard, caller scope | One authorized Document or Collection Y.Doc. | +| SSE connection | deployment instance, Workspace, authorized Spaces, last revision | Scoped catalog invalidations. | +| MCP tool | token-derived Workspace and Space scope, target shard | Authorized catalog or content operation. | +| Persistence and audit | Workspace, Space where applicable, shard | Independently keyed durable state. | + +Path segments, query values, and Yjs room names are selectors only. They may +choose among contexts the server has already authorized for that caller, but +they never grant authority. The browser receives its bootstrap configuration +from the server; it does not invent a Workspace identity. + +## 6. Persistence and recovery + +Catalog writes are immediately transactional SQLite writes. Yjs content retains +its CRDT-first write path and is persisted as shard-scoped snapshots, with the +update-log versus periodic-snapshot decision revisited from #31's recovery and +write-churn measurements. + +There is no false atomicity claim between a catalog transaction and an arbitrary +later Yjs snapshot. Operations that change both catalog and content must define +their recovery protocol explicitly: record an operation ID, make retries +idempotent, and recover on restart from the catalog operation/outbox state plus +the relevant shard snapshot. Deleting or moving a Document must not publish its +catalog event until its required durable transition is complete. + +Each shard supports lazy loading and flushing. It may unload only when it has no +authorized live connections, no active holds requiring it, and no unflushed +state. The catalog remains small and resident enough to route requests and +serve the sidebar without opening content shards. + +## 7. Migration + +The Phase-0 workspace migrates into one Workspace with one default Space. The +migration preserves: + +- Document, Collection, record, and relation IDs. +- Document hierarchy and order. +- Collection schema, rows, and embedded-view references. +- Token grants, actor attribution, audit history, and snapshot recovery data. + +Migration is versioned, idempotent, observable, and tested against a copied +representative database. A failed migration leaves the original recovery unit +intact and never partially exposes a catalog that points to unavailable shards. + +## 8. Measurements that decide approval + +#31 must record, for a representative seeded workspace and realistic concurrent +browser/MCP workload: + +- Per-Document and per-Collection encoded state size. +- Catalog size, catalog refresh payload, and SSE reconnect/resync behavior. +- Cold load, initial Yjs sync, update-apply latency, and fan-out bytes per + active shard. +- Server/client memory, CPU, event-loop delay, snapshot/restart time, and + recovery behavior. +- The cost of an all-Yjs navigation catalog compared with committed catalog + writes plus SSE invalidation. + +The approval decision must state whether the proposed Document/Collection +granularity is accepted, whether Collections need an initial partition rule, +the retained event window, snapshot cadence, and any measured thresholds that +require further partitioning or compaction. + +## 9. Required verification + +- A client connected to one content shard never receives another shard's Yjs + updates or Awareness state. +- A Space-scoped SSE subscriber never receives an event that reveals an + inaccessible title, hierarchy, or existence signal. +- A missed SSE event causes a revision-aware catalog resync, not a stale UI. +- UI and MCP operations enforce the same catalog, Space, and shard scope. +- A cross-document agent hold works only for records the caller can access and + does not expose active editing state outside the relevant shard. +- Existing data migrates losslessly into the default Space and survives restart. +- Two local deployment instances remain isolated even when they run the same + Compendium version concurrently. From 5b149772ef06427e272ac393eb15b83674c4f935 Mon Sep 17 00:00:00 2001 From: Brylie Christopher Oxley Date: Sun, 30 Aug 2026 16:54:07 +0300 Subject: [PATCH 2/3] docs: refine workspace sharding contracts --- docs/specifications/workspace-sharding.md | 152 +++++++++++++++++----- 1 file changed, 116 insertions(+), 36 deletions(-) diff --git a/docs/specifications/workspace-sharding.md b/docs/specifications/workspace-sharding.md index 7111dc5..8ce7032 100644 --- a/docs/specifications/workspace-sharding.md +++ b/docs/specifications/workspace-sharding.md @@ -1,6 +1,7 @@ # Workspace catalog and CRDT sharding -**Status:** Proposed — requires the capacity baseline from #31 before approval. +**Status:** Proposed — #31's capacity baseline is complete; design approval is +pending review of the contracts in this document. **Depends on:** [`prd.md`](../prd.md), [`architecture.md`](./architecture.md), [`persistence.md`](./persistence.md), [`collaboration.md`](./collaboration.md), @@ -20,8 +21,8 @@ not as the universal transport for every reactive interface element. - A **Document shard** is one Y.Doc containing one Document's ordered blocks and rich text. - A **Collection shard** is one Y.Doc containing one Collection's schema and - rows. It is the initial unit; #31 determines whether a later row partition is - needed and under which measured conditions. + rows. It is the initial unit; a later row partition requires a measured + threshold and a separate design decision. - The browser uses a small server-sent-events feed for catalog changes. SSE delivers invalidations; it never becomes a second source of truth or a replacement for collaborative content sync. @@ -29,24 +30,39 @@ not as the universal transport for every reactive interface element. Document or Collection shard. This is intentionally a proposed decision, not an implementation commitment. -#31 must establish the Phase-0 baseline and identify whether a one-Y.Doc-per -Document/Collection design meets the expected operating envelope. The approval -record for #112 must state which assumptions survived that measurement. +The completed [#31 capacity baseline](../benchmarks/crdt-capacity-baseline-2026-08-30.md) +supports a small daily Phase-0 workspace and establishes a global-state +escalation boundary. Its document-size projection supports this shard direction, +but does not substitute for real shard-aware transport measurements. The +approval record for #112 must state which assumptions survived that measurement +and which must be proved by #113. ## 2. Terms and boundaries -| Term | Meaning | Authority | -| ------------------- | --------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------- | -| Deployment instance | One locally configured Compendium process and database target. It separates dogfooding, development, and test environments. | Server configuration; see #111. | -| Workspace | The top-level durable knowledge environment within a deployment instance. | Server-authorized request context. | -| Space | An organizational and permission boundary within one Workspace. A person switches Spaces to work on different projects. | Workspace catalog. | -| Catalog | The SQLite metadata projection needed to route, list, search, and authorize content without opening every content shard. | Catalog service and committed database transactions. | -| Content shard | The Y.Doc for one Document or Collection. | Authorized Yjs and service-layer mutations. | +| Term | Meaning | Authority | +| ------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------- | +| Deployment instance | Server-owned instance identity, persistence target, and event scope. One or more coordinated processes with all three values belong to the same instance. | Server configuration; see #111. | +| Workspace | The top-level durable knowledge environment within a deployment instance. | Server-authorized request context. | +| Space | An organizational and permission boundary within one Workspace. A person switches Spaces to work on different projects. | Workspace catalog. | +| Catalog | The SQLite metadata projection needed to route, list, search, and authorize content without opening every content shard. | Catalog service and committed database transactions. | +| Content shard | The Y.Doc for one Document or Collection. | Authorized Yjs and service-layer mutations. | A deployment instance is not a Workspace, and a Space is not a tenant. This keeps local operational isolation (#111), single-user multi-space organization (#6), and eventual multi-user authentication (#3) independently evolvable. +Phase 0 and the first #113 implementation run one application process per +Deployment instance. Starting independent processes against the same instance +identity and SQLite target is unsupported until the coordination contract below +is implemented; separate personal, development, and test instances must use +distinct server-owned identities and persistence targets. A future scale-out +deployment still represents one Deployment instance only when every process +shares the same identity, catalog/outbox database, authorization policy, and +event backplane. It must use durable shard ownership leases, transactional +outbox claims, and a deployment-scoped event backplane; independent mutable +in-memory replicas are invalid. #111 owns the startup diagnostics and isolation +tests, while #113 must not silently introduce a second writer. + ## 3. Ownership model ### 3.1 Workspace catalog @@ -58,12 +74,20 @@ and authorize a Workspace without subscribing to its content shards: - Document and Collection identities, titles, owning Space, hierarchy/order, and shard locator. - Space-level access grants and the derived scopes assigned to access tokens. +- A workspace-wide record locator with a unique `(workspace_id, record_id)` + key, owning-shard locator, and record kind. - Catalog revision and committed event records. Document titles and hierarchy are catalog fields. They are not duplicated into a Document Y.Doc and derived later, because the sidebar needs them before opening the Document. Record IDs remain stable global identities; a title is never an -identity or authorization key. +identity or authorization key. Creating a record reserves its supplied or +generated ID in the workspace-wide locator transaction before inserting it into +the target shard. A duplicate `(workspace_id, record_id)` is rejected; a retry +with the same operation ID returns its original result rather than creating a +second record. The hold coordinator resolves this locator before using its +`(workspaceId, recordId)` key, so a hold cannot conflate records in different +shards. The existing `snapshots`, `access_tokens`, and `audit_log` storage must become explicitly scoped by Workspace and, where applicable, Space and shard. A @@ -107,6 +131,23 @@ authorize caller and target scope → publish the committed event to active SSE subscribers ``` +Operations that also move, create, or delete content use a durable operation +row rather than publishing directly from the initial catalog transaction: + +```text +pending_content → content_durable → publishable → published +``` + +The catalog transaction records the operation ID, affected shards, intended +catalog revision, and a non-publishable outbox row. The shard transition writes +the operation ID into its durable snapshot/recovery marker; only then does one +transaction mark the operation and its outbox row `publishable`. An outbox +consumer may claim and emit only `publishable` rows. Restart recovery resumes a +pending operation idempotently by inspecting those durable markers, never by +assuming a catalog row proves the content transition completed. This applies to +move and delete events as well as creation, so a client cannot be invalidated +toward absent or stale content. + SSE is driven by the committed catalog write, never by the periodic persistence of a Y.Doc snapshot. A snapshot callback would be delayed, might batch unrelated content changes, and cannot safely identify which navigation metadata changed. @@ -114,16 +155,22 @@ content changes, and cannot safely identify which navigation metadata changed. The SSE endpoint emits compact invalidations, for example: ```text -id: 184 +id: opaque-authorized-stream-cursor event: catalog-changed -data: {"workspaceId":"…","revision":184,"documents":["…"],"spaces":["…"]} +data: {"documents":["…"],"spaces":["…"]} ``` The client treats this event as a hint to fetch authoritative catalog state. It does not mutate its catalog cache solely from event payload. On reconnect, the -client supplies `Last-Event-ID` (or a known revision); if retained events no -longer cover that revision, the server sends a resync signal and the client -reloads the authorized catalog snapshot. +client supplies an opaque `Last-Event-ID` cursor. The cursor is scoped to the +Deployment instance, Workspace, and an authorization-scope fingerprint; it +maps server-side to the highest catalog revision delivered to that authorized +stream, but never exposes a raw workspace revision or event ID. If the scope +has changed, the cursor fails validation, retained source events no longer +cover its hidden revision, or filtering makes replay coverage ambiguous, the +server sends `catalog-resync` and the client reloads the authorized catalog +snapshot. Inaccessible Space changes therefore create neither an observable +cursor gap nor an existence signal. An SSE subscriber is authorized for a Workspace and a set of Spaces before it is registered. Filtering happens before an event is emitted, not after it @@ -131,9 +178,9 @@ reaches the browser. A permission revocation closes or resynchronizes affected subscriptions immediately. For the initial single-process deployment, an in-process publisher may drain -committed outbox rows immediately. Multiple application processes require a -shared outbox consumer/event backplane and shard ownership; independent mutable -in-memory replicas are invalid. +only `publishable` outbox rows immediately. Multiple application processes +require the deployment coordination rules in §2; independent mutable in-memory +replicas are invalid. ## 5. Trusted routing @@ -161,11 +208,11 @@ update-log versus periodic-snapshot decision revisited from #31's recovery and write-churn measurements. There is no false atomicity claim between a catalog transaction and an arbitrary -later Yjs snapshot. Operations that change both catalog and content must define -their recovery protocol explicitly: record an operation ID, make retries -idempotent, and recover on restart from the catalog operation/outbox state plus -the relevant shard snapshot. Deleting or moving a Document must not publish its -catalog event until its required durable transition is complete. +later Yjs snapshot. Operations that change both catalog and content use the +durable operation state in §4: record an operation ID, make retries idempotent, +persist a shard completion marker, and recover on restart from the operation, +outbox, and relevant shard snapshot. Deleting or moving a Document must not +publish its catalog event until its required durable transition is complete. Each shard supports lazy loading and flushing. It may unload only when it has no authorized live connections, no active holds requiring it, and no unflushed @@ -174,8 +221,31 @@ serve the sidebar without opening content shards. ## 7. Migration -The Phase-0 workspace migrates into one Workspace with one default Space. The -migration preserves: +The Phase-0 workspace migrates into one Workspace with one default Space. A +versioned migration first loads the legacy snapshot into an isolated Y.Doc and +records a manifest keyed by `(legacy_snapshot_id, migration_version)`. It then +deterministically creates: + +- one catalog row and Document shard for each legacy Document, with its ordered + block records and Y.Text content; +- one catalog row and Collection shard for each legacy Collection, with its + schema and rows; and +- workspace-wide record-locator entries for every migrated block and row. + +The migration copies IDs exactly: `parentId`, record order, relation targets, +page-link targets, and collection-view references retain their original IDs and +are resolved only after the corresponding target locator exists. Each target +shard is snapshotted with the migration operation ID and content checksum before +the manifest records it durable. The catalog exposes the new workspace only +after every planned target has a durable marker and the locator manifest is +complete. Re-running a migration reuses verified targets from that manifest and +creates no duplicate rows, records, or snapshots. + +The original mixed Y.Doc snapshot and its audit/recovery metadata remain a +read-only legacy recovery unit until all target checksums, IDs, references, +token grants, and restart recovery checks pass and an explicit retention policy +permits removal. A failed migration leaves that unit intact and does not expose +a partial catalog. The migration preserves: - Document, Collection, record, and relation IDs. - Document hierarchy and order. @@ -183,13 +253,13 @@ migration preserves: - Token grants, actor attribution, audit history, and snapshot recovery data. Migration is versioned, idempotent, observable, and tested against a copied -representative database. A failed migration leaves the original recovery unit -intact and never partially exposes a catalog that points to unavailable shards. +representative database. ## 8. Measurements that decide approval -#31 must record, for a representative seeded workspace and realistic concurrent -browser/MCP workload: +The completed #31 baseline records the Phase-0 global-workspace envelope. #113 +must rerun the same representative seeded workspace and realistic concurrent +browser/MCP workload with real shards, recording: - Per-Document and per-Collection encoded state size. - Catalog size, catalog refresh payload, and SSE reconnect/resync behavior. @@ -212,9 +282,19 @@ require further partitioning or compaction. - A Space-scoped SSE subscriber never receives an event that reveals an inaccessible title, hierarchy, or existence signal. - A missed SSE event causes a revision-aware catalog resync, not a stale UI. +- An opaque SSE cursor cannot reveal an inaccessible Space change; a changed or + ambiguous authorized scope causes a catalog resync. - UI and MCP operations enforce the same catalog, Space, and shard scope. - A cross-document agent hold works only for records the caller can access and does not expose active editing state outside the relevant shard. -- Existing data migrates losslessly into the default Space and survives restart. -- Two local deployment instances remain isolated even when they run the same - Compendium version concurrently. +- A duplicate caller-supplied record ID in two shards of one Workspace is + rejected, while independent Workspaces remain isolated. +- Crash/restart recovery never emits a move/delete catalog event before its + target shard transition is durable, and retries remain idempotent. +- Existing data migrates losslessly into the default Space and survives restart: + fixtures cover mixed Documents/Collections, IDs, embedded views, links, + recovery metadata, and a re-run after partial completion. +- #111 proves two local deployment instances remain isolated even when they run + the same Compendium version concurrently; a later scale-out implementation + proves lease, outbox-claim, and event-backplane coordination within one + Deployment instance. From 9099cc9d3716b22ba4b39809b6661a65efdc4e17 Mon Sep 17 00:00:00 2001 From: Brylie Christopher Oxley Date: Sun, 30 Aug 2026 17:03:26 +0300 Subject: [PATCH 3/3] docs: close catalog publication gaps --- docs/specifications/workspace-sharding.md | 63 ++++++++++++++++------- 1 file changed, 44 insertions(+), 19 deletions(-) diff --git a/docs/specifications/workspace-sharding.md b/docs/specifications/workspace-sharding.md index 8ce7032..a3447a7 100644 --- a/docs/specifications/workspace-sharding.md +++ b/docs/specifications/workspace-sharding.md @@ -138,15 +138,30 @@ row rather than publishing directly from the initial catalog transaction: pending_content → content_durable → publishable → published ``` -The catalog transaction records the operation ID, affected shards, intended -catalog revision, and a non-publishable outbox row. The shard transition writes -the operation ID into its durable snapshot/recovery marker; only then does one -transaction mark the operation and its outbox row `publishable`. An outbox -consumer may claim and emit only `publishable` rows. Restart recovery resumes a -pending operation idempotently by inspecting those durable markers, never by -assuming a catalog row proves the content transition completed. This applies to -move and delete events as well as creation, so a client cannot be invalidated -toward absent or stale content. +The initial catalog transaction records the operation ID, affected shards, +hidden desired catalog state, and a non-publishable outbox row. It does not +allocate a public catalog revision or alter the published catalog projection. +The shard transition writes the operation ID into its durable +snapshot/recovery marker; only then does one transaction promote the desired +state into the published catalog, allocate the next public Workspace revision, +and mark its outbox row `publishable`. An outbox consumer may claim and emit +only `publishable` rows. + +Page loads, catalog APIs, MCP routing reads, and SSE all use the published +catalog projection only. While an operation is `pending_content`, a create and +its reserved record IDs are unrouteable; a move continues to serve its previous +published locator/hierarchy; and a delete continues to serve its previous +published state. A failed or recovered operation either promotes atomically or +leaves that prior public state intact. Restart recovery resumes a pending +operation idempotently by inspecting durable markers, never by assuming an +internal catalog intent proves the content transition completed. This prevents a +client from being routed to absent or stale content. + +Public revisions are allocated only at that promotion transaction and are +strictly ordered per Workspace. An outbox consumer delivers a contiguous prefix +of publishable public revisions; it cannot advance a stream past an earlier +unresolved operation. If durable recovery cannot provide that prefix, the +consumer emits `catalog-resync` rather than advancing a cursor across the gap. SSE is driven by the committed catalog write, never by the periodic persistence of a Y.Doc snapshot. A snapshot callback would be delayed, might batch unrelated @@ -175,7 +190,11 @@ cursor gap nor an existence signal. An SSE subscriber is authorized for a Workspace and a set of Spaces before it is registered. Filtering happens before an event is emitted, not after it reaches the browser. A permission revocation closes or resynchronizes affected -subscriptions immediately. +subscriptions immediately. A move or permission change that crosses an +authorization boundary sends a generic `catalog-resync` to subscribers whose +authorized catalog may lose or gain the entry. That signal contains no Document +ID, destination Space, title, hierarchy, or inaccessible metadata; the +subscriber reloads its authorized catalog snapshot to remove or add entries. For the initial single-process deployment, an in-process publisher may drain only `publishable` outbox rows immediately. Multiple application processes @@ -187,13 +206,13 @@ replicas are invalid. Every boundary resolves a server-authorized context before reading or mutating state: -| Boundary | Required context | Result | -| ------------------------- | ---------------------------------------------------------------- | -------------------------------------------- | -| Page load and catalog API | deployment instance, Workspace, authorized Spaces | Catalog snapshot or a scoped subset. | -| Yjs WebSocket upgrade | deployment instance, Workspace, content shard, caller scope | One authorized Document or Collection Y.Doc. | -| SSE connection | deployment instance, Workspace, authorized Spaces, last revision | Scoped catalog invalidations. | -| MCP tool | token-derived Workspace and Space scope, target shard | Authorized catalog or content operation. | -| Persistence and audit | Workspace, Space where applicable, shard | Independently keyed durable state. | +| Boundary | Required context | Result | +| ------------------------- | ---------------------------------------------------------------- | ---------------------------------------------- | +| Page load and catalog API | deployment instance, Workspace, authorized Spaces | Published catalog snapshot or a scoped subset. | +| Yjs WebSocket upgrade | deployment instance, Workspace, content shard, caller scope | One authorized Document or Collection Y.Doc. | +| SSE connection | deployment instance, Workspace, authorized Spaces, last revision | Scoped catalog invalidations. | +| MCP tool | token-derived Workspace and Space scope, target shard | Authorized catalog or content operation. | +| Persistence and audit | Workspace, Space where applicable, shard | Independently keyed durable state. | Path segments, query values, and Yjs room names are selectors only. They may choose among contexts the server has already authorized for that caller, but @@ -258,8 +277,8 @@ representative database. ## 8. Measurements that decide approval The completed #31 baseline records the Phase-0 global-workspace envelope. #113 -must rerun the same representative seeded workspace and realistic concurrent -browser/MCP workload with real shards, recording: +must rerun a workspace seeded with representative data and a realistic +concurrent browser/MCP workload with real shards, recording: - Per-Document and per-Collection encoded state size. - Catalog size, catalog refresh payload, and SSE reconnect/resync behavior. @@ -282,8 +301,14 @@ require further partitioning or compaction. - A Space-scoped SSE subscriber never receives an event that reveals an inaccessible title, hierarchy, or existence signal. - A missed SSE event causes a revision-aware catalog resync, not a stale UI. +- Pending content operations are absent from routing reads; a move/delete + retains its previous published state until its successor is durable. +- Public outbox delivery cannot advance an SSE cursor past an unresolved earlier + Workspace operation; recovery produces a contiguous prefix or resync. - An opaque SSE cursor cannot reveal an inaccessible Space change; a changed or ambiguous authorized scope causes a catalog resync. +- A move out of scope, move into scope, and permission change each trigger a + generic scoped resync without naming an inaccessible target or destination. - UI and MCP operations enforce the same catalog, Space, and shard scope. - A cross-document agent hold works only for records the caller can access and does not expose active editing state outside the relevant shard.