Skip to content

fix(authz): a failed post-write owner step leaves committed entries with no artifact owner, bricking the entire scope with no supported recovery #1677

Description

@jcs130

Summary

Memory entries can be committed to the artifact heads/manifest while the corresponding
artifact-owner rows are missing. From that moment the entire scope — every read and
every write — fails with HTTP 503 artifact_owner_pending, and there is no supported way
to repair it from outside (the flush path hits the same predicate and self-locks).

Environment

PowerContext 1.0.0 (pip install), self-hosted server, SQLite persistence (sqlite-vec),
local embedding + generation providers, WSL2. Single process, several scopes.

Evidence

  1. Invariant check heads ∪ manifest − owners on the live DB reported exactly 2 entry ids
    with no owner row, both present in heads and manifest.
  2. POST /v1/memory/search against the affected scope failed in 4/4 search modes with
    artifact_owner_pending (not a channel-specific problem).
  3. Gate location (verified in the installed source):
    • server/authz/casbin.py:99 — sets reason = "artifact-owner-pending" when no owner matches
      (decision is allowed = 0, applied to reads and writes alike).
    • server/app.py:3944, 3955, 3990, 4012, 4345 — raise AccessUnavailableError("artifact_owner_pending").
    • _COLLECTION_CONTENT_OPERATIONS (server/app.py:4294), checked by
      require_scope_content_ready (server/app.py:4325, invoked at :4371) — this is why a
      single missing row takes out collection-wide content operations, not just one entry.
  4. Recovery: no CLI or HTTP entry point to establish an owner row was found; the documented flush
    route returns the same pending error (self-locking). We repaired the scope by calling
    RelationalAccessRepository.establish_artifact_owner offline with the server stopped — i.e. by
    reaching into an internal API. That worked (+2 revisions, no restart needed), but it is not a
    path a user can be expected to find.

Hypothesis (NOT verified — please check)

We suspect owner rows are written by a step that runs after entries are committed, so any
failure/timeout between the two steps leaves the invariant permanently broken. We could not confirm
the ordering from the code; treating it as a hypothesis, not a finding.

Requests

  1. Make the owner write atomic with the entry commit (or write owners first, idempotently), so
    owners ⊇ heads ∪ manifest holds at every commit boundary.
  2. Provide a supported repair command for an already-broken scope (e.g. a pc subcommand or an
    admin endpoint that re-establishes owners from heads ∪ manifest).
  3. Consider degrading to read-only rather than 503-ing every read for the whole scope when only
    owner metadata is missing.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions