Skip to content

Report artifact index status to the knowledge base - #204

Merged
daniilperkin merged 2 commits into
devfrom
feature/KB-updates-final
Sep 25, 2026
Merged

daniilperkin merged 2 commits into
devfrom
feature/KB-updates-final

Conversation

@daniilperkin

@daniilperkin daniilperkin commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Part of the Knowledge Base v1 upgrade (SprintStartProject/Wiki#303 epic). Counterparts: SprintStartProject/sprintstart-backend#258, SprintStartProject/sprintstart-frontend#264.

Why

The Knowledge Base now shows, per artifact card, whether the AI assistant's index holds that artifact. The backend has no view of that; the AI's ingestion metadata store does (status, chunk_count, updated_at per artifact id). This adds one read-only route the backend can ask in a single batch per visible page.

What changed

  • GET /api/v1/ingest/status?artifact_ids=a&artifact_ids=b (src/api/routes/ingest_status.py, registered after ingest_run) returns 200 {items: [{artifact_id, status, updated_at, chunk_count}]}, one item per distinct id in first-seen order.
  • Status mapping: completed becomes indexed; processing, failed, deindexed pass through; an id with no record is unknown with null fields; an unrecognised stored status is also unknown but keeps its timestamps.
  • Limits: more than 100 ids is a 422 (FastAPI's standard validation body, counted before de-duplication); a blank id is a 422; no ids is 200 {items: []}. The backend caps at 100 before calling.
  • Reads only the metadata store, never Chroma: IngestionMetadataStore.get_artifacts does one batched lookup (chunked below SQLite's parameter limit).
  • No auth, like every other route on this service (including POST /ingest/sync); same Depends(get_ingestion_metadata_store) pattern.
  • Separate commit: .gitignore ignores the per-developer Hermes .worktrees/ and .worktreeinclude.

Known gap (documented, not fixed)

If a deindex fails in the vector store, retrieval already hides the artifact but its metadata record still says completed until the backend's retry marks it deindexed. In that window this route reports indexed. The frontend's chip tooltip is worded accordingly ("reports the index record", not "the assistant will find it"). Accepted for v1; folding pending removals into the status is tracked in #205.

Checks

ruff format --check (214 files) · ruff check · pyright src/ 0 errors (run as uv run python -m pyright src/; the plain uv run pyright trampoline fails on this Windows host before and after the change) · uv run python -m pytest 963 passed, 8 skipped (13 new: 9 route tests in tests/api/test_ingest_status.py, 4 store tests in tests/ingestion/test_metadata_store.py).

Agent-driven work on this repository runs in git worktrees checked out
under `.worktrees/` inside the main checkout, and each developer keeps a
`.worktreeinclude` file listing the untracked, machine-local files (for
example `.env`) that must be copied into a new worktree so the service can
start there.

Both are per-developer state, never project state. A worktree directory is
a full second checkout of the repository: if git saw it, `git status` would
list thousands of files and a careless `git add .` would commit a nested copy
of the code. `.worktreeinclude` names local secrets files by path and says
nothing that another developer's machine needs -- committing it would push
one person's layout onto everyone and invite copying real `.env` files
around.

Ignoring them here rather than in each developer's `.git/info/exclude`
means a fresh clone is safe by default: nobody has to remember a local
setup step before the first agent run.
The knowledge base is getting a chip on every artifact card that says
whether the chatbot can answer from that artifact. The only record of
that lives in this service: the ingestion metadata store keeps one
ArtifactRecord per artifact with the outcome of its last ingestion. The
backend only tracks AI sync per ingestion run, so a run-level chip would
be wrong for every artifact of a partially failed batch. This adds the
read the backend needs to show the real per-artifact state.

GET /api/v1/ingest/status?artifact_ids=a&artifact_ids=b answers
{"items": [{"artifact_id", "status", "updated_at", "chunk_count"}]}:
one item per distinct requested id, in the order the ids were first
requested, so the caller can walk the answer against its page without
building a lookup of its own. Shape and names are the contract pinned
with the backend caller, snake_case like every response of this service.

Status is renamed for the reader rather than passed through: completed
becomes indexed, because that is what the chip means, and "completed"
would leak a pipeline word into the product. processing, failed and
deindexed keep their names. An id with no record is unknown, with null
updated_at and chunk_count: never ingested, older than the store, or not
ours at all, and the chip must not claim anything about it. A recorded
status the route does not know how to name reads as unknown too, rather
than being guessed at.

The request is capped at 100 ids -- one visible page, the same cap the
backend enforces on its own callers -- through FastAPI's list
validation, so an oversized batch is a 422 with the standard detail/loc
body, like every other input error here. The cap counts ids as sent,
before de-duplication, so it bounds the request itself. A blank id is a
422 as well: `artifact_ids=` can only ever produce a meaningless unknown.
No ids at all is not an error and returns an empty list, so a caller
whose page filtered down to nothing never has to special-case this
service, and never lands in its "AI unreachable" fallback for it.

The lookup is a new IngestionMetadataStore.get_artifacts(ids): a single
SELECT ... WHERE id IN (...) returning a dict keyed by id, instead of a
get_artifact call per id, so a page of 100 costs one statement under the
store lock rather than 100. It splits at 900 bound parameters, below the
999 ceiling of SQLite builds older than 3.32, so it stays safe for any
later caller while every request this route accepts is exactly one query.

The route is a pure read. It depends on the metadata store alone, never
touches Chroma, and writes nothing; a test pins that the connection's
change counter and the corpus revision do not move. chunk_count is the
recorded figure, not a live Chroma count, which would add a vector-store
round trip per artifact for what is a decoration.

Like every route of this internal service it carries no auth of its own,
mirroring /ingest/sync: it is reachable only by the backend, which
applies its project-access check and drops ids the caller may not see
before it asks.

Known gap, left as is on purpose: when a deindex fails in the vector
store, /ingest/sync leaves a durable revocation tombstone and the record
keeps saying completed until the backend's retry succeeds and marks it
deindexed. In that window retrieval already hides the artifact while
this route reports it indexed. Folding revocations into the status is a
separate decision and is not guessed at here.

Tests cover the completed-to-indexed mapping, every other status,
unknown ids, request order with duplicates, exactly 100 ids accepted and
101 rejected, blank and missing ids, the no-write guarantee, and the
batch method: hits and misses, one query for a full page, and splitting
above the parameter limit.
@daniilperkin
daniilperkin merged commit 7c78404 into dev Sep 25, 2026
4 checks passed
@daniilperkin
daniilperkin deleted the feature/KB-updates-final branch September 25, 2026 08:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants