Report artifact index status to the knowledge base - #204
Merged
Merged
Conversation
Agent-driven work on this repository runs in git worktrees checked out under `.worktrees/` inside the main checkout, and each developer keeps a `.worktreeinclude` file listing the untracked, machine-local files (for example `.env`) that must be copied into a new worktree so the service can start there. Both are per-developer state, never project state. A worktree directory is a full second checkout of the repository: if git saw it, `git status` would list thousands of files and a careless `git add .` would commit a nested copy of the code. `.worktreeinclude` names local secrets files by path and says nothing that another developer's machine needs -- committing it would push one person's layout onto everyone and invite copying real `.env` files around. Ignoring them here rather than in each developer's `.git/info/exclude` means a fresh clone is safe by default: nobody has to remember a local setup step before the first agent run.
The knowledge base is getting a chip on every artifact card that says
whether the chatbot can answer from that artifact. The only record of
that lives in this service: the ingestion metadata store keeps one
ArtifactRecord per artifact with the outcome of its last ingestion. The
backend only tracks AI sync per ingestion run, so a run-level chip would
be wrong for every artifact of a partially failed batch. This adds the
read the backend needs to show the real per-artifact state.
GET /api/v1/ingest/status?artifact_ids=a&artifact_ids=b answers
{"items": [{"artifact_id", "status", "updated_at", "chunk_count"}]}:
one item per distinct requested id, in the order the ids were first
requested, so the caller can walk the answer against its page without
building a lookup of its own. Shape and names are the contract pinned
with the backend caller, snake_case like every response of this service.
Status is renamed for the reader rather than passed through: completed
becomes indexed, because that is what the chip means, and "completed"
would leak a pipeline word into the product. processing, failed and
deindexed keep their names. An id with no record is unknown, with null
updated_at and chunk_count: never ingested, older than the store, or not
ours at all, and the chip must not claim anything about it. A recorded
status the route does not know how to name reads as unknown too, rather
than being guessed at.
The request is capped at 100 ids -- one visible page, the same cap the
backend enforces on its own callers -- through FastAPI's list
validation, so an oversized batch is a 422 with the standard detail/loc
body, like every other input error here. The cap counts ids as sent,
before de-duplication, so it bounds the request itself. A blank id is a
422 as well: `artifact_ids=` can only ever produce a meaningless unknown.
No ids at all is not an error and returns an empty list, so a caller
whose page filtered down to nothing never has to special-case this
service, and never lands in its "AI unreachable" fallback for it.
The lookup is a new IngestionMetadataStore.get_artifacts(ids): a single
SELECT ... WHERE id IN (...) returning a dict keyed by id, instead of a
get_artifact call per id, so a page of 100 costs one statement under the
store lock rather than 100. It splits at 900 bound parameters, below the
999 ceiling of SQLite builds older than 3.32, so it stays safe for any
later caller while every request this route accepts is exactly one query.
The route is a pure read. It depends on the metadata store alone, never
touches Chroma, and writes nothing; a test pins that the connection's
change counter and the corpus revision do not move. chunk_count is the
recorded figure, not a live Chroma count, which would add a vector-store
round trip per artifact for what is a decoration.
Like every route of this internal service it carries no auth of its own,
mirroring /ingest/sync: it is reachable only by the backend, which
applies its project-access check and drops ids the caller may not see
before it asks.
Known gap, left as is on purpose: when a deindex fails in the vector
store, /ingest/sync leaves a durable revocation tombstone and the record
keeps saying completed until the backend's retry succeeds and marks it
deindexed. In that window retrieval already hides the artifact while
this route reports it indexed. Folding revocations into the status is a
separate decision and is not guessed at here.
Tests cover the completed-to-indexed mapping, every other status,
unknown ids, request order with duplicates, exactly 100 ids accepted and
101 rejected, blank and missing ids, the no-write guarantee, and the
batch method: hits and misses, one query for a full page, and splitting
above the parameter limit.
This was referenced Sep 24, 2026
daniilperkin
marked this pull request as ready for review
September 24, 2026 11:01
daniilperkin
marked this pull request as draft
September 24, 2026 11:05
daniilperkin
marked this pull request as ready for review
September 24, 2026 11:05
Afif-del
approved these changes
Sep 24, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of the Knowledge Base v1 upgrade (SprintStartProject/Wiki#303 epic). Counterparts: SprintStartProject/sprintstart-backend#258, SprintStartProject/sprintstart-frontend#264.
Why
The Knowledge Base now shows, per artifact card, whether the AI assistant's index holds that artifact. The backend has no view of that; the AI's ingestion metadata store does (
status,chunk_count,updated_atper artifact id). This adds one read-only route the backend can ask in a single batch per visible page.What changed
GET /api/v1/ingest/status?artifact_ids=a&artifact_ids=b(src/api/routes/ingest_status.py, registered afteringest_run) returns200 {items: [{artifact_id, status, updated_at, chunk_count}]}, one item per distinct id in first-seen order.completedbecomesindexed;processing,failed,deindexedpass through; an id with no record isunknownwith null fields; an unrecognised stored status is alsounknownbut keeps its timestamps.200 {items: []}. The backend caps at 100 before calling.IngestionMetadataStore.get_artifactsdoes one batched lookup (chunked below SQLite's parameter limit).POST /ingest/sync); sameDepends(get_ingestion_metadata_store)pattern..gitignoreignores the per-developer Hermes.worktrees/and.worktreeinclude.Known gap (documented, not fixed)
If a deindex fails in the vector store, retrieval already hides the artifact but its metadata record still says
completeduntil the backend's retry marks itdeindexed. In that window this route reportsindexed. The frontend's chip tooltip is worded accordingly ("reports the index record", not "the assistant will find it"). Accepted for v1; folding pending removals into the status is tracked in #205.Checks
ruff format --check(214 files) ·ruff check·pyright src/0 errors (run asuv run python -m pyright src/; the plainuv run pyrighttrampoline fails on this Windows host before and after the change) ·uv run python -m pytest963 passed, 8 skipped (13 new: 9 route tests intests/api/test_ingest_status.py, 4 store tests intests/ingestion/test_metadata_store.py).