Skip to content

fix(interview): accept adapter streams and bind them to their agent - #1788

Merged
eldad-caura-ai merged 1 commit into
mainfrom
fix/interview-adapter-streams
Oct 2, 2026
Merged

eldad-caura-ai merged 1 commit into
mainfrom
fix/interview-adapter-streams

Conversation

@eldad-caura-ai

Copy link
Copy Markdown
Member

Audit finding

Addresses part of M-86 (POST /interview/submit trusted node_id), and a regression that merged caura PR #1775 introduced in the same check.

Change

  • A UUID node_id must still name a fleet node of the tenant (404 otherwise), as in fix: apply the same identity, trust, fleet and visibility rules across write paths, lifecycle and audit #1775.
  • Any other node_id is an adapter stream and is accepted again.
  • An agent credential, or an install credential, may continue an adapter stream only when the agent it resolved to last advanced that stream's watermark. Otherwise the route returns 409 and persists nothing. 409 rather than 403, because caura-interviewer skips one transcript on a 409 but aborts its whole run on a 403. Tenant credentials keep tenant-wide authority, as on the other write gates.
  • read_watermark_state returns the cursor and the last-advancing agent in one primary read; read_watermark keeps its contract.

Not in this PR

  • Fleet-node windows. A same-tenant agent credential can still submit for another fleet node, capped at 1,000,000 past its watermark per call. Closing that depends on binding node heartbeats and command receipt to the node's credential (M-85), which needs a migration and a policy decision. It is tracked separately.
  • A reserved main credential still names its subject, the existing known gap until reserved_agent_id_policy=reject.

Verification

  • Tests first: the first push carries only the new tests, so CI can show them failing against unfixed main. The fix is then amended into the same single commit.
  • No repository scripts, dependency installs, tests or builds were run on Eldad's Mac under its safety hold. CI runs the suites.

One signed-off commit. Do not merge or enable auto-merge; Eldad merges after review.

🤖 Generated with Claude Code

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Claude Code Review — skipped: draft PR

caura-interviewer keys one interview stream per transcript, as cc:<machine>:<session> or cursor:..., and registers no fleet node, so the fleet-node check added in #1775 refused every window it sent. A UUID node_id must still name a fleet node of the tenant; any other node_id is an adapter stream again. An agent or install credential may continue an adapter stream only when the agent it resolved to last advanced it (409 otherwise), so it cannot jump another agent's cursor. Tenant credentials keep tenant-wide authority.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: eldad-caura <eldad@caura.ai>
@eldad-caura-ai
eldad-caura-ai force-pushed the fix/interview-adapter-streams branch from af85fed to cbe0909 Compare October 2, 2026 22:44
@eldad-caura-ai
eldad-caura-ai marked this pull request as ready for review October 2, 2026 22:55
@eldad-caura-ai
eldad-caura-ai requested a review from a team as a code owner October 2, 2026 22:55
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

🤖 Review by Claude Code

Claude Code Review ✅ No issues found.


Reviewed by claude-sonnet-5-5 · cost $0.261577

@eldad-caura-ai

Copy link
Copy Markdown
Member Author

@erni-a ready for human review. CI, CodeQL, DCO and the Claude review (no issues found) all pass on the single commit cbe0909. The tests-only push failed all 5 new tests against unfixed main (run 37072329502). This restores caura-interviewer streams that #1775 started refusing, and release PR #1738 already lists #1775, so it should merge before that release. Not auto-merging; Eldad merges.

🤖 Posted by Claude Code

@eldad-caura-ai
eldad-caura-ai merged commit 3930964 into main Oct 2, 2026
18 checks passed
@eldad-caura-ai
eldad-caura-ai deleted the fix/interview-adapter-streams branch October 2, 2026 23:18
@caura-deploy-bot caura-deploy-bot Bot mentioned this pull request Oct 2, 2026
erni-a pushed a commit that referenced this pull request Oct 3, 2026
🤖 I have created a release *beep* *boop*
---


<details><summary>backend: 3.21.0</summary>

##
[3.21.0](backend-v3.20.1...backend-v3.21.0)
(2026-10-03)


### Features

* **contradiction:** add a per-tenant switch to turn contradiction
detection off (SIDE-58)
([#1761](#1761))
([8ce0fff](8ce0fff))
* **search:** per-request recall_boost/entity_boost opt-outs on REST
/search ([#1763](#1763))
([9a6eb4e](9a6eb4e))
* **stats:** report pending background work and a settled flag on GET
/memories/stats ([#1768](#1768))
([5ea8f80](5ea8f80))


### Bug Fixes

* apply the same identity, trust, fleet and visibility rules across
write paths, lifecycle and audit
([#1775](#1775))
([b58759f](b58759f))
* **audit:** give the audit flusher's storage-slot acquire its own
budget (oss-0927-m-04)
([#1741](#1741))
([ac33b26](ac33b26))
* **ci:** isolate review tools from runner credentials
([#1784](#1784))
([5bc489e](5bc489e))
* **ci:** validate review-memory authors before model capture
([#1783](#1783))
([2273911](2273911))
* **embed:** record the re-embed give-up that reported success
(oss-0924-m-05) ([#1733](#1733))
([d9ee8a5](d9ee8a5))
* **enrich:** record the enrichment give-up that reported success
(oss-0927-m-02) ([#1736](#1736))
([94a142b](94a142b))
* **events:** refuse a forge dry_run instead of silently running for
real ([#1737](#1737))
([2e10e33](2e10e33))
* **graph:** preserve relations with surviving evidence
([#1787](#1787))
([37f8f15](37f8f15))
* **graph:** validate relation errors and tenant-scope overlap seeds
([#1782](#1782))
([17da42f](17da42f))
* **interview:** accept adapter streams and bind them to their agent
([#1788](#1788))
([3930964](3930964))
* lifecycle dedup cadence, bulk write ordering, fresh reads, skill fleet
scope, PII policy on edits
([#1772](#1772))
([f5983a0](f5983a0))
* lifecycle, storage, write-path and configuration reliability
([#1776](#1776))
([adca395](adca395))
* **lifecycle:** settle the embed-backfill topic on one spelling
([#1739](#1739))
([586b249](586b249))
* **llm:** refuse anthropic structured output instead of silently faking
it (oss-0915-m-01)
([#1742](#1742))
([fab8174](fab8174))
* low-batch (oss-0909-l-02, l-03, l-04)
([#1744](#1744))
([77abe9e](77abe9e))
* **plugin:** align tool parameters and document requests with REST
([#1779](#1779))
([b3ef98d](b3ef98d))
* **plugin:** bound credential provisioning and reject API redirects
([#1780](#1780))
([8bf503d](8bf503d))
* **plugin:** honor auto-write opt-out for conversation persistence
([#1778](#1778))
([ad8f8c3](ad8f8c3))
* **plugin:** preserve keystone truncation and normalize tool results
([#1781](#1781))
([1adae88](1adae88))
* **plugin:** refuse to send the API key over plain HTTP to non-loopback
hosts (oss-0917-m-01)
([#1745](#1745))
([4981f45](4981f45))
* **ratchet:** a new file inherits old-name lines only from files the
change deletes ([#1785](#1785))
([71b8883](71b8883))
* **ratchet:** read a JSON key's $comment marker on the line below it
([#1732](#1732))
([a196921](a196921))
* respect memory visibility in lifecycle passes, keep distinct entities
apart, bound Google LLM calls
([#1771](#1771))
([bed4935](bed4935))
* **search:** caller-named top_k beats profile/tenant default top_k
([#1764](#1764))
([47c2796](47c2796))
* **search:** honour explicit top_k on recent_context; expose retrieval
strategy header ([#1762](#1762))
([b6dde15](b6dde15))
* **settings:** let a null unset a search.default_profile knob
([#1765](#1765))
([0e6f4ab](0e6f4ab))
* **storage:** compile the stats breakdown filter without psycopg bind
casts ([#1790](#1790))
([11d5d13](11d5d13))
* **tasks:** record the three known-open give-ups; exclude the audit one
(oss-0927-m-03) ([#1746](#1746))
([2f6c247](2f6c247))
* **tests:** let unit-marked tests run without a database
([#1747](#1747))
([51542cb](51542cb))
* tighten fleet command validation, installer URL handling and settings
storage ([#1769](#1769))
([f96768c](f96768c))


### Dependencies

* bump the uv-minor-patch group across 4 directories with 9 updates
([#1770](#1770))
([0194ce6](0194ce6))
* update sqlalchemy[asyncio] requirement from &lt;2.1,&gt;=2.0.51 to
&gt;=2.0.51,&lt;2.2 in /core-storage-api
([#1758](#1758))
([1e9d73f](1e9d73f))


### Documentation

* **embedding:** stop telling operators the nightly sweep repairs
unembedded rows (oss-0927-h-01)
([#1740](#1740))
([b6cc308](b6cc308))
* **embedding:** the gateway's None is not queued for a backfill sweep
(oss-0927-m-01) ([#1735](#1735))
([5e1179e](5e1179e))
* update Eldad's GitHub handle to
[@eldad-caura-ai](https://github.com/eldad-caura-ai)
([#1767](#1767))
([d148bb6](d148bb6))
</details>

<details><summary>plugin: 2.23.4</summary>

##
[2.23.4](plugin-v2.23.3...plugin-v2.23.4)
(2026-10-03)


### Bug Fixes

* apply the same identity, trust, fleet and visibility rules across
write paths, lifecycle and audit
([#1775](#1775))
([b58759f](b58759f))
* **plugin:** align tool parameters and document requests with REST
([#1779](#1779))
([b3ef98d](b3ef98d))
* **plugin:** bound credential provisioning and reject API redirects
([#1780](#1780))
([8bf503d](8bf503d))
* **plugin:** honor auto-write opt-out for conversation persistence
([#1778](#1778))
([ad8f8c3](ad8f8c3))
* **plugin:** preserve keystone truncation and normalize tool results
([#1781](#1781))
([1adae88](1adae88))
* **plugin:** refuse to send the API key over plain HTTP to non-loopback
hosts (oss-0917-m-01)
([#1745](#1745))
([4981f45](4981f45))
* tighten fleet command validation, installer URL handling and settings
storage ([#1769](#1769))
([f96768c](f96768c))
</details>

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

Signed-off-by: release-please[bot] <release-please[bot]@users.noreply.github.com>
Co-authored-by: caura-deploy-bot[bot] <265395343+caura-deploy-bot[bot]@users.noreply.github.com>
eldad-caura-ai added a commit that referenced this pull request Oct 3, 2026
## Audit finding

Addresses **M-85**: fleet node identity was the caller-supplied
`node_name`. Any agent or install credential in a tenant could heartbeat
as another node, receive and ack its queued commands (deploy payloads
included, which an agent credential may not queue itself), report
results for any command by UUID, and list every command's payload and
result.

## Change

Each fleet node is now bound to the credential that heartbeats it.

- **Principal.** A gateway-verified agent key binds as `agent:<id>` and
an install credential as `install:<uuid>`. Every other credential is
tenant-wide (`tenant`): tenant, user and admin keys, the shared
`CAURA_API_KEY`, standalone mode, and an `X-Agent-ID` the caller
asserted itself. OSS standalone and shared-key deployments therefore see
no change. An install credential that arrives without its install UUID
is refused (403 `INSTALL_UUID_MISSING`) rather than bound to a shared
placeholder.
- **Heartbeat.** `fleet_upsert_node` decides in its `ON CONFLICT ...
WHERE`, against the locked row. A new or unbound node takes the caller's
principal. A node bound to another narrow credential refuses a narrow
caller with 403 `FLEET_NODE_BOUND_TO_OTHER_CREDENTIAL`, and the refused
heartbeat changes nothing and delivers nothing. A tenant-wide credential
is always admitted and takes the node back.
- **Command result and listing.** A narrow credential reports on, and
lists, only commands of nodes bound to it. Another node's command is the
same 404 as a missing one.
- **Release.** `POST /api/v1/fleet/nodes/{node_id}/release` (tenant
credentials only) moves a node to a new credential, for example after
rotating its key. With `bind_agent_id` or `bind_install_uuid` the node
is bound to that credential in the same statement, so no other
credential can claim it in between. Bare, it clears the binding and the
next heartbeat binds the node.
- **Migration 055** adds nullable `fleet_nodes.owner_principal`, with no
backfill. Existing nodes bind on their first heartbeat after deploy.

## Trade-offs (approved policy)

- Existing nodes bind on first heartbeat, and that heartbeat also
receives the node's pending commands. Between deploy and each existing
node's first heartbeat, a narrow credential in the same tenant that
heartbeats first can claim the node and take what is queued. The
tenant's own next heartbeat takes the node back; a node whose own
credential is narrow is rebound with a release that names it. Nothing
recorded which credential an existing node uses, and binding every
existing node to the tenant instead would lock out nodes that heartbeat
with agent keys or install credentials until an operator rebinds each
one.
- Moving a node to a different narrow credential takes a release that
names it. A bare release reopens the first-heartbeat window for that
node.
- Deploy `core-storage-api` (migration 055) before `core-api`, as the
pipeline already does. A new `core-api` sends `owner_principal`, which
an older storage cannot write.

## Not in this PR

- Interview fleet-node windows (the rest of M-86) follow in a separate
PR, now that caura PR #1788 has merged: a window for a fleet node must
cite a delivered `interview_request`, used once. It relies on this PR's
listing filter, which stops an agent credential reading another node's
pending command ids.
- The enterprise dashboard has no button for the release endpoint yet.

## Verification

- Storage tests cover bind, refuse-without-change, refresh, reclaim,
legacy unbound rows, caller-omitted binding, release (including tenant
scoping and binding straight to a named credential), and the result and
listing filters. Route tests cover the drain attack, over-refusal guards
for own nodes and asserted headers, install credentials (including one
with no install UUID), release authorization, release that binds a named
agent key or install credential with no window for a squatter, at most
one bind target, and 404 for an unknown node.
- Tests first: the tests-only push ([run
37074834988](https://github.com/caura-ai/caura/actions/runs/37074834988/job/111062277857))
failed all 10 attack tests in `tests/test_fleet_node_binding.py` against
unfixed `main`, with 8459 passing. The drain test's response carried the
victim node's queued command. The two over-refusal guards passed, as
they must without the fix. That step stopped the job before the storage
suite ran. The fix is now amended into the same single commit.
- No repository scripts, dependency installs, tests or builds were run
on Eldad's Mac under its safety hold. CI runs the suites.

One signed-off commit. Do not merge or enable auto-merge; Eldad merges
after review.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Signed-off-by: eldad-caura <eldad@caura.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
eldad-caura-ai added a commit that referenced this pull request Oct 3, 2026
…to it

An agent or install credential could submit an interview window for any fleet node of its tenant by naming the node and any command_id, which advanced that node's watermark and wrote its job doc (M-86, fleet-node half; #1788 bound adapter streams). A fleet node's window from an agent or install credential must now cite an interview_request the scheduler queued for that node and the node's heartbeat delivered. Storage claims the request in one conditional UPDATE, so each request admits one window, and the node's own result report still closes the command. Tenant credentials keep tenant-wide authority. The plugin already cites the delivered command's id, so its flow is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: eldad-caura <eldad@caura.ai>
eldad-caura-ai added a commit that referenced this pull request Oct 3, 2026
…to it

An agent or install credential could submit an interview window for any fleet node of its tenant by naming the node and any command_id, which advanced that node's watermark and wrote its job doc (M-86, fleet-node half; #1788 bound adapter streams). A fleet node's window from an agent or install credential must now cite an interview_request the scheduler queued for that node and the node's heartbeat delivered. Storage claims the request in one conditional UPDATE, so each request admits one window, and the node's own result report still closes the command. Tenant credentials keep tenant-wide authority. The plugin already cites the delivered command's id, so its flow is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: eldad-caura <eldad@caura.ai>
eldad-caura-ai added a commit that referenced this pull request Oct 3, 2026
…to it

An agent or install credential could submit an interview window for any fleet node of its tenant by naming the node and any command_id, which advanced that node's watermark and wrote its job doc (M-86, fleet-node half; #1788 bound adapter streams). A fleet node's window from an agent or install credential must now cite an interview_request the scheduler queued for that node and the node's heartbeat delivered. Storage claims the request in one conditional UPDATE, so each request admits one window, and the node's own result report still closes the command. Tenant credentials keep tenant-wide authority. The plugin already cites the delivered command's id, so its flow is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: eldad-caura <eldad@caura.ai>
eldad-caura-ai added a commit that referenced this pull request Oct 3, 2026
…to it

An agent or install credential could submit an interview window for any fleet node of its tenant by naming the node and any command_id, which advanced that node's watermark and wrote its job doc (M-86, fleet-node half; #1788 bound adapter streams). A fleet node's window from an agent or install credential must now cite an interview_request the scheduler queued for that node and the node's heartbeat delivered. Storage claims the request in one conditional UPDATE, so each request admits one window, and the node's own result report still closes the command. Tenant credentials keep tenant-wide authority. The plugin already cites the delivered command's id, so its flow is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: eldad-caura <eldad@caura.ai>
eldad-caura-ai added a commit that referenced this pull request Oct 3, 2026
…to it (#1792)

## Audit finding

Finishes **M-86**. caura PR #1788 bound adapter streams (`cc:` /
`cursor:` node ids) to the agent that advances them. Fleet-node windows
were still open: an agent or install credential could submit a window
for any fleet node of its tenant by naming the node and any
`command_id`. That advanced the node's watermark, which only moves
forward, and wrote its job doc.

## Change

- **Gate.** A window for a fleet node (UUID `node_id`) from an agent or
install credential must cite the `interview_request` the scheduler
queued for that node, and the node's heartbeat must have delivered it
(`acked`). Anything else gets 403 `INTERVIEW_REQUEST_REQUIRED` before
the watermark or job doc is touched. This is the same "agent or install
credential" test #1788 applies to adapter streams.
- **Used once.** Storage claims the request in one conditional UPDATE:
this tenant, this node, `interview_request`, `acked`, no `result` yet.
The claim writes a marker into `result`, so a second window citing the
same id is refused. A used id stops being secret, because the job doc
and watermark record it. The node's own result report still overwrites
the marker and closes the command.
- **Unchanged.** Tenant credentials keep tenant-wide authority, as they
do for adapter streams. The plugin already submits once per delivered
command and cites its id (`plugin/src/heartbeat.ts`), and the scheduler
only looks at `pending` commands, so neither changes. A failed submit is
retried by the scheduler as a fresh command, as before.

## Why this relies on caura PR #1789

The request id is the credential here. Since #1789, an agent or install
credential can only list or receive commands for nodes bound to it, so
it cannot read another node's pending `interview_request` before that
node uses it.

## Verification

- Route tests cover an uninvited agent and an uninvited install, a
queued but undelivered request, one window per request, a request for
another node, and a guard that tenant credentials still submit directly.
The existing attribution test now cites a delivered request.
- Storage tests cover claim once, undelivered, wrong node, wrong tenant,
non-interview command, an already completed request, and the node's
result report closing a claimed request.
- Tests first: the tests-only push ([run
37126305662](https://github.com/caura-ai/caura/actions/runs/37126305662/job/111212158891))
failed exactly the five attack tests against `main`, with 8644 passing.
In each one the unfixed route answered `accepted` and moved the node's
watermark. The tenant guard and the updated attribution test passed.
That step stopped the job before the storage suite ran. The fix is now
amended into the same single commit.
- No repository scripts, dependency installs, tests or builds were run
on Eldad's Mac under its safety hold. CI runs the suites.

One signed-off commit. Do not merge or enable auto-merge; Eldad merges
after review.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Signed-off-by: eldad-caura <eldad@caura.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants