Skip to content

node status/spawn look for connection.json one level above where the broker writes it, reporting a live node as STOPPED #1575

Description

@khaliqgant

Summary

A node started with --state-dir <X> writes its connection.json to <X>/state/, but the CLI resolves it at <X>/connection.json. So passing the same --state-dir the node was started with makes every local broker command conclude the broker is dead.

The node is fine. The CLI just looks one directory too shallow.

Reproduction

Node launched by launchd via a start script containing:

agent-relay node up --broker-name sf-mini --state-dir /Users/khaliqgant/.agentworkforce/relay/sf-mini-node

Using that exact path:

$ agent-relay node status --state-dir /Users/khaliqgant/.agentworkforce/relay/sf-mini-node
Status: STOPPED

$ agent-relay node agent spawn codex --name <x> --cwd <y>
No running broker found (/Users/khaliqgant/.agentworkforce/relay/connection.json does not exist).

Adding /state:

$ agent-relay node status --state-dir /Users/khaliqgant/.agentworkforce/relay/sf-mini-node/state
Status: RUNNING
Mode: broker (stdio)
PID: 62446
Agents: 1
Node delivery: CONNECTED
Node: sf-mini (node_203549044126121984)

Confirmed on disk — the only connection.json is at <X>/state/connection.json; <X>/connection.json does not exist.

Why this is worth fixing

  • It makes a healthy node look dead to its own operator. Status: STOPPED on a running broker is the single most likely trigger for restarting a node — and restarting a fleet node kills every spawned agent on it, since they are children of the broker. The remedy the wrong answer invites is destructive.
  • It silently disables local spawn. node agent spawn fails with "No running broker found" naming a path the operator never chose, so the message points away from the real cause. On the affected machine this was misread first as a broker-version problem, then as a placement failure.
  • The error names the default path, not the requested one. Even though --state-dir was passed, the message reports ~/.agentworkforce/relay/connection.json, which hides that the flag was honoured at all.

Notes

node agent attach accepts --state-dir and documents it as auto-discovered; node agent spawn has no --state-dir option at all, so there is no way to point it at a non-default broker even knowing the correct path. Working around this required cd-ing into the state directory.

Suggested direction

  • Resolve connection.json consistently: either the broker writes it at <state-dir>/connection.json, or every reader looks under <state-dir>/state/. One convention, applied on both sides.
  • Accept both for a release, preferring the deeper path, so existing nodes keep working.
  • Give node agent spawn a --state-dir flag matching attach.
  • When resolution fails, report the path actually searched and the flag that produced it, and distinguish "no broker is running" from "cannot find this broker's connection file". Those are different facts with different remedies, and only one of them warrants a restart.

Observed on CLI 11.7.1 against a broker pinned to 11.6.7 via BROKER_BINARY_PATH, though the path mismatch looks independent of the version skew.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions