Skip to content

manager: a spawned process can still be reported failed_to_start #14

Description

@kp2pml30

Found by a review pass over feat/rework-manager-api. Verified in the code.

What

implementation/src/manager/run.rs:2297 (proc.spawn()? succeeds) — after that point, a failure writing execution data to the child's stdin does anyhow::bail!, so run_genvm_process returns Err and supervise_genvm publishes failed_to_start (run.rs:2341, :2108).

Scenario

Start an executor with extra_args: ["--bad-flag"] and an execution-data body large enough not to fit the pipe buffer. Clap exits and closes stdin while the manager is still writing. The select! then races:

  • child.wait() wins → the run is reported finished;
  • the stdin task's EPIPE wins → the run is reported failed_to_start.

Why it matters

The protocol defines the distinction by whether a process was created (docs/website/src/impl-spec/appendix/manager-socket.rst:223), and the process was created in both branches. tests/system/manager-socket/test.py:526 expects finished, so the test is load-bearing on losing the race.

Suggested fix

Record "spawned" state as soon as proc.spawn() returns, and classify every later failure as finished regardless of which arm of the select! fires.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions