Skip to content

fix(cron): keep scheduled jobs single-owner across desktop instances - #3175

Merged
nonoqing merged 1 commit into
GCWing:mainfrom
nonoqing:fix/cron-multi-instance-scheduling-lease
Sep 22, 2026
Merged

nonoqing merged 1 commit into
GCWing:mainfrom
nonoqing:fix/cron-multi-instance-scheduling-lease

Conversation

@nonoqing

Copy link
Copy Markdown
Collaborator

Summary

Scheduled jobs now have one owner per user data directory. Two running desktop
instances no longer both schedule the same job, and no longer retry a trigger
the other instance already delivered.

Type and Areas

Type: regression fix

Areas: Rust core (Agent Kernel scheduled-job service, services-core)

Motivation / Impact

The user data root is not keyed by bundle identifier, so the dev and release
apps share one data/cron/jobs.json while macOS treats them as independent
single instances. Both schedulers fired the same trigger; the loser then
retried the same turn id every five seconds (Dialog turn ID is already active or completed) until the job was permanently stuck as overdue — 442
consecutive failures in the reported case.

What changes:

  • The instance holding the scheduling lease schedules jobs and is the only
    writer of jobs.json; a second instance runs as standby.
  • In standby, creating, updating, deleting, or manually running a job returns an
    explicit error instead of silently overwriting the owner's state, and job
    lists are read from disk so they report the owner's state rather than a stale
    startup snapshot. This is a visible behavior change: editing scheduled jobs in
    the second instance now says which instance owns them.
  • When the owner exits, the standby takes over within 30 seconds and adopts the
    store before scheduling from it.
  • A trigger that turns out to be consumed already converges instead of retrying
    forever.

Verification

All commands were run on base e60f92a01:

cargo test -p openbitfun-core --no-default-features --features agent-runtime,scheduled-jobs,git --lib service::cron
  -> 22 passed
cargo test -p openbitfun-core --no-default-features --features agent-runtime,scheduled-jobs,git --lib agentic::tools::implementations::cron_tool
  -> 11 passed
cargo test -p openbitfun-agent-runtime --features agent-runtime --test agent_session_contracts scheduled_job
  -> 11 passed
cargo test -p openbitfun-services-core --no-default-features --features local-storage --test exclusive_file_lease_contracts
  -> 7 passed (includes real cross-process exclusion and release after an abnormal process exit)
cargo test -p openbitfun-services-core --no-default-features --features local-storage --test session_write_lock_contracts
  -> 10 passed
cargo check -p openbitfun-services-core --no-default-features
cargo check -p openbitfun-services-core --no-default-features --features runtime-ownership
cargo check -p openbitfun-core --no-default-features --features agent-runtime,scheduled-jobs,git --lib
pnpm run check:core-boundaries

Manual check: no end-to-end run of two desktop instances was performed; the
multi-instance behavior above is covered by the lease's cross-process contract
tests plus the unit tests of the ownership state, not by a two-window run.

Reviewer Notes

Design:

  • ExclusiveFileLease is a directory-scoped sibling of session::write_lock,
    reusing the internal FileLock. Ownership is the OS lock, so a leftover lock
    file never implies ownership and a crashed owner releases it automatically.
  • CronService takes the lease before it touches the store, and publishes
    ownership only after adopting the store from disk, so a takeover cannot
    schedule from a stale in-memory map.
  • A lease that cannot be created at all (as opposed to being held) fails open
    with an error log: scheduled jobs keep working exactly as before instead of
    the feature going silently dead because one lock file is unavailable.
  • The platform contention check moved into file_lock as is_lock_contention
    and now relies on ErrorKind::WouldBlock, which std maps from the
    EAGAIN/EWOULDBLOCK a contended flock reports; the previous explicit
    libc::EAGAIN branch was redundant and did not compile in a
    runtime-ownership-only build.

Known limits, deliberately out of scope:

  • Other files under the user data root (config/app.json, token usage) still
    have no cross-process write protection.
  • An older binary does not know about the lease, so mixed-version instances
    stay unprotected.
  • Takeover latency is up to 30 seconds, because only process exit releases the
    lease and there is no cross-process notification.

AI assistance: the change was written with an AI coding agent. Testing level:
fully tested for the touched Rust modules (commands above); not exercised with
two live desktop instances.

Checklist

  • This PR is focused and does not include secrets, temporary prompts, generated scratch files, or unrelated artifacts.
  • Relevant verification is recorded above, or skipped checks are explained.
  • User-facing strings, docs, and locales are updated where applicable.

The second instance's refusal error is service-level text today; a dedicated
UI/i18n state for "another instance owns scheduled jobs" is a follow-up, so the
last box is left unchecked on purpose.

Two desktop instances share one user data directory: the user root is not
keyed by bundle identifier, so the dev and release apps each ran their own
scheduler over the same jobs.json. Both fired the same trigger, and the
loser retried a turn id the winner had already consumed every five seconds
until the job was permanently stuck as overdue.

Scheduling now has one owner per user data directory:

- services-core gains ExclusiveFileLease, a typed cross-process lease over
  one resource path that reuses the existing FileLock. A second holder gets
  InUse, a leftover lock file never implies ownership, and an abnormal
  process exit releases the lease.
- CronService takes that lease before touching the store. The holder
  schedules jobs and is the only writer of jobs.json. A leapless instance
  becomes standby: it refuses job changes with an explicit error, skips turn
  lifecycle bookkeeping, and reads jobs from disk so it reports the owner's
  state instead of a stale startup snapshot.
- A standby re-checks the lease every 30 seconds and, on takeover, adopts
  the store from disk before scheduling from it. Ownership is published only
  after that adoption, so a takeover cannot schedule from a stale map.
- A lease that cannot be created at all fails open with an error log, which
  keeps the pre-lease behavior instead of silently disabling scheduled jobs.
- The cron tool reports the standby state in its list result, so an agent
  does not present another instance's job list as the state that will run.

Delivery also converges when a trigger was consumed already: that outcome
now clears the pending trigger and stops retrying instead of failing
forever.

Verified on base e60f92a:
- cargo test -p openbitfun-core --no-default-features --features agent-runtime,scheduled-jobs,git --lib service::cron (22 passed)
- cargo test -p openbitfun-core --no-default-features --features agent-runtime,scheduled-jobs,git --lib agentic::tools::implementations::cron_tool (11 passed)
- cargo test -p openbitfun-agent-runtime --features agent-runtime --test agent_session_contracts scheduled_job (11 passed)
- cargo test -p openbitfun-services-core --no-default-features --features local-storage --test exclusive_file_lease_contracts (7 passed)
- cargo test -p openbitfun-services-core --no-default-features --features local-storage --test session_write_lock_contracts (10 passed)
- cargo check -p openbitfun-core --no-default-features --features agent-runtime,scheduled-jobs,git --lib
- pnpm run check:core-boundaries

Not covered: other files under the user data root (config/app.json, token
usage) still have no cross-process write protection, and an older binary
does not know about this lease, so mixed-version instances stay unprotected.
No test builds a full CronService, so the service wiring is covered by lease
unit tests and contract tests, not end to end.

Co-authored-by: bitfun-ai <bitfun-ai@users.noreply.github.com>
@nonoqing
nonoqing merged commit a9250ba into GCWing:main Sep 22, 2026
14 checks passed
@nonoqing
nonoqing deleted the fix/cron-multi-instance-scheduling-lease branch September 22, 2026 01:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant