Skip to content

fix(usage): a single failed usage-accounting write tears down the entire round; the default SQLite busy timeout is easily exhausted by concurrent writers #1679

Description

@jcs130

Summary

On a SQLite-backed deployment with concurrent work, the daily-usage UPSERT can hit
database is locked. That exception aborts the whole generation round — all model work already
paid for is lost, and the server log shows tracebacks naming the usage table. Bookkeeping should
not be able to fail a user-visible operation.

Evidence

  1. Server tracebacks naming the usage table with database is locked during rounds that were doing
    concurrent work (several scopes active, long embedding/generation calls in flight).
  2. After raising the busy timeout (POWERCONTEXT_SERVER_DATABASE_BUSY_TIMEOUT_MS=60000) the same
    workload produced 0 database is locked errors and 0 tracebacks over the following runs
    (~3,000 model calls' worth).
  3. Default in the installed source: builtin/persistence/sqlite/profile.py:50 —
    busy_timeout_ms: int = Field(default=5_000, ge=0), applied at profile.py:131
    (PRAGMA busy_timeout = ...). Under SQLite's single-writer model, 5 s is not much once one
    writer holds the lock across a slow model call.

Requests

  1. Make usage accounting best-effort: record it in its own short transaction and log-and-continue on
    failure rather than failing the round.
  2. If it must participate in the round, take the write lock early (before the expensive model calls)
    or batch accounting so the lock is held briefly.
  3. Document, or raise, the busy-timeout default for multi-writer deployments.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions