Skip to content

Document recommended observability and alerting #64

Description

@flavorjones

Placeholder to review; not sure this needs doing, but capturing it so it isn't lost.

Check that the docs cover how we recommend operators do observability and alerting. The Rails World talk describes it as three signals (logging from the Rails app, metrics tracked by the Supervisor, logging from the Supervisor) and this guidance:

  • Alarm on cell availability. hotcell_up, per host. The app polls describe on the control socket and exports 1 or 0. A deploy that misses a role, a missing group, or a crashed supervisor shows here first.
  • Alarm on queue depth. hotcell_queued and hotcell_queue_high_water, from metrics on the control socket. High water near queue_size means no headroom for a spike; capacity failures in steady state mean under-provisioned.
  • Alarm on available disk space. Scratch disk used per host, from the node exporter on the scratch mount. A full disk reaches the caller as ENOSPC.
  • Monitor throughput and failures by outcome. Requests per minute by outcome: ok, unreadable, capacity, killed (deadline, memory, fsize, crashed). Any shift away from ok is the early warning.
  • Health checks. The Docker HEALTHCHECK and the Rails /up/hotcell endpoint both ask the control socket, so they stay green even when the app cannot use the work socket.

If the docs already say this, close.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions