From 10807aed24365ff60f96d4ed79c35ccf8710337c Mon Sep 17 00:00:00 2001 From: Mike Dalessio Date: Wed, 30 Sep 2026 16:17:03 -0400 Subject: [PATCH 1/6] Tidy the README and CHANGELOG for v0.6.0 The Upgrading note that `yabeda-hotcell` writes no log line misled readers once `HotCell::LogSubscriber` shipped. The README's "Logs" section covered only the cell's logs. Rename it "Cell logs", move `HotCell::LogSubscriber` into an "Application logs" section beside it, rewrite the Upgrading notes as action-and-rationale pairs and add one for the health controllers, and shorten entries that repeated the README. --- CHANGELOG.md | 17 +++++++++-------- README.md | 40 +++++++++++++++++++++++----------------- 2 files changed, 32 insertions(+), 25 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index c5c3295..279d343 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -18,27 +18,28 @@ gem but are what an operator runs against their own image. Some actions that application developers should consider taking when upgrading from an earlier version: -* If the application has its own Yabeda integration for HotCell, replace it with the `yabeda-hotcell` gem. The gem uses the same metric names. It does not write a log line for each call. -* Replace a cell's copies of `examples/operations/echo.rb` and `reopen.rb` with `require "hot_cell/health_operations"`, and point the application's clients at `health.echo` and `health.reopen`. See the README's "Rails healthcheck". -* Remove the application's own log line for the `perform.hot_cell` event. `HotCell::LogSubscriber` now writes one. See the README's "Per-call telemetry". +* Replace the application's own Yabeda integration for HotCell with the `yabeda-hotcell` gem. The README's "Metrics collection" lists the gem's metrics, to compare with the application's dashboards and alerts. +* Remove the application's own log line for each HotCell call, whether a `perform.hot_cell` subscriber or the application's Yabeda integration writes it. `HotCell::LogSubscriber` now writes that line, as the README's "Application logs" describes. +* Replace the application's own HotCell health endpoints with `HotCell::HealthController` and `HotCell::DiagnosticsController`. The README's "Rails healthcheck" shows the routes and how to put the diagnostics endpoint behind authentication. +* Replace a cell's copies of `examples/operations/echo.rb` and `reopen.rb` with `require "hot_cell/health_operations"`, and point the application's clients at `health.echo` and `health.reopen`. The README's "Rails healthcheck" explains what the round trips prove. ### HotCell::Server #### Added * `hot_cell/health_operations` defines `health.echo` and `health.reopen`, the round trips an application calls to prove it can use a cell's work socket. A cell serves them only if it requires the file. -* The `hotcell.describe` response now includes `server_version`. This value is the version of the `hotcell-server` gem that the cell runs. You do not need a shell in the container to find it. (#21) +* The `hotcell.describe` response includes `server_version`, the `hotcell-server` version the cell runs, so finding it no longer takes a shell in the container. (#21) #### Improved -* The supervisor now forks a worker into each free slot at boot. When the supervisor reaps a worker that served a request, it forks a replacement immediately. A request that finds a waiting worker does not wait for `fork`. A request that waits in the queue still waits for `fork`. The supervisor also runs `Process.warmup` one time, at boot, before the first fork. +* A request that finds an idle worker no longer waits for `fork`. The supervisor forks a worker into each free slot at boot, and forks a replacement as soon as it reaps a worker that served a request. A request that waits in the queue still waits for `fork`. The supervisor runs `Process.warmup` once, at boot, before the first fork. ### HotCell::Client #### Added -* `HotCell::LogSubscriber` writes one `info` line to the Rails log for each call. The line has the cell, the operation, the code and both durations. It also has the byte counts when the client can measure them. A failed call also has the cause and the `stderr` when it has them. When an exception interrupts the call, the line has the exception's class in place of the code and the cell's duration. The railtie attaches it. Without Rails, require `hot_cell/log_subscriber`, call `HotCell::LogSubscriber.attach_to :hot_cell`, and set `ActiveSupport::LogSubscriber.logger`. -* Added public and private health check controllers. `HotCell::HealthController` answers `OK` or `FAIL` from each registered cell's control socket without taking a worker. `HotCell::DiagnosticsController` returns every check as JSON, including `health.echo` and `health.reopen` round trips, and inherits from the class named by `HotCell.diagnostics_controller_parent`. Your application adds the routes; see the README's "Rails healthcheck". +* `HotCell::LogSubscriber` writes one `info` line to the Rails log for each call, with the cell, the operation, the outcome, both durations and, when measured, the byte counts. The railtie attaches it. See the README's "Application logs". +* `HotCell::HealthController` and `HotCell::DiagnosticsController` serve a cell healthcheck without a controller of the application's own. The health endpoint answers `OK` or `FAIL` from each registered cell's control socket without taking a worker, so it can be public. The diagnostics endpoint returns every check as JSON, including the `health.echo` and `health.reopen` round trips, and belongs behind authentication. The application adds the routes; see the README's "Rails healthcheck". #### Fixed @@ -49,7 +50,7 @@ Some actions that application developers should consider taking when upgrading f #### Added -* New gem `yabeda-hotcell`. It publishes Yabeda metrics for the application side. Call `Yabeda::HotCell.install!` one time at boot. The gem counts each call by cell, operation, code and cause, and measures the time the cell spent. On each scrape, it reads the counters of each registered cell. See the README's "Metrics collection". +* New gem `yabeda-hotcell` publishes Yabeda metrics for every call, and gauges from each registered cell's counters on every scrape. Call `Yabeda::HotCell.install!` once at boot. See the README's "Metrics collection". ### ActiveStorage::HotCell::Client diff --git a/README.md b/README.md index 020fb3e..bfeccd5 100644 --- a/README.md +++ b/README.md @@ -25,7 +25,8 @@ applied, and the results are returned or written to the output file. * [Using the Active Storage operations](#using-the-active-storage-operations) * [Using custom operations](#using-custom-operations) - [Observability](#observability) - * [Logs](#logs) + * [Cell logs](#cell-logs) + * [Application logs](#application-logs) * [Metrics collection](#metrics-collection) * [Per-call telemetry](#per-call-telemetry) * [Container healthcheck](#container-healthcheck) @@ -538,13 +539,27 @@ its class, clamped to the cell's exactly as the shipped ones are. ## Observability Some strategies that are working for us to monitor HotCell, which we recommend you add to your -application. (Some of these things may show up more-fully-formed in a future release.) +application. -### Logs +### Cell logs The cell writes one JSON object per event to stdout, so whatever ships your container logs ships these too. Alert on the presence of `worker.crashed`, and on `worker.killed` by cause. +### Application logs + +In a Rails application, `HotCell::LogSubscriber` writes one `info` line per call to the Rails log: + +``` + HotCell (41.2ms) {"cell":"images","operation":"active_storage.transformers.image.vips","code":"ok","perform_ms":38,"duration_ms":41.2,"bytes_in":20480,"bytes_out":8192} +``` + +A failed call adds `cause` and `stderr` when it has them. A call interrupted by an exception, such as the +application's own request timeout, logs the exception's class in place of the code. To turn the line off, call +`HotCell::LogSubscriber.detach_from :hot_cell` in an initializer. Without Rails, require +`hot_cell/log_subscriber`, call `HotCell::LogSubscriber.attach_to :hot_cell`, and set +`ActiveSupport::LogSubscriber.logger`. + ### Metrics collection The `yabeda-hotcell` gem publishes [Yabeda](https://github.com/yabeda-rb/yabeda) metrics. Add it to the @@ -569,19 +584,10 @@ scraped process must be on the cell's own host. Watch `queued`, `queue_high_wate ### Per-call telemetry -Subscribe to the `perform.hot_cell` Active Support Notification for logging, metrics, or both. It -fires on every call, success or failure, and it is the only signal that survives a dead cell -- an -unreachable socket comes back as code `unavailable`, so the primary alarm belongs here. - -In a Rails application, `HotCell::LogSubscriber` already writes one `info` line per call to the Rails log: - -``` - HotCell (41.2ms) {"cell":"images","operation":"active_storage.transformers.image.vips","code":"ok","perform_ms":38,"duration_ms":41.2,"bytes_in":20480,"bytes_out":8192} -``` - -A failed call adds `cause` and `stderr` when it has them. A call interrupted by an exception, such as the -application's own request timeout, logs the exception's class in place of the code. To turn the line off, call -`HotCell::LogSubscriber.detach_from :hot_cell` in an initializer. +The `perform.hot_cell` Active Support Notification fires on every call, success or failure, and it is +the only signal that survives a dead cell -- an unreachable socket comes back as code `unavailable`, so +the primary alarm belongs here. `HotCell::LogSubscriber` and `yabeda-hotcell` both subscribe to it, and +an application can subscribe to it for anything else. ### Container healthcheck @@ -592,7 +598,7 @@ remember to use this. ### Rails healthcheck -hotcell-client ships two controllers. Your application adds the routes. +`hotcell-client` ships two controllers. Your application adds the routes. `HotCell::HealthController` asks each registered cell for `describe` and `metrics` over its control socket. It returns `OK` with a 200 when at least one cell is registered and every cell answers, and `FAIL` with a From af6d69cfea44639a6be9c365e3b7871ed993831a Mon Sep 17 00:00:00 2001 From: Mike Dalessio Date: Wed, 30 Sep 2026 16:22:05 -0400 Subject: [PATCH 2/6] Recommend alerts for a HotCell deployment The README's observability section described each signal but not what to alert on, and its advice was spread across three sections. Add a "Recommended alerts" section from the guidance in #64, and point at `docs/LOGS.md` and `docs/TUNING.md` for the rest. `docs/TUNING.md` said to watch `queue_high_water` near `queue_size`, but that value resets only at boot, so name `queued` and a rise in `queue_high_water` instead. [Fix #64] --- README.md | 43 ++++++++++++++++++++++++++++++++++--------- docs/TUNING.md | 2 +- 2 files changed, 35 insertions(+), 10 deletions(-) diff --git a/README.md b/README.md index bfeccd5..9f562bc 100644 --- a/README.md +++ b/README.md @@ -25,6 +25,7 @@ applied, and the results are returned or written to the output file. * [Using the Active Storage operations](#using-the-active-storage-operations) * [Using custom operations](#using-custom-operations) - [Observability](#observability) + * [Recommended alerts](#recommended-alerts) * [Cell logs](#cell-logs) * [Application logs](#application-logs) * [Metrics collection](#metrics-collection) @@ -538,13 +539,38 @@ its class, clamped to the cell's exactly as the shipped ones are. ## Observability -Some strategies that are working for us to monitor HotCell, which we recommend you add to your -application. +HotCell's signals are the cell's own log, the counters the cell reports on its control socket, and +the application's record of every call. The alerts below are the ones we recommend, and the sections +after them say where each signal comes from. + +### Recommended alerts + +- **Cell availability.** Alert when the `up` gauge is 0 or absent for any cell on any host. A deploy that + missed a role, an application without the cell's group, or a dead supervisor shows here first. A host + without `HOTCELL_ROOT` reports no `up` at all, and its calls raise `HotCell::CellNotConfigured`. +- **Failed calls.** Alert on the `requests` counter by `code`. `unavailable` means the cell is down, + restarting, or unreachable, and it is recorded even when the cell cannot answer. Any other shift away + from `ok` is the early warning. +- **Queue headroom.** Alert when `queued` nears the cell's `queue_size`, when `queue_high_water` rises + toward it, or when `capacity` appears in steady state. Each means the cell is under-provisioned. + `queue_size` is configuration, not a metric, and `queue_high_water` resets only at boot, so alert on + its rise. A rising `cancelled` means callers gave up waiting. +- **Scratch space.** Alert on free space on each host's scratch: `node_filesystem_avail_bytes` from the + node exporter for a disk-backed scratch or, for a tmpfs, the container's memory usage against the + tmpfs `size=`. A full scratch fails + every request that needs it, and a write that fails inside libvips comes back `unreadable`, a permanent + verdict against the file (see [docs/IMAGEMAGICK.md](docs/IMAGEMAGICK.md)). + [Where scratch lives](docs/DEPLOYMENT.md#where-scratch-lives) covers the layouts. +- **Cell errors.** Alert on any `ERROR` event in the cell log, such as `worker.crashed` or + `worker.unforkable`, which should never happen. Alert on a rise in the `killed` gauge by cause: a + single kill for `memory` or `fsize` is the cell rejecting a hostile file. + +[What to watch](docs/TUNING.md#what-to-watch) adds the signals for tuning a cell's limits. ### Cell logs The cell writes one JSON object per event to stdout, so whatever ships your container logs ships -these too. Alert on the presence of `worker.crashed`, and on `worker.killed` by cause. +these too. [docs/LOGS.md](docs/LOGS.md) lists every event and field. ### Application logs @@ -579,15 +605,14 @@ scrape the gem polls `cell.metrics` from every registered cell and sets the gaug `queued`, `queue_high_water`, `cancelled`, `killed` (by `cause`), and `uptime_seconds`. The control socket answers even while the work socket is saturated, and it is host-local, so the -scraped process must be on the cell's own host. Watch `queued`, `queue_high_water`, `cancelled`, and -`killed` by cause. +scraped process must be on the cell's own host. ### Per-call telemetry -The `perform.hot_cell` Active Support Notification fires on every call, success or failure, and it is -the only signal that survives a dead cell -- an unreachable socket comes back as code `unavailable`, so -the primary alarm belongs here. `HotCell::LogSubscriber` and `yabeda-hotcell` both subscribe to it, and -an application can subscribe to it for anything else. +The `perform.hot_cell` Active Support Notification fires in the application on every call, success or +failure, so it still reports a dead cell: an unreachable socket comes back as code `unavailable`. +`HotCell::LogSubscriber` and `yabeda-hotcell` both subscribe to it, and an application can subscribe to +it for anything else. ### Container healthcheck diff --git a/docs/TUNING.md b/docs/TUNING.md index fcf1196..c02154a 100644 --- a/docs/TUNING.md +++ b/docs/TUNING.md @@ -143,7 +143,7 @@ safe to decide per caller. | `killed_by` by cause | the only legitimate reason to tighten a limit | | `queued_ms` p95 rising, `perform_ms` p95 flat | the cell needs more workers, not faster ones | | `perform_ms` p95 rising | the work got more expensive; check for a library upgrade | -| `queue_high_water` near `queue_size` | no headroom left | +| `queued` near `queue_size`, or `queue_high_water` rising toward it | no headroom left; `queue_high_water` resets only at boot | | `capacity` above zero in steady state | under-provisioned | | `unavailable` | the cell is down, restarting, or unreachable | | `unreadable` rate | worth watching after a toolchain upgrade | From d456cc47fdca31ad691a34b7e53e8bed140d8fbd Mon Sep 17 00:00:00 2001 From: Mike Dalessio Date: Wed, 30 Sep 2026 16:30:51 -0400 Subject: [PATCH 3/6] Tighten the README prose changed since v0.5.0 Give each sentence a subject that can do what its verb says, and cut words that carried nothing. --- README.md | 48 ++++++++++++++++++++++++------------------------ 1 file changed, 24 insertions(+), 24 deletions(-) diff --git a/README.md b/README.md index 9f562bc..69f718a 100644 --- a/README.md +++ b/README.md @@ -119,8 +119,8 @@ never evaluates a byte of image data -- it hands the accepted connection itself the worker reads them. A **worker** is a child the supervisor forks before a request needs it: one per slot at boot, and a -replacement as soon as a worker that served is reaped. Every untrusted byte is touched there and nowhere -else. It applies the cell's resource limits before touching the socket, serves +replacement as soon as the supervisor reaps one that served. Every untrusted byte is touched there and +nowhere else. It applies the cell's resource limits before touching the socket, serves `max_requests_per_worker` requests, and exits without running finalizers. A **slot** is the numbered workspace a worker borrows. It holds one directory per request, which is @@ -539,33 +539,33 @@ its class, clamped to the cell's exactly as the shipped ones are. ## Observability -HotCell's signals are the cell's own log, the counters the cell reports on its control socket, and -the application's record of every call. The alerts below are the ones we recommend, and the sections -after them say where each signal comes from. +Monitor HotCell through the cell's log, the counters the cell reports on its control socket, and the +application's record of every call. We recommend the alerts below. The sections after them say where +each signal comes from. ### Recommended alerts -- **Cell availability.** Alert when the `up` gauge is 0 or absent for any cell on any host. A deploy that - missed a role, an application without the cell's group, or a dead supervisor shows here first. A host - without `HOTCELL_ROOT` reports no `up` at all, and its calls raise `HotCell::CellNotConfigured`. -- **Failed calls.** Alert on the `requests` counter by `code`. `unavailable` means the cell is down, - restarting, or unreachable, and it is recorded even when the cell cannot answer. Any other shift away - from `ok` is the early warning. +- **Cell availability.** Alert when the `up` gauge is 0 or absent for any cell on any host. It reads 0 + first when a deploy missed a role, when the application lacks the cell's group, or when the supervisor is + dead. On a host without `HOTCELL_ROOT`, the gauge is absent and calls raise + `HotCell::CellNotConfigured`. +- **Failed calls.** Alert on the `requests` counter by `code`. The application records `unavailable` + when the cell is down, restarting, or unreachable. Any shift away from `ok` is an early warning. - **Queue headroom.** Alert when `queued` nears the cell's `queue_size`, when `queue_high_water` rises toward it, or when `capacity` appears in steady state. Each means the cell is under-provisioned. - `queue_size` is configuration, not a metric, and `queue_high_water` resets only at boot, so alert on - its rise. A rising `cancelled` means callers gave up waiting. + `queue_size` is configuration, not a metric. `queue_high_water` resets only at boot, so alert on its + rise. A rising `cancelled` means callers gave up waiting. - **Scratch space.** Alert on free space on each host's scratch: `node_filesystem_avail_bytes` from the node exporter for a disk-backed scratch or, for a tmpfs, the container's memory usage against the - tmpfs `size=`. A full scratch fails - every request that needs it, and a write that fails inside libvips comes back `unreadable`, a permanent - verdict against the file (see [docs/IMAGEMAGICK.md](docs/IMAGEMAGICK.md)). + tmpfs `size=`. A full scratch fails every request that needs it. A write that fails inside libvips + gets `unreadable` from the cell, a permanent verdict against the file (see + [docs/IMAGEMAGICK.md](docs/IMAGEMAGICK.md)). [Where scratch lives](docs/DEPLOYMENT.md#where-scratch-lives) covers the layouts. - **Cell errors.** Alert on any `ERROR` event in the cell log, such as `worker.crashed` or `worker.unforkable`, which should never happen. Alert on a rise in the `killed` gauge by cause: a single kill for `memory` or `fsize` is the cell rejecting a hostile file. -[What to watch](docs/TUNING.md#what-to-watch) adds the signals for tuning a cell's limits. +[What to watch](docs/TUNING.md#what-to-watch) lists the signals for tuning a cell's limits. ### Cell logs @@ -580,11 +580,11 @@ In a Rails application, `HotCell::LogSubscriber` writes one `info` line per call HotCell (41.2ms) {"cell":"images","operation":"active_storage.transformers.image.vips","code":"ok","perform_ms":38,"duration_ms":41.2,"bytes_in":20480,"bytes_out":8192} ``` -A failed call adds `cause` and `stderr` when it has them. A call interrupted by an exception, such as the -application's own request timeout, logs the exception's class in place of the code. To turn the line off, call -`HotCell::LogSubscriber.detach_from :hot_cell` in an initializer. Without Rails, require -`hot_cell/log_subscriber`, call `HotCell::LogSubscriber.attach_to :hot_cell`, and set -`ActiveSupport::LogSubscriber.logger`. +For a failed call, the line adds `cause` and `stderr` when they exist. For a call interrupted by an +exception, such as the application's own request timeout, the line has the exception's class in place of +the code. To turn the line off, call `HotCell::LogSubscriber.detach_from :hot_cell` in an initializer. +Without Rails, require `hot_cell/log_subscriber`, call `HotCell::LogSubscriber.attach_to :hot_cell`, and +set `ActiveSupport::LogSubscriber.logger`. ### Metrics collection @@ -599,8 +599,8 @@ gem "yabeda-hotcell" Yabeda::HotCell.install! ``` -The metrics are in the `hotcell` group. The `requests` counter counts every call, tagged with `cell`, -`operation`, `code` and `cause`, and the `perform` histogram measures the time the cell spent. On each +The metrics are in the `hotcell` group. The `requests` counter counts every call by `cell`, `operation`, +`code` and `cause`, and the `perform` histogram measures the time the cell spent. On each scrape the gem polls `cell.metrics` from every registered cell and sets the gauges `up`, `running`, `queued`, `queue_high_water`, `cancelled`, `killed` (by `cause`), and `uptime_seconds`. From 6a0366d273f8d300d4778497310f26e86e9ff347 Mon Sep 17 00:00:00 2001 From: Mike Dalessio Date: Wed, 30 Sep 2026 16:36:34 -0400 Subject: [PATCH 4/6] State that hotcell-client defines no healthcheck routes The README and CHANGELOG said "Your application adds the routes", which described the reader rather than the gem. Say that `hotcell-client` defines the controllers and does not define routes for them. --- CHANGELOG.md | 2 +- README.md | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 279d343..7b29fa6 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -39,7 +39,7 @@ Some actions that application developers should consider taking when upgrading f #### Added * `HotCell::LogSubscriber` writes one `info` line to the Rails log for each call, with the cell, the operation, the outcome, both durations and, when measured, the byte counts. The railtie attaches it. See the README's "Application logs". -* `HotCell::HealthController` and `HotCell::DiagnosticsController` serve a cell healthcheck without a controller of the application's own. The health endpoint answers `OK` or `FAIL` from each registered cell's control socket without taking a worker, so it can be public. The diagnostics endpoint returns every check as JSON, including the `health.echo` and `health.reopen` round trips, and belongs behind authentication. The application adds the routes; see the README's "Rails healthcheck". +* `HotCell::HealthController` and `HotCell::DiagnosticsController` check each registered cell. `HotCell::HealthController` uses only the control socket and returns `OK` or `FAIL`. It takes no worker. `HotCell::DiagnosticsController` also sends `health.echo` and `health.reopen` on the work socket. It takes a worker for each round trip. It returns each result as JSON. Its default superclass, `ActionController::Base`, has no authentication. `hotcell-client` does not define routes for these controllers. See the README's "Rails healthcheck". #### Fixed diff --git a/README.md b/README.md index 69f718a..21df08c 100644 --- a/README.md +++ b/README.md @@ -623,7 +623,7 @@ remember to use this. ### Rails healthcheck -`hotcell-client` ships two controllers. Your application adds the routes. +`hotcell-client` defines two controllers. It does not define routes for them. `HotCell::HealthController` asks each registered cell for `describe` and `metrics` over its control socket. It returns `OK` with a 200 when at least one cell is registered and every cell answers, and `FAIL` with a From a5450f3f74358f8d3f5db902673ddc120800c5be Mon Sep 17 00:00:00 2001 From: Mike Dalessio Date: Wed, 30 Sep 2026 16:39:54 -0400 Subject: [PATCH 5/6] Tell the application how to route the healthcheck controllers The README and CHANGELOG said what `hotcell-client` does not do. Say what the application should do: add a route for each controller, and put the diagnostics route behind authentication. --- CHANGELOG.md | 2 +- README.md | 3 ++- 2 files changed, 3 insertions(+), 2 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 7b29fa6..f5a3f8e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -39,7 +39,7 @@ Some actions that application developers should consider taking when upgrading f #### Added * `HotCell::LogSubscriber` writes one `info` line to the Rails log for each call, with the cell, the operation, the outcome, both durations and, when measured, the byte counts. The railtie attaches it. See the README's "Application logs". -* `HotCell::HealthController` and `HotCell::DiagnosticsController` check each registered cell. `HotCell::HealthController` uses only the control socket and returns `OK` or `FAIL`. It takes no worker. `HotCell::DiagnosticsController` also sends `health.echo` and `health.reopen` on the work socket. It takes a worker for each round trip. It returns each result as JSON. Its default superclass, `ActionController::Base`, has no authentication. `hotcell-client` does not define routes for these controllers. See the README's "Rails healthcheck". +* `HotCell::HealthController` and `HotCell::DiagnosticsController` check each registered cell. `HotCell::HealthController` uses only the control socket and returns `OK` or `FAIL`. It takes no worker. `HotCell::DiagnosticsController` also sends `health.echo` and `health.reopen` on the work socket. It takes a worker for each round trip. It returns each result as JSON. To use the controllers, add a route for each to the application. Put the diagnostics route behind authentication: in an initializer, set `HotCell.diagnostics_controller_parent` to the name of an authenticated controller class, or subclass `HotCell::DiagnosticsController`, authenticate in the subclass, and route to the subclass. See the README's "Rails healthcheck". #### Fixed diff --git a/README.md b/README.md index 21df08c..f3f5161 100644 --- a/README.md +++ b/README.md @@ -623,7 +623,8 @@ remember to use this. ### Rails healthcheck -`hotcell-client` defines two controllers. It does not define routes for them. +`hotcell-client` defines two controllers, `HotCell::HealthController` and `HotCell::DiagnosticsController`. +Add a route for each to the application's `config/routes.rb`, as the example below shows. `HotCell::HealthController` asks each registered cell for `describe` and `metrics` over its control socket. It returns `OK` with a 200 when at least one cell is registered and every cell answers, and `FAIL` with a From dc5c5ab896ef1ccccdeeee75938b2cdb0fbaf69f Mon Sep 17 00:00:00 2001 From: Mike Dalessio Date: Wed, 30 Sep 2026 16:41:56 -0400 Subject: [PATCH 6/6] Rewrite the docs added since v0.5.0 in Simplified Technical English The README and CHANGELOG text added since v0.5.0 used marketing verbs, described the reader instead of the software, and put several statements in one sentence. Rewrite it in ASD-STE100: one statement per sentence, active voice, and instructions in the imperative. --- CHANGELOG.md | 30 +++++++++++++-------------- README.md | 58 ++++++++++++++++++++++++++-------------------------- 2 files changed, 44 insertions(+), 44 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index f5a3f8e..f67ff80 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -18,17 +18,17 @@ gem but are what an operator runs against their own image. Some actions that application developers should consider taking when upgrading from an earlier version: -* Replace the application's own Yabeda integration for HotCell with the `yabeda-hotcell` gem. The README's "Metrics collection" lists the gem's metrics, to compare with the application's dashboards and alerts. -* Remove the application's own log line for each HotCell call, whether a `perform.hot_cell` subscriber or the application's Yabeda integration writes it. `HotCell::LogSubscriber` now writes that line, as the README's "Application logs" describes. -* Replace the application's own HotCell health endpoints with `HotCell::HealthController` and `HotCell::DiagnosticsController`. The README's "Rails healthcheck" shows the routes and how to put the diagnostics endpoint behind authentication. -* Replace a cell's copies of `examples/operations/echo.rb` and `reopen.rb` with `require "hot_cell/health_operations"`, and point the application's clients at `health.echo` and `health.reopen`. The README's "Rails healthcheck" explains what the round trips prove. +* Replace the application's Yabeda integration for HotCell with the `yabeda-hotcell` gem. The README's "Metrics collection" lists the gem's metrics for comparison with the application's dashboards and alerts. +* Remove the log line that the application writes for each HotCell call, in a `perform.hot_cell` subscriber or in its Yabeda integration. `HotCell::LogSubscriber` writes this line now; see the README's "Application logs". +* Replace the application's HotCell health endpoints with `HotCell::HealthController` and `HotCell::DiagnosticsController`. The README's "Rails healthcheck" shows the routes and the authentication for the diagnostics route. +* Replace the cell's copies of `examples/operations/echo.rb` and `reopen.rb` with `require "hot_cell/health_operations"`, and change the application's clients to call `health.echo` and `health.reopen`. See the README's "Rails healthcheck". ### HotCell::Server #### Added -* `hot_cell/health_operations` defines `health.echo` and `health.reopen`, the round trips an application calls to prove it can use a cell's work socket. A cell serves them only if it requires the file. -* The `hotcell.describe` response includes `server_version`, the `hotcell-server` version the cell runs, so finding it no longer takes a shell in the container. (#21) +* `hot_cell/health_operations` defines the `health.echo` and `health.reopen` operations. An application calls them to make sure that it can use the cell's work socket. A cell accepts them only when one of its operation files requires `hot_cell/health_operations`. Otherwise, the cell answers `unsupported`. +* The `hotcell.describe` response includes `server_version`, the version of `hotcell-server` that the cell runs. Use it to find the version without a shell in the container. (#21) #### Improved @@ -38,42 +38,42 @@ Some actions that application developers should consider taking when upgrading f #### Added -* `HotCell::LogSubscriber` writes one `info` line to the Rails log for each call, with the cell, the operation, the outcome, both durations and, when measured, the byte counts. The railtie attaches it. See the README's "Application logs". -* `HotCell::HealthController` and `HotCell::DiagnosticsController` check each registered cell. `HotCell::HealthController` uses only the control socket and returns `OK` or `FAIL`. It takes no worker. `HotCell::DiagnosticsController` also sends `health.echo` and `health.reopen` on the work socket. It takes a worker for each round trip. It returns each result as JSON. To use the controllers, add a route for each to the application. Put the diagnostics route behind authentication: in an initializer, set `HotCell.diagnostics_controller_parent` to the name of an authenticated controller class, or subclass `HotCell::DiagnosticsController`, authenticate in the subclass, and route to the subclass. See the README's "Rails healthcheck". +* `HotCell::LogSubscriber` writes one `info` line to the Rails log for each call. The line contains the cell, the operation, the outcome and both durations. It also contains the byte counts when the client measures them. The railtie attaches the subscriber. See the README's "Application logs". +* `HotCell::HealthController` and `HotCell::DiagnosticsController` check each registered cell. `HotCell::HealthController` uses only the control socket and returns `OK` or `FAIL`. It takes no worker. `HotCell::DiagnosticsController` also sends `health.echo` and `health.reopen` on the work socket. It takes a worker for each round trip. It returns each result as JSON. To use the controllers, add a route for each to the application. Put the diagnostics route behind authentication. To do this, set `HotCell.diagnostics_controller_parent` in an initializer to the name of an authenticated controller class. Alternatively, subclass `HotCell::DiagnosticsController`, authenticate in the subclass, and route to the subclass. See the README's "Rails healthcheck". #### Fixed -* The `perform.hot_cell` event now carries `cell` and `operation` when an exception escapes the call. Previously, a subscriber saw neither. +* The `perform.hot_cell` event now contains `cell` and `operation` when an exception escapes the call. Previously, the event contained neither. * The client now returns `capacity` when a full cell closes the connection before the client finishes sending the request. Previously, the client returned `unavailable` for the broken pipe. ### Yabeda::HotCell #### Added -* New gem `yabeda-hotcell` publishes Yabeda metrics for every call, and gauges from each registered cell's counters on every scrape. Call `Yabeda::HotCell.install!` once at boot. See the README's "Metrics collection". +* New gem `yabeda-hotcell`. It records Yabeda metrics for each call. On each scrape, it also sets gauges from the counters of each registered cell. Call `Yabeda::HotCell.install!` once at boot. See the README's "Metrics collection". ### ActiveStorage::HotCell::Client #### Fixed -* An application that uses only the Vips transformer no longer needs the `mini_magick` gem. Previously, `require "active_storage/hot_cell/client"` raised `LoadError` when `mini_magick` was not installed. The gem now loads `Transformers::Image::Magick` when the application first names it. +* `require "active_storage/hot_cell/client"` now works without the `mini_magick` gem. Previously, it raised `LoadError` when `mini_magick` was not installed. The gem now loads `Transformers::Image::Magick` when the application first references that constant. An application that uses only the Vips transformer can remove `mini_magick`. ### ActiveStorage::HotCell::Server #### Improved -* The transform operations write their output straight to its final path, saving a file rename on every transform. This needs image_processing 2.2.0, which the gemspec now requires. (#4) +* The transform operations write their output directly to its final path. This removes one file rename from each transform. This change requires image_processing 2.2.0, which the gemspec now specifies. (#4) ### Tooling #### Changed -* The example cell serves the gem's `health.echo` and `health.reopen` in place of its own `example.echo` and `example.reopen`. +* The example cell accepts `health.echo` and `health.reopen` from `hot_cell/health_operations`. Its own `example.echo` and `example.reopen` operations are removed. #### Fixed -* `bin/conformance` no longer fails intermittently at "offered overload answers capacity" against a healthy cell. A worker still cleaning up after the previous check could take one of the places the overload check fills, so the cell never filled. The check now waits for the cell to go idle first. -* `bin/example-image` and `bin/load` no longer pass their inputs through a shell. Before, a checkout path that contained shell syntax ran as a command in `bin/example-image`. A `bin/load` scenario, duration or thread count that contained shell syntax ran as a command in the driver container. Now the scripts pass these values as arguments. (#33) +* `bin/conformance` no longer fails intermittently at "offered overload answers capacity" against a healthy cell. Previously, a worker that was still cleaning up after the previous check could take a place that the overload check fills, so the cell did not fill. The check now waits until the cell is idle. +* `bin/example-image` and `bin/load` pass their inputs as arguments, not through a shell. Previously, shell syntax in a checkout path ran as a command in `bin/example-image`. Shell syntax in a `bin/load` scenario, duration or thread count ran as a command in the driver container. (#33) ## v0.5.0 / 2026-09-09 diff --git a/README.md b/README.md index f3f5161..4ec51a1 100644 --- a/README.md +++ b/README.md @@ -539,15 +539,15 @@ its class, clamped to the cell's exactly as the shipped ones are. ## Observability -Monitor HotCell through the cell's log, the counters the cell reports on its control socket, and the -application's record of every call. We recommend the alerts below. The sections after them say where -each signal comes from. +Use these signals to monitor HotCell: the cell's log, the counters that the cell reports on its control +socket, and the application's record of each call. Set the alerts below. The sections after the alerts +tell where each signal comes from. ### Recommended alerts - **Cell availability.** Alert when the `up` gauge is 0 or absent for any cell on any host. It reads 0 first when a deploy missed a role, when the application lacks the cell's group, or when the supervisor is - dead. On a host without `HOTCELL_ROOT`, the gauge is absent and calls raise + dead. On a host without `HOTCELL_ROOT`, the gauge is absent. Calls on that host raise `HotCell::CellNotConfigured`. - **Failed calls.** Alert on the `requests` counter by `code`. The application records `unavailable` when the cell is down, restarting, or unreachable. Any shift away from `ok` is an early warning. @@ -555,9 +555,9 @@ each signal comes from. toward it, or when `capacity` appears in steady state. Each means the cell is under-provisioned. `queue_size` is configuration, not a metric. `queue_high_water` resets only at boot, so alert on its rise. A rising `cancelled` means callers gave up waiting. -- **Scratch space.** Alert on free space on each host's scratch: `node_filesystem_avail_bytes` from the - node exporter for a disk-backed scratch or, for a tmpfs, the container's memory usage against the - tmpfs `size=`. A full scratch fails every request that needs it. A write that fails inside libvips +- **Scratch space.** Alert on free space on the scratch of each host. For a disk-backed scratch, use + `node_filesystem_avail_bytes` from the node exporter. For a tmpfs, compare the container's memory usage + with the tmpfs `size=`. A full scratch fails every request that needs it. A write that fails inside libvips gets `unreadable` from the cell, a permanent verdict against the file (see [docs/IMAGEMAGICK.md](docs/IMAGEMAGICK.md)). [Where scratch lives](docs/DEPLOYMENT.md#where-scratch-lives) covers the layouts. @@ -583,13 +583,13 @@ In a Rails application, `HotCell::LogSubscriber` writes one `info` line per call For a failed call, the line adds `cause` and `stderr` when they exist. For a call interrupted by an exception, such as the application's own request timeout, the line has the exception's class in place of the code. To turn the line off, call `HotCell::LogSubscriber.detach_from :hot_cell` in an initializer. -Without Rails, require `hot_cell/log_subscriber`, call `HotCell::LogSubscriber.attach_to :hot_cell`, and -set `ActiveSupport::LogSubscriber.logger`. +Without Rails, require `hot_cell/log_subscriber`. Then call `HotCell::LogSubscriber.attach_to :hot_cell`. +Then set `ActiveSupport::LogSubscriber.logger`. ### Metrics collection -The `yabeda-hotcell` gem publishes [Yabeda](https://github.com/yabeda-rb/yabeda) metrics. Add it to the -application's `Gemfile`, and install it once at boot: +The `yabeda-hotcell` gem records HotCell metrics in [Yabeda](https://github.com/yabeda-rb/yabeda). Add +the gem to the application's `Gemfile`. Call `Yabeda::HotCell.install!` once at boot: ```ruby # Gemfile @@ -599,20 +599,19 @@ gem "yabeda-hotcell" Yabeda::HotCell.install! ``` -The metrics are in the `hotcell` group. The `requests` counter counts every call by `cell`, `operation`, -`code` and `cause`, and the `perform` histogram measures the time the cell spent. On each -scrape the gem polls `cell.metrics` from every registered cell and sets the gauges `up`, `running`, -`queued`, `queue_high_water`, `cancelled`, `killed` (by `cause`), and `uptime_seconds`. +The metrics are in the `hotcell` group. The `requests` counter counts each call by `cell`, `operation`, +`code` and `cause`. The `perform` histogram measures the time that the cell used. On each scrape, the +gem gets `cell.metrics` from each registered cell. It sets these gauges: `up`, `running`, `queued`, +`queue_high_water`, `cancelled`, `killed` (by `cause`), and `uptime_seconds`. -The control socket answers even while the work socket is saturated, and it is host-local, so the -scraped process must be on the cell's own host. +The control socket answers when the work socket is saturated. The control socket is local to its host. +Thus the scraped process must be on the same host as the cell. ### Per-call telemetry -The `perform.hot_cell` Active Support Notification fires in the application on every call, success or -failure, so it still reports a dead cell: an unreachable socket comes back as code `unavailable`. -`HotCell::LogSubscriber` and `yabeda-hotcell` both subscribe to it, and an application can subscribe to -it for anything else. +The application sends the `perform.hot_cell` Active Support Notification for each call, successful or +failed. Thus the notification reports a dead cell: an unreachable socket gives the code `unavailable`. +`HotCell::LogSubscriber` and `yabeda-hotcell` subscribe to it. To record other data, subscribe to it. ### Container healthcheck @@ -627,19 +626,20 @@ remember to use this. Add a route for each to the application's `config/routes.rb`, as the example below shows. `HotCell::HealthController` asks each registered cell for `describe` and `metrics` over its control socket. -It returns `OK` with a 200 when at least one cell is registered and every cell answers, and `FAIL` with a -503 otherwise. These calls take no worker, so the endpoint can be public, like `/up`. +It returns `OK` with a 200 when at least one cell is registered and each cell answers. Otherwise, it +returns `FAIL` with a 503. These calls take no worker. Thus you can make the endpoint public, like `/up`. `HotCell::DiagnosticsController` returns the result of every check as JSON, with a 503 if any check fails. Along with `describe` and `metrics`, it sends `health.echo` and `health.reopen` over the work socket. Each -of those takes a worker, so put this endpoint behind authentication. +round trip takes a worker. Put this endpoint behind authentication. -The control socket carries no file descriptors, so `describe` and `metrics` succeed even when your -application cannot use the work socket. Only the round trips exercise the work socket. A cell without the -shared group passes `health.echo` and fails `health.reopen` with `EACCES`. The cell serves both operations -only if one of its operation files has `require "hot_cell/health_operations"`. +Of the cell's two sockets, only the work socket carries file descriptors. Thus `describe` and `metrics` +succeed when the application cannot use the work socket. Only the round trips test the work socket. A +cell without the shared group passes `health.echo` and fails `health.reopen` with `EACCES`. Add +`require "hot_cell/health_operations"` to one of the cell's operation files. Without it, the cell +answers `unsupported` for both operations. -Set the diagnostics controller's superclass in an initializer, then add the routes: +Set the diagnostics controller's superclass in an initializer. Then add the routes: ```ruby # config/initializers/hotcell.rb