Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 16 additions & 15 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,61 +18,62 @@ gem but are what an operator runs against their own image.

Some actions that application developers should consider taking when upgrading from an earlier version:

* If the application has its own Yabeda integration for HotCell, replace it with the `yabeda-hotcell` gem. The gem uses the same metric names. It does not write a log line for each call.
* Replace a cell's copies of `examples/operations/echo.rb` and `reopen.rb` with `require "hot_cell/health_operations"`, and point the application's clients at `health.echo` and `health.reopen`. See the README's "Rails healthcheck".
* Remove the application's own log line for the `perform.hot_cell` event. `HotCell::LogSubscriber` now writes one. See the README's "Per-call telemetry".
* Replace the application's Yabeda integration for HotCell with the `yabeda-hotcell` gem. The README's "Metrics collection" lists the gem's metrics for comparison with the application's dashboards and alerts.
* Remove the log line that the application writes for each HotCell call, in a `perform.hot_cell` subscriber or in its Yabeda integration. `HotCell::LogSubscriber` writes this line now; see the README's "Application logs".
* Replace the application's HotCell health endpoints with `HotCell::HealthController` and `HotCell::DiagnosticsController`. The README's "Rails healthcheck" shows the routes and the authentication for the diagnostics route.
* Replace the cell's copies of `examples/operations/echo.rb` and `reopen.rb` with `require "hot_cell/health_operations"`, and change the application's clients to call `health.echo` and `health.reopen`. See the README's "Rails healthcheck".

### HotCell::Server

#### Added

* `hot_cell/health_operations` defines `health.echo` and `health.reopen`, the round trips an application calls to prove it can use a cell's work socket. A cell serves them only if it requires the file.
* The `hotcell.describe` response now includes `server_version`. This value is the version of the `hotcell-server` gem that the cell runs. You do not need a shell in the container to find it. (#21)
* `hot_cell/health_operations` defines the `health.echo` and `health.reopen` operations. An application calls them to make sure that it can use the cell's work socket. A cell accepts them only when one of its operation files requires `hot_cell/health_operations`. Otherwise, the cell answers `unsupported`.
* The `hotcell.describe` response includes `server_version`, the version of `hotcell-server` that the cell runs. Use it to find the version without a shell in the container. (#21)

#### Improved

* The supervisor now forks a worker into each free slot at boot. When the supervisor reaps a worker that served a request, it forks a replacement immediately. A request that finds a waiting worker does not wait for `fork`. A request that waits in the queue still waits for `fork`. The supervisor also runs `Process.warmup` one time, at boot, before the first fork.
* A request that finds an idle worker no longer waits for `fork`. The supervisor forks a worker into each free slot at boot, and forks a replacement as soon as it reaps a worker that served a request. A request that waits in the queue still waits for `fork`. The supervisor runs `Process.warmup` once, at boot, before the first fork.

### HotCell::Client

#### Added

* `HotCell::LogSubscriber` writes one `info` line to the Rails log for each call. The line has the cell, the operation, the code and both durations. It also has the byte counts when the client can measure them. A failed call also has the cause and the `stderr` when it has them. When an exception interrupts the call, the line has the exception's class in place of the code and the cell's duration. The railtie attaches it. Without Rails, require `hot_cell/log_subscriber`, call `HotCell::LogSubscriber.attach_to :hot_cell`, and set `ActiveSupport::LogSubscriber.logger`.
* Added public and private health check controllers. `HotCell::HealthController` answers `OK` or `FAIL` from each registered cell's control socket without taking a worker. `HotCell::DiagnosticsController` returns every check as JSON, including `health.echo` and `health.reopen` round trips, and inherits from the class named by `HotCell.diagnostics_controller_parent`. Your application adds the routes; see the README's "Rails healthcheck".
* `HotCell::LogSubscriber` writes one `info` line to the Rails log for each call. The line contains the cell, the operation, the outcome and both durations. It also contains the byte counts when the client measures them. The railtie attaches the subscriber. See the README's "Application logs".
* `HotCell::HealthController` and `HotCell::DiagnosticsController` check each registered cell. `HotCell::HealthController` uses only the control socket and returns `OK` or `FAIL`. It takes no worker. `HotCell::DiagnosticsController` also sends `health.echo` and `health.reopen` on the work socket. It takes a worker for each round trip. It returns each result as JSON. To use the controllers, add a route for each to the application. Put the diagnostics route behind authentication. To do this, set `HotCell.diagnostics_controller_parent` in an initializer to the name of an authenticated controller class. Alternatively, subclass `HotCell::DiagnosticsController`, authenticate in the subclass, and route to the subclass. See the README's "Rails healthcheck".

#### Fixed

* The `perform.hot_cell` event now carries `cell` and `operation` when an exception escapes the call. Previously, a subscriber saw neither.
* The `perform.hot_cell` event now contains `cell` and `operation` when an exception escapes the call. Previously, the event contained neither.
* The client now returns `capacity` when a full cell closes the connection before the client finishes sending the request. Previously, the client returned `unavailable` for the broken pipe.

### Yabeda::HotCell

#### Added

* New gem `yabeda-hotcell`. It publishes Yabeda metrics for the application side. Call `Yabeda::HotCell.install!` one time at boot. The gem counts each call by cell, operation, code and cause, and measures the time the cell spent. On each scrape, it reads the counters of each registered cell. See the README's "Metrics collection".
* New gem `yabeda-hotcell`. It records Yabeda metrics for each call. On each scrape, it also sets gauges from the counters of each registered cell. Call `Yabeda::HotCell.install!` once at boot. See the README's "Metrics collection".

### ActiveStorage::HotCell::Client

#### Fixed

* An application that uses only the Vips transformer no longer needs the `mini_magick` gem. Previously, `require "active_storage/hot_cell/client"` raised `LoadError` when `mini_magick` was not installed. The gem now loads `Transformers::Image::Magick` when the application first names it.
* `require "active_storage/hot_cell/client"` now works without the `mini_magick` gem. Previously, it raised `LoadError` when `mini_magick` was not installed. The gem now loads `Transformers::Image::Magick` when the application first references that constant. An application that uses only the Vips transformer can remove `mini_magick`.

### ActiveStorage::HotCell::Server

#### Improved

* The transform operations write their output straight to its final path, saving a file rename on every transform. This needs image_processing 2.2.0, which the gemspec now requires. (#4)
* The transform operations write their output directly to its final path. This removes one file rename from each transform. This change requires image_processing 2.2.0, which the gemspec now specifies. (#4)

### Tooling

#### Changed

* The example cell serves the gem's `health.echo` and `health.reopen` in place of its own `example.echo` and `example.reopen`.
* The example cell accepts `health.echo` and `health.reopen` from `hot_cell/health_operations`. Its own `example.echo` and `example.reopen` operations are removed.

#### Fixed

* `bin/conformance` no longer fails intermittently at "offered overload answers capacity" against a healthy cell. A worker still cleaning up after the previous check could take one of the places the overload check fills, so the cell never filled. The check now waits for the cell to go idle first.
* `bin/example-image` and `bin/load` no longer pass their inputs through a shell. Before, a checkout path that contained shell syntax ran as a command in `bin/example-image`. A `bin/load` scenario, duration or thread count that contained shell syntax ran as a command in the driver container. Now the scripts pass these values as arguments. (#33)
* `bin/conformance` no longer fails intermittently at "offered overload answers capacity" against a healthy cell. Previously, a worker that was still cleaning up after the previous check could take a place that the overload check fills, so the cell did not fill. The check now waits until the cell is idle.
* `bin/example-image` and `bin/load` pass their inputs as arguments, not through a shell. Previously, shell syntax in a checkout path ran as a command in `bin/example-image`. Shell syntax in a `bin/load` scenario, duration or thread count ran as a command in the driver container. (#33)

## v0.5.0 / 2026-09-09

Expand Down
110 changes: 71 additions & 39 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,9 @@ applied, and the results are returned or written to the output file.
* [Using the Active Storage operations](#using-the-active-storage-operations)
* [Using custom operations](#using-custom-operations)
- [Observability](#observability)
* [Logs](#logs)
* [Recommended alerts](#recommended-alerts)
* [Cell logs](#cell-logs)
* [Application logs](#application-logs)
* [Metrics collection](#metrics-collection)
* [Per-call telemetry](#per-call-telemetry)
* [Container healthcheck](#container-healthcheck)
Expand Down Expand Up @@ -117,8 +119,8 @@ never evaluates a byte of image data -- it hands the accepted connection itself
the worker reads them.

A **worker** is a child the supervisor forks before a request needs it: one per slot at boot, and a
replacement as soon as a worker that served is reaped. Every untrusted byte is touched there and nowhere
else. It applies the cell's resource limits before touching the socket, serves
replacement as soon as the supervisor reaps one that served. Every untrusted byte is touched there and
nowhere else. It applies the cell's resource limits before touching the socket, serves
`max_requests_per_worker` requests, and exits without running finalizers.

A **slot** is the numbered workspace a worker borrows. It holds one directory per request, which is
Expand Down Expand Up @@ -537,18 +539,57 @@ its class, clamped to the cell's exactly as the shipped ones are.

## Observability

Some strategies that are working for us to monitor HotCell, which we recommend you add to your
application. (Some of these things may show up more-fully-formed in a future release.)

### Logs
Use these signals to monitor HotCell: the cell's log, the counters that the cell reports on its control
socket, and the application's record of each call. Set the alerts below. The sections after the alerts
tell where each signal comes from.

### Recommended alerts

- **Cell availability.** Alert when the `up` gauge is 0 or absent for any cell on any host. It reads 0
first when a deploy missed a role, when the application lacks the cell's group, or when the supervisor is
dead. On a host without `HOTCELL_ROOT`, the gauge is absent. Calls on that host raise
`HotCell::CellNotConfigured`.
- **Failed calls.** Alert on the `requests` counter by `code`. The application records `unavailable`
when the cell is down, restarting, or unreachable. Any shift away from `ok` is an early warning.
- **Queue headroom.** Alert when `queued` nears the cell's `queue_size`, when `queue_high_water` rises
toward it, or when `capacity` appears in steady state. Each means the cell is under-provisioned.
`queue_size` is configuration, not a metric. `queue_high_water` resets only at boot, so alert on its
rise. A rising `cancelled` means callers gave up waiting.
- **Scratch space.** Alert on free space on the scratch of each host. For a disk-backed scratch, use
`node_filesystem_avail_bytes` from the node exporter. For a tmpfs, compare the container's memory usage
with the tmpfs `size=`. A full scratch fails every request that needs it. A write that fails inside libvips
gets `unreadable` from the cell, a permanent verdict against the file (see
[docs/IMAGEMAGICK.md](docs/IMAGEMAGICK.md)).
[Where scratch lives](docs/DEPLOYMENT.md#where-scratch-lives) covers the layouts.
- **Cell errors.** Alert on any `ERROR` event in the cell log, such as `worker.crashed` or
`worker.unforkable`, which should never happen. Alert on a rise in the `killed` gauge by cause: a
single kill for `memory` or `fsize` is the cell rejecting a hostile file.

[What to watch](docs/TUNING.md#what-to-watch) lists the signals for tuning a cell's limits.

### Cell logs

The cell writes one JSON object per event to stdout, so whatever ships your container logs ships
these too. Alert on the presence of `worker.crashed`, and on `worker.killed` by cause.
these too. [docs/LOGS.md](docs/LOGS.md) lists every event and field.

### Application logs

In a Rails application, `HotCell::LogSubscriber` writes one `info` line per call to the Rails log:

```
HotCell (41.2ms) {"cell":"images","operation":"active_storage.transformers.image.vips","code":"ok","perform_ms":38,"duration_ms":41.2,"bytes_in":20480,"bytes_out":8192}
```

For a failed call, the line adds `cause` and `stderr` when they exist. For a call interrupted by an
exception, such as the application's own request timeout, the line has the exception's class in place of
the code. To turn the line off, call `HotCell::LogSubscriber.detach_from :hot_cell` in an initializer.
Without Rails, require `hot_cell/log_subscriber`. Then call `HotCell::LogSubscriber.attach_to :hot_cell`.
Then set `ActiveSupport::LogSubscriber.logger`.

### Metrics collection

The `yabeda-hotcell` gem publishes [Yabeda](https://github.com/yabeda-rb/yabeda) metrics. Add it to the
application's `Gemfile`, and install it once at boot:
The `yabeda-hotcell` gem records HotCell metrics in [Yabeda](https://github.com/yabeda-rb/yabeda). Add
the gem to the application's `Gemfile`. Call `Yabeda::HotCell.install!` once at boot:

```ruby
# Gemfile
Expand All @@ -558,30 +599,19 @@ gem "yabeda-hotcell"
Yabeda::HotCell.install!
```

The metrics are in the `hotcell` group. The `requests` counter counts every call, tagged with `cell`,
`operation`, `code` and `cause`, and the `perform` histogram measures the time the cell spent. On each
scrape the gem polls `cell.metrics` from every registered cell and sets the gauges `up`, `running`,
`queued`, `queue_high_water`, `cancelled`, `killed` (by `cause`), and `uptime_seconds`.
The metrics are in the `hotcell` group. The `requests` counter counts each call by `cell`, `operation`,
`code` and `cause`. The `perform` histogram measures the time that the cell used. On each scrape, the
gem gets `cell.metrics` from each registered cell. It sets these gauges: `up`, `running`, `queued`,
`queue_high_water`, `cancelled`, `killed` (by `cause`), and `uptime_seconds`.

The control socket answers even while the work socket is saturated, and it is host-local, so the
scraped process must be on the cell's own host. Watch `queued`, `queue_high_water`, `cancelled`, and
`killed` by cause.
The control socket answers when the work socket is saturated. The control socket is local to its host.
Thus the scraped process must be on the same host as the cell.

### Per-call telemetry

Subscribe to the `perform.hot_cell` Active Support Notification for logging, metrics, or both. It
fires on every call, success or failure, and it is the only signal that survives a dead cell -- an
unreachable socket comes back as code `unavailable`, so the primary alarm belongs here.

In a Rails application, `HotCell::LogSubscriber` already writes one `info` line per call to the Rails log:

```
HotCell (41.2ms) {"cell":"images","operation":"active_storage.transformers.image.vips","code":"ok","perform_ms":38,"duration_ms":41.2,"bytes_in":20480,"bytes_out":8192}
```

A failed call adds `cause` and `stderr` when it has them. A call interrupted by an exception, such as the
application's own request timeout, logs the exception's class in place of the code. To turn the line off, call
`HotCell::LogSubscriber.detach_from :hot_cell` in an initializer.
The application sends the `perform.hot_cell` Active Support Notification for each call, successful or
failed. Thus the notification reports a dead cell: an unreachable socket gives the code `unavailable`.
`HotCell::LogSubscriber` and `yabeda-hotcell` subscribe to it. To record other data, subscribe to it.

### Container healthcheck

Expand All @@ -592,22 +622,24 @@ remember to use this.

### Rails healthcheck

hotcell-client ships two controllers. Your application adds the routes.
`hotcell-client` defines two controllers, `HotCell::HealthController` and `HotCell::DiagnosticsController`.
Add a route for each to the application's `config/routes.rb`, as the example below shows.

`HotCell::HealthController` asks each registered cell for `describe` and `metrics` over its control socket.
It returns `OK` with a 200 when at least one cell is registered and every cell answers, and `FAIL` with a
503 otherwise. These calls take no worker, so the endpoint can be public, like `/up`.
It returns `OK` with a 200 when at least one cell is registered and each cell answers. Otherwise, it
returns `FAIL` with a 503. These calls take no worker. Thus you can make the endpoint public, like `/up`.

`HotCell::DiagnosticsController` returns the result of every check as JSON, with a 503 if any check fails.
Along with `describe` and `metrics`, it sends `health.echo` and `health.reopen` over the work socket. Each
of those takes a worker, so put this endpoint behind authentication.
round trip takes a worker. Put this endpoint behind authentication.

The control socket carries no file descriptors, so `describe` and `metrics` succeed even when your
application cannot use the work socket. Only the round trips exercise the work socket. A cell without the
shared group passes `health.echo` and fails `health.reopen` with `EACCES`. The cell serves both operations
only if one of its operation files has `require "hot_cell/health_operations"`.
Of the cell's two sockets, only the work socket carries file descriptors. Thus `describe` and `metrics`
succeed when the application cannot use the work socket. Only the round trips test the work socket. A
cell without the shared group passes `health.echo` and fails `health.reopen` with `EACCES`. Add
`require "hot_cell/health_operations"` to one of the cell's operation files. Without it, the cell
answers `unsupported` for both operations.

Set the diagnostics controller's superclass in an initializer, then add the routes:
Set the diagnostics controller's superclass in an initializer. Then add the routes:

```ruby
# config/initializers/hotcell.rb
Expand Down
2 changes: 1 addition & 1 deletion docs/TUNING.md
Original file line number Diff line number Diff line change
Expand Up @@ -143,7 +143,7 @@ safe to decide per caller.
| `killed_by` by cause | the only legitimate reason to tighten a limit |
| `queued_ms` p95 rising, `perform_ms` p95 flat | the cell needs more workers, not faster ones |
| `perform_ms` p95 rising | the work got more expensive; check for a library upgrade |
| `queue_high_water` near `queue_size` | no headroom left |
| `queued` near `queue_size`, or `queue_high_water` rising toward it | no headroom left; `queue_high_water` resets only at boot |
| `capacity` above zero in steady state | under-provisioned |
| `unavailable` | the cell is down, restarting, or unreachable |
| `unreadable` rate | worth watching after a toolchain upgrade |
Expand Down
Loading