How should compatibility failures be handled during rolling deployments? #473
Replies: 4 comments 4 replies
|
I think we should explore the options in more detail before making the decision, because each option has different implications for retry semantics, ordering, observability, and mixed-version deployments. Common prerequisitesRegardless of the selected option, the implementation should introduce explicit exceptions for:
These failures should be explicitly separated from exceptions thrown after a handler has been invoked. This gives us two distinct categories:
This distinction should exist independently of how compatibility state is retained or retried. I would also prefer to avoid database schema changes such as adding columns or tables. The existing status column stores textual values, so adding another record status does not require a database migration. It is still an application-model change that needs defined upgrade behavior. Option D: Lazy or metadata-first record materializationRepositories currently materialize every payload before returning records to the scheduler. One incompatible record can therefore prevent earlier compatible records from being processed: With metadata-first materialization, payloads are resolved individually during processing: This allows the scheduler to identify the exact incompatible record and process compatible predecessors:
The disadvantage is that this changes the materialization semantics of Option A: Instance-local compatibility exclusionsIn simple terms, Option A works as follows:
Advantages
Disadvantages
The current PR limits exclusions to a fixed number of entries. With more incompatible entries than the capacity and a small processing batch, the scheduler could repeatedly rediscover evicted exclusions and fail to reach otherwise compatible keys: This can result in starvation rather than only additional queries or log entries. Option A with Option DCombining Option A with metadata-first materialization would allow compatible predecessors to be processed before an incompatible record is encountered. It would also allow compatibility handling to respect
However, this combination may introduce efficiency concerns. The scheduler could repeatedly fetch metadata for many locally excluded records while processing only a small number of compatible records. Avoiding that would require careful repository query, paging, and exclusion behavior. Therefore, I am not yet sure whether the additional complexity of combining Option A with lazy materialization is justified. Option B: Persist a dedicated compatibility stateOption B introduces a dedicated When an instance cannot load the payload type, deserialize the payload or context, or find the registered handler, it marks the exact record as Normal processing queries omit This option requires metadata-first materialization. The scheduler must receive enough record metadata to identify and update the problematic record even when its payload cannot be deserialized. With eager materialization, deserialization can fail before the scheduler receives any records. An The existing status column stores textual values, so adding Healing strategiesA healing process is required to make 1. Reactivate all incompatible recordsWhen an instance acquires a partition, it changes all UPDATE outbox_record
SET status = 'NEW'
WHERE status = 'INCOMPATIBLE'
AND partition IN (...)The normal scheduler then reevaluates each record using its regular batching, ordering, transaction, and concurrency controls. If a record is still incompatible, it is marked 2. Fully evaluate every incompatible recordWhen an instance acquires a partition, the healing process loads each
Only records that pass all checks are changed back to This strategy requires reading and deserializing every incompatible record during healing. The normal scheduler subsequently processes the reactivated records. 3. Perform metadata preflightWhen an instance acquires a partition, the healing process retrieves the distinct payload type names and handler IDs used by The instance then determines:
Records whose payload type and handler are both available are changed back to Conceptually, the update is: UPDATE outbox_record
SET status = 'NEW'
WHERE status = 'INCOMPATIBLE'
AND partition IN (...)
AND handler_id NOT IN (:unavailableHandlerIds)
AND record_type NOT IN (:unavailablePayloadTypes)The unavailable predicates would only be added when their corresponding sets are non-empty. This is a metadata preflight rather than complete compatibility validation. Payload and context deserialization remain record-specific and are performed by the normal scheduler. A record with an available class and handler but incompatible serialized data may therefore transition to The implementation must account for database parameter limits when the unavailable sets are large and provide equivalent behavior for JDBC, JPA, and MongoDB. Advantages
Disadvantages
Option C: Reuse the delivery failure stateOption C reuses the existing There are several possible variants. Option C.1: Treat compatibility like an ordinary delivery failureThis is the simplest variant and is close to the current missing-handler behavior:
Advantages
Disadvantages
Option C.2: Use negative
|
| Concern | Option A: local exclusions | Option B: INCOMPATIBLE status |
Option C: reuse failure state |
|---|---|---|---|
| Consumes delivery retries | No | No | Yes, or changes counter semantics |
| Database DDL required | No | No, uses the existing textual status column | No |
| Persisted visibility | No | Strong | Partial |
| Recovery after ownership change | Automatic through local state reset | Requires healing | Not reliable after FAILED |
| Permanent incompatibility visible | Only indirectly | Yes | Yes after FAILED |
| Additional lifecycle | No | Yes, healing lifecycle | Variant-dependent, compatibility retry for some variants |
| Mixed-version application risk | Low | Medium | Medium |
| Implementation complexity | Medium | Medium | Medium |
| Risk of premature permanent failure | No | No | High |
Current view
My current conclusions are:
- Explicit compatibility exceptions should be introduced regardless of the selected option.
- Compatibility failures should not consume delivery retries because no delivery was attempted.
- Option C should be rejected because it either consumes delivery retries or overloads existing fields with incompatible semantics.
- The main decision is therefore between Option A and Option B.
Option A is simpler and avoids changing the persisted record lifecycle, but it provides weak observability and requires careful handling of exclusion capacity, query behavior, and starvation.
Option B provides explicit persisted state and more direct record-level behavior, but it introduces a healing lifecycle, requires metadata-first materialization.
If the priority for 1.10.0 is the smallest behavioral correction under a strict expand-and-contract deployment contract, Option A may be the more practical starting point.
If persistent operational visibility and explicit management of permanently incompatible records are requirements, Option B provides a clearer model, but it should be treated as a larger architectural change.
|
After reflecting further on the alternatives, I think we may be solving a much larger problem than #464 actually requires. Producer and Consumer ResponsibilitiesNamastack Outbox contains both sides of an asynchronous contract:
This remains true even though both sides are implemented by the same library and often run inside the same application. During a rolling deployment, producers and consumers from different application versions coexist while sharing the same persisted records. The same compatibility rules therefore apply as in other event-driven systems. Before a producer writes a new message contract, every consumer that may receive that message must be able to understand it. For Namastack Outbox, that contract includes:
This is why expand-and-contract is a required deployment contract:
If an instance produces a record before every possible consumer can understand it, the deployment has violated that contract. The library should fail safely in this situation, but it does not need to preserve normal throughput or provide a complete secondary lifecycle for an invalid deployment. I think we sometimes lose sight of this producer/consumer relationship when discussing the library internally. We then start treating compatibility failures as something the library must fully manage, even though the primary solution is the same as in any other event-driven architecture: consumers must become compatible before producers use the new contract. A Simpler Compatibility SafeguardThe essential defect in the current behavior is narrow:
I propose addressing exactly this problem:
Later records for the affected key would not be processed during that attempt. Other record keys already selected in the same batch could complete normally. Instance-Wide Compatibility CooldownWithout additional handling, an incompatible record could be selected and logged during every polling cycle. With a two-second polling interval, that could produce a database query and warning every two seconds until the instance is replaced. We can avoid this without maintaining payload-type, handler-ID, or record-key exclusion sets. When the scheduler encounters a compatibility failure, it can set a single instance-local timestamp: The current batch is allowed to finish, but that scheduler instance does not query for another batch during the cooldown. After the cooldown, it probes again:
This uses one local timestamp rather than collections of incompatible identifiers. It does not introduce persisted retry state, another record lifecycle, repository query changes, or database schema changes. ConsequencesThis deliberately prioritizes safety over availability after a deployment contract violation:
I consider the temporary loss of throughput acceptable here. Applications that require uninterrupted processing of all compatible keys during a rollout already have a supported solution: follow the documented expand-and-contract sequence. ScopeI would keep arbitrary payload and context deserialization failures outside this change. They may indicate version incompatibility, corrupted data, or another serialization problem. They already occur before handler execution and therefore do not need to consume handler delivery retries, but poison-record management deserves a separate decision if it becomes a concrete requirement. With this approach, we could remove the following parts of the current proposal:
The result would be a much smaller change focused on the actual correctness issue: a missing handler must not consume a delivery retry when delivery never started. I therefore propose this minimal compatibility handling with a scheduler-wide local cooldown as the direction for 1.10.0. Persisted incompatibility states, healing, metadata-first materialization, and poison-record handling can remain separate proposals if concrete requirements emerge. |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Context
During a rolling deployment, different application versions can process records from the same outbox.
A newer instance may persist a record whose payload type or handler is not available on an older instance that currently owns the corresponding partition.
This creates two compatibility failures before handler execution:
The problem is described in #464. PR #470 contains a working implementation of one possible approach. The implementation should be treated as a concrete proposal and source of practical evidence, rather than as a decision that has already been finalized.
Decision to Make
We need to decide how Namastack Outbox should classify, retain, retry, and expose compatibility failures during mixed-version deployments.
In particular:
Terminology
Delivery failure
A registered handler was invoked and threw an exception.
Compatibility failure
Handler execution could not begin because the current instance cannot load the persisted payload type or cannot find the persisted handler.
This distinction is important because a compatibility failure does not prove that delivery itself has failed.
Required Semantics
The selected solution should address the following:
record when doing so could violate ordering.
Deployment Assumption
A library cannot reliably determine whether application code is:
The proposed deployment contract therefore follows an expand-and-contract model:
For removal:
This assumption is part of the decision and may be challenged in this discussion.
Considered Options
Option A: Instance-local compatibility exclusions
This is the approach currently implemented in PR #470.
failureCount.NEW.instance-local data structure.
Advantages
Disadvantages
is corrected.
during the first encounter with an incompatible record.
Option B: Persist a dedicated compatibility state
Persist a separate status, retry timestamp, failure category, or compatibility counter.
Advantages
state.
Disadvantages
state.
temporary or permanent.
Option C: Reuse the delivery failure state
Increment
failureCountand process compatibility through the existing retry policy.Variants include interpreting negative counts or recognizing a compatibility category from
failureReason.Advantages
Disadvantages
FAILEDbefore its handler has ever been invoked.Option D: Lazy or metadata-first record materialization
Load record metadata first and defer payload resolution until the payload is accessed.
This could allow compatible predecessors to be processed before the first incompatible record is encountered.
Advantages
Disadvantages
OutboxRecord.considerations.
This option may therefore be complementary to the other options rather than a replacement for them.
Proposed Direction
For version 1.10.0, adopt Option A together with the documented expand-and-contract deployment contract.
The reasons are:
Accept the following limitations:
Lazy or metadata-first materialization should be evaluated separately because it addresses repository materialization and predecessor isolation rather than the core failure semantics.
Observability should be designed separately through the planned SPI so that this decision does not prematurely define that public API.
Questions for Review
NEWwithout consumingfailureCount?contract?
Decision Process
All reactions