Conversation
/_replicationStatus fell back to the cfg's target state while no replicator had published status, so a replication reported running before anything started. - If this is assigned to this node, return starting if it hasn't get made it to running yet. - add starting to activeOnly=true filter - Add unassigned to the openapi docs
Redocly previews |
Contributor
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The empty target state is treated as running during startup but remains uncovered and incorrectly reported by the new fallback.
Review effort: Balanced
Findings: 1
What changed in this PR
Updates ISGR status reporting so pending replications report starting instead of prematurely reporting running.
Changes:
- Reports assigned, not-yet-started replications as
starting. - Includes
startingreplications inactiveOnlyresults. - Updates tests and OpenAPI documentation.
| File | Description |
|---|---|
docs/api/components/schemas.yaml |
Documents the unassigned state. |
docs/api/components/parameters.yaml |
Clarifies activeOnly behavior. |
db/sg_replicate_cfg.go |
Corrects pending replication status handling. |
db/sg_replicate_cfg_test.go |
Tests pre-start and active-only status behavior. |
💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| require.NoError(t, err) | ||
| require.NoError(t, mgr.StartLocalNode(localNodeUUID, nil)) | ||
|
|
||
| for _, targetState := range []string{ReplicationStateRunning, ReplicationStateStopped} { |
- An unset target state starts like running, so report it as starting too. - activeOnly now includes reconnecting replications. - Cover replications assigned to another node, live replicator states, the activeOnly REST query, and upserts while starting. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Create replications with UpsertReplication and let rebalance assign them, with the other node joining first where it owns the replication. - Move the unset target state case into its own test. - Run activeOnly against a real passive node, so running and reconnecting both come from a replicator, and drop the test that forced states. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Contributor
|
Droid finished @torcolvin's task —— View job The status fallback and activeOnly behavior changes look correct and are backed by focused tests for the startup window and reconnecting behavior. One small OpenAPI wording mismatch remains: the |
bbrks
previously approved these changes
Oct 2, 2026
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
torcolvin
enabled auto-merge (squash)
October 2, 2026 15:33
bbrks
approved these changes
Oct 2, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

CBG-5838
TestRequireReplicatorStoppedBeforeUpsertfailed in CI with a 400 on the upsert that follows the stop:Cause
GetReplicationStatusfalls back to the cfg's target state while no replicator has published status, so a replication reportsrunningbefore anything has started. The wait forrunningreturned on its first poll.RefreshReplicationCfgreads the cfg, then starts assigned replications while holdingactiveReplicatorsLock. The test carried on with that start still in flight, stopped the replication, and the start landed afterwards on the pre-stop cfg. The upsert then found a running replicator. The 20.7 ms the failing PUT took is it waiting onactiveReplicatorsLockwhile that start held it.Fix
Report
startingfor that stub. The target state stays the answer for every other state, becausestoppedanderrorare true when no replicator has reported, andrunningis not.Two things already agree with this:
PUT _replicationStatus/{id}?action=startanswersstartingthroughtransitionStateNamewhenever the current state differs from the target.UpsertReplicationgates onstopped, so the rejection of a config change during a start is unchanged.A wait for
runningnow needs either a local replicator in that state or a status document, which only a running replicator writes. No start can be in flight when the wait returns, because the refresh holdsactiveReplicatorsLockacrossStartand the status read needs the read lock.Behaviour changes
GET _replicationStatusand the_clusterresponse reportstartinginstead ofrunningfor a replication that no replicator has published status for.?activeOnly=truereturnsstartingas well asrunning, so a replication does not drop out of the listing while it starts. The spec description of the parameter is updated to match.PUT _replicationStatus/{id}?action=startanswersstartinginstead ofrunningin that window.runningnow waits for a replicator instead of returning at once.The status document is still the only cross-node channel, so a replication whose status write fails, or whose document expired, reads
startingon other nodes until the next successful publish. That trades a silent falserunningfor a visible falsestarting.Tests
TestReplicationStatusBeforeReplicatorStartsadds a replication assigned to the local node, never runs a refresh, and asserts the reported status. It fails on the old code withexpected: "starting", actual: "running"and has no timing dependence.Verified by delaying
RefreshReplicationCfgbetween its cfg read and the start, which widens the window deterministically: the endpoint reportedrunning3/3 before the change andstarting3/3 after, andTestRequireReplicatorStoppedBeforeUpsertpasses 5/5 under that delay at 8.1 s per run. Without the delay it passes 100/100 under-race. The fullrest/replicatortestpackage, thedbreplication tests and therestreplication tests pass.unassignedis added to theISGRReplicationStateenum in the API spec. The code already returned it.🤖 Generated with Claude Code