Repository navigation
Replies: 1 comment
|
Hi @Oparop, thanks for trying xyOps, and for the question! xySat does not have an autonomous scheduling mode. It is specifically designed to be a job runner, while scheduling and job dispatch remain the responsibility of the active primary conductor. For production resilience, we recommend running multiple conductors. xyOps supports a primary conductor with hot standby peers and automatic primary election. If the primary conductor becomes unavailable, another conductor can take over scheduling and coordination. xySat also maintains awareness of the conductor cluster and can reconnect to the newly elected primary. On the worker side, we recommend running multiple xySat servers, placing them in a group, and targeting your events at that group rather than at one specific server. When launching a job, xyOps only considers workers that are currently online and eligible. Any server that has dropped offline is automatically omitted when selecting the target. This scenario is also handled very well by the job queue. If you configure both Max Concurrent Jobs and a nonzero Max Queue Limit on the event, queuing is enabled. If no eligible target server is available when the job is launched, the job is automatically placed into the queue and waits for a suitable server to become available. Once a worker reconnects, the job is dequeued and launched normally. There is another layer of resilience for jobs that have already started. If xySat loses its connection to the conductor while running a job, the job continues running locally. xySat holds its progress updates in memory until it can reconnect to the primary conductor, even if a different conductor has become primary in the meantime. If the job completes while disconnected, xySat also handles that as well. It retains the completion status, job data, and output files locally until it can reconnect and successfully send the final completion request to a conductor. If you need scheduled runs to be recovered after a period when no conductor was available, you can also add a Catch-Up trigger modifier. This causes missed schedule occurrences to be evaluated when scheduling resumes. So the system is designed to be quite resilient, but that resilience is provided in layers: multiple conductors for scheduler availability, server groups and queues for worker availability, and xySat's local buffering for jobs already in progress. With those pieces configured appropriately, temporary worker, network, or conductor outages should not result in lost work. Help this helps. |
Uh oh!
There was an error while loading. Please reload this page.
Hi,
I've been testing xyOps and I really like the Conductor/Satellite architecture. However, I ran into an issue regarding resilience:
When the Conductor is unable to reach a Satellite (network split, Conductor down, temporary outage, etc.), the jobs that are scheduled on that Satellite simply don't run anymore. It seems the Satellite fully relies on the Conductor to trigger scheduled jobs, rather than having its own local scheduling capability.
Is there a way to configure a Satellite to run in an "autonomous" mode — i.e., if it loses communication with the Conductor for a certain period (heartbeat timeout, etc.), it would start evaluating and triggering its own scheduled jobs locally, based on the last known job definitions/schedule it received from the Conductor?
This would be very useful for production setups where network reliability between Conductor and Satellites isn't guaranteed, and where missing scheduled jobs during an outage is not acceptable.
All reactions