Skip to content

[Feature Request]: Global "Maintenance Mode" toggle (pause scheduler + monitors + alerts together) #441

Description

@para6666

Is there an existing feature request for this?

  • I have searched the existing issues

What is your feature request?

Problem:

Right now there is no single switch to put a whole xyOps instance (or a server group) into maintenance. Scheduler, Alerts, and Monitors are three separate systems, and each behaves differently:

Scheduler: has a global Pause switch, but it only stops automatic event triggers (schedule/interval/single/plugin).
Alerts: each Alert definition has its own enabled flag. You have to turn it off one by one. Alerts also have no category of their own - the groups field only says which servers the alert applies to, not a group of alerts. So you can't disable a set of alerts together.
Monitors: a Monitor has no enabled flag at all. It only has display, which just hides the graph in the UI. The monitor keeps running every minute no matter what.
Servers: setting enabled to false removes the server from job selection, but the server stays online and keeps sending metrics, so Monitors and Alerts keep running on it.
Result: during a real maintenance window (e.g. patching, planned downtime), pausing the scheduler is not enough. Alerts keep firing unless you disable every single alert by hand, and Monitors keep collecting data no matter what you do.

Suggestion

This is just an idea, not a hard requirement - it would just be nice to have. Something like a "Maintenance Mode" with two levels could help:

  1. Global: one toggle (next to the existing scheduler pause) that also stops alerts from firing/running actions, and optionally pauses monitor sampling for the whole instance.

  2. Per server / server group: a "Maintenance" flag on a Server or Group that:

    • Removes the server from job target selection (same as enabled false today).
    • Stops alerts from firing (or at least stops their actions) for that server/group.
    • Optionally pauses monitor sampling for that server/group, or at least marks the data points as "maintenance" so they don't mess up the history/graphs.
  3. Grouping for Alerts and Monitors themselves: a category/tag field on Alert and Monitor definitions (separate from the existing groups field, which is about servers, not about grouping the alerts/monitors) would let you:

    • Turn off a whole category of alerts at once (e.g. all "disk" alerts, or everything tied to one application).
    • Pause a whole category of monitors at once.
    • Have one button in the UI ("Disable Category") instead of clicking through every single alert or monitor one at a time.

It would also be great if all of the above (global toggle, per-server/group flag, category enable/disable) could be set through the API as well, not just the UI.

Code of Conduct

  • I agree to follow this project's Code of Conduct

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions