Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,7 +37,7 @@ alt="Meshery Logo" width="70%" /></picture></a><br /><br /></p>

MeshSync is Meshery's event-driven, continuous discovery and synchronization engine. It ensures that the configuration and operational state of Kubernetes (and any supported Meshery platform) are known to Meshery Server. When deployed into a Kubernetes cluster, MeshSync runs as a custom controller under the control of [Meshery Operator](https://docs.meshery.io/concepts/architecture/operator) and publishes resource changes over Meshery Broker (NATS).

MeshSync runs in one of two modes: **nats** (default - publishes Kubernetes resource events to NATS) and **file** (writes deduplicated cluster snapshots to disk, with no NATS or CRD dependency). Run `meshsync --help` for input parameters.
MeshSync runs in one of two output modes: **broker** (default - publishes Kubernetes resource events to Meshery Broker over NATS) and **file** (writes deduplicated cluster snapshots to disk, with no NATS or CRD dependency). Run `meshsync --help` for input parameters.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch. Swept the remaining mentions in ed83861: the -outputNamespaces / -outputResources flag help text in main.go and the determineUseCRDFlag comment in pkg/lib/meshsync/meshsync.go (which also referenced a nonexistent channel output mode).


## Documentation

Expand Down
Empty file removed all-certmgs.yaml
Empty file.
139 changes: 0 additions & 139 deletions cert-manager-cainjector-c7d4dbdd9-jlstt.yaml

This file was deleted.

171 changes: 0 additions & 171 deletions cert-manager-webhook-847d7676c9-89rtq.yaml

This file was deleted.

2 changes: 1 addition & 1 deletion docs/agent-instructions/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,7 @@ main.go --parses CLI flags--> pkg/lib/meshsync.Run(...)

- `main.go` parses flags (`-output`, `-outputFile`, `-outputNamespaces`, `-outputResources`, `-stopAfter`) and calls `pkg/lib/meshsync.Run(...)`.
- `meshsync.Handler` (`meshsync/meshsync.go`) holds the config, logger, broker handle, dynamic informer factory, kube client, channel pool, output writer, and output-filtration config. `meshsync.New(...)` wires them together and derives the cluster ID via `pkg/utils.GetClusterID`.
- `GetDynamicInformer` builds a `dynamicinformer.DynamicSharedInformerFactory` filtered by a label selector derived from the config's informer blacklist.
- `GetDynamicInformer` builds a `dynamicinformer.DynamicSharedInformerFactory`. Resource filtering happens in the watch-list config (`internal/config/crd_config.go` decides which informers get registered); the factory's list-options hook (`GetListOptionsFunc`) is a deliberate no-op.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated in ed83861. Corrected both spots in docs/design/fd5-periodic-reconciliation.md: the reconcile-LIST parenthetical and the scale-mitigation bullet now state that GetListOptionsFunc is a deliberate no-op and that reconcile scope comes from which informer pipelines are registered from the watch-list config.


## Discovery Pipeline (`internal/pipeline`)

Expand Down
4 changes: 2 additions & 2 deletions docs/design/fd5-periodic-reconciliation.md
Original file line number Diff line number Diff line change
Expand Up @@ -107,7 +107,7 @@ Implemented as the third branch above: for each reconciled GVR, take the key set
- **`meshsync/meshsync.go`**: `Handler` gets no new fields — a reconcile loop is a *behavior* of `Run()`, not new handler state, keeping with `pipeline.New`'s per-call-fresh philosophy the architecture doc already mandates.
- **`meshsync/handlers.go`** — the core of the feature:
- **New unexported function `(h *Handler) reconcileLoop(intervalStr string, stopCh channels.StopChannel)`**, started as a goroutine from `Run()` (alongside the existing `startAndTrackDiscovery()` call, `handlers.go:57-58`) only when the resolved interval is non-zero. On each tick (a `time.Ticker`, jittered per §"scale" below), it iterates `h.stores` (the same map `startDiscovery` populates at `discovery.go:48` — `h.stores[gvr-name] = cache.Store`) and, for each GVR present, performs:
1. A fresh, filtered LIST against `h.kubeClient.DynamicKubeClient` for that GVR (respecting the same `GetListOptionsFunc` blacklist label selector already used for informer construction, `meshsync/meshsync.go:34-53`, so reconcile never re-lists resource kinds the informer itself was configured to exclude).
1. A fresh, filtered LIST against `h.kubeClient.DynamicKubeClient` for that GVR (reconcile iterates only the GVRs whose informer pipelines were registered from the watch-list config; `GetListOptionsFunc` is a deliberate no-op, so there is no list-time label filtering to rely on).
2. Diff against `store.List()`/`store.ListKeys()` per §3(b)/(c).
3. Publish only drifted objects through the *existing* `internal/pipeline.RegisterInformer.publishItem`-equivalent path — since `publishItem` is a method on `RegisterInformer` (pipeline-internal, not exported), the cleanest reuse is to export a small helper in `internal/pipeline` (e.g. `pipeline.PublishDrift(log, outputWriter, clusterID, outputFiltration, cfg PipelineConfig, obj *unstructured.Unstructured, evtype broker.EventType) error`) that wraps the existing `model.ParseList` + `checkMustSkip` + `outputWriter.Write` sequence (lines 93-111 of `internal/pipeline/handlers.go`), called from both `RegisterInformer.publishItem` (refactored to delegate to it) and the new reconcile loop — avoiding duplicating the skip/filter/write logic in two places, which the CLAUDE.md "consistency with existing patterns... migrate the existing code accordingly" directive requires rather than a second parallel implementation.
- Guard the whole loop on the `channels.Stop` channel exactly like `WatchCRDs` does (`handlers.go:359-361`) for clean shutdown.
Expand Down Expand Up @@ -151,7 +151,7 @@ The default-off interval (§3(d)) **is** the feature flag — no separate boolea

## 8. Risks, Failure Modes, Perf/Scale

- **Re-list storm on large clusters.** A naive "list every configured GVR every tick" is the primary named risk. Mitigations designed in: (1) default-off (§3(d)); (2) enforced floor interval (5m minimum, §3(d)); (3) intra-tick staggering across GVRs with bounded concurrency (§4); (4) fleet-wide jitter via the existing `jitter()` helper (§4) so many clusters' MeshSync instances don't synchronize; (5) reconcile scope is **exactly** the informer's existing configured GVR set filtered through `GetListOptionsFunc`'s blacklist label selector (§4) — it never lists more than what's already being watched, and it respects `-outputResources` (via `outputFiltration.ResourceSet`, already applied at the pipeline-construction/write layer) so a restricted deployment reconciles only its restricted resource set. **`-outputNamespaces` is a write-time filter, not an informer/LIST-time scope** (§2) — a genuinely namespace-scoped reconcile (listing only within the allowed namespaces at the API-server level, cheaper than a cluster-wide LIST filtered client-side) is not achievable without deeper informer-factory changes (`NewFilteredDynamicSharedInformerFactory` takes one cluster-wide-or-single-namespace argument, not a set) — flagged explicitly as an out-of-scope enhancement in §11, with the interim mitigation being that write-time filtering still means a namespace-restricted deployment publishes no more drift than it would publish live events, even though its LIST cost is unreduced.
- **Re-list storm on large clusters.** A naive "list every configured GVR every tick" is the primary named risk. Mitigations designed in: (1) default-off (§3(d)); (2) enforced floor interval (5m minimum, §3(d)); (3) intra-tick staggering across GVRs with bounded concurrency (§4); (4) fleet-wide jitter via the existing `jitter()` helper (§4) so many clusters' MeshSync instances don't synchronize; (5) reconcile scope is **exactly** the informer's existing configured GVR set as registered from the watch-list config (§4; `GetListOptionsFunc` is a deliberate no-op) — it never lists more than what's already being watched, and it respects `-outputResources` (via `outputFiltration.ResourceSet`, already applied at the pipeline-construction/write layer) so a restricted deployment reconciles only its restricted resource set. **`-outputNamespaces` is a write-time filter, not an informer/LIST-time scope** (§2) — a genuinely namespace-scoped reconcile (listing only within the allowed namespaces at the API-server level, cheaper than a cluster-wide LIST filtered client-side) is not achievable without deeper informer-factory changes (`NewFilteredDynamicSharedInformerFactory` takes one cluster-wide-or-single-namespace argument, not a set) — flagged explicitly as an out-of-scope enhancement in §11, with the interim mitigation being that write-time filtering still means a namespace-restricted deployment publishes no more drift than it would publish live events, even though its LIST cost is unreduced.
- **Dedup/RV-suppression interaction failing silently.** If the reconcile loop's own RV-diff (§3(b)) has a bug and always treats "unchanged" as "changed," it defeats the entire purpose and re-floods Server every tick. Test plan (§9) includes an explicit "steady state produces zero publishes" assertion, not just a "drift produces N publishes" assertion — the null case is equally load-bearing.
- **Missed-delete false positives from partial/paginated or rate-limited LIST failures.** A LIST that fails midway or is truncated must never be diffed as "everything not returned is deleted" — a failed/partial LIST must abort that GVR's reconcile pass for that tick (log + error code, per §4) rather than falling through to the diff, or every transient API-server hiccup becomes a false mass-delete storm published to Server. This is the single most dangerous failure mode in the design and must be a hard early-return, tested explicitly (§9).
- **Stale reconstructed object on missed-delete.** The synthesized Delete event (§3(c)) carries the last-known Store copy, not a live cluster read (the object no longer exists to read) — downstream consumers relying on delete-event payload freshness (there are none identified in Server's `Delete` handling, which only uses it for a keyed row delete, `models/meshsync_events.go:242-246`) are unaffected, but this should be called out in the PR description as an intentional, unavoidable characteristic.
Expand Down
4 changes: 2 additions & 2 deletions main.go
Original file line number Diff line number Diff line change
Expand Up @@ -94,14 +94,14 @@ func parseFlags() {
&outputNamespacesString,
"outputNamespaces",
"",
"k8s namespaces for which limit the output, comma separated list f.e. \"default,agile-otter\", applicable for both nats and file output mode",
"k8s namespaces for which limit the output, comma separated list f.e. \"default,agile-otter\", applicable for both broker and file output mode",
)
var outputResourcesString string
flag.StringVar(
&outputResourcesString,
"outputResources",
"",
"k8s resources for which limit the output, comma separated case insensitive list of k8s resources, f.e. \"pod,deployment,service\", applicable for both nats and file output mode",
"k8s resources for which limit the output, comma separated case insensitive list of k8s resources, f.e. \"pod,deployment,service\", applicable for both broker and file output mode",
)
flag.DurationVar(
&stopAfterDuration,
Expand Down
4 changes: 2 additions & 2 deletions pkg/lib/meshsync/meshsync.go
Original file line number Diff line number Diff line change
Expand Up @@ -370,8 +370,8 @@ func determineUseCRDFlag(
log logger.Handler,
kubeClient *mesherykube.Client,
) bool {
// if output mode is not nats generally it is not expected to have CRD present in cluster.
// theoretically CRD could be present even in file, channel output mode.
// if output mode is not broker generally it is not expected to have CRD present in cluster.
// theoretically CRD could be present even in file output mode.
// hence check if CRD are present in the cluster,
// and only skip them if it is not present.
crd, errGetMeshsyncCRD := config.GetMeshsyncCRD(kubeClient.DynamicKubeClient)
Expand Down
Loading