Skip to content

Latest commit

 

History

History
613 lines (495 loc) · 39.2 KB

File metadata and controls

613 lines (495 loc) · 39.2 KB

VCF Operations integration (research track)

Status: lab-verified, and the first slice is implemented. Written 2026-08-31 from the API specifications, verified the same day against a live VCF 9.1 lab, and shipped as working plugins immediately after (see What is implemented). Every question the checklist asked is answered in Lab verification results, which supersedes any guess earlier in this document.

Why this doc exists

vCheck asks vCenter "what is the configuration right now?" and every plugin applies its own hardcoded threshold. VCF Operations has already done the analysis: time-series, dynamic thresholds, symptom correlation, capacity forecasting, and a Broadcom-maintained diagnostics rule engine.

Querying Operations for raw inventory would be pointless duplication. Querying it for verdicts is the entire value of this track. That reframe is what the rest of this document builds on.

Where the specs are

The API specifications are not in this repo. They are published at https://github.com/vmware/vcf-api-specs (bundle 9.1.0.0-25372366). The bundle is ~97 MB, mostly the vSphere WSDL tree, and it does not belong in this repository. The paths below are relative to its specifications/ folder.

Three services, not one

specifications/vcf-operations/ covers three separate services with three different base paths and three different auth headers. Conflating them is the first trap.

Spec file Base path Auth header Surface
vcf-operations-openapi.json /suite-api Authorization: OpsToken <token> 343 paths, 587 schemas
log-management-openapi.json /api/v2 X-JWT-Token: <jwt> 24 paths (log search, forwarders, agents)
realtime-metrics/realtime-metrics-openapi.yaml /api/v1 Authorization: Bearer <jwt> 4 paths, PromQL, Prometheus-compatible

Plus, adjacent and relevant later: fleet-lcm/fleet-lcm-openapi.yaml (38 paths: upgrade plans, component status, release versions) and sddc-manager/sddc-manager-openapi.json.

The auth chain

Documented in the description block of realtime-metrics-openapi.yaml (roughly lines 20 to 55). That is the only place in the whole bundle where the header format is spelled out, so it is recorded here:

  1. POST /suite-api/api/auth/token/acquire with {"username","password","authSource"} returns {"token","validity","expiresAt","roles"}. Validity is extended on each call and reset to 6 hours from the last call.
  2. GET /suite-api/api/integrations/services with Authorization: OpsToken <token> returns the registered services. Find the entry whose type is VCF_VODAP and take its service key.
  3. POST /suite-api/api/auth/token/exchange with {"serviceKeys":["<key>"]} and the OpsToken header returns a JWT.
  4. Use that JWT as Authorization: Bearer <jwt> for realtime metrics, and as X-JWT-Token for log management.

Note the header prefix is OpsToken, not the vRealizeOpsToken used by Aria Operations 8.x. Verified on the 9.1 lab: OpsToken works (see the results below).

The PowerShell SDK changes the plan

VMware.Sdk.Vcf.Ops ships as part of VCF.PowerCLI 9.1, which this project already mandates. It is generated from exactly the suite-api spec above.

Measured 2026-08-31 (these numbers are real):

Fact Value
Module version 13.5.0.25380678 (also 13.4.0.24798382 present)
Cmdlet count 822 (505 Invoke-*, 313 Initialize-*, plus Connect/Disconnect/Get/Set)
Import time, fresh process 0.80s on macOS, but 4.8 to 5.4s on Windows (measured 2026-08-31; 0.11s on a re-import inside the same process)
Auto-loaded by VMware.VimAutomation.Core? No. It is a separate, explicit import.

On Windows the fresh-process import is ~5s, not 0.8s: roughly 5% of the ~110s full-run budget spent before a single call is made. That settles the design question: the import must sit behind the same conditional as the rest of this track, and must not happen at all when no Operations server is configured.

Decision: use the SDK for /suite-api, hand-roll REST only for the other two services. The SDK has no log-search and no PromQL cmdlets (checked: only log forwarding and log configuration cmdlets exist, Invoke-VcfOpsGetLogForwardingConfiguration and friends). Log search and realtime metrics therefore need Invoke-RestMethod plus the JWT exchange above.

Verified cmdlet signatures

Confirmed by Get-Command against 13.5.0.25380678 on 2026-08-31, and since called against the lab (see the results below).

Connect-VcfOpsServer [-Server] <string[]> [-Credential <pscredential>] [-User <string>]
    [-Password <securestring>] [-AuthSource <string>] [-Protocol <string>] [-Port <int>]
    [-NotDefault] [-IgnoreInvalidCertificate]
# second parameter set additionally takes -VcfOAuthSecurityContext <VcfOAuthSecurityContext>

Invoke-VcfOpsGetServicesInfo            [-Server] [-AsInvokeRestMethodRequest]
Invoke-VcfOpsQueryFindings              [-FindingsQuery <FindingsQuery>] [-Page] [-PageSize] [-SortBy] [-SortOrder]
Invoke-VcfOpsQueryAlert                 -AlertQuery <AlertQuery> [-Page] [-PageSize]
Invoke-VcfOpsGetMatchingResources       -ResourceQuery <ResourceQuery> [-Page] [-PageSize]
Invoke-VcfOpsQueryLatestStatsOfResources -LatestStatQuery <LatestStatQuery>
Invoke-VcfOpsGetReclaimData             -Id <guid> -Reason <string> [-ShowExcluded <bool>] [-Page] [-PageSize]

Query bodies are built with Initialize-* cmdlets:

Initialize-VcfOpsFindingsFilter [-Severities <List[string]>] [-Categories <List[string]>]
    [-ResourceKinds <List[string]>] [-AdapterKinds <List[string]>] [-FindingTypes <List[string]>]
    [-ResourceIds <List[guid]>] [-RuleUuids <List[guid]>] [-FromOccurrenceTime <long>]
    [-Capabilities <List[string]>] [-RefreshTypes <List[string]>]

Initialize-VcfOpsFindingsQuery [-VarFilter <FindingsFilter>]

Initialize-VcfOpsalertquery [-ActiveOnly <bool>] [-AlertCriticality <List[...]>]
    [-AlertStatus <List[...]>] [-ResourceQuery <ResourceQuery>] [-ResourceKind <string>] ...

Initialize-VcfOpsresourcequery [-ResourceKind <List[string]>] [-ResourceHealth <List[ResourceHealth]>]
    [-ResourceStatus <List[ResourceDataCollectionStatus]>] [-AdapterInstanceId <List[guid]>] ...

Initialize-VcfOpslateststatquery [-ResourceId <List[guid]>] [-StatKey <List[string]>]
    [-MaxSamples <int>] [-CurrentOnly <bool>]

Two generator quirks that will cost time if not known in advance:

  • The filter property of FindingsQuery is exposed as -VarFilter, not -Filter.
  • Casing is inconsistent because it follows the spec's schema names: Initialize-VcfOpsalertquery and Initialize-VcfOpsresourcequery are lowercase, while Initialize-VcfOpsFindingsQuery and Initialize-VcfOpsFindingsFilter are PascalCase. PowerShell resolves cmdlet names case-insensitively, so this only matters when grepping or tab-completing.
  • Several parameters are typed as generated enums (AlertQuery+AlertCriticalityEnum, ResourceDataCollectionStatus). Verified on the lab: plain strings coerce fine. Initialize-VcfOpsFindingsFilter -Severities @('CRITICAL','WARNING') is accepted.

Three more quirks, all found the hard way on the lab (2026-08-31):

  • Invoke-VcfOpsQueryAlert returns its payload on _Alerts, not Alerts. $a.Alerts is silently empty while $a.PageInfo.TotalCount says 75. Read $a._Alerts. Invoke-VcfOpsQueryFindings does not do this; its payload is on Findings. Do not assume the property name; check it.
  • Invoke-VcfOpsGetReclaimData is broken. It serialises its -Id <guid> into the literal string Variant,10,Version,4, so the server answers 400 Cannot convert "Variant,10,Version,4" to uuid. This happens whether a [guid] or a string is passed. Use raw REST for reclamation (see the results section for the working call).
  • Invoke-VcfOpsGetStatKeysOfResources takes -ResourceId, not -Id (unlike Invoke-VcfOpsGetReclaimData, which takes -Id).
  • Disconnect-VcfOpsServer has no -Confirm parameter, so the usual -Confirm:$false throws. Just call it bare.

-AsInvokeRestMethodRequest on any Invoke-* cmdlet returns the request instead of executing it. That is the fastest way to learn the exact wire shape, and a good debugging aid on the lab.

Proposed architecture

Plugins/00 Initialize/005 Connection Plugin for VCF Operations.ps1   # $PluginTags = "core"
Plugins/85 Operations/...                                           # $PluginTags = "..., ops"
VcfOps.ps1                                                          # dot-sourced by the engine, like Charts.ps1
  • VcfOps.ps1 holds the shared client concerns: paging with a hard cap, retry, the JWT exchange for the two non-suite-api services, and the resource-id join helper. Dot-sourced by the engine exactly the way Charts.ps1 already is, so this introduces no module infrastructure the repo does not already have.
  • 005 Connection Plugin for VCF Operations.ps1 is a sibling of the vCenter connection plugin. It connects once and pre-collects a small set of globals ($OpsFindings, $OpsAlerts, $OpsResourceIndex, $OpsHealth). Every Operations plugin consumes those, per the existing key-reuse rule in CONTRIBUTING.md. No plugin re-queries.
  • New tag ops added to the $PluginTags vocabulary so Profiles.psd1 can include or drop the entire track in one line.

Two non-negotiable design rules

1. Graceful absence. No Operations server configured, or the connect fails, sets $OpsAvailable = $false, and every Operations plugin returns immediately. Users without VCF Operations must get today's report byte for byte. This track is strictly additive.

2. Read-only by construction. That SDK exposes 505 Invoke-* cmdlets, including Invoke-VcfOpsDeleteResource, Invoke-VcfOpsDeleteAdapterInstance, and adapter monitoring-state start/stop. The wrapper in VcfOps.ps1 must expose only read verbs, and no plugin should call an SDK cmdlet directly. A health report must never mutate the thing it is reporting on.

The technical crux: the resource join (solved)

Verified on the lab, 2026-08-31. Every vSphere resource in Operations carries its vCenter identity in resourceKey.resourceIdentifiers, and two of those identifiers are flagged isPartOfUniqueness = True:

Identifier Value seen Unique vCheck-side equivalent
VMEntityObjectID vm-6258, host-28, domain-c9, datastore-15 yes $vm.ExtensionData.MoRef.Value
VMEntityVCID the vCenter's InstanceUuid (a GUID) yes (Get-View ServiceInstance).Content.About.InstanceUuid
VMEntityInstanceUUID the VM's instance UUID (a GUID) no $vm.ExtensionData.Config.InstanceUuid
VMEntityName display name no $vm.Name

The VMEntityVCID observed on all 339 VM resources was byte-identical to the connected vCenter's own InstanceUuid, and the same identifier scheme applies uniformly to HostSystem, ClusterComputeResource, Datastore and Datacenter, not just VMs.

So the join key is the pair (VMEntityVCID, VMEntityObjectID), and the helper in VcfOps.ps1 should build a hashtable keyed on MoRef, filtered to this run's vCenter:

# index: MoRef -> Operations resource, scoped to the vCenter this run is about
$thisVcId = (Get-View ServiceInstance).Content.About.InstanceUuid
foreach ($r in $resourceList) {
    $ids   = $r.ResourceKey.ResourceIdentifiers
    $vcid  = ($ids | Where-Object { $_.IdentifierType.Name -eq 'VMEntityVCID' }).Value
    if ($vcid -ne $thisVcId) { continue }      # drop other vCenters in the fleet
    $moref = ($ids | Where-Object { $_.IdentifierType.Name -eq 'VMEntityObjectID' }).Value
    $index[$moref] = $r
}

Do not join on name: 339 Operations VM resources exist against a vCenter whose VM list is dominated by transient VKS/Supervisor pod VMs, and name lookups missed every sampled one. MoRef is the only reliable key.

A second index is required, and it is not optional. Alerts do not carry a resource name or a MoRef: an alert has only ResourceId, the Operations GUID. Rendering an alert against a vCenter object therefore needs ResourceId -> resource -> VMEntityObjectID, i.e. the $OpsResourceIndex global must be keyed both ways. Building it for all five vSphere kinds cost 1.1s for 350 objects on the lab.

Candidate plugin set

Ordered by value against effort. Endpoint paths are from the spec; every cmdlet in this table has now been called against the lab (2026-08-31) except the two raw-REST rows (log search and realtime metrics) and the rightsizing/dynamic-threshold endpoints. What each one actually returned on the lab is in Lab verification results. In short: alerts 75, findings 0, reclamation 136+1, capacity stats present, self-health all OK.

Plugin Endpoint / cmdlet Why it cannot exist today
Operations self-health GET /api/deployment/node/services/info, /node/status, Invoke-VcfOpsGetServicesInfo A report that trusts Operations must first check Operations is healthy. Service health enum: UNKNOWN, INVALID, OK, WARNING, ERROR.
Collection blind spots POST /api/resources/query filtered on resourceStatus, Invoke-VcfOpsGetMatchingResources Operations knows what it has stopped watching (DOWN, NO_DATA_RECEIVING, COLLECTOR_DOWN). Nothing else in the report does.
Diagnostic findings POST /api/diagnostics/findings/query, Invoke-VcfOpsQueryFindings Broadcom-maintained best-practice rules with severity, category, ruleDescription, affectedObjectsCount. One POST renders a section that improves itself as Broadcom ships new rules, with no code change here.
Active alerts by criticality POST /api/alerts/query, Invoke-VcfOpsQueryAlert Correlated causation with contributing symptoms, rather than 40 independent threshold trips. alertLevel: UNKNOWN, NONE, INFORMATION, WARNING, IMMEDIATE, CRITICAL, AUTO.
Capacity time remaining POST /api/resources/stats/latest/query, Invoke-VcfOpsQueryLatestStatsOfResources A real forecast, versus the naive projection in Plugins/20 Cluster/060 Capacity Planning.ps1. Stat keys must be discovered on a live system, see checklist.
Reclaimable waste, with cost GET /api/optimization/datacenters/{id}/reclaim/resources, Invoke-VcfOpsGetReclaimData Idle VMs, orphaned disks, powered-off waste, each carrying costSaving, reclaimableCpu, reclaimableMemory, reclaimableDiskSpace. A health report with a currency figure on it reads very differently to a manager.
Rightsizing GET /api/optimization/datacenters/{id}/rightsize/resources Oversized VMs, allocatedCpu/allocatedMemory versus recommendation.
Anomalies vs dynamic threshold POST /api/resources/stats/dt/query "This host is abnormal for itself" beats any global constant in a # Start of Settings block.
Fleet certificate expiry POST /api/fleet-management/certificate-management/certificates/query Covers every VCF appliance, not just the ESXi hosts that Plugins/30 Host/120 Host Certificate Expiration Check.ps1 probes by TLS handshake.
Log error bursts log-management POST /api/v2/logs/search (JWT, raw REST) Log signal has never been in the report.
Realtime spot check realtime-metrics GET /api/v1/query (PromQL, JWT, raw REST) 20s-granularity metrics, new in VCF 9.

The three moves worth doing for their own sake

These are the reasons to do this track at all, as opposed to adding another table of alerts.

1. Push vCheck findings back into Operations

The API has POST /api/events/bulk and POST /api/resources/{id}/stats, both supporting a push adapter kind. vCheck can write its own verdicts onto the matching Operations resources.

Consequences: vCheck's config-compliance findings appear on the object timeline inside Operations, can drive Operations alerting and notification, and become historical. vCheck stops being a dead-end email and becomes a custom data source. It also gives vCheck the historical memory it structurally lacks (the trend of "how many problems do we have" over time) without vCheck ever needing a database of its own.

This is the one item here that writes, so it contradicts the read-only rule above and must be opt-in, off by default, and clearly separated in the code.

2. Report the disagreements

For any object both systems see, compare verdicts. vCheck says the datastore is fine at 78% used; Operations says 12 days to full. Neither system alone produces that row, and it is plausibly the highest-signal row in the whole report. Depends entirely on the resource join.

3. Let Operations supply the thresholds

Instead of $WarningDays = 60 hardcoded per plugin, pull the dynamic threshold band for the same metric and report outliers against each object's own baseline.

Policy tensions to resolve before coding

  • Fleet scope collides with the one-run-per-vCenter policy (resolved). The project policy is no ELM/AllLinked and one vCheck run per vCenter. VCF Operations is fleet-wide by design, so an Operations query can silently pull in objects belonging to vCenters this run is not about. The VMEntityVCID identifier solves this: filter every resource to VMEntityVCID -eq (Get-View ServiceInstance).Content.About.InstanceUuid and the fleet collapses to exactly the vCenter this run is about, with no adapter-instance bookkeeping. (The VMwareAdapter Instance resource also carries VCURL = <vcenter fqdn> as its unique identifier, which is a usable second route, but the VCID filter is cheaper and exact.) The lab has a single vCenter, so this filter is correct-by-construction there rather than demonstrated against a multi-vCenter fleet, but the identifier is present and matches. Sections that are genuinely fleet-wide (Operations self-health, certificate expiry, the password-expiry alerts) must be labelled fleet-wide instead of filtered.
  • Licensing and deployment. The whole track presumes a licensed, deployed VCF Operations instance. The lab used for verification has one (see the results below).
  • Credentials. A second set of credentials is now in play. This should land on the v9.1 roadmap item "SecretManagement + VCF token auth" rather than inventing a second VCHECK_PASSWORD-style environment variable. Note Connect-VcfOpsServer already accepts -VcfOAuthSecurityContext, which is the same mechanism that roadmap item names.
  • Report length. vCheck's premise is "only problems". A findings feed can be long. It needs the same severity gate and row cap as everything else.

Lab verification checklist

Executed 2026-08-31 against a VCF Operations 9.1 lab instance. All answers are recorded in Lab verification results below. Kept here because it is still the right sequence to re-run against any other Operations instance. Note the snippet below is the original plan and contains two calls that do not work as written (the reclaim reasons and -Confirm:$false); the results section has the corrected forms.

Run these in order against the target environment. Each step is cheap and each one de-risks the step after it.

# 0. Does the lab even have VCF Operations, and does the module resolve?
Get-Module -ListAvailable VMware.Sdk.Vcf.Ops | Select-Object Name, Version
Import-Module VMware.Sdk.Vcf.Ops

# 1. Auth. Self-signed certs are expected in the lab, hence -IgnoreInvalidCertificate.
$cred = Get-Credential            # local Operations user, or the SSO account
Connect-VcfOpsServer -Server '<vcf-ops-fqdn>' -Credential $cred -IgnoreInvalidCertificate

# 2. Cheapest real call. Confirms the token works end to end.
Invoke-VcfOpsGetServicesInfo | Select-Object -ExpandProperty service |
    Select-Object name, health, details

# 3. Findings. Note -VarFilter, not -Filter.
$filter = Initialize-VcfOpsFindingsFilter -Severities @('CRITICAL','WARNING')
$query  = Initialize-VcfOpsFindingsQuery -VarFilter $filter
$f = Invoke-VcfOpsQueryFindings -FindingsQuery $query -PageSize 100
$f.findings | Select-Object severity, category, ruleName, affectedObjectsCount |
    Sort-Object severity | Format-Table -AutoSize
# If -Severities rejects plain strings, that answers the enum-coercion question.
# Also record: what do the severity and category values actually look like?

# 4. THE JOIN. Highest-risk unknown. Can an Operations resource be tied back to a vCenter object?
$rq = Initialize-VcfOpsresourcequery -ResourceKind @('VirtualMachine')
$r  = Invoke-VcfOpsGetMatchingResources -ResourceQuery $rq -PageSize 5
$r.resourceList[0].resourceKey | ConvertTo-Json -Depth 6
#   -> look for resourceIdentifiers containing the vCenter MoRef (vm-123) or instanceUUID.
#      Whatever key is present decides how the join helper in VcfOps.ps1 is written.
#      RECORD THE ANSWER IN THIS DOC.

# 5. Stat keys are runtime data and are NOT in the spec. Discover the real capacity keys.
$id = $r.resourceList[0].identifier
Invoke-RestMethod -Method Get -Uri "https://<vcf-ops-fqdn>/suite-api/api/resources/$id/statkeys" `
    -Headers @{ Authorization = "OpsToken <token>" } -SkipCertificateCheck |
    Select-Object -ExpandProperty 'resource-type-attributes' |
    Where-Object key -match 'timeremaining|capacity' | Select-Object key, name
# Anything previously written about 'summary|timeremaining' is a GUESS until this runs.

# 6. Reclamation needs a datacenter id, which may not exist if nobody configured custom DCs.
Invoke-RestMethod -Method Get -Uri "https://<vcf-ops-fqdn>/suite-api/api/resources/customdatacenters" `
    -Headers @{ Authorization = "OpsToken <token>" } -SkipCertificateCheck
# If empty: does Invoke-VcfOpsGetReclaimData accept a normal datacenter resource id instead?

# 7. Cost of the whole thing, against the ~110s run budget.
Measure-Command { Import-Module VMware.Sdk.Vcf.Ops }        # expect ~0.8s
Measure-Command { Invoke-VcfOpsQueryFindings -FindingsQuery $query -PageSize 100 }

Disconnect-VcfOpsServer -Server '<vcf-ops-fqdn>' -Confirm:$false

Lab verification results (2026-08-31)

Every open question is answered. Run against a lab VCF Operations instance at version 9.1.0.0, paired with a single VCF 9.1 vCenter.

  • Does the lab have a licensed VCF Operations instance at all? Yes, a full stack: an analytics node, a cloud proxy (collector), SDDC Manager, an Orchestrator, plus Operations for Networks (platform and collector). 26 adapter instances are registered, including VMWARE for the vCenter, a vSAN adapter, NSXTAdapter, VcfAdapter, SupervisorAdapter and NETWORK_INSIGHT.

  • Does Connect-VcfOpsServer work, and which -AuthSource? Yes, with the local Operations user and no -AuthSource at all (the default local source). -IgnoreInvalidCertificate is required in the lab. Connect costs 0.49s and needs no separate token handling.

  • Do plain strings coerce into the generated enum parameters? Yes.

  • What identifier joins an Operations resource to a vCenter object? (VMEntityVCID, VMEntityObjectID); see the join section. This was the highest-risk unknown and it works.

  • What are the real capacity/time-remaining stat keys? The guess summary|timeremaining was wrong. The real namespace is OnlineCapacityAnalytics|… and the key set is per resource kind (263 stat keys on a cluster, 289 on a host, 88 on a datastore):

    Stat key Cluster Host Datastore Lab value (cluster)
    OnlineCapacityAnalytics|timeRemaining yes yes yes 366 (days; 366 is the cap)
    OnlineCapacityAnalytics|capacityRemainingPercentage yes yes yes 24.42
    OnlineCapacityAnalytics|{cpu,mem,diskspace}|demand|timeRemaining yes yes no 366
    OnlineCapacityAnalytics|{cpu,mem,diskspace}|demand|capacityRemaining yes yes no cpu 49366.8, disk 21757.4
    OnlineCapacityAnalytics|{…}|demand|recommendedSize yes yes no rightsizing input
    OnlineCapacityAnalytics|diskspace|total|timeRemaining no no yes
    reclaimable|cost, reclaimable|idle_vms|{cost,cpu,mem,diskspace}, reclaimable|poweredOff_vms|{cost,diskspace}, reclaimable|vm_snapshots|{cost,diskspace} yes yes no reclaimable|cost = 3.48
    reclaimable|orphaned_disk|{cost,diskspace} no no yes

    Clusters and hosts additionally carry …|timeRemainingWithCommit / …|recommendedSizeWithCommit (the same figures including committed-but-unprovisioned demand). Discover, never hardcode: GET /suite-api/api/resources/{id}/statkeys, and note the response property is stat-key, not the resource-type-attributes guessed earlier here.

  • Do custom datacenters exist, and what id does the reclamation endpoint want? GET /api/resources/customdatacenters returns 0: none are configured. It does not matter, because the endpoint accepts a plain Datacenter resource id (the Operations resource id of an ordinary vCenter datacenter). The reason values in this document were wrong too; the server enumerates the real ones in its 400 body: POWERED_OFF, IDLE, SNAPSHOT, ORPHANED_DISK (not IDLE_VMS / POWERED_OFF_VMS / SNAPSHOTS). The working call, which must be raw REST because the SDK cmdlet mangles the id:

    GET /suite-api/api/optimization/datacenters/{datacenterResourceId}/reclaim/resources?reason=ORPHANED_DISK&pageSize=50
    -> { pageInfo, reclaimRightsizeResources: [ { id, name, costSaving, reclaimableDiskSpace, duration } ] }
    

    Real lab output: 136 orphaned disks (costSaving 0.0032 each, vSAN paths) and 1 powered-off VM worth costSaving 3.48, the same 3.48 that reclaimable|cost reports on the cluster. IDLE and SNAPSHOT were both empty. Each call is 0.05 to 0.16s.

  • What is the real per-call latency, and does the track fit the performance budget? Yes, comfortably: the calls are nearly free and the import is the whole cost.

    Step Cost
    Import-Module VMware.Sdk.Vcf.Ops 5.37s
    Connect-VcfOpsServer 0.49s
    Resource index, 350 objects across 5 kinds 1.10s
    Invoke-VcfOpsGetServicesInfo 0.73s
    Invoke-VcfOpsQueryFindings (pageSize 100) 0.20s
    Invoke-VcfOpsQueryAlert (pageSize 500, 75 alerts) 0.11s
    Invoke-VcfOpsGetMatchingResources VMs (339) 0.54s
    Total for the whole spike 8.54s

    Against the ~110s budget that is ~8%, of which 63% is the module import. Everything after the import is sub-second, so the plugin count is not the thing to economise on; the import is.

  • Is the header prefix really OpsToken and not vRealizeOpsToken? OpsToken works. So does vRealizeOpsToken: 9.1 still accepts the legacy 8.x prefix. Token acquire takes 0.28s and returns a 74-character token; roles came back empty for the local admin user.

What the lab actually has to report on

This is the part that changes the plugin ordering, so it is worth stating plainly.

  • Invoke-VcfOpsQueryFindings returns TotalCount = 0, unfiltered, not merely for CRITICAL/WARNING. The diagnostics rule engine has produced nothing on this lab. The call works; there is simply nothing in it. A findings-led first slice would have rendered an empty section, so findings cannot be the lead plugin.

  • Invoke-VcfOpsQueryAlert returns 75 active alerts (65 WARNING, 8 CRITICAL, 2 IMMEDIATE; all ACTIVE / OPEN). That is the real verdict feed, and it is dominated by things vCheck has no equivalent for:

    Count Alert definition
    41 Virtual Machine is violating the vSphere Security Configuration Guide (vSphere 8+)
    4 vSphere Distributed Port Group is violating the Security Configuration Guide
    3 ESXi Host is violating the Security Configuration Guide
    2 Objects are not receiving data from adapter instance
    2 VCF Roles definition updated at identity broker
    1 each VM experiencing disk read latency; VCF Operations password expired; VCF Operations for Networks password expired; vCenter Server violating the Security Configuration Guide

    47 of the 75 joined cleanly to a vCenter object via ResourceId -> resource -> MoRef (for example the vCenter appliance VM and the Operations cloud proxy VM). The other 28 are either non-vSphere kinds absent from the five-kind index (distributed port groups, the vCenter object itself) or genuinely Operations-internal (password expiry, identity-broker notifications). Widening the index to the port-group and vCenter kinds recovers most of the remainder; the Operations-internal ones belong in a fleet-wide section with no object column.

  • Collection blind spots are real and large. Of 339 VM resources, 263 have an empty collection status and GREY health, 76 are DATARECEIVING; health is GREEN on 74 and RED on 2. Note the observed enum is DATARECEIVING, no underscore, whereas the spec's resource-data-collection-status list below says DATA_RECEIVING. Match on the observed value, and treat empty as its own case. The 263 are mostly transient VKS/Supervisor pod VMs Operations has stopped collecting; a blind-spot plugin must exclude those or it will report 263 rows of noise on day one.

Revised first slice, given the above

The original plan led with diagnostic findings. On this lab that renders nothing. Lead with alerts instead: the same shape of work, but it produces 47 joined rows immediately.

  1. 005 Connection Plugin for VCF Operations.ps1: conditional import, connect, build the two-way $OpsResourceIndex filtered on VMEntityVCID, set $OpsAvailable.
  2. One plugin rendering active alerts joined to vCenter objects, severity-gated and row-capped, with a donut by AlertLevel per the $Chart contract.
  3. Operations self-health from Invoke-VcfOpsGetServicesInfo (all 7 services OK on the lab: CASA, LOCATOR, COLLECTOR, ADMINUI, API, UI, ANALYTICS). It is cheap, and it gates trusting the rest.

Findings stay in the plugin set, but as a section that is empty until Broadcom's rules fire.

What is implemented

Shipped 2026-08-31, starting with the three items above. Everything not listed in this section is still proposal.

File Role
VcfOps.ps1 Read-only client, dot-sourced by the engine like Charts.ps1. Every SDK call goes through Invoke-VcfOpsRead, which refuses any cmdlet not on an explicit allowlist. Holds the paging caps, the retry, the identifier helper and the resource index.
Plugins/00 Initialize/005 Connection Plugin for VCF Operations.ps1 Conditional import + connect, then pre-collects $OpsServices, $OpsResourceIndex and $OpsAlerts. Sets $OpsAvailable / $OpsConnectError.
Plugins/85 Operations/010 VCF Operations Health.ps1 Per-service health, plus a row when Operations is configured but unreachable.
Plugins/85 Operations/020 VCF Operations Active Alerts.ps1 Active alerts joined to vCenter objects, severity-gated, row-capped, severity donut.
Plugins/85 Operations/030 VCF Operations Monitoring Coverage.ps1 vCenter objects Operations is not watching, per object type.
Plugins/85 Operations/040 VCF Operations Capacity Forecast.ps1 Capacity headroom and time remaining from `OnlineCapacityAnalytics
Plugins/85 Operations/050 VCF Operations Reclaimable Capacity.ps1 Reclaimable waste with cost, via raw REST.
Plugins/85 Operations/060 VCF Operations Disagreements.ps1 Datastores where this report's static threshold and Operations' forecast disagree.
Profiles.psd1 New ops profile; the plugins are tagged health, ops, slow.

Two gates keep it free when it is not wanted. No $VcfOpsServer configured → the connection plugin returns before importing anything. Configured, but the active profile selected no Operations plugin → it also returns, which is what keeps quick quick (verified: the plugin started and finished inside the same second on a quick run with Operations configured).

The index is built both ways on purpose. .ByMoRef is the join for plugins that start from a $VM/$VMH object; .ById is the join for alerts, which carry only an Operations GUID. Objects whose VMEntityVCID belongs to another vCenter are indexed ById but flagged Foreign and never exposed ByMoRef, because a MoRef is only unique within one vCenter and mixing fleets there would silently mis-attribute findings. The alerts plugin drops foreign alerts and labels genuinely unscoped ones (Operations-internal, or object kinds outside the indexed set).

End-to-end lab result (-PluginProfile ops, VCF Operations 9.1.0.0): 76 active alerts rendered (47 joined to vCenter objects: 43 VMs and 4 hosts; 28 labelled fleet-wide), sorted severity-first, with a working severity donut, no unresolved CIDs, and the health section correctly empty because all seven services were OK. Added ~6s to the run.

The second batch (030 / 040 / 050)

030 Monitoring Coverage started as the "collection blind spots" idea in the plugin table above, and the lab reshaped it. The blunt version, report resources Operations has stopped collecting from, is pure noise: all 266 such resources on the lab are objects that no longer exist in vCenter at all (271 Operations resources reference a dead MoRef). Filtering them through the MoRef join leaves zero rows, which is the correct answer and a useless plugin. Inverting the question is what makes it useful: which objects in this vCenter is Operations not watching? Measured on the lab: hosts 3/3, clusters 1/1, datastores 4/4, all DATARECEIVING; VMs 69 of 680 (10.1%). That VM figure is VKS/Supervisor churn rather than a real gap, which is why the VM check ships off by default while the stable object types are checked at 100%.

040 Capacity Forecast uses one batched latest-stats call for every cluster, host and datastore (0.1s for the lab). Observed: the management cluster at 24.4% capacity remaining, hosts 45 to 53%, datastores 55 to 99.6%, and every timeRemaining pinned at 366, which is Operations' ceiling and reads as "more than a year", not "366 days".

050 Reclaimable Capacity is the row that reads differently to a manager: 113 items, 1668.1 GB, 154.03 cost saving on the lab (mostly orphaned vSAN disks, plus a powered-off VM that joins cleanly back to vm-445). Full vSAN paths are far too wide for an email table, so the file and its immediate folder are shown instead of the path.

Two more SDK shapes, measured: Invoke-VcfOpsQueryLatestStatsOfResources wants [guid[]] resource ids and returns Values → .ResourceId / .StatList.Stat → .StatKey.Key / .Data (an array; take the last sample). It does not use the _-prefixed property quirk that Invoke-VcfOpsQueryAlert does, so check per cmdlet rather than assuming.

060 Disagreements is move #2 above, built for datastores (the case the move describes). For each datastore it holds this report's static verdict (free % against the same 15% threshold 40 Datastore/010 uses) beside Operations' forecast, and reports only where the two differ. Both directions matter: Operations is more concerned is the dangerous one (plenty of free space, but the growth rate says otherwise), while vCheck is more concerned flags a fixed threshold firing on a datastore that has been flat for months, a candidate for an exception rather than a recurring alert. Free % is computed from CapacityGB/FreeSpaceGB rather than the PercentFree VIProperty, so the plugin does not depend on the connection plugin defining it. Every lab datastore agrees under the shipped thresholds, so the section is correctly empty there; the disagreement path was exercised by tightening the Operations threshold to 60%, which produced the expected row for the vSAN datastore (60.1% free versus 55.5% remaining). Clusters and hosts are the obvious extension, but the vCheck-side cluster verdict lives in 20 Cluster/060 Capacity Planning as QuickStats arithmetic that would have to be duplicated, so it is worth doing deliberately rather than by copy-paste.

Still not built (in rough value order): diagnostic findings (works, but the lab has none), rightsizing, dynamic-threshold anomalies, fleet certificate expiry, log-management search and realtime PromQL metrics (both raw REST with the JWT exchange), and the two remaining "moves worth doing" above: pushing vCheck findings back into Operations, and letting Operations supply the thresholds.

First slice

Superseded by Revised first slice. The plan below was written before the lab run; it leads with findings, of which the lab has none. Kept for the acceptance criteria, which still stand, and which the lab run has now largely met in advance: auth works, the join works, and the cost is measured at 8.54s.

One spike, one plugin. Connect via Connect-VcfOpsServer, call Invoke-VcfOpsGetServicesInfo (self-health) and Invoke-VcfOpsQueryFindings, render the findings as a severity-sorted table with a donut by category using the $Chart contract from CHARTS.md.

Acceptance: it proves auth works, proves the SDK is usable inside the dot-sourced plugin scope, measures the real runtime cost, and produces a genuinely useful report section on its own. If the join in step 4 also works, everything else in this document is mechanical.

Enum values

Observed on the lab (2026-08-31): these win over the spec

Enum Observed values
resource-data-collection-status DATARECEIVING (no underscore, unlike the spec below) and empty string. 76 and 263 of 339 VM resources respectively.
resource-health GREEN (74), GREY (263), RED (2)
alert.alertLevel WARNING (65), CRITICAL (8), IMMEDIATE (2)
alert.status / alert.controlState ACTIVE (75) / OPEN (75)
service.health / service.name OK on all 7: CASA, LOCATOR, COLLECTOR, ADMINUI, API, UI, ANALYTICS
reclamation reason POWERED_OFF, IDLE, SNAPSHOT, ORPHANED_DISK (the server lists these in its 400 body)
AuditFinding.severity / .category not observable: the lab has zero findings

From the spec

For filters, in case string coercion fails (it does not; plain strings were accepted).

Enum Values
resource-health GREEN, YELLOW, ORANGE, RED, GREY
resource-state STOPPED, STARTING, STARTED, STOPPING, UPDATING, FAILED, MAINTAINED, MAINTAINED_MANUAL, REMOVING, NOT_EXISTING, NONE, UNKNOWN
resource-data-collection-status NONE, ERROR, UNKNOWN, DOWN, DATA_RECEIVING, OLD_DATA_RECEIVING, NO_DATA_RECEIVING, NO_PARENT_MONITORING, COLLECTOR_DOWN
alert.alertLevel UNKNOWN, NONE, INFORMATION, WARNING, IMMEDIATE, CRITICAL, AUTO
alert.status NEW, ACTIVE, UPDATED, CANCELED
alert.controlState OPEN, ASSIGNED, SUSPENDED, SUPPRESSED
service.health UNKNOWN, INVALID, OK, WARNING, ERROR
service.name UI, ADMINUI, CASA, ANALYTICS, COLLECTOR, API, LOCATOR, UNKNOWN

AuditFinding fields: ruleName, ruleDescription, ruleUuid, severity, category, findingType, affectedObjectsCount, lastObservedTimeInMillis, refreshMode, capabilities. The severity and category value sets are not enumerated in the spec, so they have to be observed on a live system.