Skip to content

feat(observability): integrate Ansible nodes with Grafana Cloud Application Observability #55

Description

@altanoruc

Objective

Provide an opt-in Grafana Cloud Application Observability integration for Ansible-deployed deCDN nodes, without exposing telemetry endpoints publicly or committing Grafana credentials.

Repository findings

  • The Ansible role already renders [observability] with metrics_bind, metrics_port, and an optional otlp_endpoint; its node metrics (9090) and admin RPC remain loopback-only.
  • The node service supports a separate 0600 environment file for secrets, but Grafana credentials must not be added to the node process environment unnecessarily.
  • The Helm chart already offers a private metrics ClusterIP service and optional ServiceMonitor; this issue targets the VM/Ansible path first.

Proposed implementation

  • Add an opt-in decdn_grafana_cloud_enabled: false integration to the Ansible role or a dedicated observability role.
  • Install and manage Grafana Alloy as a hardened systemd service.
  • Configure Alloy to:
    • scrape http://127.0.0.1:<decdn_metrics_port>/metrics and remote-write Prometheus metrics to Grafana Cloud;
    • receive the node's OTLP export on loopback and forward it to Grafana Cloud with the required authenticated OTLP exporter;
    • attach stable resource/metric labels, at minimum service.name=decdn-node, node identity, region, and environment;
    • avoid high-cardinality identifiers and define trace sampling before enabling fleet-wide export.
  • Set decdn_otlp_endpoint to Alloy's loopback receiver only when the integration is enabled.
  • Store Grafana Cloud endpoint, instance/tenant ID, and API token in a dedicated host-provisioned 0600 secret file. Do not commit them, render them into node.toml, or place them in the deCDN node environment.
  • Keep all listeners loopback-only and make no nftables changes beyond existing node exposure.
  • Document Grafana Cloud stack setup, minimum token scopes, secret provisioning, enablement, rollback, and cost/cardinality controls.

Acceptance criteria

  • Default deployment remains unchanged and installs no Alloy/Grafana components.
  • With the feature enabled and secrets provisioned, ansible-playbook --check and a normal deploy complete successfully.
  • Node Prometheus metrics arrive in Grafana Cloud with the required labels.
  • Node OTLP data reaches Grafana Cloud through Alloy; the node itself has no Grafana Cloud credentials.
  • No Grafana, OTLP, or metrics port is publicly reachable; metrics scrape and OTLP receiver bind to loopback.
  • The deployment fails clearly when the integration is enabled but required secret values are absent or malformed.
  • Molecule/template coverage verifies default-off behavior, generated Alloy config, secret-file permissions, loopback binds, and decdn_otlp_endpoint wiring.
  • Decide in a follow-up whether the Helm chart needs a separate Alloy/OTel Collector integration; retain its existing ServiceMonitor behavior unchanged in this issue.

References

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Fields

Priority

None yet

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions