A telemetry cost & cardinality profiler for SigNoz: it finds the waste, generates the fix, and proves it drops zero error traces first.
31 / 31 error traces kept · every slow trace kept · the error-span panel stayed flat — and only then the number: −95.0% span storage measured on a live ingester (honest counterweight: only ~6% on an error-storm day — the policy refuses to drop errors) · 14.1M spans profiled in 895 ms · 41 tests, 45 subtests, 7 packages · read-only by design — it exits non-zero rather than emit a policy that would drop one error trace · MIT
The first waste it ever found was ClickHouse's own system.metric_log — ~80 merges a minute, 6.2 GiB peaks, OOM-killing the very scan that discovered it.
▶ Watch it work (3:42): https://youtu.be/lZKU3A7t5GM
One real scan, rendered as the bill it is — with the sampling policy explorer replaying 21 recorded positions over 168,711 of our own traces. No install, no SigNoz, no network calls.
See it work · The 15-minute tour · Quickstart · Profilers · Architecture · Status · Compatibility · Learn
One command, 895 ms, against a live SigNoz instance holding 14.1M spans: six profilers, a ranked bill in GB/month and dollars, and a replay of the recommended sampling policy over 75,586 real traces that keeps every error and slow trace. Nothing was written to the cluster.
Why that is fast, and what it costs in precision. The profilers read aggregates rather than rows — day-sliced
GROUP BYqueries with top-N limits (internal/store/clickhouse.go) — and attribute and series cardinality comes from ClickHouse'suniqCombined, an approximate distinct count. So cardinality figures are estimates, and the GB/month numbers are a transparent ingest model rather than an invoice. The safety replay is the exact one: its input is one row per trace,GROUP BY trace_idacross the whole window, which is why it can claim every error trace by count.
Cutting telemetry blindly is scary — drop the wrong span and you are blind during the next incident. TELELENS is observability for your observability: it ranks the waste with hard numbers, generates the exact OpenTelemetry Collector config that fixes it, and proves by replay that the fix keeps every error and slow trace before you ever apply it. It writes files to out/; a human reviews the diff and applies it.
The Savings Tracker dashboard is the other half of the proof — ingest falls off a cliff when the generated config lands, while the error-span panel stays flat.
Measured, not projected — applied to a live SigNoz ingester, then reverted and the revert verified:
| Result | Measured | Evidence |
|---|---|---|
| Injected error traces kept | 31 / 31 (live) · 40 / 40 (fixtures) | assets/live-apply-measure-m4-evidence.md |
| Simulator replayed over real last-24h traces | 162,760 traces — SAFE | assets/live-simulator-m3-evidence.txt |
| Error-spans-per-service dashboard panel after the config landed | flat — no drop | assets/screenshots/m5-03-savings-tracker.png |
| Only then: span storage cut on healthy known-volume traffic | −95.0% | assets/live-apply-measure-m4-evidence.md |
| Honest counterweight: droppable on an error-storm day | ~6% (the policy refuses to drop errors) | same |
Console capture behind the recording: assets/live-scan-2026-07-25.txt; assets/README.md indexes the rest.
Copy-pasteable, in order. Steps 1–6 need no SigNoz, no Docker, no network — everything runs against the committed fixture corpus; the live path (7–8) is optional.
1. Clone and build — one static binary, one direct non-stdlib dependency.
git clone https://github.com/vinayaksonthalia/telelens
cd telelens
go build -o telelens ./cmd/telelensExpect: no output;
./telelens --helplistsscan,simulate,report,generate.
2. Prove the suite is green — offline.
go test ./...Expect:
okfor 7 packages, includingTestSimulateSafetyInvariantininternal/analyze.
3. Find the waste — the ranked bill, in colour, in well under a second.
./telelens scan --fixturesExpect: six profilers, a severity-coloured table, then
Total identified savings: 130.2 GB/month ≈ $39.06/monthand ascan completedline withoutputs in out/.
4. Prove the fix is safe — the policy is replayed trace by trace before anything is applied.
./telelens simulate --fixtures --sample-pct 5 --latency-ms 750Expect:
error traces: 40 / 40 kept,slow traces: 55 / 55 kept,verdict: SAFE. One dropped error or slow trace prints UNSAFE and exits non-zero (how the invariant is enforced).
5. Read the generated fix — this is the artifact you would actually merge.
./telelens report --findings out/findings.json # re-renders out/waste-report.md
./telelens generate --findings out/findings.json # re-renders the collector fragment + casting patch
sed -n '40,60p' out/collector-fragment.yamlExpect:
wrote out/collector-fragment.yaml/wrote out/casting-patch.yaml, and afilter/drop_unread_metricsblock shipped deliberately commented out under a review banner — the review step is the product (why).
6. Take the agent surface — the same findings as stable JSON on stdout, progress on stderr.
./telelens scan --fixtures --json | jq '.findings[] | select(.category=="quality")'Expect: well-formed JSON objects;
NO_COLOR=1 ./telelens scan --fixtures | catdegrades cleanly too.
7. (Optional) Point it at your own SigNoz — read-only: ClickHouse SELECTs and SigNoz GETs only.
cp .env.example .env # CLICKHOUSE_HTTP_URL, SIGNOZ_API_URL, SIGNOZ_API_KEY, COST_PER_GB
set -a; source .env; set +a
./telelens scan --window 7 --cost-per-gb 0.30Expect:
telelens scan · source: clickhouse (…, window=7d)and a real ranked bill. Reaching ClickHouse and the least-privilegetelelens_rouser: DOCS § Path 2.
8. (Optional) Import the dashboards and guardrails.
export SIGNOZ_URL=http://localhost:8080 SIGNOZ_API_KEY=...
./dashboards/import.shExpect: 3 dashboards + 3 alert rules created, including the Savings Tracker shown above (what is in the pack).
Requires Go ≥ 1.24. One static binary; the only direct non-stdlib dependency is charmbracelet/lipgloss.
git clone https://github.com/vinayaksonthalia/telelens
cd telelens
go build ./cmd/telelens
# Demo mode — no SigNoz needed. Scans the committed "Noisy Neighborhood"
# fixture corpus (every classic waste pattern, recorded as query results):
./telelens scan --fixtures
# Outputs land in out/:
# waste-report.md ranked findings with evidence, GB/mo and $
# findings.json machine-readable report
# collector-fragment.yaml annotated otelcol processors (tail_sampling, filter, transform)
# casting-patch.yaml the same fix as a Foundry ingester.config.data patchOnce the repo is public, go install github.com/vinayaksonthalia/telelens/cmd/telelens@latest installs the CLI directly. Re-render a previous scan without re-profiling via telelens report / telelens generate (step 5). Full flag and exit-code reference: DOCS § Command reference.
| Profiler | Findings |
|---|---|
| traces | volume+bytes by service · attribute cardinality offenders (user.id as a span attribute) · jumbo attributes (4 KB db.statement on every span) · near-duplicate success spans (tail-sampling candidates) |
| logs | Drain-style template mining ("this one DEBUG line is half your log bytes") · per-service severity distribution · exact-duplicate line detection |
| metrics | per-metric series cardinality · which label is the bomb (user_id → 210k series) · stale/silent metrics |
| usage-xref | metrics written but read by no dashboard or alert — diffed against GET /api/v1/dashboards + GET /api/v1/rules, walked across schema versions |
| quality | missing service.name · missing/non-standard/miscased log severity · non-conformant metric names — unqueryable data is waste at any size |
| ecosystem | cost attribution for modern workloads: prices AI-agent telemetry (gen_ai.* spans + tokens) and browser-RUM sessions (session.id spans, KB/session); flags any unbounded ID used as a metric label as structural at ANY series count |
Pricing is deliberately transparent — uncompressed-ingest bytes × 30 days × --cost-per-gb; unpriceable findings ship as directional (model). Every profiler query lives in internal/store/, and together they exercise all six SigNoz Query Builder showcase capabilities (mapping).
The product is a loop: find the waste, price it, prove the fix is safe on your own
traces, hand a human the config, then watch the guardrails. Numbers below are the
live-verified ones from assets/.
%%{init: {'flowchart': {'rankSpacing': 40, 'nodeSpacing': 30, 'wrappingWidth': 340}}}%%
flowchart TD
CH[("ClickHouse :8123<br/>SELECT only")] --> SCAN
API[("SigNoz API :8080<br/>GET only")] --> SCAN
SCAN["1 · scan — six profilers, read-only<br/>26 findings in 3.74 s"]
SCAN --> PRICE["2 · price — uncompressed GB × 30 days × your cost per GB<br/>risks that cannot be priced ship as directional, not padded"]
PRICE --> SIM{"3 · prove — replay the policy over<br/>162,760 of your own traces"}
SIM -->|"UNSAFE — one error trace would be lost"| STOP["exit 1, nothing is generated"]
SIM -->|"SAFE — 31 / 31 error traces kept, every slow trace kept"| GEN["4 · generate → out/ — collector-fragment.yaml + casting-patch.yaml<br/>the drop-metrics block ships commented out, under a review banner"]
GEN --> REVIEW{{"a human reads the diff — TELELENS applies nothing"}}
REVIEW --> CAST["the operator runs foundryctl cast<br/>−95.0% span storage measured (~6% on an error-storm day)"]
CAST --> GUARD["5 · guardrails — 3 dashboards + 3 alert rules<br/>ingest falls off a cliff, the error-span panel stays flat"]
GUARD -.->|"next scan window"| SCAN
Read-only by design (the product invariant). Profilers issue only ClickHouse SELECTs and SigNoz GETs; the live client refuses non-SELECT statements in code, and .env.example documents the least-privilege telelens_ro ClickHouse user that enforces it at the database. The only writes the profiler performs are files in out/; the one write path in the project is the opt-in dashboards/import.sh, which POSTs the dashboard pack and guardrail rules you explicitly ask for. (Foundry is SigNoz's official deployment tool — github.com/SigNoz/foundry; foundryctl cast applies a casting, its config.)
| Scope | Status |
|---|---|
| P0 — all six profilers, ranked waste report (md+json), config generator (tail_sampling / filter / transform), casting patch, fixture mode, offline test suite | ✅ Done |
P1 — sampling simulator with safety invariant, Telemetry Bill dashboard pack + guardrail alerts, usage cross-reference, --json agent/MCP surface, design-system CLI |
✅ Done |
| Live-verified (Jul 18) — full scan of a real mixed-workload instance (3.74 s, 26 findings); simulator over ALL 162,760 last-24h traces (SAFE); generated config applied to a live ingester measuring −95.0% span storage with zero error-trace loss on healthy known-volume traffic (on an error-storm day the same policy can only drop ~6% — it refuses to touch errors), then reverted; 3 dashboards + 3 alert rules via API; MCP data-quality hook demoed with a real agent | ✅ Evidence in assets/ |
P2 / roadmap — web panel, --explain LLM narration, telelens verify automated savings assertion, --import flag, otelcol validate in CI, multi-cluster scan |
❌ Not done |
| Known limits — pricing is a transparent ingest model, not a cloud invoice; the simulator replays per-trace summaries, not packet-level collector behaviour; SigNoz v0.13x schemas with startup drift detection; heavy queries time-sliced + memory-guarded; meter buckets are hourly, so sub-hour A/B measurements count storage rows instead |
Live-verified against self-hosted SigNoz v0.132.2 (signoz_index_v3 / logs_v2 / time_series_v4, with startup drift detection). SigNoz Cloud does not expose ClickHouse to customers, so live mode is self-hosted only — Cloud users run the fixtures demo, which exercises every profiler and generator against the recorded corpus.
Uninstall: delete the telelens binary and the out/ directory — there is no other state. Drop the telelens_ro ClickHouse user if you created one; remove imported dashboards and alerts from the SigNoz UI or its APIs.
DOCS.md — operator manual: live setup, rollback runbook, command reference, the agent/MCP hook, honest caveats. learning/README.md — teaching curriculum: pipeline walkthrough, SigNoz deep-dive, trade-offs, FAQ, design rationale, bug-hunt diary. LEARNINGS.md — what the live stack actually taught us. Licensed MIT.
Contributions welcome — see CONTRIBUTING.md.
AI coding assistants were used during development. The profilers, config generator, sampling simulator, and dashboards are original work; the analysis itself is deterministic SQL and arithmetic — there is no LLM in the cost loop — and every live number in this README is backed by evidence under assets/.


