Self-hosted observability in one binary. Analytics, error tracking, APM, logs, session replay, monitoring, feature flags, experiments — with three of those surfaces working together in one install:
- AI query assistant on the SQL explorer. English in, SQL out (your LLM key, your cost). Every call logged back to the LLM-tracing table.
- Incident-mode markers. When an alert fires, every time-series chart overlays a translucent vertical band for the window.
- Scheduled SQL exports to any S3-compatible bucket (S3, R2, MinIO).
Two processes (Observe + Nucleus). 91 MB idle measured on the reference 4-core host — see BENCHMARKS.md.
brew install useteploy/tap/observegit clone https://github.com/useteploy/teploy-observe.git
cd teploy-observe
docker compose upThe Compose file runs the published image. To test local source changes, build
the root Dockerfile explicitly.
Open http://localhost:3000. First visit lands on the setup wizard — there's
no default password to change later, since none is set until you choose one
there.
Downloads the installer from the latest release (not the mutable main
branch) and verifies its SHA-256 against the release's checksums.txt
before executing it — the script itself already checksum+signature-verifies
the observe binary it installs, but nothing verified the script itself
until now:
(
set -e
curl -fsSLO https://github.com/useteploy/teploy-observe/releases/latest/download/install.sh
curl -fsSLO https://github.com/useteploy/teploy-observe/releases/latest/download/checksums.txt
grep " install.sh\$" checksums.txt > checksum.txt
if command -v sha256sum >/dev/null 2>&1; then sha256sum -c checksum.txt || exit 1; else shasum -a 256 -c checksum.txt || exit 1; fi
sh install.sh
)The script generates a random admin password and prints it on completion;
it is also stored in /etc/observe/observe.env and rotatable from
Settings → Users.
The direct installer verifies the release's SHA256 through an Ed25519-signed
checksums.txt before installing anything. It fails closed if the signature
or archive hash is invalid. OBSERVE_HEALTH_URL is read by the install script
only, for the health poll it runs after restarting the existing service
(default http://127.0.0.1:3000/healthz). It does not configure
observe upgrade, which derives its readiness URL from OBSERVE_ADDR; pass
--health-url instead for a non-default endpoint.
Use the manager that installed Observe:
# Direct Linux/systemd install
sudo observe upgrade
sudo observe upgrade --version v1.2.3
# --service <unit> systemd unit name (default: observe.service)
# --health-url <url> readiness URL for custom service configuration,
# derived from OBSERVE_ADDR by default
# Homebrew
brew upgrade useteploy/tap/observe
# Docker Compose
docker compose pull && docker compose up -d
# Pin a release with OBSERVE_VERSION=1.2.3The direct updater authenticates and stages the release while the current server remains online. It then asks systemd to stop Observe gracefully, atomically replaces the binary, and requires three healthy responses from the exact new version. A failed start or readiness check restores and restarts the previous version automatically. There is a brief restart window; telemetry senders should retain their normal retry policy.
git clone https://github.com/useteploy/teploy-observe.git
cd teploy-observe
go build ./cmd/observe # neutron-go is vendored; no network setupYou also need a Nucleus database binary — see the Docker compose file for the exact image and version.
The tested path from a fresh clone to a query-visible first event is the
script sdk/e2e/first_event_test.sh (needs podman; uses its own
fixtures on ports 55446/38080 and cleans up after itself):
git clone https://github.com/useteploy/teploy-observe.git
cd teploy-observe
bash sdk/e2e/first_event_test.shIt boots a Nucleus fixture, builds and starts Observe, creates a site-scoped API key, proves keyless ingest is refused with 401, sends a first event from each SDK (browser, sentry-shim, Go, Python), and polls the stats API until those records are query-visible. Every step above is exercised by that script (verified passing on this branch); it is not part of the normal test suites because it needs podman and exclusive ports.
The manual equivalent, condensed from the script:
- Run: start Nucleus, then
OBSERVE_NUCLEUS_URL=... OBSERVE_ADMIN_PASSWORD=... observe. Sign in, or drive the API:POST /api/v1/auth/loginwith the admin credentials. - API key:
POST /api/v1/sites/default/keyswith the admin JWT (Settings > API keys in the UI) — a site-scoped telemetry key (obs_..., shown once). - First event: send one with any SDK or a plain POST to
/api/v1/events/batchwith theX-API-Keyheader. - See it:
GET /api/v1/stats/events?site_id=default(JWT auth) lists the event once it has flushed.
Failure modes: a 401 on ingest means the key is missing/wrong (keys are
site-scoped; browser keys are telemetry-only by design); /healthz 503
with nucleus: ... means the engine connection failed — check
OBSERVE_NUCLEUS_URL first. Durable-ingest posture and all counters are
on /healthz (see the WAL notes under Platform below).
Tested paths (installation → credential → first event → verification, per SDK) and dual-write recipes are in docs/sdk/MIGRATION.md. The per-API compatibility tables — supported, changed semantics, intentional no-op, unsupported — are the authority: docs/sdk/COMPATIBILITY.md. The rule for every recipe: run both integrations side by side, compare, and only remove the old one after the Observe data has proven itself.
Declared limits, fixture-tested scale points, and where to read them live:
- Measured throughput (reference host, Intel N5000 4-core/4GB —
BENCHMARKS.md): 75 req/s analytics pageviews at 45 ms
p95; 58 req/s OTLP traces (2 spans/request) at 66 ms p95; error events
are rate-limited per site (default 1000 events/s,
OBSERVE_RATE_LIMIT). Memory 91 MB idle / 124 MB after the full bench. The CI-enforced ingest budget (weekly, GitHub-hosted runner) is 10,000 events/s sustained with p95 < 50 ms — docs/operations/perf-budget.md. - Query admission (refusals are labeled 429/504, never truncated answers): per-query row budget 1M, wall-time 30 s, 8 global / 4 per-site concurrent expensive queries, funnel/retention window clamp 186 days — knobs and refusal codes in docs/operations/capacity.md.
- Metrics ingest cardinality: 20k distinct series per site, 20k data points per request (413 past — exporters must split), label truncation with counters; the series registry is bounded and in-memory.
- Ingest durability: WAL-backed group commit (durable mode by
default; a 200 means fsynced). The disk high-water (512 MB across WAL
segments by default) refuses new events with a retryable 503 in durable
mode;
OBSERVE_WAL_LOSSY=trueopts into the declared loss budget instead. Counters on/healthz. - Single install: two processes, one binary each. There is no
clustering, no multi-node storage; scale is a single instance plus the
budgets above. Ingest can be exposed separately from the dashboard via
OBSERVE_INGEST_ADDR. - No PromQL: metrics are queried via the structured query API; PromQL is an explicit future decision (see "Metrics query semantics" below).
- Pageviews, visitors, sessions, bounce rate, duration.
- Top pages, referrers, UTM tracking, channel classification.
- Browser, OS, device, country, language breakdowns.
- Custom events with property drill-down.
- Funnels, retention cohorts, user journeys, goals.
- Real-time active visitors.
- Cookie-free (no cookies are set; visitor identity is derived, and all data stays on your server). Whether that satisfies your jurisdiction's obligations is your compliance review to make, not a claim of ours.
- Automatic grouping (MD5 of type + in-app frames).
- Stack trace viewer with source-map support.
- Full-text search across messages (BM25).
- Issue status (open / resolved / ignored), release health, breadcrumbs.
- Error-to-session cross-correlation.
- OTLP ingest over HTTP for all three signals — traces, metrics and logs — in
both wire formats (
application/x-protobuf, which is what OTLP exporters send by default, andapplication/json). gRPC is not served; point an exporter at HTTP transport or put a Collector in front. - Service list with RED metrics, waterfall + flame-graph views, dependency map, p50/p95/p99 latency.
- Level, service, trace-id correlation, full-text search.
- Pipelines (JSON parse, regex extract, rename, mask, sample).
- Two recorders, one ingest and player. The structural recorder
(
observe-replay.js) ships bounded full-DOM snapshots (5,000 nodes, depth 32) plus mouse / click / scroll interactions, re-snapshotted periodically — mutations are counted, not captured. The delta recorder (observe-replay-delta.js, rrweb 2.1.6 bundled and wrapped in Observe's sanitizer) records full snapshots plus incremental DOM mutations (characterData included), so playback shows DOM changes between keyframes. Both ride the same v2 batch transport; sessions are indistinguishable server-side. - Privacy boundary (both recorders, audited allowlist policy): attribute
allowlist on every node, form controls / scripts / head subtrees /
contenteditable/data-observe-blockregions replaced by opaque placeholders, input text never recorded, image srcs sanitized to origin+path and re-loaded through the replay asset proxy. The player re-sanitizes at play time (second independent layer) inside asandbox="allow-same-origin"iframe with a strict CSP — replayed script tags stay inert. - Playback with timeline scrubbing, click heatmaps, and error correlation. Sessions recorded by either tracker replay through the appropriate player automatically (delta sessions via the rrweb Replayer class; keyframe sessions via the structural renderer).
- Uptime HTTP monitors with response-time tracking.
- Cron heartbeat monitors with missed-check detection.
- Track model calls (tokens, cost, latency).
- Cost estimation for GPT / Claude / Gemini.
- The AI query assistant dogfoods this — every generated-SQL call writes a row.
- Feature flags (boolean + multivariate, rollout %, user targeting).
- A/B experiments (frequentist p-value + Bayesian probability-to-beat).
- Surveys, custom dashboards with panels.
- RBAC enforced — JWT carries a role claim (
admin/editor/viewer). Writes require editor or admin; destructive config routes require admin. - Ingest is WAL-backed with durable acknowledgment (group commit) —
accepted events are mirrored to
$OBSERVE_QUEUE_DIRwhen the queue is available, and (the default durable mode) an ingest 200 is sent only after the batch's WAL frame has been fsynced: requests admitted within one commit window (default 25 ms,OBSERVE_WAL_GROUP_COMMIT_MAX_DELAY) share a single fsync, so a hard crash loses nothing that was acknowledged. Crash recovery replays fsynced-but-uncheckpointed records since the last checkpoint. When the WAL's disk high-water (OBSERVE_WAL_MAX_TOTAL_BYTES) is reached, durable mode REFUSES new events with a retryable 503 + Retry-After until the checkpoint advances — uncheckpointed segments are never deleted. Operators who prefer the older availability-first behavior can setOBSERVE_WAL_LOSSY=true: acknowledgments return without an fsync, the periodic sync (500 ms) bounds the declared loss budget (up to one sync interval of acked events, plus segments a high-water breach then deletes — loudly counted), and/healthzsurfaces the counters (events.acceptedvsevents.durably_acked,wal.unsynced_events). The full acknowledgment/commit contract is specified indocs/O01_DURABLE_INGEST_ADR.md. - Error ingest is durable and idempotent — error records ride the
same WAL machinery (their own
errorsqueue in$OBSERVE_QUEUE_DIR): an errors 200 means the record's frame is fsynced, a storage failure leaves it PENDING (retried, never dropped), and crash recovery replays uncheckpointed records. SDKs that send a stableevent_idget idempotent application: a retry of the same id + payload is acknowledged once and applied once ({ok:true, deduped:true}), the same id with a different payload is rejected 409, and the same guarantees hold across restarts via theerror_inboxledger. Counters live at/healthzundererrors(accepted, durably_acked, applied, deduped, quarantined, conflicting_id, pending). - Per-site rate limiting — each site has its own token bucket. One
noisy site can't starve a quiet one. Admin-editable via
PUT /api/v1/sites/{id}/ratelimit. - Alerting (threshold per metric, cooldown, silence). Alert-fire auto-opens an incident marker.
- Integrations (Jira, GitHub, PagerDuty, Slack, email) + webhooks.
- SSO / SAML, email digests, data export (CSV/JSON).
- SQL query explorer with lexer-guarded read-only enforcement
(rejects
/* comment */ INSERT ...and stacked statements). POST /api/v1/query/explainreturns the Nucleus plan.
<!-- Analytics -->
<script defer src="https://your-observe.com/t/observe.js"
data-site-id="YOUR_SITE_ID"></script>
<!-- Error tracking -->
<script defer src="https://your-observe.com/t/observe-errors.js"
data-site-id="YOUR_SITE_ID"></script>
<!-- Session replay (structural keyframes) -->
<script defer src="https://your-observe.com/t/observe-replay.js"
data-site-id="YOUR_SITE_ID"></script>
<!-- Session replay (delta recording via rrweb — DOM mutations captured;
same attributes, 30 s keyframes via data-checkout-interval) -->
<script defer src="https://your-observe.com/t/observe-replay-delta.js"
data-site-id="YOUR_SITE_ID"></script>
<!-- Feedback widget -->
<script defer src="https://your-observe.com/t/observe-feedback.js"
data-site-id="YOUR_SITE_ID"></script>observe.track("signup", { plan: "pro" });
observe.revenue(49.99, "USD", { product: "annual" });
observe.trackVitals();observeErrors.captureException(error);
observeErrors.captureMessage("Something went wrong");
observeErrors.addBreadcrumb({ type: "user", category: "click", message: "Button" });| Language | Package | Install |
|---|---|---|
| Python | teploy-observe |
pip install teploy-observe |
| Go | sdk/go |
go get github.com/useteploy/teploy-observe/sdk/go |
| Variable | Default | Description |
|---|---|---|
OBSERVE_ADDR |
:3000 |
Listen address. Keep on localhost/tailnet when publishing ingest. |
OBSERVE_WEBHOOK_ALLOW_CIDRS |
(unset) | Networks webhook delivery may reach despite being private, as CIDRs (100.64.0.0/10, 10.0.0.0/8); a bare IP means that address alone. Self-hosted fleets live on a tailnet, which the SSRF guard blocks by design — without this an alert can never reach a self-hosted receiver. Applies to webhook delivery ONLY, never to integrations or uptime monitoring. Link-local (169.254.169.254 cloud metadata), multicast and the unspecified address stay blocked whatever you declare. Hostnames are refused: allowing by name would hand back DNS rebinding. |
OBSERVE_INGEST_ADDR |
(unset) | Optional second bind address serving ONLY telemetry-write endpoints (e.g. :3001). This is the port to expose publicly; the dashboard does not listen on it. |
OBSERVE_PUBLIC_URL |
(unset) | External base URL (https://observe.example.com) used for SSO metadata and generated links. Falls back to the request's Host header, which a client can spoof — set it whenever the instance is reachable by a name. |
OBSERVE_NUCLEUS_URL |
postgres://localhost:5432/observe |
Nucleus connection. |
OBSERVE_JWT_SECRET |
(random) | JWT signing secret — set in prod to persist sessions across restarts. |
OBSERVE_SECRET_KEY |
(unset) | Master key for encrypting stored secrets (LLM API key, S3/R2 credentials) at rest; required to configure those features. |
OBSERVE_ADMIN_USER |
admin |
Bootstrap admin username. |
OBSERVE_ADMIN_PASSWORD |
(unset) | Bootstrap admin password; unset means no default — the /setup wizard creates the account on first visit. |
OBSERVE_SESSION_SALT |
(random) | Session-ID hashing salt — set it to keep session/visitor IDs stable across restarts. |
OBSERVE_DEMO_MODE |
(unset) | Set to true to lock the deployment to a read-only public demo (write ops on /api/v1/* return 403). |
OBSERVE_SEED_DEMO |
(unset) | Set to true for first-boot demo seeding (off by default; also on when demo mode is set). |
OBSERVE_RATE_LIMIT |
1000 |
Default per-site events/sec. Per-site overrides via API. |
OBSERVE_TRUSTED_PROXIES |
(unset) | Comma-separated CIDRs/IPs whose X-Forwarded-For / X-Real-Ip are trusted for client-IP extraction. Empty trusts none (peer address) so clients can't spoof their IP to evade per-IP rate limiting. |
OBSERVE_BUFFER_SIZE |
100000 |
Max buffered events in memory. |
OBSERVE_FLUSH_SIZE |
500 |
Flush threshold (events). |
OBSERVE_FLUSH_INTERVAL_MS |
2000 |
Flush threshold (time). |
OBSERVE_DATA_DIR |
./data |
Root dir for WAL, queue, local state. |
OBSERVE_QUEUE_DIR |
$OBSERVE_DATA_DIR/queue |
Ingest WAL directory. |
OBSERVE_WAL_MAX_SEGMENT_BYTES |
67108864 |
Per-segment cap: appends past it roll to a fresh numbered WAL segment. |
OBSERVE_WAL_MAX_TOTAL_BYTES |
536870912 |
Disk high-water across all WAL segments. Durable mode (default): a roll that would breach it refuses admission with a retryable 503 until the checkpoint advances (wal.refused_high_water at /healthz). Lossy mode: the oldest segment is dropped (loudly, counted) — those events lose their crash-recovery copy but ingestion keeps accepting work. |
OBSERVE_WAL_GROUP_COMMIT_MAX_DELAY |
25 |
Durable mode: bound of the group-commit window in ms — ingest 200s wait for one shared WAL fsync per window instead of an fsync per event. |
OBSERVE_WAL_LOSSY |
(unset) | Set to true (or 1) to opt into lossy mode: acknowledgments return without an fsync (declared loss budget: one 500 ms sync interval plus high-water headroom) and the disk high-water deletes instead of refusing. Default is durable mode. |
OBSERVE_AUDIT_KEY |
Audit-chain HMAC key: base64 (>=32 decoded bytes) or raw (>=32 bytes). Unset uses the persistent generated key at $OBSERVE_DATA_DIR/audit.key. |
|
OBSERVE_AUDIT_KEYRING |
Comma-separated id:base64key historical audit keys kept for verification through rotation (startup logs the ready-to-paste entry for the file key). |
|
OBSERVE_REPLAY_ASSET_HOSTS |
Comma-separated hostnames the replay asset proxy may fetch images from (empty disables proxied replay images). | |
OBSERVE_REQUIRE_WAL |
(unset) | Set to true (or 1) to refuse to start when WAL-backed ingestion durability is unavailable, instead of degrading to memory-only. |
OBSERVE_RAW_RETENTION_DAYS |
30 |
Raw event retention. Also the window over which visitor counts are exact from raw events; past it they are counted from the sessions table (90 days), and past both the dashboard says which window the figure covers. |
OBSERVE_HOURLY_RETENTION_DAYS |
365 |
Hourly rollup retention. |
OBSERVE_LLM_RETENTION_DAYS |
30 |
LLM trace retention, including stored prompts/completions. Must be at least 1. Cleanup runs daily; historical model-price catalog entries are retained. Logical expiry does not guarantee database files immediately shrink. |
OBSERVE_LOG_ROUTES |
0 |
Set to 1 to print route table at boot. |
OBSERVE_SMTP_HOST |
SMTP server for email reports. | |
OBSERVE_SMTP_PORT |
587 |
SMTP port. |
OBSERVE_SMTP_USER |
SMTP username. | |
OBSERVE_SMTP_PASS |
SMTP password. | |
OBSERVE_SMTP_FROM |
From email address. | |
TEPLOY_NAV_DASH_URL |
URL of your Teploy Dash dashboard. When set, it appears in the top-left cross-product switcher. | |
TEPLOY_NAV_SHIP_URL |
URL of your Teploy Ship dashboard. When set, it appears in the top-left cross-product switcher. |
Optional. When OBSERVE_OIDC_ISSUER and OBSERVE_OIDC_CLIENT_ID are set, the
login page offers an SSO button and Observe acts as an OpenID Connect relying
party (authorization-code flow with PKCE), minting its normal JWT after the IdP
authenticates the user. Password login stays available as the break-glass path.
Register https://<your-observe-host>/api/v1/auth/oidc/callback as the redirect
URI with your provider. When SSO is enabled, the first-run open-access grace
period is disabled (authentication becomes required).
| Variable | Default | Description |
|---|---|---|
OBSERVE_OIDC_ISSUER |
IdP issuer URL (discovery base, e.g. https://your-org.okta.com). Required to enable SSO. |
|
OBSERVE_OIDC_CLIENT_ID |
OAuth client ID. Required to enable SSO. | |
OBSERVE_OIDC_CLIENT_SECRET |
OAuth client secret. Omit for a public (PKCE-only) client. | |
OBSERVE_OIDC_REDIRECT_URL |
(derived) | Callback URL. Derived from the request Host when unset; set explicitly behind a proxy that rewrites Host. Must end in /api/v1/auth/oidc/callback. |
OBSERVE_OIDC_SCOPES |
openid profile email |
Space/comma-separated scopes (openid always included). Add groups for group-based role mapping. |
OBSERVE_OIDC_LABEL |
Single sign-on |
Text on the SSO button. |
OBSERVE_OIDC_USERNAME_CLAIM |
preferred_username |
Claim used as the username (falls back to email, then sub). |
OBSERVE_OIDC_ROLE_CLAIM |
teploy_role |
Claim carrying the role directly (admin/editor/viewer). Checked first. |
OBSERVE_OIDC_GROUPS_CLAIM |
groups |
Claim listing the user's groups, used when no direct role claim matches. |
OBSERVE_OIDC_ADMIN_GROUP |
Group whose members become admin. |
|
OBSERVE_OIDC_EDITOR_GROUP |
Group whose members become editor. |
|
OBSERVE_OIDC_VIEWER_GROUP |
Group whose members become viewer. |
|
OBSERVE_OIDC_DEFAULT_ROLE |
viewer |
Role for an authenticated user matching no role claim or group (least privilege). |
Role resolution order: a recognized teploy_role claim wins; otherwise groups
are matched (admin > editor > viewer); otherwise the default role. SSO users are
not stored in the admin_users table — their role comes fresh from the IdP on
every login.
Any OIDC provider works. Two are worth calling out because if you already run Teploy you probably already run one of them, so SSO costs you no new software.
Forgejo (or Gitea) is a full OIDC provider. Its discovery document
advertises openid profile email groups and a groups claim.
- Register an OAuth2 application — Site Administration → Applications for an
org-wide one, or user Settings → Applications for a personal one. Set the
redirect URI to
https://<your-observe-host>/api/v1/auth/oidc/callback. - Point Observe at it:
OBSERVE_OIDC_ISSUER=https://forgejo.example.com
OBSERVE_OIDC_CLIENT_ID=<client id>
OBSERVE_OIDC_CLIENT_SECRET=<client secret>
OBSERVE_OIDC_SCOPES="openid profile email groups"
OBSERVE_OIDC_ADMIN_GROUP=platform:owners
OBSERVE_OIDC_EDITOR_GROUP=platform:deployers- Request
groupsexplicitly. It is not in the default scopes, and without it no group matches, so every user lands onOBSERVE_OIDC_DEFAULT_ROLE. - Forgejo emits one entry per org (
platform) and one per team (platform:deployers). Group comparison is exact and case-sensitive, so copy the names as Forgejo spells them. - Forgejo cannot mint a custom claim, so leave
ROLE_CLAIMat its default and map roles by group. - Each dashboard needs its own OAuth2 application because the redirect URIs differ, but all three can map against the same orgs and teams.
OpenBao also serves OIDC (identity/oidc/provider), which is convenient if
you already run it for teploy secret --provider openbao. Create a provider,
an assignment, and a client, then use the provider's discovery URL as the
issuer:
OBSERVE_OIDC_ISSUER=https://openbao.example.com/v1/identity/oidc/provider/teployMap roles with a scope template that emits a groups array (matched as above),
or one that emits a teploy_role string — OpenBao can produce a custom claim,
so the direct role claim is available here and takes precedence over groups.
| Method | Path | Description |
|---|---|---|
| POST | /api/v1/events |
Ingest analytics event. |
| POST | /api/v1/events/batch |
Ingest batch of events. |
| POST | /api/v1/errors |
Ingest error event. |
| POST | /api/v1/logs |
Ingest log entry. |
| POST | /v1/traces |
OTLP trace ingestion (protobuf or JSON). |
| POST | /v1/metrics |
OTLP metric ingestion (protobuf or JSON). |
| POST | /v1/logs |
OTLP log ingestion (protobuf or JSON). |
| POST | /api/v1/llm/ingest |
Ingest LLM trace. |
| POST | /api/v1/infra/report |
Host metrics. |
| POST | /api/v1/replays |
Session replay events. |
| POST | /api/v1/feedback |
User feedback. |
| Method | Path | Role | Description |
|---|---|---|---|
| GET | /api/v1/stats/* |
any | Overview, timeseries, pages, referrers, journeys, correlations, retention. |
| GET | /api/v1/issues |
any | Error issue list. |
| GET | /api/v1/traces/* |
any | Service RED metrics, search, waterfall. |
| GET | /api/v1/logs/search |
any | Log search. |
| POST | /api/v1/query |
editor+ | SQL explorer (read-only, lexer-guarded). |
| POST | /api/v1/query/explain |
editor+ | Return the Nucleus plan. |
| POST | /api/v1/ai/query |
editor+ | NL → SQL via configured LLM. |
| GET/PUT | /api/v1/ai/config |
admin | Instance LLM provider / key. |
| GET | /api/v1/incidents |
any | List / filter incidents. |
| POST | /api/v1/incidents |
editor+ | Declare manual incident. |
| POST | /api/v1/incidents/{id}/close |
editor+ | Close. |
| GET/POST/DELETE | /api/v1/exports/scheduled |
admin | Scheduled SQL exports to S3. |
| POST | /api/v1/sites |
admin | Create site. |
| DELETE | /api/v1/sites/{id} |
admin | Delete site. |
| PUT | /api/v1/sites/{id}/ratelimit |
admin | Set per-site events/sec cap. |
| POST | /api/v1/platform/* |
admin | Users, alert rules, webhooks. |
Observe does not implement PromQL. Metrics are queried via Observe's
structured query API (GET /api/v1/metrics/query and
GET /api/v1/metrics/series: metric name, exact-match label filters,
group-by, step, and the reducers last|avg|sum|min|max|rate|p50|p95|p99).
PromQL support is an explicit future decision, not an implied
compatibility; if it ever lands it will wrap the upstream Prometheus
engine, never a hand-rolled subset.
Reducer semantics are pinned by hand-computed reference tables
(internal/metrics/o07_reference_test.go):
rateis computed per series before any cross-series aggregation; a collapsed query returns the sum of per-series rates. On cumulative counters a value decrease is a counter reset — a new epoch counted under the restart-at-zero assumption (the Prometheusrate()convention), never a negative rate. Bucket values are time-weighted (increase / covered span), so mixed scrape intervals and gaps weigh samples by duration. Delta-temporality counters are direct: the bucket sum divided by the bucket length.p50/p95/p99interpolate inside the crossing histogram bucket between its explicit boundaries (the official cumulative-histogram convention); cumulative histograms are differenced per series first.- Every point carries
estimate+method: histogram quantiles are alwaysestimate: true(interpolated from bucketed data); rate buckets that invoked the restart-at-zero assumption are labeled estimates naming the assumption; everything else is labeledestimate: false, method: "exact".
Browser / SDKs
|
v
Observe (Go, ~26MB) --- JWT / API key auth, RBAC middleware
| per-site rate limiter
| disk-backed ingest WAL
| background jobs (rollups, retention,
| exports, alerts)
| embedded SPA dashboard
| AI query assistant (admin-supplied LLM key)
|
v pgwire
Nucleus (Rust, ~32MB) --- SQL + multi-model (KV, columnar, FTS,
vector, doc, graph, time-series)
MergeTree, ReplacingMergeTree
WAL, TLS, query cache
Two processes. No Redis, no Kafka, no ClickHouse, no ZooKeeper.
Server: FSL-1.1-MIT (Functional Source License) — any use is permitted
except offering a competing product, and each version automatically becomes
MIT two years after its release. See LICENSE.
SDKs (sdk/browser, sdk/sentry-shim, sdk/python, sdk/go):
MIT. Each SDK subdirectory has its own LICENSE.
See CONTRIBUTING.md for contribution guidelines and SECURITY.md for
vulnerability reports.