diff --git a/CHANGELOG.md b/CHANGELOG.md index 97114856..012b62db 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,6 +5,75 @@ All notable changes to this project are documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). +## [Unreleased] + +### Added + +- **Value correlation (D13).** A fifteenth statistical detector, `value_correlation`, + adapted from AMiner's `VariableCorrelationDetector` and intra-record: for a field pair it + mines implication rules `A = x ⇒ B = y` — an antecedent value with at least + `stat_correlation_min_support` reference events whose dominant consequent accounts for at + least `stat_correlation_rule_confidence` of them, both directions — and reports a window in + which the rule's violation rate rises: a 2×2 G-test of conforming against violating events + between the reference and the window, one Benjamini–Hochberg pool per run, an effect floor + of `stat_correlation_min_ratio` on the violation-rate ratio. Only rises are reported; a rule + that appears is a proportion shift. Both frames: `rule-g-test` mines from the baseline + window and tests each suspect window; `self-rule-g-test` mines from the timeline and tests + each leave-one-out slice against the rest. Pairs come from an explicit field list or the + recommender's top `stat_correlation_auto_fields` categorical fields, capped at + `stat_correlation_max_pairs` with a warning; each pair is one `GROUP BY a, b` scan capped + at `stat_correlation_max_rows_per_pair` rows. Findings carry the rule as mined, both + sides' counts and violation rates, and the consequent value that most often took the + rule's place; the allowlist key is the combo one. Gate entry (two categorical fields, a + sliceable span in the self frame), wizard card, evidence figure, agent knob + (`rule_confidence`), seven settings with registry specs, run snapshot and + `docs/ANOMALY_DETECTION.md` §17. The demo case asserts the contractor's + `user ⇒ home workstation` rule breaking on every host the intrusion visits, in both frames. + +- **Time-of-day habit (D12).** A fourteenth statistical detector, `time_of_day`, adapted + from AMiner's `PathValueTimeIntervalDetector`: per (field, value) it cuts the day into + `bucket_minutes`-wide wall-clock buckets (15–240 minutes, default 60) read in an explicit + IANA `timezone` (default `UTC`, `stat_habit_timezone` for the site), learns the value's + habit — the buckets holding at least `stat_habit_min_bucket_count` reference occurrences, + for values with at least `stat_habit_min_baseline` of them — and reports an occurrence in + any other bucket, scored by the circular distance in hours to the nearest habitual one. + Interval cadence measures the gap between arrivals; this reads the hour on the wall, and + the two are independent (a nightly job that moves keeps its cadence and breaks its + habit). Both frames: `habit` learns from the baseline window and scores each suspect + window; `self-habit` takes the value's own busy buckets across the timeline and scores + its thin ones. The zone and resolution are snapshotted into the persisted run and carried + on every finding, since the same instant is a different hour elsewhere. Auto field + selection follows the novelty recommender and the timeline's field overrides; the + allowlist key is `(field, value)`. Gate entry (always offered in the self frame), wizard + card with a bucket choice and a zone box, a day-strip evidence figure, agent knobs, five + settings with registry specs and `docs/ANOMALY_DETECTION.md` §16 ship with it. The demo + case gains a one-off manual afternoon backup run (the self-frame signal) and moves the + contractor's lateral movement onto the jump host at 03:00, a host whose baseline logons + are an administrator's office hours; the nightly backup's move to 03:40 is the benign hit. + +- **Transition speed (D15).** A thirteenth statistical detector, `transition_time`, adapted + from AMiner's `MinimalTransitionTimeDetector`: per ordered value pair of a series field + it learns the fastest a stream ever moved from one value to the next and reports a + transition that undercuts that floor by at least `stat_transition_min_ratio` (default + 2×) — one account on two hosts seconds apart, a session skipping states. Transitions are + one step of the sequence detectors' n-gram assembly, per source and per value of a new + `partition_field` knob (the identifier whose moves are timed; rows without it are left + out rather than pooled), so two users' interleaved logons never read as one actor. Both + frames from day one: with a baseline the floor is the baseline window's fastest + transition of the pair (`min-transition`); without one it is the pair's next-fastest + transition anywhere on the timeline (`self-min-transition`, leave-one-out by + construction). A floor is learned from at least `stat_transition_min_transitions` (3) + transitions, a zero floor is skipped and counted in a warning, the per-source candidate + cap `stat_transition_max_candidates` (2000) keeps the fastest pairs and discloses itself, + and score is `1 − observed / reference`. The finding carries the pair, the stream that + made the move, both durations, which floor was used and the speed-up; the allowlist key + is `(series_field, "a → b")` in both frames. Gate, wizard card, evidence figure, agent + tool, persisted-run snapshot (`partition_field`, `min_transitions`) and + `docs/ANOMALY_DETECTION.md` §15 ship with it; the analysis cache moves to version 5. The + demo case gains an administrator's routine jump-host hop as the floor and the + contractor's wmic call landing on `FILE-01` two seconds later as the signal, asserted in + both frames. + ## [1.19.7] — 2026-09-15 ### Added diff --git a/CLAUDE.md b/CLAUDE.md index ef989a91..7b9465cc 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -26,7 +26,7 @@ trail (`api/routers/auth.py`, `admin.py`, `deps.py`). - `CONCEPT.md` / `MODEL_REFINEMENT.md` — product vision and the Case/Source/Timeline/Event/ Artifact data model. Read before touching the model; rarely changes. - `TECH_STACK.md` — backing-service decision record (*why*, not *what's shipped*). -- `ANOMALY_DETECTION.md` — reference for all fourteen analysis tools actually running +- `ANOMALY_DETECTION.md` — reference for all seventeen analysis tools actually running (statistical detectors, Sigma runner, log templates, semantic similarity), plus the baseline/disposition model. Update alongside any detector change in the same commit. - `AGENT.md` — the optional AI investigation agent (design invariants, MCP tools, provider @@ -253,7 +253,7 @@ instead of rebuilding. does not strand its verdict bar a screen below the claim; `ToolsSheet` is four tabs (Scope, Methods, Signatures, Explore) rather than one scroll, so a thousand-row template list cannot bury the baseline picker; `method-registry.ts` is the - single description of all twelve methods, including the prose that used to live in a + single description of all fifteen methods, including the prose that used to live in a Method tab and each method's optional `railFloor`, a presentation-only bar on the ranked feed whose held-back count is always disclosed; the sheet's method mode runs a method with the analyst's own knob values, which is what keeps the analysis diff --git a/README.md b/README.md index 6de7c371..f4963ede 100644 --- a/README.md +++ b/README.md @@ -69,7 +69,7 @@ resolved on the machine you deploy to. ClickHouse, time histogram with anomaly overlays, keyset pagination with jump-to-time, tag/comment annotations with bulk apply, saved views, and streaming CSV/JSONL export that keeps the forensic columns. -- **Anomaly detection** — fourteen analysis tools: twelve statistical detectors over +- **Anomaly detection** — seventeen analysis tools: fifteen statistical detectors over ClickHouse needing no embeddings, a Sigma rule runner, and semantic similarity search over local embeddings. Each is documented method by method, scores against explicit baseline-vs-suspect windows, and yields findings whose confirm/dismiss disposition @@ -109,7 +109,7 @@ feel like, and the Case/Timeline model here is descended from it. That is the co invite, and three axes are where we think we are already the better place to run an investigation: -- **Detection is the workflow, not an add-on** — fourteen analysis tools in the box, each +- **Detection is the workflow, not an add-on** — seventeen analysis tools in the box, each scoring against an analyst-declared baseline and carrying a verdict that survives re-scans, so triage accumulates instead of being redone. - **Provenance goes all the way down** — not just "this file was imported": a finding is diff --git a/docs/AGENT.md b/docs/AGENT.md index 615f0623..7f35e432 100644 --- a/docs/AGENT.md +++ b/docs/AGENT.md @@ -716,6 +716,14 @@ detector runs without a baseline (D18) took it to **43,891 over 35 tools**, ceil unchanged — the first draft landed at 44,197 and the rule was applied: the tool's docstring was rewritten compact (−306 chars net) rather than the ceiling moved. +The 1.20 detectors (2026-09-16: `transition_time`'s `partition_field`, `time_of_day`'s +`bucket_minutes` and `timezone`) took the first draft to 44,421 and the rule was applied +again: the docstring now names each detector once and each knob in a few words, and the +`timezone` length constraints came off the schema (the runner validates the zone anyway) — +**43,832 over 35 tools**, ceiling unchanged, below the D11 figure. `value_correlation`'s +`rule_confidence` (D13) pushed it over once more; the same treatment (no schema-level +bounds, a terser docstring) lands it at **43,845 over 35 tools**. + Detector findings additionally reduce their inline example event in the **model's copy** to `event_id` + truncated `message` (`_deflate_findings` — on the turn that motivated it: 33.7k → 15.7k tokens); diff --git a/docs/ANOMALY_DETECTION.md b/docs/ANOMALY_DETECTION.md index ef29c5a1..25ebda37 100644 --- a/docs/ANOMALY_DETECTION.md +++ b/docs/ANOMALY_DETECTION.md @@ -7,7 +7,7 @@ This document covers every detector actually running in the codebase today. If a detector described here changes (formula, default, field name), update this file and the "Method" tab copy in the same commit. -There are fourteen independent analysis tools in Vestigo: +There are seventeen independent analysis tools in Vestigo: 1. [Value novelty](#1-value-novelty-rare--first-seen-values) — rare/new field values, single field or [combinations](#value-combinations-the-value_combo-variant) (ClickHouse, no ML) 2. [Frequency anomalies](#2-frequency-anomalies-volume-spikes--silences) — volume spikes/silences (ClickHouse, no ML) @@ -23,13 +23,16 @@ There are fourteen independent analysis tools in Vestigo: 12. [Repeating sequences](#12-repeating-sequences-motif-mining) — recurring time-ordered n-grams of a field's values, ranked by support and cadence regularity; the discovery/mining complement of detector 9 (ClickHouse, no ML) 13. [Sigma rule runner](#13-sigma-rule-runner-signature-matching) — signature matching: community/custom Sigma rules compiled to ClickHouse predicates (ClickHouse, no ML) 14. [Log templates](#14-log-templates-structural-line-clustering) — structural clustering of raw lines into templates, so rare *shapes* surface without naming a field (ClickHouse, no ML) +15. [Transition speed](#15-transition-speed-value-to-value-moves-faster-than-ever-seen) — a stream reaching the next value of a field faster than that pair was ever reached before (ClickHouse, no ML) +16. [Time-of-day habit](#16-time-of-day-habit-values-at-an-hour-they-never-keep) — a value occurring at a wall-clock hour it has no habit of, in an explicit timezone (ClickHouse, no ML) +17. [Value correlation](#17-value-correlation-field-to-field-rules-that-break) — an implication rule between two fields of the same event (`A = x ⇒ B = y`) whose violation rate rises in a window (ClickHouse + a real significance test, no ML) All but the eleventh are **statistical or rule-based**: pure counting, arithmetic and predicate matching over already-ingested events — no machine learning, no network calls, working the instant ingestion finishes. The eleventh needs an explicit embedding step first. -Code: `src/vestigo/db/anomaly_stats.py` (detectors 1–10, 12, 14), +Code: `src/vestigo/db/anomaly_stats.py` (detectors 1–10, 12, 14–17), `src/vestigo/db/similarity.py` (detector 11), `src/vestigo/sigma/` (detector 13). UI: `frontend/src/components/analysis/`. @@ -61,6 +64,9 @@ teaches the tooling instead of the work. What each tool has to find: | 12 | Repeating sequences | The beacon motif (proxy request + matching firewall allow in the same second), alongside a benign nightly-backup motif. | | 13 | Sigma | Four case-scoped rules ship inside the case: encoded PowerShell, suspicious service install, wmic remote process creation, failed-logon burst. | | 14 | Log templates | ~40 syslog templates, with an `unattended-upgrade` shape that appears only in the suspect window. | +| 15 | Transition speed | The contractor account on `FILE-01` two seconds after its wmic call from `JUMP-01`, timed per account over `attr:computer_name`, against an administrator whose routine jump-host hop takes the same pair 20–90 s. | +| 16 | Time-of-day habit | The nightly backup program at 03:xx and 04:xx after three weeks of 02:xx runs (benign, the same move interval cadence sees), and the contractor on `JUMP-01` at 03:00, a host whose baseline logons are an administrator's office hours. Without a baseline, the one manual afternoon backup run twelve hours from the program's nightly slot. | +| 17 | Value correlation | The rule `user = m.okonkwo ⇒ computer_name = WKS-004`, which holds over three weeks of the contractor's logons and breaks on every host the intrusion takes the account to — the jump host, the file server, two finance workstations. Without a baseline, the same rule against the rest of the timeline, slice by slice. | `tests/test_demo_detector_coverage_clickhouse.py` asserts that each of these actually returns findings. If a retuned threshold silences one of them, that @@ -351,15 +357,15 @@ checked, and the UI must never present the two the same way. | Method | Precondition | Setting | |---|---|---| -| `value_novelty`, `timestamp_order`, `log_template`, `entropy` | none — always offered | — | -| `value_combo` | ≥2 categorical fields | — | +| `value_novelty`, `timestamp_order`, `log_template`, `entropy`, `time_of_day` | none — always offered | — | +| `value_combo`, `value_correlation` | ≥2 categorical fields | — | | `numeric_range` | ≥1 field whose sampled values are ≥90 % numeric | `analysis_gate_min_numeric_ratio` | | `charset` | ≥1 field above the enum-like ceiling | `analysis_gate_max_enum_distinct` | | `frequency` | span of at least the minimum number of seconds | `analysis_gate_min_frequency_buckets` | | `interval_periodicity` | enough events per series value to fit a cadence | `analysis_gate_min_interval_periods` | -| `sequence_novelty` | a series field with ≥2 distinct values | `analysis_gate_min_series_distinct` | -| `proportion_shift`, `value_distribution_drift` | a span of more than one instant (self frame: it is cut into slices) | `stat_self_slices` | -| any of those four in the **baseline frame** | an active baseline definition | — (reported as `needs_setup`) | +| `sequence_novelty`, `transition_time` | a series field with ≥2 distinct values (one value yields no ordering and no transition) | `analysis_gate_min_series_distinct` | +| `proportion_shift`, `value_distribution_drift`, `value_correlation` | a span of more than one instant (self frame: it is cut into slices) | `stat_self_slices` | +| any of those six, or `time_of_day`, in the **baseline frame** | an active baseline definition | — (reported as `needs_setup`) | Three rows encode a distinction worth stating, because each was once drawn wrong: @@ -2538,6 +2544,335 @@ entry if that need grows past a browser flag. --- +## 15. Transition speed (value-to-value moves faster than ever seen) + +**What it answers:** "Did something get from *here* to *there* faster than it ever +has?" The AMiner `MinimalTransitionTimeDetector` analog (roadmap D15). One account +logging on to two hosts three seconds apart, a session skipping from `created` to +`closed` with nothing in between, a workflow reaching its last state in a fraction of +its usual time — each value is ordinary, each ordering may even be ordinary, but the +*speed* is not. Event sequences (§9) own the ordering axis; this detector owns the +time between two consecutive values. + +**How it works.** A **transition** is one step `a → b` between two consecutive events +of one *stream* whose `series_field` values differ (same-value repeats are not +transitions). The stream is the source, further split by the **partition field** when +one is set — the identifier whose movements are being timed, `attr:user` for host +moves, a session id for state moves. Without it every source is a single stream, and +two users' interleaved logons read as one actor moving between their hosts, which is +rarely the question. Transitions are one step of the sequence detectors' n-gram +assembly (`_ngram_inner_sql`, n = 2), so they inherit its deterministic ordering +(effective timestamp, then record order), the once-per-source scan discipline, and the +guarantee that a pair never spans a window boundary. The duration is +`dateDiff('millisecond', previous, current)`. Per ordered pair the detector learns a +**floor** — the fastest that pair was ever reached in the reference — and flags a +transition that undercuts it by at least the speed-up factor (`min_ratio`, default 2): +a floor learned from a handful of transitions is not precise to the second, so +"slightly faster" is not a finding. + +**Two frames.** + +| | Self (`self-min-transition`) | Baseline (`min-transition`) | +|---|---|---| +| Reference | the pair's **next-fastest** transition anywhere in the scope | the pair's fastest transition in the baseline window | +| Learned from | at least `stat_transition_min_transitions` transitions of the pair in the scope (default 3) | at least that many in the baseline window | +| Flags | the pair's single fastest transition, when it undercuts the next-fastest by `min_ratio`× | each suspect window's fastest transition of the pair, when it undercuts the baseline floor by `min_ratio`× | +| Zero floor | two instantaneous transitions: skipped, counted in a warning | a baseline floor of zero: skipped, counted in a warning | + +The self frame is leave-one-out by construction: every other transition of the pair is +at least as slow as the second-fastest, so the fastest is judged against everything +else that pair ever did. Two equally fast transitions vouch for each other and nothing +is flagged. A zero floor cannot be undercut — with second-resolution timestamps, +zero-length transitions are routine — so such pairs are skipped rather than scored, and +the run says how many. + +A zero-length *observation* is only as precise as the timestamps it is the difference +of. In a source that records sub-second timestamps (any event with a non-zero +millisecond part) a 0 ms gap is genuinely instant and judged as such. In a source whose +every timestamp is a whole second — Plaso CSV, syslog — `12:00:05 → 12:00:05` says only +that the move took *under a second*: judged as instant it would undercut any positive +floor (`0 × min_ratio` beats everything) and rank first, though it may have taken +0.99 s against a 1 s floor. So there it is judged and scored at its **one-second +bound**: flagged only when one second still undercuts the floor by `min_ratio`×, scored +`1 − 1 / reference_seconds`, with `speedup` a lower bound. Pairs the bound holds back +are counted in a warning. The resolution is probed per source, only for a source that +produced a zero-length candidate, and the probe stops at its first sub-second +timestamp. A zero-length finding carries `timestamp_resolution` (`second` / +`sub-second`) and `observed_upper_bound_seconds` (`1.0`, or `null` when instant). + +Per source: the `stat_transition_max_candidates` fastest pairs (baseline frame: per +suspect window) are fetched fastest-first, with a warning when the cap is hit — the cap +keeps the fastest, which are the ones the question is about; in the baseline frame the +floor is then learned for exactly those candidate pairs. On a multi-source scope the +floor is the minimum over every source's reference and counts are summed, so a pair +that is slow in one source and fast in another is judged against the fast one. + +**Score = 1 − observed / reference** (observed at its bound, above), in `[0, 1]`: 1.0 is an instantaneous transition, +0.5 is exactly twice as fast as the floor, and nothing under `1 − 1/min_ratio` is +reported. The representative event is the **arriving** event of the fastest transition +(the `b` side), and `first_seen` is its timestamp. Findings carry `observed_seconds`, +`reference_seconds`, `reference_kind` (`next-fastest` / `baseline-min`), `speedup` +(reference ÷ observed; `null` when the observation is instant), `count` (the pair's +transitions in the window, or in the scope), `baseline_count` (the transitions the +floor was learned from) and, when a partition field was set, `partition_field` / +`partition_value` — the stream that made the move, which is how the analyst finds the +actor. The baseline frame adds `window_label`/`window_start`/`window_end`; the self +frame adds `scope_transitions` and no window keys. + +**Parameters.** + +- `series_field` (request, default `artifact`) — the field whose consecutive values + form transitions. Shared with the frequency and sequence detectors' group-by; any + field token works, including `attr:` and mapped canonical fields. As with event + sequences, a source whose every row carries one artifact value has no transitions + under the default — pick the field that moves (`attr:computer_name`, a state field). +- `partition_field` (request, default unset) — the stream key. Rows without a value + for it are left out rather than lumped into one anonymous stream. Snapshotted into + the persisted `DetectorRun` as `partition_field`. +- `min_ratio` (request) / `VESTIGO_STAT_TRANSITION_MIN_RATIO` (server default 2.0, + must exceed 1) — the speed-up floor. Snapshotted as `min_ratio`. +- `VESTIGO_STAT_TRANSITION_MIN_TRANSITIONS` (default 3) — the learning floor. + Snapshotted as `min_transitions`. +- `VESTIGO_STAT_TRANSITION_MAX_CANDIDATES` (default 2000) — the per-source candidate + cap; hitting it attaches a warning. + +**Allowlist key:** `(series_field, "a → b")` — the same shape as event sequences, in +both frames, so one **Normal** verdict on a pair covers it wherever it recurs. + +### Caveats + +- **The floor is only as good as the reference.** A baseline in which a pair occurred + three times has a floor that is the fastest of three draws; the learning floor keeps + one-off pairs out, and `min_ratio` keeps "a bit faster" out, but on a thin baseline + the flagged speed-ups are as much about the baseline's poverty as about the suspect + window. Prefer longer baselines, or the self frame, which learns from everything. +- **Interleaved streams manufacture speed.** Without a partition field, a multi-writer + source (one Windows log carrying every account) times the gap between *any* two + consecutive events with different values, whoever produced them — and two accounts + on two hosts in the same second reads as an impossible move. Set the partition field + to the identifier whose moves are meant, as the demo case does with `attr:user`. +- **Timestamp resolution bounds the question.** Second-resolution sources produce + zero-length transitions routinely; a pair whose floor is zero is skipped with a + warning rather than scored, and a floor of one second is compared against + observations that are themselves rounded. Millisecond sources make the detector far + sharper. +- **Only the fastest transition per pair is reported** (per window in the baseline + frame). `count` says how many transitions of the pair the window held; it does not + say how many undercut the floor. The Explorer drill on the pair shows all of them. +- A fast transition is **not malicious by itself** — a scripted deployment legitimately + touches ten hosts in ten seconds. Rank for triage; mark the routine pair Normal once. + +## 16. Time-of-day habit (values at an hour they never keep) + +**What it answers:** "Does this value keep hours — and did it break them?" The AMiner +`PathValueTimeIntervalDetector` analog (roadmap D12). A backup job that runs at 02:15, +a service account that works office hours, a user who has never logged on after +eight: each has a daily *habit*, and an occurrence outside it is worth a look whatever +the value itself is. Interval cadence (§8) measures the gap between arrivals; this +detector reads the hour on the wall. The two are independent: the demo's backup keeps +its once-a-day cadence perfectly when it moves from 02:15 to 03:40, and breaks its +habit. + +**How it works.** The day is cut into `1440 / bucket_minutes` wall-clock buckets +(default 60, so 24 hourly buckets) read in an explicit **IANA timezone** (default +`UTC`; `stat_habit_timezone` sets a site's zone once). The zone is not a +presentation choice: the same instant is 03:40 in Berlin and 01:40 in London, so the +run records it in `DetectorRun.params` and every finding carries it, or the run could +not be reproduced. Per (field, value) the detector counts occurrences per bucket in +the reference, and the value's **habit** is the set of buckets holding at least +`stat_habit_min_bucket_count` (3) of them — learned only for values with at least +`stat_habit_min_baseline` (20) reference occurrences, since a handful of events say +nothing about hours kept. An occurrence in any other bucket is flagged, one finding +per (value, bucket) — per suspect window in the baseline frame — with the bucket's +count, and scored by the **circular distance in hours** to the nearest habitual +bucket: 23:xx is one hour from a 00:xx habit, not twenty-three. + +**Two frames.** + +| | Self (`self-habit`) | Baseline (`habit`) | +|---|---|---| +| Reference | the value's own occurrences across the whole scope | the value's occurrences in the baseline window | +| Habit | its buckets holding ≥ `min_bucket_count` scope occurrences | its buckets holding ≥ `min_bucket_count` baseline occurrences | +| Flags | every *thin* bucket — under the floor, hence not habitual — against the busy ones | each suspect window's occurrences outside the baseline habit, habitual or not | +| Catches | one manual afternoon run of a nightly job | a nightly job that moved; an account active at 03:00 for the first time | + +The self frame is leave-one-out in the only sense that matters here: a value that +occurs once at 15:00 and fifty times at 02:00 is judged by the fifty. What it cannot +see is a *repeated* new hour — a bucket that fills past the floor becomes habitual by +definition — which is what a baseline is for: the same fifty-plus-eight becomes "eight +occurrences at 03:xx against a baseline habit of 02:xx" once the analyst declares the +window. + +One scan per field: per (value, bucket, window) counts with the first occurrence, +restricted to the `stat_habit_max_candidates_per_field` (500) highest-volume values +(cap → warning; lower than the other per-field caps because each value returns one +row per occupied bucket and window). Auto field selection is the novelty +recommender's categorical set, steered by the timeline's +[field overrides](#declaring-which-fields-a-method-reads). Values below the learning +floor are counted in a warning rather than silently skipped. + +**Score = hours to the nearest habitual bucket**, in whole multiples of the bucket +width: a 60-minute resolution scores 1, 2 … 12; a 15-minute one 0.25 upward. The +representative event is the first occurrence in the bucket (within the window, or the +scope), and `first_seen` is its timestamp. Findings carry `bucket`, `bucket_label` +(`"03:00–04:00"`), `bucket_minutes`, `timezone`, `count`, `baseline_count` (the +reference occurrences the habit was learned from), `habit_buckets` with their labels, +`nearest_habit` / `nearest_habit_label` and `distance_hours`; the baseline frame adds +`window_label`/`window_start`/`window_end`, the self frame `scope_occurrences`. + +**Parameters.** + +- `fields` (request, default auto) — the recommender's categorical fields, as for + value novelty; steered by field overrides, bypassed by an explicit list. +- `bucket_minutes` (request) / `VESTIGO_STAT_HABIT_BUCKET_MINUTES` (server default + 60) — one of 15, 30, 60, 120, 180, 240; each divides the day so the last bucket ends + at midnight. Snapshotted as `bucket_minutes`. +- `timezone` (request) / `VESTIGO_STAT_HABIT_TIMEZONE` (server default `UTC`) — an + IANA zone name, validated against the host's zoneinfo database and a strict token + pattern before it is inlined into SQL (ClickHouse takes a zone as a constant). + Snapshotted as `timezone`. +- `VESTIGO_STAT_HABIT_MIN_BASELINE` (default 20) — reference occurrences a value + needs before it has a habit. +- `VESTIGO_STAT_HABIT_MIN_BUCKET_COUNT` (default 3) — reference occurrences a bucket + needs to be habitual. +- `VESTIGO_STAT_HABIT_MAX_CANDIDATES_PER_FIELD` (default 500) — the per-field + candidate cap; hitting it attaches a warning. + +**Allowlist key:** `(field, value)` — a value declared Normal is normal at any hour. +There is no per-bucket key on purpose: "the backup may run at 03:40 now" is a change +to the baseline definition, not a verdict on a finding. + +### Caveats + +- **The zone is the claim.** A run in `UTC` over a site that works in `Asia/Tokyo` + reports office hours as a night-time habit and a 03:00 logon as ordinary. Set + `stat_habit_timezone` for the site, or the knob per run; the finding says which zone + it read, so a reader can tell. +- **Habits need volume.** Twenty occurrences over a three-week baseline is the floor, + not a comfortable sample; a value with thirty occurrences spread thinly across the + day has a patchy habit, and the empty buckets between its busy ones will flag. The + bucket floor (3) is what keeps one stray reference occurrence from becoming a habit, + and it also means a bucket with two reference occurrences is *not* habit — read + `habit_buckets` before reading the distance. +- **Round-the-clock values have no habit to break.** A value present in every bucket + is never flagged, which is correct: an all-hours process has no hour it does not + keep. High-volume fields (a busy user, a chatty host) mostly look like this, and the + detector earns its keep on the low-volume, scheduled and role-bound values around + them. +- **Weekends and holidays are not modelled.** A weekday-only habit is still a + time-of-day habit, and a Saturday occurrence at 10:00 is inside it. Day-of-week is + a different question (`time:day_of_week` in Visualize answers it), deliberately not + folded in here. +- **Clock skew moves the hour.** Occurrences are bucketed on the corrected timestamp + (W2), so a source with a declared offset is read at its corrected wall-clock hour, + as every other time-derived query reads it. +- An off-hours occurrence is **not malicious by itself** — a manual run, a time-zone + traveller, a shifted maintenance window. Rank for triage; mark the value Normal, or + move the baseline, once it is explained. + +--- + +## 17. Value correlation (field-to-field rules that break) + +**What it answers:** "One field used to decide another — when did that stop being true?" +The AMiner `VariableCorrelationDetector` analog (roadmap D13), intra-record: two fields of +the *same event*, unlike the sequence detectors, which relate consecutive events. An +account that always logs on to its own workstation, a status code that always follows a +given action, a service name that always runs from one path — each is an **implication +rule** `A = x ⇒ B = y`, and the event where it fails is often the event that matters even +when `x` and `y` are each ordinary. Value combos (§1) find a *pair* that is rare; this +finds a *rule* that broke. + +**How it works.** For a field pair `(A, B)` the detector counts events per `(a, b)` value +pair in the reference and per window. In each direction, an antecedent value `x` with at +least `stat_correlation_min_support` (20) reference events forms a rule with its dominant +consequent `y` when `y` accounts for at least `stat_correlation_rule_confidence` (0.95) +of them — nineteen in twenty, so a user with two home workstations has no rule and a +user with one has. A rule is **broken** in a window when the share of `x` events whose +`B` is not `y` rises: a 2×2 G-test of conforming against violating events between the +reference and the window (the same log-likelihood ratio proportion shift uses), one +Benjamini–Hochberg pool over every (rule, window) test in the run, and an effect floor +of `stat_correlation_min_ratio` (2) on the violation-rate ratio — 4 % to 6 % is not a +break however significant. Only rises are reported: a rule that *appears* in a window is +a value whose share changed, which proportion shift already owns, and a rule that +tightens is not a finding. + +**Two frames.** + +| | Self (`self-rule-g-test`) | Baseline (`rule-g-test`) | +|---|---|---| +| Rules mined from | the whole scope | the baseline window | +| Window | each of `stat_self_slices` equal time slices | each suspect window | +| Reference | the other slices (leave-one-out) | the baseline window | +| Catches | a rule that breaks in one stretch of the timeline | a rule that breaks after the incident start | + +In the self frame a rule mined from the whole scope already contains its own violations, +so a rule broken everywhere is not a rule and is never tested — which is right, since +nothing in the scope says it should have held. A rule broken in one stretch keeps its +confidence over the scope and fails against the complement of that stretch. + +**Which pairs.** An explicit `fields` list is scanned as every pair among them; auto mode +takes the recommender's top `stat_correlation_auto_fields` (6) categorical fields, steered +by the timeline's [field overrides](#declaring-which-fields-a-method-reads), and pairs +them (15 pairs). More than `stat_correlation_max_pairs` (20) pairs are truncated with a +warning naming the count — name fewer fields to choose which. Each pair is one +`GROUP BY a, b` scan capped at `stat_correlation_max_rows_per_pair` (5000) +highest-volume value pairs (cap → warning: a rule whose antecedent lives in the tail is +not tested). Fields with many distinct values (identifiers) are not recommended in the +first place, and pairing two of them is the one way to make this detector expensive. + +**Score = the G statistic**, like proportion shift. The representative event is the +**first violating occurrence** in the window, and `first_seen` is its timestamp. Findings +carry `fields` (`[antecedent, consequent]`) and `values` (`[x, y]`), `confidence` and +`support` (the rule as mined), `count` and `violations` (the window), `baseline_count` and +`baseline_violations` (the reference side, whichever frame), both violation rates and +their `rate_ratio`, `top_violator` with its count — the consequent value that most often +took `y`'s place, usually the answer to "where did it go instead?" — and `g_statistic`, +`p_value`, `q_value`. The baseline frame adds `window_label`/`window_start`/`window_end`; +the self frame the slice keys (`slice_index`, `rest_slices`). + +**Parameters.** + +- `fields` (request, default auto) — two or more; every pair among them is scanned. +- `fdr_q` (request) / `VESTIGO_STAT_CORRELATION_FDR_Q` (0.05) — the BH ceiling. +- `min_ratio` (request) / `VESTIGO_STAT_CORRELATION_MIN_RATIO` (2.0) — the violation-rate + ratio floor. +- `rule_confidence` (request) / `VESTIGO_STAT_CORRELATION_RULE_CONFIDENCE` (0.95) — the + consequent share that forms a rule. Snapshotted into the persisted `DetectorRun`. +- `min_support` (request) / `VESTIGO_STAT_CORRELATION_MIN_SUPPORT` (20) — reference + events an antecedent value needs. Snapshotted. +- `VESTIGO_STAT_CORRELATION_AUTO_FIELDS` (6), `VESTIGO_STAT_CORRELATION_MAX_PAIRS` (20), + `VESTIGO_STAT_CORRELATION_MAX_ROWS_PER_PAIR` (5000) — the pair and row caps above. + +**Allowlist key:** the combo one — `(A,B)` joined with `,` and `x␟y` joined with the +combo value separator — so **Mark normal** on a broken rule suppresses that rule in both +frames and wherever it recurs, and a rule declared normal from a combo row is the same +key. + +### Caveats + +- **Confidence is not causation, and the direction is mined both ways.** `user ⇒ host` + and `host ⇒ user` are different rules with different supports; a shared workstation + has no `host ⇒ user` rule while each of its users may keep a `user ⇒ host` one. Read + `fields` to see which direction broke. +- **A rule the reference already breaks a little is judged on the rise.** The G-test + compares rates, so a 2 % baseline violation rate that becomes 40 % is a strong finding + and one that becomes 3 % is not, whatever the counts. +- **Thin antecedents have no rules.** Twenty reference events is the floor; a value seen + a dozen times cannot form a rule however consistent, and appears in no finding. Lower + `min_support` for a short baseline, knowing that a rule mined from twenty events has a + 95 % confidence that is one violation wide. +- **The self frame cannot see a rule broken from the start.** Mined over the scope, a + rule that never held is not a rule; the baseline frame, mined over a declared normal + period, is where "it held before the incident and not after" lives. +- **Identifier pairs are the cost.** Two high-cardinality fields produce a pair table the + row cap truncates; the warning says so, and the recommender keeps identifiers out of + auto mode for this reason. +- A broken rule is **not malicious by itself** — a new laptop, a reassigned service, a + migration. Rank for triage; mark the rule Normal once it is explained. + +--- + ## Dispositions and normality (implementation notes) Analyst verdicts on findings live in one `finding_dispositions` table with the @@ -2567,7 +2902,10 @@ Every successful scan (`GET .../anomalies` with the default `persist=true`, and always for `tag_anomalies`) writes a `DetectorRun` row: the request params it ran with — fields, `series_field`, thresholds, `baseline_id`, resolved windows, `windows_hash`, `dispositions_hash`, the per-source clock-skew offsets in -effect, entropy's `variant`, and for a self-frame run of the slice methods the +effect, entropy's `variant`, transition speed's `partition_field` and +`min_transitions`, time-of-day's `bucket_minutes` and `timezone`, value +correlation's `rule_confidence` and `min_support`, and for a self-frame run of +the slice methods the `slices` payload, its `slices_hash` and the resolved self settings (`self_slices`, `pause_ratio`, `min_span_seconds`, `sequence_rarity_floor`) — plus the serialized result, and returns its id as `run_id`. Rows diff --git a/docs/PROGRESS.md b/docs/PROGRESS.md index f61ba5a9..4871dbbf 100644 --- a/docs/PROGRESS.md +++ b/docs/PROGRESS.md @@ -4,8 +4,149 @@ Append-only session log — what changed and why, newest first. This file keeps sessions only; older ones live in git history, and every release is summarized in `CHANGELOG.md`. Plans belong in `ROADMAP.md`, not here. -Last updated: 2026-09-15 (v1.19.7; session 238 — review findings on the D18/D19/D11 -branch: the self frame's complement, the per-slice scan budget, three disclosure gaps). +Last updated: 2026-09-16 (1.20 in progress; sessions 239–241 — transition speed D15, +time-of-day habit D12 and value correlation D13, the 1.20 detector cluster). + +## Session 241 — 2026-09-16: value correlation (D13) + +The third and last cheap AMiner analog, `value_correlation`, in the same one-commit shape. + +**Rules, not pairs.** Value combos already find a rare `(x, y)`. What this detector adds is +the *rule*: an antecedent value `x` with at least 20 reference events whose dominant +consequent `y` covers at least 95 % of them, mined in both directions per field pair, and +tested per window with the 2×2 G-test proportion shift already has — conforming against +violating events, reference against window, one BH pool, a 2× floor on the violation-rate +ratio. The roadmap's design problem was field-pair explosion, and the answer is three caps +that each disclose themselves: auto mode pairs the recommender's top six categorical +fields (15 pairs), more than 20 pairs are truncated with a warning naming the count, and +each pair's `GROUP BY a, b` is capped at 5000 highest-volume rows with a warning that a +tail antecedent was not tested. Identifiers never enter auto mode, which is what keeps the +pair table small in the common case. + +**Only rises, only breaks.** A rule that appears in a window is a value whose share rose, +which proportion shift owns; a rule that tightens is not a finding. And the self frame +mines its rules over the whole scope, so a rule broken from the start is not a rule and is +never tested — stated in the reference section as the frame's limit rather than hidden. +The finding carries `top_violator`, the consequent value that most often took `y`'s +place, because "where did it go instead?" is the first question an analyst asks of a +broken rule and the data to answer it was already in the scan. + +**The demo needed nothing.** Every human account has one or two home workstations +(`_home_hosts`); the contractor has one, so `m.okonkwo ⇒ WKS-004` holds over 6,500 +baseline logons at confidence ~1.0 and breaks on the jump host, the file server and two +finance workstations. Two-home users never form a rule (confidence ~0.5), the +administrator's hop is two events a day against 280 and keeps their rule intact, and in +the self frame the contractor's violations sit in the last five of 24 slices. Asserted in +both frames with `fields=["attr:user", "attr:computer_name"]`. + +**The agent schema budget, again.** Adding `rule_confidence` to `run_anomaly_detector` was +absorbed by the compact docstring from session 240; the measured total is recorded in +`AGENT.md`. + +## Session 240 — 2026-09-16: time-of-day habit (D12) + +The second 1.20 detector, `time_of_day`, in the same shape as D15: one commit with its +gate entry, params model, agent knobs, run snapshot, settings, method card, finding type, +evidence figure, reference section and demo signal. + +**The timezone is the design decision, and it is a knob plus a setting, not a timeline +attribute.** The roadmap's one requirement was an explicit zone stamped into +`DetectorRun.params`. Inventing a per-timeline zone would be a data-model change +(`MODEL_REFINEMENT.md` territory) for a single detector, so the zone is `stat_habit_timezone` +(server default `UTC`, set once for a site) overridable per run, validated against zoneinfo +and a strict token pattern, then inlined into SQL — ClickHouse takes a zone as a constant, +so it cannot be bound — and recorded on the run and on every finding. `_time_fields.py` +already pins its `toHour` to `'UTC'` for the same reason this detector cannot leave the +zone implicit: the server's zone can change under a stored run. + +**Habit = buckets with enough reference mass; distance is circular.** A bucket is habitual +when it holds at least `stat_habit_min_bucket_count` (3) reference occurrences, and only a +value with at least `stat_habit_min_baseline` (20) of them has a habit at all. Score is +the circular distance in hours to the nearest habitual bucket, in multiples of the bucket +width, so 23:xx sits one hour from a 00:xx habit. The self frame is the same rule over the +whole scope: every thin bucket is, by definition, not habitual, and is scored against the +busy ones — which catches a one-off manual run and cannot catch a *repeated* new hour, and +the reference section says so rather than pretending otherwise; that is what the baseline +frame is for. + +**One scan per field, bounded by value.** Rows are per (value, bucket, window), so the +per-field cap is 500 values rather than the usual 2000: at a 15-minute resolution with +four suspect windows a value can return 480 rows. The candidate set is the highest-volume +values, selected in a subquery over the same predicate. + +**The demo's habits.** `walk()` gives every hour of the day a non-zero weight, so every +high-volume value (any human account, any home workstation) is habitual round the clock — +correct behaviour, and it means the demo signals had to come from low-volume scheduled or +role-bound streams. Baseline frame: the nightly backup program at 03:xx/04:xx against +three weeks of 02:xx (benign, deliberately the same move interval cadence already sees), +and the contractor on `JUMP-01` at 03:00 — the lateral-movement leg now visits the jump +host first, at night, and the jump host's baseline logons are the administrator's hop from +session 239, all office hours. Self frame: one manual backup run on a baseline afternoon, +twelve hours from the program's nightly slot; without it the self frame's only hit was two +runs that happened to spill past 04:00, which is the kind of RNG accident this file's +demo-coverage rule exists to replace with a fabricated signal. + +**Environment note.** This machine cannot run rootless podman, so PostgreSQL 16 and the +pinned ClickHouse 26.6.1.1193 run user-space from `~/.local/share/vestigo-devstack/`; the +`embeddings` extra was installed to match CI's `--all-extras`. Qdrant is still absent, so +five case-delete tests (`test_stories_api`, `test_rbac_api`, `test_demo_api`) 502 here and +only here; they are unrelated to this work and pass in CI. + +## Session 239 — 2026-09-16: transition speed (D15), the first 1.20 detector + +The 1.20 cluster is the three cheap AMiner analogs left on the roadmap — D15, D12, D13 — +landed one commit each on `feat/1.20`, in that order because each reuses machinery the +previous fortnight touched. This session is D15, `transition_time`. + +**What it is.** A transition is one step `a → b` between consecutive events of one stream +whose series-field values differ; the detector learns each pair's fastest transition and +reports one that undercuts it by `min_ratio`×. The stream is the source, split by a new +`partition_field` (the identifier whose moves are timed), which is the knob that makes +the question meaningful: without it a Windows log carrying every account times the gap +between *any* two consecutive logons, and two users on two hosts in the same second reads +as an impossible move. Rows without a partition value are left out rather than pooled +into one anonymous stream, for the same reason. + +**Built on `_ngram_inner_sql`, not beside it.** A transition is an n-gram of length two +with `gram[1] != gram[2]`, so the assembly, the per-source scan discipline, the +record-order tie-breaks and the window-boundary guarantee are the sequence detectors'. +The helper gained an optional `partition_col` (added to every `PARTITION BY`, `None` +keeps the sequence detectors' SQL shape) and now emits `pkey` and the arriving event's +`last_eid` — the representative event is the `b` side, the one that arrived too soon. +Adding columns to a shared subquery is what moved `CACHE_VERSION` to 5, not the new +method id, which could not collide with an old key on its own. + +**Both frames from day one (D18 is a rule now, not a migration).** Baseline: +`min-transition`, the floor is the pair's fastest baseline-window transition over at +least `min_transitions` of them, learned only for the candidate pairs the suspect scan +surfaced (Query B is bound to Query A's grams). Self: `self-min-transition`, the floor is +the pair's *next-fastest* transition anywhere in the scope — `groupArraySorted(2)` per +source, the two smallest merged across sources — which is leave-one-out by construction, +and two equally fast transitions vouch for each other. A zero floor is skipped in both +frames and counted in a warning: with second-resolution timestamps zero-length +transitions are routine, and "faster than instant" is not a claim. The self mode was +first named `loo-min-transition`; the demo coverage test's `self-`/`rare-` prefix +convention is the right one, so it is `self-min-transition` like its siblings. + +**The demo had no such signal.** The baseline frame found nothing on the demo case: the +contractor's lateral moves are minutes apart while every account's random alternation +between its own home hosts produces baseline floors of seconds. Per this file's own rule +the fabricated signal was strengthened rather than the assertion: one administrator now +has a routine jump-host hop (JUMP-01 by RDP, FILE-01 by network logon 20–90 s later, +twice a working day, kept tight so their own workstation logons rarely fall between), and +each wmic remote process creation during lateral movement now logs the contractor on to +FILE-01 two seconds later — which from JUMP-01 is the administrator's pair at a tenth of +the time. Both frames assert on it. + +**Surface.** Gate entry (same `series_distinct ≥ 2` floor as sequences: one value has no +transition; `needs_setup` in the baseline frame without a baseline), `_TransitionTimeParams` +(`series_field`, `partition_field`, `min_ratio`), the `/anomalies` and tag endpoints and +the agent tool gain `partition_field`, the persisted run snapshots `partition_field` and +`min_transitions`, three settings with registry specs. Frontend: the thirteenth method +card (`Gauge` icon, a "Stream" field knob), `TransitionTimeFinding`, a two-bar evidence +figure labelled by `reference_kind`, verdict/normalize/subject cases, and the mode sets in +`finding-frame.ts`. Docs: `ANOMALY_DETECTION.md` §15, the tool list, demo table and gate +table; README and CLAUDE.md counts. ## Session 238 — 2026-09-15: review of the D18/D19/D11 branch (PR #377) diff --git a/docs/ROADMAP.md b/docs/ROADMAP.md index a9efd763..84b63c73 100644 --- a/docs/ROADMAP.md +++ b/docs/ROADMAP.md @@ -12,11 +12,11 @@ from that day. **Milestone 10 — AI agent log investigation — is the 2.0 thru everything below**; the numbered list orders the remaining 1.x work by payoff-per-effort: 1. **A12** local transform tools — no design round, no OPSEC gate. -2. **D12** / **D13** / **D15** — cheap detectors reusing existing SQL machinery. -3. **W8** query-time field extraction — makes bespoke unstructured logs first-class. -4. **A8** external MCP toolsets — needs its own design round (policy, not plumbing). -5. **D10** / **D16** — heaviest lifts, last of the detector line. -6. **Milestone 11** external processors — P1 (the protocol doc) gates the rest; the +2. **W8** query-time field extraction — makes bespoke unstructured logs first-class. +3. **A8** external MCP toolsets — needs its own design round (policy, not plumbing). +4. **D10** / **D16** — heaviest lifts, last of the detector line (the cheap three, D15, D12 + and D13, shipped in 1.20). +5. **Milestone 11** external processors — P1 (the protocol doc) gates the rest; the Hayabusa engine half lives in `overcuriousity/hayabusa-processor`. Milestones 2–3 are polish, picked up opportunistically. Milestone 9 is additive work on @@ -141,31 +141,13 @@ burns its numbers out of that file**; the migration is done when the file is `{} Detectors adapted from [ait-aecid/logdata-anomaly-miner](https://github.com/ait-aecid/logdata-anomaly-miner), constrained to be **field-agnostic** and SQL-explainable per the forensic-reproducibility -requirement. D1–D9, `proportion_shift` and `sequence_motif` shipped — `ANOMALY_DETECTION.md` +requirement. D1–D9, D12, D13, D15, `proportion_shift` and `sequence_motif` shipped — `ANOMALY_DETECTION.md` is each detector's contract, updated in the same commit as any detector change. Every item below is incomplete until the frontend half lands with it: a plain-language method explanation, the SQL/params visible on the finding, disposition + allowlist wiring. A detector whose reasoning an analyst cannot read does not count as shipped. -**Low effort, high value:** - -- [ ] **D12 — Time-of-day habit** (`PathValueTimeIntervalDetector`): per value, learn which - times of day it occurs at in the baseline, flag suspect-window occurrences outside that - habit. Distinct from `interval_periodicity`, which measures inter-arrival gaps. Bucket by - `toHour`/`toMinute`, score by distance to the nearest occupied bucket. Needs an explicit - **timezone** decision stamped into `DetectorRun.params`, or the run is not reproducible. -- [ ] **D13 — Cross-field value correlation** (`VariableCorrelationDetector`): learn which - field-value pairs co-occur *within the same event*, flag violations. Intra-record, unlike - D10. Reuses `GROUP BY a, b` plus the G-test and Benjamini–Hochberg pool that - `proportion_shift` has. Field-pair explosion is the design problem: needs a preselection - rule and a candidate cap in the `HEAVY_SCAN_SETTINGS` family, honestly reported. -- [ ] **D15 — Impossible-speed transitions** (`MinimalTransitionTimeDetector`): learn the - minimum observed time between consecutive values of a field per identifier, flag a - suspect-window transition faster than the baseline ever saw. `find_sequence_novelty`'s - `lagInFrame` partitions already produce the pairs; this is a `min(dateDiff)` over the - same shape. Score = `1 − (observed / learned_min)`. - **High effort, high value:** - [ ] **D10 — Event correlation rules** (`EventCorrelationDetector`): mine baseline diff --git a/frontend/src/api/anomalies.ts b/frontend/src/api/anomalies.ts index f4b5a42b..205c4ff5 100644 --- a/frontend/src/api/anomalies.ts +++ b/frontend/src/api/anomalies.ts @@ -21,7 +21,7 @@ export interface LogTemplatesParams { } export interface AnomalyParams { - detector?: "value_novelty" | "value_combo" | "frequency" | "timestamp_order" | "numeric_range" | "charset" | "entropy" | "proportion_shift" | "interval_periodicity" | "sequence_novelty" | "sequence_motif" | "value_distribution_drift"; + detector?: "value_novelty" | "value_combo" | "frequency" | "timestamp_order" | "numeric_range" | "charset" | "entropy" | "proportion_shift" | "interval_periodicity" | "sequence_novelty" | "sequence_motif" | "value_distribution_drift" | "transition_time" | "time_of_day" | "value_correlation"; /** Comma-separated field tokens for value_novelty, e.g. "artifact,display_name,attr:user_agent" */ fields?: string; /** Field to group frequency series / build event sequences by */ @@ -46,6 +46,14 @@ export interface AnomalyParams { group_field?: string; /** sequence_novelty / sequence_motif only: break an n-gram when consecutive events are more than this many seconds apart. Omit for no gap bound. */ max_gap_seconds?: number; + /** transition_time only: the stream whose transitions are timed (e.g. "attr:user"). Omit for one stream per source. */ + partition_field?: string; + /** time_of_day only: width of the wall-clock buckets (15, 30, 60, 120, 180 or 240 minutes). Omit for the server default. */ + bucket_minutes?: number; + /** time_of_day only: IANA zone the clock is read in (e.g. "Europe/Berlin"). Omit for the server default. */ + timezone?: string; + /** value_correlation only: share of an antecedent's reference events one consequent value must account for to form a rule. */ + rule_confidence?: number; /** ID of a saved baseline definition (baseline range + suspect windows). Omit for self-baseline. */ baseline_id?: string; limit?: number; @@ -102,7 +110,7 @@ export const anomaliesApi = { sourceId: string, eventId: string, body: { - detector: "value_novelty" | "value_combo" | "frequency" | "timestamp_order" | "numeric_range" | "charset" | "entropy" | "proportion_shift" | "interval_periodicity" | "sequence_novelty" | "sequence_motif" | "value_distribution_drift"; + detector: "value_novelty" | "value_combo" | "frequency" | "timestamp_order" | "numeric_range" | "charset" | "entropy" | "proportion_shift" | "interval_periodicity" | "sequence_novelty" | "sequence_motif" | "value_distribution_drift" | "transition_time" | "time_of_day" | "value_correlation"; content: string; details: Record; /** diff --git a/frontend/src/api/types.ts b/frontend/src/api/types.ts index f32edd20..22da2934 100644 --- a/frontend/src/api/types.ts +++ b/frontend/src/api/types.ts @@ -786,6 +786,164 @@ export interface SequenceMotifFinding { confirmed_other_scope?: boolean; } +/** + * One value-to-value transition faster than its learned floor, from the + * transition_time detector (D15). `details.method` is `min-transition` + * (the floor is the baseline window's fastest transition of the pair, and + * `details` carries `window_*` keys) or `self-min-transition` (the floor is + * the pair's next-fastest transition anywhere in the timeline, with + * `scope_transitions` and no window keys). `reference_kind` names which. + */ +export interface TransitionTimeFinding { + type: "transition_time"; + /** Field token whose consecutive values form the transition (e.g. "attr:computer_name"). */ + field: string; + /** [from, to]. */ + values: string[]; + /** "from → to" — display form and the allowlist key. */ + value: string; + /** The stream the transition was timed within (e.g. "attr:user"); null = per source. */ + partition_field: string | null; + /** That stream key's value on the flagged transition; null when unpartitioned. */ + partition_value: string | null; + /** The fastest transition of this pair in the window (or the timeline), in seconds. */ + observed_seconds: number; + /** The floor it undercut, in seconds. */ + reference_seconds: number; + reference_kind: "baseline-min" | "next-fastest"; + /** + * reference_seconds ÷ observed_seconds; null when the observation is instant. + * A zero gap between whole-second timestamps is judged at its one-second + * bound (`details.timestamp_resolution === "second"`), so this is then a + * lower bound. + */ + speedup: number | null; + /** Transitions of this pair in the window (baseline frame) or the timeline (self). */ + count: number; + /** Transitions the floor was learned from. */ + baseline_count: number; + /** 1 − observed ÷ reference (observed at its bound, see `speedup`); 1.0 = instantaneous. */ + score: number; + /** Timestamp of the arriving event of the fastest transition. */ + first_seen: string | null; + event_id: string | null; + event: Event | null; + details: Record; + /** Present (true) only when the request passed `include_dismissed`. */ + dismissed?: boolean; + /** Present (true) when a confirmed disposition covers this finding's event. */ + confirmed?: boolean; + /** + * Present (true) when the only confirmed verdict on this event was reached + * under a *different* comparison. The claim stands, but not for this scope — + * so the row is marked rather than badged, and Confirm stays live. + */ + confirmed_other_scope?: boolean; +} + +/** + * One value occurring at a time of day it has no habit of, from the + * time_of_day detector (D12). `details.method` is `habit` (the habit was + * learned from the baseline window; `details` carries `window_*` keys) or + * `self-habit` (the value's own busy buckets across the timeline, with + * `scope_occurrences` and no window keys). The bucket resolution and the + * IANA zone the clock was read in are on every finding — the same wall-clock + * hour in two zones is two different claims. + */ +export interface TimeOfDayFinding { + type: "time_of_day"; + field: string; + value: string; + /** The offending wall-clock bucket: index, "HH:MM–HH:MM" label, width and zone. */ + bucket: number; + bucket_label: string; + bucket_minutes: number; + timezone: string; + /** Occurrences of the value in this bucket (in the suspect window, or the timeline). */ + count: number; + /** Reference occurrences the habit was learned from. */ + baseline_count: number; + /** The habitual bucket indexes, ascending; the nearest one and its label. */ + habit_buckets: number[]; + nearest_habit: number; + nearest_habit_label: string; + /** Circular distance to the nearest habitual bucket, in hours. */ + distance_hours: number; + /** = distance_hours — used for ranking. */ + score: number; + /** First occurrence in the bucket. */ + first_seen: string | null; + event_id: string | null; + event: Event | null; + details: Record; + /** Present (true) only when the request passed `include_dismissed`. */ + dismissed?: boolean; + /** Present (true) when a confirmed disposition covers this finding's event. */ + confirmed?: boolean; + /** + * Present (true) when the only confirmed verdict on this event was reached + * under a *different* comparison. The claim stands, but not for this scope — + * so the row is marked rather than badged, and Confirm stays live. + */ + confirmed_other_scope?: boolean; +} + +/** + * One implication rule `A = x ⇒ B = y` broken in a window, from the + * value_correlation detector (D13). `fields` is [antecedent, consequent] and + * `values` is [x, y], the combo shape, so the allowlist key is the combo one. + * `details.method` is `rule-g-test` (mined from the baseline window, tested + * per suspect window; `window_*` keys) or `self-rule-g-test` (mined from the + * timeline, each leave-one-out slice tested against the rest; `slice_index`, + * `rest_slices`). The `baseline_*` fields hold the reference side either way. + */ +export interface ValueCorrelationFinding { + type: "value_correlation"; + fields: string[]; + values: string[]; + /** "x ⇒ y" — display form. */ + value: string; + /** Share of the antecedent's reference events that carried y. */ + confidence: number; + /** Antecedent events the rule was mined from. */ + support: number; + /** Antecedent events in the window or slice under test. */ + count: number; + /** Of those, the ones whose consequent was not y. */ + violations: number; + /** The reference side: antecedent events and violations in the baseline window, or the other slices. */ + baseline_count: number; + baseline_violations: number; + violation_rate: number; + baseline_violation_rate: number; + /** violation_rate ÷ baseline_violation_rate (0.5-smoothed when the reference has none). */ + rate_ratio: number; + /** The most common violating consequent value in the window, and its count. */ + top_violator: string; + top_violator_count: number; + g_statistic: number; + p_value: number; + /** Benjamini–Hochberg adjusted p-value across every test in the run. */ + q_value: number; + /** = g_statistic — used for ranking. */ + score: number; + /** First violating occurrence in the window. */ + first_seen: string | null; + event_id: string | null; + event: Event | null; + details: Record; + /** Present (true) only when the request passed `include_dismissed`. */ + dismissed?: boolean; + /** Present (true) when a confirmed disposition covers this finding's event. */ + confirmed?: boolean; + /** + * Present (true) when the only confirmed verdict on this event was reached + * under a *different* comparison. The claim stands, but not for this scope — + * so the row is marked rather than badged, and Confirm stays live. + */ + confirmed_other_scope?: boolean; +} + export type AnomalyFinding = | ValueNoveltyFinding | ValueComboFinding @@ -798,7 +956,10 @@ export type AnomalyFinding = | IntervalPeriodicityFinding | SequenceNoveltyFinding | SequenceMotifFinding - | DistributionDriftFinding; + | DistributionDriftFinding + | TransitionTimeFinding + | TimeOfDayFinding + | ValueCorrelationFinding; export interface AnomaliesResponse { status: "ok" | "no_data" | "insufficient_data"; @@ -1004,7 +1165,10 @@ export interface AnomalyMarker { | "proportion_shift" | "interval_periodicity" | "sequence_novelty" - | "value_distribution_drift"; + | "value_distribution_drift" + | "transition_time" + | "time_of_day" + | "value_correlation"; /** Raw structured finding data — stored verbatim on the persisted annotation. */ rawDetails: Record; /** End of the anomalous window, for frequency findings — enables a range highlight. */ diff --git a/frontend/src/components/analysis/FindingEvidence.tsx b/frontend/src/components/analysis/FindingEvidence.tsx index 434ca45e..9b6a47f3 100644 --- a/frontend/src/components/analysis/FindingEvidence.tsx +++ b/frontend/src/components/analysis/FindingEvidence.tsx @@ -172,6 +172,57 @@ function NovelChars({ value, novel }: { value: string; novel: string[] }) { ); } +/** + * The day as a strip of buckets: the habitual ones in the reference neutral, + * the offending one in the anomaly accent, the rest empty. Every cell comes + * from the finding (`habit_buckets`, `bucket`, `bucket_minutes`); nothing is + * drawn for buckets the payload says nothing about. + */ +function DayStrip({ + bucket, + habit, + bucketMinutes, + bucketLabel, + timezone, +}: { + bucket: number; + habit: number[]; + bucketMinutes: number; + bucketLabel: string; + timezone: string; +}) { + const n = Math.max(1, Math.floor(1440 / bucketMinutes)); + const habitual = new Set(habit); + const caption = `${bucketLabel} ${timezone}; habitual buckets ${habit.length}`; + return ( +
+
+ {Array.from({ length: n }, (_, i) => ( + + ))} +
+
+ 00:00 + + {bucketLabel} {timezone} + + 24:00 +
+
+ ); +} + /** The n-gram, oldest → newest, so the *order* is what the eye reads. */ function Ngram({ values }: { values: string[] }) { return ( @@ -326,6 +377,62 @@ export function FindingEvidence({ finding }: { finding: MethodResult }) { case "sequence_novelty": case "sequence_motif": return ; + case "value_correlation": { + // The claim is a violation rate against a reference rate, both counted. + const self = findingMode(finding) === "self-rule-g-test"; + return ( + + ); + } + case "time_of_day": + return ( + + ); + case "transition_time": { + // The claim is one duration against one floor, both measured; which + // floor is in the label, since the two answer different questions. A + // zero gap between whole-second timestamps means "under a second", and + // was judged at that bound — the figure shows the bound, not an instant. + const bounded = finding.details.timestamp_resolution === "second"; + return ( + + ); + } case "timestamp_order": return (
diff --git a/frontend/src/components/analysis/FindingGroup.tsx b/frontend/src/components/analysis/FindingGroup.tsx index 534d7161..e430970f 100644 --- a/frontend/src/components/analysis/FindingGroup.tsx +++ b/frontend/src/components/analysis/FindingGroup.tsx @@ -50,9 +50,9 @@ interface Props { * Rows this group holds that no sweep method produces — today, Sigma hits in * the Named-techniques group. * - * A slot rather than a thirteenth entry in `METHODS`: that registry is pinned - * by tests to exactly the twelve ids `db/analysis_plan.py` plans for and the - * twelve param sets `api/routers/analysis.py` accepts, and Sigma is neither + * A slot rather than a sixteenth entry in `METHODS`: that registry is pinned + * by tests to exactly the fifteen ids `db/analysis_plan.py` plans for and the + * fifteen param sets `api/routers/analysis.py` accepts, and Sigma is neither * planned nor run through the findings endpoint. */ extraRows?: React.ReactNode; diff --git a/frontend/src/components/analysis/MethodKnobForm.tsx b/frontend/src/components/analysis/MethodKnobForm.tsx index 92a5c541..b636e2d4 100644 --- a/frontend/src/components/analysis/MethodKnobForm.tsx +++ b/frontend/src/components/analysis/MethodKnobForm.tsx @@ -44,7 +44,8 @@ export function buildParams( } const value = (raw[knob.param] ?? "").trim(); if (!value) continue; - out[knob.param] = knob.kind === "number" ? Number(value) : value; + const numeric = knob.kind === "number" || (knob.kind === "choice" && knob.numeric); + out[knob.param] = numeric ? Number(value) : value; } return out; } @@ -129,6 +130,16 @@ export function knobHelp(knob: MethodKnob): string { return "How many consecutive events form one sequence. Three is a good default."; case "max_gap_seconds": return "Break a sequence when consecutive events are farther apart than this."; + case "partition_field": + return "Whose moves are timed: transitions are measured within one value of this field, such as one account. Without it every source is a single stream."; + case "rule_confidence": + return "How consistently one value must imply the other before it counts as a rule. 0.95 means nineteen times in twenty."; + case "min_support": + return "How many reference events a value needs before a rule is learned from it."; + case "bucket_minutes": + return "How finely the day is cut. One hour tells a nightly job from a daytime one; fifteen minutes tells 02:15 from 02:45."; + case "timezone": + return "The zone the clock is read in, as an IANA name such as Europe/Berlin. Recorded on the run, since the same instant is a different hour elsewhere."; case "field": return "The text field to cluster into templates. Usually the message."; case "order": diff --git a/frontend/src/components/analysis/detector-registry.ts b/frontend/src/components/analysis/detector-registry.ts index c69d412a..e8aa1d71 100644 --- a/frontend/src/components/analysis/detector-registry.ts +++ b/frontend/src/components/analysis/detector-registry.ts @@ -11,8 +11,11 @@ */ import { Activity, + Clock, + Gauge, Hash, Layers, + Link2, ListOrdered, Percent, Replace, @@ -32,6 +35,9 @@ export type DetectorId = | "interval" | "drift" | "sequence" + | "transition" + | "habit" + | "correlation" | "order" | "range" | "charset" @@ -59,6 +65,7 @@ export const DETECTOR_CATEGORIES: { id: DetectorCategory; label: string }[] = [ export const DETECTORS: DetectorMeta[] = [ { id: "novelty", detector: "value_novelty", icon: Hash, label: "Rare values", hint: "Rare or first-seen field values", category: "values", scoreUnit: "surprise" }, { id: "combo", detector: "value_combo", icon: Layers, label: "Value combos", hint: "Rare combinations of fields", category: "values", scoreUnit: "surprise" }, + { id: "correlation", detector: "value_correlation", icon: Link2, label: "Value correlation", hint: "Field-to-field rules that break", category: "values", scoreUnit: "G" }, { id: "range", detector: "numeric_range", icon: Ruler, label: "Numeric range", hint: "Values outside a learned band", category: "values", scoreUnit: "× band" }, { id: "charset", detector: "charset", icon: Type, label: "Charset novelty", hint: "Never-seen characters", category: "values", scoreUnit: "surprise" }, { id: "entropy", detector: "entropy", icon: Shuffle, label: "Entropy outliers", hint: "Random or degenerate strings", category: "values", scoreUnit: "× band" }, @@ -66,8 +73,10 @@ export const DETECTORS: DetectorMeta[] = [ { id: "shift", detector: "proportion_shift", icon: Percent, label: "Proportion shift", hint: "Value shares that change between windows", category: "volume", scoreUnit: "G" }, { id: "interval", detector: "interval_periodicity", icon: Timer, label: "Interval cadence", hint: "Broken heartbeats and new beaconing", category: "volume", scoreUnit: "−log₁₀ p" }, { id: "drift", detector: "value_distribution_drift", icon: Replace, label: "Distribution drift", hint: "Whole-field value-mix changes between windows", category: "volume", scoreUnit: "−log₁₀ p" }, + { id: "habit", detector: "time_of_day", icon: Clock, label: "Time-of-day habit", hint: "Values at an hour they never keep", category: "volume", scoreUnit: "h off habit" }, { id: "order", detector: "timestamp_order", icon: Rewind, label: "Timestamp order", hint: "Timestamps running backwards", category: "volume", scoreUnit: "s skew" }, { id: "sequence", detector: "sequence_novelty", icon: ListOrdered, label: "Event sequences", hint: "Never-seen or rare event orderings (n-grams)", category: "sequences", scoreUnit: "surprise" }, + { id: "transition", detector: "transition_time", icon: Gauge, label: "Transition speed", hint: "Value-to-value moves faster than ever seen", category: "sequences", scoreUnit: "1 − obs/ref" }, ]; export const DETECTORS_BY_ID = Object.fromEntries(DETECTORS.map((d) => [d.id, d])) as Record< diff --git a/frontend/src/components/analysis/method-registry.ts b/frontend/src/components/analysis/method-registry.ts index e8ffa36f..90a49d6d 100644 --- a/frontend/src/components/analysis/method-registry.ts +++ b/frontend/src/components/analysis/method-registry.ts @@ -21,9 +21,12 @@ */ import { Activity, + Clock, FileText, + Gauge, Hash, Layers, + Link2, ListOrdered, Percent, Replace, @@ -46,6 +49,9 @@ export type MethodId = | "interval_periodicity" | "timestamp_order" | "sequence_novelty" + | "transition_time" + | "time_of_day" + | "value_correlation" | "log_template"; export type EvidenceClass = "named" | "statistical" | "exploration"; @@ -96,6 +102,11 @@ export interface MethodKnob { placeholder: string; /** `kind: "choice"` only — the options, first one the default. */ options?: { value: string; label: string }[]; + /** + * `kind: "choice"` only — the option values are numbers, sent as such. The + * API's int `Literal` for them refuses the string a `