diff --git a/dev-analytics/tasks/lessons.md b/dev-analytics/tasks/lessons.md index e567791..4986f5f 100644 --- a/dev-analytics/tasks/lessons.md +++ b/dev-analytics/tasks/lessons.md @@ -68,6 +68,7 @@ Pre-seeded with the mistakes most likely on this codebase, based on architecture - `[seed]` Used "average lead time" in a UI label when the metric is actually `..._MEDIAN` → label must match the metric type suffix. (×0) - Derived a new metric's bot policy by copying the nearest existing doc → derive it from the attribution predicate instead. Matching on a numeric account ID (PRs, reviews, issues) excludes bots by construction, so a `NOT LIKE '%[bot]%'` clause there is a no-op; the commit predicate's email branch is satisfiable by an automation account using the user's address, so the five commit queries filter explicitly. Fixed on `116-bot-exclusion-code-doc-alignment`, 2026-09-14. (×1) - Corrected a metric doc's attribution claim in the notes bullet while the Definition paragraph and `## Formula` block still stated the opposite → a doc states its attribution three times; grep the whole file before editing one of them. All three issue docs still say "project-level, no author filter" while `IssueRepository` filters on `creator_github_id` / `assignee_github_id`; whether the doc or the code is wrong is still open. 2026-09-14. (×1) +- Corrected a metric doc's attribution bullet while its Definition, Formula block and §8.3 controlled-change test still described the pre-attribution behaviour → a metric doc states its scope in four places, and a fix that reaches one of them leaves the file arguing with itself. The three issue docs drifted from commit 24e4a13 until 2026-09-14. When attribution changes, grep the doc for every restatement, and check whether the controlled-change test could even detect the change — the old ones passed under both readings. (×1) ### Tooling & environment diff --git a/docs/metrics/daily-issues-closed.md b/docs/metrics/daily-issues-closed.md index 2c268d3..f7e0751 100644 --- a/docs/metrics/daily-issues-closed.md +++ b/docs/metrics/daily-issues-closed.md @@ -9,7 +9,7 @@ ## Definition -The number of issues closed in the repositories and Jira projects the user is subscribed to on a given calendar day. Like `DAILY_ISSUES_CREATED`, this is a project-level metric: it measures throughput at the project scope rather than attributing closures to a specific developer. A consistently positive gap between issues closed and issues created signals a healthy, backlog-reducing project. +The number of issues assigned to the user that were closed on a given calendar day, across the repositories and Jira projects they have registered or are subscribed to. A GitHub issue counts when the user's numeric GitHub account ID is its assignee; a Jira issue counts when the user's Jira `accountId` is its assignee. Closures are credited to the assignee rather than to whoever performed the close — neither source records the closing actor — so an issue someone else closes on the user's behalf counts for the user, and an issue the user closes for someone else does not. Read against `DAILY_ISSUES_CREATED`, the two series describe different roles: issues the user raised versus issues assigned to them. Their difference is a rough balance of intake against completion for one person, not a personal backlog figure. ## Formula @@ -17,15 +17,18 @@ The number of issues closed in the repositories and Jira projects the user is su FOR each calendar day D in [from, to): daily_issues_closed(D) = COUNT(*) - FROM issues - WHERE repository_id IN :repoIds - AND closed_at IS NOT NULL - AND closed_at >= D 00:00:00 UTC - AND closed_at < D+1 00:00:00 UTC + FROM issues i + LEFT JOIN jira_project_repo_mappings rm ON rm.jira_project_id = i.jira_project_id + WHERE COALESCE(i.repository_id, rm.repository_id) IN :repoIds + AND ( (i.source = 'GITHUB' AND i.assignee_github_id = user.githubUserId) + OR (i.source = 'JIRA' AND i.assignee_account_id = user.jiraAccountId) ) + AND i.closed_at IS NOT NULL + AND i.closed_at >= D 00:00:00 UTC + AND i.closed_at < D+1 00:00:00 UTC ``` - Attribution: GitHub issues by `assignee_github_id = user.githubUserId`; Jira issues by `assignee_account_id = user.jiraAccountId`. Closures are credited to the assignee, not to whoever performed the close. See [author-attribution.md](author-attribution.md). -- Bot exclusion: implicit. Attribution matches on the numeric GitHub account ID alone, which no bot account shares with a user, so no bot filter is applied. See [author-attribution.md](author-attribution.md). +- Bot exclusion: implicit. Attribution matches on a numeric account identifier alone — the user's GitHub account ID or their Jira `accountId` — which no bot account shares with a user, so no bot filter is applied. See [author-attribution.md](author-attribution.md). - Time window: UTC calendar day boundaries applied to `closed_at`. See [timezone.md](timezone.md). - Only issues with `closed_at IS NOT NULL` are included. - Same `repoIds` scope and Jira-linking rules as `DAILY_ISSUES_CREATED`. @@ -33,12 +36,16 @@ FOR each calendar day D in [from, to): ## Edge cases - **No subscriptions**: early return; no snapshots written. -- **Jira issues without mapped repo**: not counted (no `repository_id`). -- **Reopened then re-closed issues**: the latest `closed_at` is used (upsert on collection overwrites the row). An issue re-closed on a different day is counted on the re-close date. +- **User with no linked identity**: a user who has neither a GitHub account ID nor a Jira `accountId` cannot match either branch of the predicate, so `DailyIssuesCalculator.calculate` returns early and writes nothing at all — not zeros. `DailyIssuesCalculator` produces `DAILY_ISSUES_CREATED` and `DAILY_ISSUES_CLOSED` together, so the early return suppresses both series. +- **Jira issues without mapped repo**: not counted. A Jira issue carries `repository_id = NULL` by construction; the `jira_project_repo_mappings` row, not the column, is what attaches it to a repository, and without one the coalesced repository is NULL. +- **No assignee**: a closed issue with no assignee, or one assigned to an account no user has linked, matches no user and is counted for nobody. Summing the metric over all users is therefore a lower bound on the project's closed-issue count, not equal to it. +- **Multiple GitHub assignees**: collection stores GitHub's single `assignee` field only. On an issue with several assignees the primary assignee is credited and the secondary ones are not counted. +- **Reopened then re-closed issues**: the latest `closed_at` is used (upsert on collection overwrites the row). An issue re-closed on a different day is counted on the re-close date. The same upsert overwrites the assignee, so attribution follows the *current* assignee: reassigning an issue that has already been counted moves it to the new assignee on the next collection run, retroactively. - **Issues closed via commit message** (`fixes #N`): `closed_at` is set by the GitHub API at merge time; the metric reflects that timestamp. +- **Multiple repos**: one snapshot per repo per day; read side sums across repos. ## Validation (thesis §8.3) -- **Expected range**: 0–15 issues/day for a healthy project; sustained 0 over multiple days may indicate planning or review phases. -- **Comparison baseline**: Jira board "Resolved" report or GitHub Issues list `state=closed` filtered by `closed:YYYY-MM-DD..YYYY-MM-DD`. -- **Controlled-change test**: close exactly 2 issues in a tracked project on a specific day → `daily_issues_closed` must equal 2 after the next collection run. +- **Expected range**: not yet re-derived. The figure of 0–15 issues/day recorded here described project-wide throughput, before attribution narrowed the metric to issues assigned to the user. A per-user range must be measured against real data before it is quoted. +- **Comparison baseline**: Jira board "Resolved" report filtered to the user as assignee, or GitHub Issues search `is:closed assignee: closed:YYYY-MM-DD..YYYY-MM-DD`. `` must be the user's *current* GitHub login: the metric matches on the numeric account ID, so a rename desynchronises the baseline. +- **Controlled-change test**: close exactly 2 issues **assigned to the user** in a tracked repository or Jira project on a specific day, and at least one issue in the same repository or project assigned to a different account → `daily_issues_closed` must equal 2 after the next collection run, not 3. diff --git a/docs/metrics/daily-issues-created.md b/docs/metrics/daily-issues-created.md index ee00d94..ba99d42 100644 --- a/docs/metrics/daily-issues-created.md +++ b/docs/metrics/daily-issues-created.md @@ -9,7 +9,7 @@ ## Definition -The number of issues created in the repositories and Jira projects the user is subscribed to on a given calendar day. This is a project-level metric: it counts all issues created in the tracked projects regardless of who created them, reflecting the workload intake rate of the user's team or project. It is not filtered by author identity. +The number of issues the user created on a given calendar day, across the repositories and Jira projects they have registered or are subscribed to. A GitHub issue counts when the user's numeric GitHub account ID is recorded as its creator; a Jira issue counts when the user's Jira `accountId` is recorded as its reporter. Issues filed by anyone else in the same projects are not counted. The metric measures how much work the user raises, not the intake rate of their team. ## Formula @@ -17,26 +17,31 @@ The number of issues created in the repositories and Jira projects the user is s FOR each calendar day D in [from, to): daily_issues_created(D) = COUNT(*) - FROM issues - WHERE repository_id IN :repoIds - AND created_at >= D 00:00:00 UTC - AND created_at < D+1 00:00:00 UTC + FROM issues i + LEFT JOIN jira_project_repo_mappings rm ON rm.jira_project_id = i.jira_project_id + WHERE COALESCE(i.repository_id, rm.repository_id) IN :repoIds + AND ( (i.source = 'GITHUB' AND i.creator_github_id = user.githubUserId) + OR (i.source = 'JIRA' AND i.reporter_account_id = user.jiraAccountId) ) + AND i.created_at >= D 00:00:00 UTC + AND i.created_at < D+1 00:00:00 UTC ``` - Attribution: GitHub issues by `creator_github_id = user.githubUserId`; Jira issues by `reporter_account_id = user.jiraAccountId`. Both sources are collected into the unified `issues` table and counted together. See [author-attribution.md](author-attribution.md). - Time window: UTC calendar day boundaries applied to `created_at`. See [timezone.md](timezone.md). -- Bot exclusion: implicit. Attribution matches on the numeric GitHub account ID alone, which no bot account shares with a user, so no bot filter is applied. Jira automation issues are reported under the automation rule's own `accountId` and are likewise never attributed to a user. See [author-attribution.md](author-attribution.md). +- Bot exclusion: implicit. Attribution matches on a numeric account identifier alone — the user's GitHub account ID or their Jira `accountId` — which no bot account shares with a user, so no bot filter is applied. Jira automation issues are reported under the automation rule's own `accountId` and are likewise never attributed to a user. See [author-attribution.md](author-attribution.md). - `repoIds`: the set of `git_repositories` linked to the user's subscriptions. For Jira, the link is through `jira_project_repo_mappings`. ## Edge cases -- **No subscriptions**: if `repoIds` is empty, `calcDailyIssues` returns early; no snapshots are written. -- **Jira issues without a mapped repo**: if a Jira project has no entry in `jira_project_repo_mappings`, its issues are not counted (they have no `repository_id`). -- **Duplicate issues**: each issue has a unique `(jira_project_id, external_id)` or `(repository_id, external_id)` composite key; the upsert on collection prevents duplicates. +- **No subscriptions**: if `repoIds` is empty, `DailyIssuesCalculator.calculate` returns early; no snapshots are written. +- **User with no linked identity**: a user who has neither a GitHub account ID nor a Jira `accountId` cannot match either branch of the predicate, so `DailyIssuesCalculator.calculate` returns early and writes nothing at all — not zeros. `DailyIssuesCalculator` produces `DAILY_ISSUES_CREATED` and `DAILY_ISSUES_CLOSED` together, so the early return suppresses both series. Linking at least one account in Settings is a precondition for this metric. +- **Jira issues without a mapped repo**: a Jira issue carries `repository_id = NULL` by construction, so the mapping row is what supplies the repository the `COALESCE` scopes on. A Jira project with no entry in `jira_project_repo_mappings` coalesces to NULL, matches no `repoIds` entry, and its issues are not counted. +- **Duplicate issues**: each issue has a unique `(jira_project_id, source_issue_key)` or `(data_source_id, source_issue_key)` composite key; the upsert on collection prevents duplicates. - **Reopened issues**: `created_at` is the original creation timestamp. Reopening does not increment the count. +- **Multiple repos**: one snapshot per repo per day; read side sums across repos. ## Validation (thesis §8.3) -- **Expected range**: 0–20 issues/day for an active project; higher values may indicate bulk imports or automation-generated issues. -- **Comparison baseline**: Jira board "Created" report or GitHub Issues list filtered by date and repository. -- **Controlled-change test**: create exactly 3 issues in a tracked project on a specific day → `daily_issues_created` must equal 3 after the next collection run. +- **Expected range**: not yet re-derived. The figure of 0–20 issues/day recorded here described project-wide intake, before attribution narrowed the metric to issues the user created. A per-user range must be measured against real data before it is quoted. +- **Comparison baseline**: Jira board "Created" report filtered to the user as reporter, or GitHub Issues list filtered by `author:` and date. `` must be the user's *current* GitHub login: the metric matches on the numeric account ID, so a rename desynchronises the baseline. +- **Controlled-change test**: create exactly 3 issues **as the user** in a tracked project on a specific day, and at least one issue as a different account → `daily_issues_created` must equal 3 after the next collection run, not 4. diff --git a/docs/metrics/issue-lead-time.md b/docs/metrics/issue-lead-time.md index aa4f565..5806a38 100644 --- a/docs/metrics/issue-lead-time.md +++ b/docs/metrics/issue-lead-time.md @@ -9,39 +9,47 @@ ## Definition -The median elapsed time in hours between an issue being created and being closed, for all issues in the repositories and Jira projects the user is subscribed to that were closed within the specified date window. This is a project-level metric — it reflects the throughput speed of the team or project as a whole, not the individual developer's resolution speed. A shorter median indicates faster issue resolution cycles. +The median elapsed time in hours between an issue being created and being closed, over the issues assigned to the user that were closed within the specified date window, across the repositories and Jira projects they have registered or are subscribed to. A GitHub issue counts when the user's numeric GitHub account ID is its assignee; a Jira issue counts when the user's Jira `accountId` is its assignee. Lead time is credited to the assignee, so the figure describes the issues this developer owns rather than the throughput of the whole project. It measures the issue's full lifetime from creation to close, not the time since it was assigned — an issue that sat in the backlog before being picked up carries that waiting time into the developer's median. A shorter median indicates faster issue resolution cycles. ## Formula ``` -FOR each closed issue I in subscribed repos/projects, closed in [from, to): - lead_time_hours(I) = FLOOR_TO_HOUR(I.closed_at - I.created_at) +FOR each closed issue I in subscribed repos/projects, closed in [from, to), + WHERE I is assigned to the user — + GitHub: I.assignee_github_id = user.githubUserId + Jira: I.assignee_account_id = user.jiraAccountId : + + lead_time_hours(I) = TRUNCATE_TO_HOUR(I.closed_at - I.created_at) issue_lead_time_hours_median = MEDIAN(lead_time_hours(I)) - grouped per repository_id + grouped per COALESCE(I.repository_id, mapped repository of I.jira_project_id) ``` - Attribution: GitHub issues by `assignee_github_id = user.githubUserId`; Jira issues by `assignee_account_id = user.jiraAccountId`. Lead time is credited to the assignee. See [author-attribution.md](author-attribution.md). -- Bot exclusion: implicit. Attribution matches on the numeric GitHub account ID alone, which no bot account shares with a user, so no bot filter is applied. See [author-attribution.md](author-attribution.md). +- Bot exclusion: implicit. Attribution matches on a numeric account identifier alone — the user's GitHub account ID or their Jira `accountId` — which no bot account shares with a user, so no bot filter is applied. See [author-attribution.md](author-attribution.md). - Window: `closed_at >= from` AND `closed_at < to+1`. `created_at` may fall before the window. -- Duration: `Duration.between(createdAt, closedAt).toHours()`. -- Grouping: one snapshot per repository. Jira issues are linked to a repository via `jira_project_repo_mappings`; issues without a `repository_id` are excluded. +- Duration: `Duration.between(createdAt, closedAt).toHours()` — the whole hours elapsed, remaining minutes discarded. Truncation is toward zero rather than downward, which differs from flooring only for the negative durations noted below. +- Grouping: one snapshot per repository. A Jira issue carries `repository_id = NULL` by construction and is attached to a repository through `jira_project_repo_mappings`; what is excluded is an issue whose coalesced repository is NULL, that is, a Jira project with no mapping. On the read side a request without a `repoId` reports the median of the per-repository medians — the same approximation applied across weeks below — while a request with a `repoId` reads only that repository's snapshots. - Saved as aggregate shape: `periodFrom` = the ISO week Monday, `periodTo` = that week Sunday. ## Edge cases - **No subscriptions / empty repoIds**: early return; no snapshot written. +- **User with no linked identity**: a user who has neither a GitHub account ID nor a Jira `accountId` cannot match either branch of the predicate, so `IssueLeadTimeCalculator.calculate` returns early and writes no snapshot at all — not a zero median. - **Issues without a mapped repository**: Jira issues not linked to a repo via `jira_project_repo_mappings` have `repository_id = NULL` and are excluded from this calculation. - **Reopened issues**: `closed_at` reflects the most recent closure (the upsert on collection overwrites). Lead time may be shorter than the true elapsed lifecycle if the issue was previously closed and reopened. - **`created_at > closed_at`**: should not occur; if it does (data quality issue), `Duration.toHours()` returns a negative value which the implementation saves as-is. These should appear as anomalous values in the UI. +- **No assignee**: a closed issue with no assignee, or one assigned to an account no user has linked, matches no user, so it contributes to no user's median. +- **Multiple GitHub assignees**: collection stores GitHub's single `assignee` field only. On an issue with several assignees the primary assignee is credited and the secondary ones are not. +- **Reassignment**: the upsert on collection overwrites the assignee, so attribution follows the *current* assignee. Reassigning an issue that has already been counted moves its lead time to the new assignee's median on the next collection run, retroactively. - **Median calculation**: same algorithm as `PR_LEAD_TIME_HOURS_MEDIAN` — sorted list, middle element for odd n, average of two middle elements for even n. ## Validation (thesis §8.3) -- **Expected range**: 24–336 hours (1–14 days) for a typical feature-tracking project; bug-fix projects tend toward the lower end. -- **Comparison baseline**: Jira "Time to Resolution" report or GitHub Issues closed date range — manually compute median of `closed_at - created_at`. -- **Controlled-change test**: create one issue and close it after exactly 72 hours → metric must equal 72 with only that issue in the window. +- **Expected range**: not yet re-derived. The figure of 24–336 hours recorded here described resolution speed across a whole project, before attribution narrowed the metric to issues assigned to the user. A per-user range must be measured against real data before it is quoted. +- **Comparison baseline**: Jira "Time to Resolution" report filtered to the user as assignee, or GitHub Issues `assignee: state=closed` over the date range — manually compute the median of `closed_at - created_at`. `` must be the user's *current* GitHub login: the metric matches on the numeric account ID, so a rename desynchronises the baseline. +- **Controlled-change test**: create one issue **assigned to the user** and close it after exactly 72 hours, and a second issue assigned to a different account closed after 12 hours → the metric must equal 72, not 42. ## Calculation grain and window resolution @@ -52,8 +60,10 @@ actually measured, and recomputing over a differently-framed range updates the s rather than adding a second window over the same days. The ISO week is the canonical grain because it matches the weekly summary job and -`COMMITS_PER_WEEK_AVG`, and because it is the smallest window over which a median is not -usually a median of one observation. +`COMMITS_PER_WEEK_AVG`, and because it is the smallest window that aggregates more than a +single day. It carries no guarantee of sample size: at per-user scope a weekly median is +often taken over very few closed issues — sometimes one — so a single week's figure reads +as an individual measurement rather than as a distribution. Reads do not require the requested window to match a stored one. A request resolves to every stored week it fully contains; where it contains none — a request narrower than one