feat(learn): learn team conventions from corrections and deliver them by retrieval - #1405
Conversation
…feedback, guidance Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- `validateCandidate` rejects unmanaged body lines, hand-edited frontmatter and duplicate ids; promote publishes the canonical serialization
- `promote` and `rollback` delete `candidate.md` so a rolled-back version is not resurrected
- `redactSecrets` `ASSIGNMENT` regex is linear; feedback and tool inputs are clipped before redaction
- lint: NFKC + `\p{Cf}`/`\p{Cc}` stripping, line separators, markdown links, URL schemes, `//host`, bare domains, emails, session ids, more shell rules; text is stored normalized
- CRLF playbooks parse; duplicate ids are re-id'd on parse
- edit budget: 1 HELPFUL per bullet, 3 EDITs and 3 REMOVEs per reflection
- distinct-feedback hash uses normalized content plus session or trajectory
- model call has a `--timeout` abort signal; `--feedback -` on a TTY fails fast
- `candidate.md`, `harmful.json` and `SKILL.md` are written via temp file and rename
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…retries, review/ci) for learn - opt-in capture (`learn.capture` / `ALTIMATE_LEARN_CAPTURE=1`) via a Bus subscription started from instance bootstrap; local-only `.altimate-code/learn/signals.jsonl` with redaction, clipping and dedupe - conservative `correctionReason` classifier; first prompt of a session is never a correction - `tool_retry` signal when the same tool fails 3+ consecutive times (one per episode) - `learn signals`, `learn signal add`, `learn reflect --session` without `--feedback`, `learn reflect --pending`; signals are consumed only after a successful reflection - auto-reflect at the end of `run` (`learn.auto_reflect` / `ALTIMATE_LEARN_AUTO=1`, model `learn.model` / `ALTIMATE_LEARN_MODEL`); stages the candidate only, never promotes, never affects the exit code - reflector prompt: user feedback is in-chat corrections (expected behaviour), prior agent behaviour is the mistake Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ished from another checkout
How coding-agent harnesses and memory frameworks bound always-on knowledge, evidence on instruction-count scaling and prompt caching, and a four-tier design (core, task-matched, archive, checks) for `altimate-code learn`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…sults Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… them Measured: a playbook holding a stale lesson next to its replacement drops held-out pass to 1/9 (the agent follows the stale one), and a weak reflector that only marks stale lessons HARMFUL lets the curator auto-remove them, losing the convention (3/9 and 4/9 vs 9/9 with a strong reflector). - anchors: lessons overlap when they share a code identifier in backticks (affix-aware, generic SQL functions ignored) - curator: an overlapping ADD/EDIT must declare `supersedes` or `coexists`; explicit supersede replaces in place and runs before fuzzy dedupe; an overlapping ADD next to a lesson marked harmful in the same reflection supersedes it implicitly - replacement step: a lesson removed for contradiction without a replacement gets one narrow model call for the corrected rule (or NONE); failures are queued in pending-replacements.jsonl and retried (5 attempts), shown in `learn show` - coexist links persist in the playbook; promote refuses undeclared overlaps - reflector prompt: rewrite a contradicted lesson; HARMFUL alone loses it End to end (Gemini agent, outdated playbook, teammate corrections, 2 iterations): weak reflector 9/9, 9/9, 9/9 held-out vs 3/9, 4/9 on the previous code; strong reflector 9/9 unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…etirement - harness: --seed-playbook for the corrections loop, REVIEWER_MAX_TURNS env - budget/: applicable/overgeneral/conflict arms, stale seed, drift scripts - research: experiments 2 and 3 results Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
From an independent review of the full `learn` diff: - cross-process lock (Flock) around candidate, harmful, history, pending and signal state; reflection re-reads and re-curates under the lock; promote commits only the candidate hash it showed - redact credential CLI arguments, connection strings, emails and SSNs in every model input (digest, feedback, existing bullets); lint rejects them in lessons - lint rejects verification-weakening lessons (skip tests/review, `commit -n`) - trajectory export includes user prompts only with `--include-prompts`, redacted - drain capture at run end and instance disposal; quarantine malformed state - consume only signals that fit in the model request - reject/rollback discard pending recovery work - external user corrections reflect without a session - accurate advice for a publish rename that collides on update - test: bootstrap and run-end leave learn untouched when learning is disabled Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… locking - stale EDIT/REMOVE/HARMFUL on a bullet changed during reflection is rejected and its signals retried - redact before clipping; cover curl --user, Authorization Basic/Bearer, scheme://user:pass@ for any scheme - verification lint catches skip/disable of tests, CI, review and hooks, respecting negation - -p/-P treated as a password only for tools where it is one (mysql, sqlcmd, bcp...) - learn lock verifies its lease before every write; 10-minute stale threshold - bootstrap does not import capture when learning is disabled Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…le rejections - redaction: attached -p for mysql/mariadb tools, line continuations, --proxy-user, key/value and YAML block credentials, NFKC + zero-width stripping - verification lint: broad targets restored, markdown stripped, clause-level negation for lists - a reflection with any stale-rejected delta keeps all its signals open - cap eviction never removes a bullet changed since the snapshot - replacement model calls run outside the lock - document the security posture (promote is the boundary; redaction and lint are guard rails) and the lock residual risk; promote prints a review notice Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ore password flags - reflection model: explicit -m > learn.model/ALTIMATE_LEARN_MODEL > the source session model > global default - verification lint: negation stops at a new clause - redact and lint mongosh/mongo -p, sshpass -p, redis-cli -a again - model test lets unrelated requests reach the real fetch Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…il corpus - each credential rule scans the text independently; matches are merged, so adding a rule cannot hide another - verification lint negates a bypass only when the negator directly precedes the verb, carried across or/nor only - guardrail-corpus.test.ts: 316 rows collected from five review rounds, asserted through redactor, lint, curate and promote validation Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… instead of parsing negation Regex negation parsing could not be right in both directions across six review rounds. Lessons mentioning skipped or disabled verification are now staged with a warning; show and promote display it, and `promote --yes` refuses flagged lessons without `--allow-flagged`. Credentials and PII are still rejected. Redaction: linear scanning (100 KB < 10 ms), tool names only at command positions and outside quotes, docker login passwords, Bearer terminology left intact. Corpus: 352 rows. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…flags A real-CLI smoke test stored `sqlcmd ... -P hunter2` unredacted in signals.jsonl because tools were only recognized at command start. Tools now count at any word occurrence unless quoted or the value of a preceding flag; Bearer in Authorization headers always redacts; flags cover validation/verification and "no need to test", CI only as the bypass object; warning scan is bounded and linear; numeric IDs are no longer redacted. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…enchmark matrix) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…, 140-char lessons - approved/candidate/versions/retired JSON snapshots under .altimate-code/learn/<name>/, atomic, locked - curator 25-bullet cap replaced by learn.max_stored (default 1000), pinned lessons never evicted - idempotent, resumable migration from a learn-managed playbook skill; malformed files quarantined - learn-managed skills excluded from auto-load; publish exports on demand - new/edited lessons <=140 chars, longer ones grandfathered; no dbt defaults; learn search Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- core tier (pinned, then net helpful) + local BM25 retrieval, frozen per session for cache safety - per-request additions appended to the user turn; file hook appends lessons for touched paths once - limits: learn.core_lessons, retrieved_lessons, budget_tokens, session_max_lessons (config + env) - shown.jsonl selection log; applied counters updated at flush, never in the prompt - projects without approved lessons are untouched Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… scheduler - enable/disable write learn.capture/auto_reflect to project config; status shows counts and limits - threshold (3 open signals), 10-min idle debounce cancelled by activity, flush-only exit - startup recovery after first idle, bounded by count/time with backoff - cross-process claims so the same signals are never reflected twice; run auto-reflect uses them Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- scope preview and confirmation (--yes for non-TTY, --dry-run sends nothing) - bounded, paged traversal of this project's root sessions; resumable, deduped extraction - batched signal writes; claimed, budgeted reflection; summary with tokens and estimated cost Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a7b80f9fb3
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
- shared config schema accepts every `learn` key (including `enabled`), so the config API no longer drops them; SDK types and OpenAPI regenerated - file lessons are delivered for tools run inside the batch tool - a first message is never classified as a correction because of a later assistant reply - published playbooks load on teammates' checkouts; the exported skill is suppressed only where a same-name local approved store exists - a successful manual reflection clears the automatic backoff - export refuses unmanaged skill directories, unexpected files and symlinks - results doc: disclose two benchmark confounds found in review (topic-switch work directory name, compression-pair lesson order), tracked in #1408 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: b808a086-9922-4555-a7ea-2e0642a31b4a) |
There was a problem hiding this comment.
All reported issues were addressed across 22 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 3e714a4919
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
- every learn write goes through one guard that rejects symlinks at the target and in any ancestor under the project root (no-follow opens); the prompt path skips delivery, CLI commands fail loudly - the kill switch also stops published learn skills from auto-loading - `learn disable` reports environment overrides that keep capture on - export recovers its own stale temps and detects edited managed exports; rollback refreshes the export - capture orders by stored message order, not id, and keeps text parts that arrive before their message; direct-feedback reflection clears backoff; shared-worktree reflections run in their originating instance Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: fdf0ff2a-92e0-47a0-8089-20905b8ea18f) |
There was a problem hiding this comment.
All reported issues were addressed across 22 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 03a0415492
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
- the kill switch also hides learn-managed skills from the skill list and the skill tool - path triggers are linted at promotion and delivery; unsafe lessons dropped - export keeps temp files it cannot prove it created; rollback restores approved.json even when the export was hand-edited, with a note - deferred startup recovery is retried when its directory opens - skill publish lookup dedupes ids repeated across pages - over-budget request notes are ignored on resume - symlink write test asserts with stat instead of fs.promises.exists Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: 9cd0e8f4-97f9-44c5-8ebb-f8475b5579c5) |
There was a problem hiding this comment.
All reported issues were addressed across 17 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
- request-note cleanup trusts only server-recorded provenance, so client text is never dropped; notes are retired immutably, which keeps the plan-layer source-structure test passing - disabled learn skills are filtered before LLM skill selection and from cached selections - cancelled reflections with open signals stay retryable; deferred recovery lookups do not consume the time budget - a skill test removes an originally unset OPENCODE_TEST_HOME Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: 1dfb67d2-ca7f-4c68-b380-b98963609772) |
- a failed signal-store read after a cancelled startup reflection refunds the recovery attempt, so open signals stay retryable - the skill selector cache is keyed by the candidate skill set, so lesson skills return after learning is re-enabled in the same process Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: 2880b9b1-d985-4faa-8baa-a7c65d1415cb) |
Resolve compaction.ts: keep client trace inheritance from main ahead of the learn replay filter for server-recorded request notes. Add the MCP listMeta and snapshot stubs that the new trace-context-loop test mock lacks. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: f3f11e80-3159-45c4-9704-219d9023afd9) |
|
🔁 @anandgupta42 — this change was promoted to the help-docs hub: https://github.com/AltimateAI/help-docs/pull/108 — please review/merge to keep the docs in sync. |
* test: [#1406] learn benchmark harness, task project and lesson pools Moved out of #1405 unchanged (from `feat/rsi-workspace-learning`), so the product PR stays reviewable. Produces the numbers in `research/rsi-workspace-learning-2026-09-30/learn-v1-results.md`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: [#1406] harden the learn benchmark harness Review fixes for research tooling; the benchmark definition (task project, seed data, `to_utc` macro and verifier checks) is unchanged so the published runs stay comparable. - safety: validate ids and paths before any deletion; fake backend binds to 127.0.0.1, debug routes need a token, authorization gaps closed - reproducibility: no machine-specific paths (dbt, CLI, worktrees); env overrides for the CLI and models; the fake backend is the default and the real SaaS needs an explicit opt-in and workspace id - scoring: runs whose agent turn did not complete no longer count; topic sessions must complete both turns in one session; analysis keeps unique directory identities - README with prerequisites (needs the `learn` features from #1405), suites and env vars; self-tests for the harness, v1bench and fake backend Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: [#1406] address the second round of harness review comments Benchmark definition (verifier checks, task project, seeds, macros) unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: [#1406] restore the published benchmark configuration; third review round Model, arm, lesson and pool defaults match the commit that produced the published runs (780168a); comparison checkouts are required env vars. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: [#1406] require an explicit workspace id for real-SaaS harness runs Removes the internal workspace id default from `run_corrections.sh` and `run_drift.sh`; SaaS mode already refuses to run without a positive `WORKSPACE_ID`, and fake-backend runs do not use it. Self-test fixtures use a neutral id. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: [#1406] combine topic-switch leak scans; safer workdir reset and baseline default - topic-switch records combine first- and second-turn leak scans - prepare_workdir deletes only a destination it created (marker + git origin) - run_corrections evaluates a fresh baseline unless one is given explicitly Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: [#1406] fix five harness review findings - publish_replace requires an explicit workspace ID (flag or WORKSPACE_ID) - bootstrap keeps training sessions from iterations 10 and later - ablation retries start from fresh learning state - topic-switch rejects unknown session IDs before touching output - leaked records make an arm incomplete and are excluded from analysis Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: [#1406] fix two regressions from the previous harness fix - ablation validates --from-loop before resetting learning state - publish_replace requires a workspace ID only for the SaaS backend Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: [#1406] reset ablation state only after preflight; scope WORKSPACE_ID to SaaS - ablation keeps prior learning state when task, user or model setup fails - publish_replace reads WORKSPACE_ID only for the SaaS backend Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Issue for this PR
No issue: new feature. Design, research and benchmark results are in
research/rsi-workspace-learning-2026-09-30/(learn-v1-plan.md,learn-v1-results.md). The benchmark harness is in #1407, which should merge after this PR.Type of change
What does this PR do?
Adds
altimate-code learn: the agent learns team conventions from corrections and gets the relevant ones back in later sessions. Local mode only; nothing is shared with a workspace unless you runlearn promote --publish.How it works
learn enable): user corrections, repeated tool failures, and review/CI feedback become redacted signals in.altimate-code/learn/.run.learn show,learn promote). Nothing reaches the agent before this step.rollback,reject,pin/unpinare available.learn.*).learn bootstrapreads past sessions;learn import-reviewsreads human review comments on merged GitHub PRs. Both show the scope and ask before sending anything.Kill switch:
learn.enabled: falseorALTIMATE_LEARN=0turns off lesson delivery, capture, auto-reflection and the TUI nudge together (default on). Explicitlearncommands still run.Behaviour change for existing users: none unless they opt in. Capture and auto-reflect default to off. Delivery runs only when a project already has an approved
approved.json; it never blocks a session (lock wait capped at ~5 s, unreadable stores skipped) and does not rewriteapproved.json(usage counters live in git-ignoredusage.json). With learning off, the TUI may show one dismissible tip after two corrections in a session (at most 3 in total;learn nudge off).Other changes:
skill publish --replaceupdates your own same-name skill published from another checkout.learnconfig:capture,auto_reflect,modeland the session-start limits are in the shared config schema; the otherlearnkeys are file/env only. The SDK is unchanged.runwaits for its own session's reflection at exit for at most 60 seconds.import-reviewskeeps only comments by owners, members and collaborators made before merge (--any-authorwidens it).Known limits (in the docs):
backend-requirements.md).How did you verify your code works?
test/altimate/learn/, including a frozen guard-rail corpus for redaction and lint, and regressions for every review finding (lock recovery, rejected-only reflections, malformed stores, redaction forms, publish adoption).main(mcp.headers ×5, HttpApi SDK, mcp status, v2 SDK error shape); CI runs it on this head.learn-v1-results.md.Screenshots / recordings
No visual UI except the one-line TUI tip.
Checklist
🤖 Generated with Claude Code
Note
Medium Risk
New prompt-injection surface (approved lessons in every session) and model calls during reflection/bootstrap, mitigated by curator lint, redaction, and mandatory promote; delivery is fail-soft but can affect agent behavior once lessons are approved.
Overview
Introduces
altimate-code learn: an opt-in pipeline that turns corrections, repeated tool failures, and review/CI feedback into human-approved lessons stored under.altimate-code/learn/, then injects relevant rules into agent sessions without a tool call.Pipeline: capture (pattern-based correction classifier + tool-retry signals, redacted) → reflect (model proposes edits; deterministic curator lints, dedupes, conflict handling via backtick anchors) → promote gate → delivery (frozen session-start “Team rules”, BM25 retrieval, per-request and per-file hooks, token budgets). Cross-process locks/claims protect concurrent reflection; bootstrap and import-reviews seed history with confirmation/dry-run.
Integration:
learnsettings added to shared config schema (packages/core); instance bootstrap starts capture/scheduler when enabled;runcan auto-reflect on exit (bounded wait). TUI nudge when capture is off;ALTIMATE_LEARNmaster switch. Docs (configure/learn.md) and GitGuardian ignores for synthetic credential tests.Reviewed by Cursor Bugbot for commit ceec300. Bugbot is set up for automated code reviews on this repo. Configure here.
Summary by CodeRabbit
New Features
learnworkflows to capture corrections and repeated tool failures, reflect on feedback, and manage lessons through review, approval, search, pinning, and rollback.skill publish --replacefor eligible name conflicts.Documentation