Skip to content

feat(learn): learn team conventions from corrections and deliver them by retrieval - #1405

Merged
anandgupta42 merged 62 commits into
mainfrom
feat/rsi-workspace-learning
Oct 5, 2026
Merged

anandgupta42 merged 62 commits into
mainfrom
feat/rsi-workspace-learning

Conversation

@anandgupta42

@anandgupta42 anandgupta42 commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Issue for this PR

No issue: new feature. Design, research and benchmark results are in research/rsi-workspace-learning-2026-09-30/ (learn-v1-plan.md, learn-v1-results.md). The benchmark harness is in #1407, which should merge after this PR.

Type of change

  • Bug fix
  • New feature
  • Refactor / code improvement
  • Documentation

What does this PR do?

Adds altimate-code learn: the agent learns team conventions from corrections and gets the relevant ones back in later sessions. Local mode only; nothing is shared with a workspace unless you run learn promote --publish.

How it works

  1. Capture (opt-in, learn enable): user corrections, repeated tool failures, and review/CI feedback become redacted signals in .altimate-code/learn/.
  2. Reflect: a model proposes lesson edits (add, edit, remove, supersede). The curator lints them, redacts secrets, dedupes, and refuses silent contradictions. Automatic after 3 signals, 10 idle minutes, or the end of run.
  3. Promote: a person reviews the diff (learn show, learn promote). Nothing reaches the agent before this step. rollback, reject, pin/unpin are available.
  4. Deliver: the harness, not a tool call, puts lessons in context. Up to 15 core and 15 retrieved lessons (BM25 over text and path triggers) form a session-start section frozen for the session, so the prompt cache holds. Per-request and per-file additions are appended to the message or tool result. All limits are configurable (learn.*).
  5. Seed: learn bootstrap reads past sessions; learn import-reviews reads human review comments on merged GitHub PRs. Both show the scope and ask before sending anything.

Kill switch: learn.enabled: false or ALTIMATE_LEARN=0 turns off lesson delivery, capture, auto-reflection and the TUI nudge together (default on). Explicit learn commands still run.

Behaviour change for existing users: none unless they opt in. Capture and auto-reflect default to off. Delivery runs only when a project already has an approved approved.json; it never blocks a session (lock wait capped at ~5 s, unreadable stores skipped) and does not rewrite approved.json (usage counters live in git-ignored usage.json). With learning off, the TUI may show one dismissible tip after two corrections in a session (at most 3 in total; learn nudge off).

Other changes:

  • skill publish --replace updates your own same-name skill published from another checkout.
  • Learn-managed skills are excluded from automatic skill loading, so lessons are not delivered twice.
  • learn config: capture, auto_reflect, model and the session-start limits are in the shared config schema; the other learn keys are file/env only. The SDK is unchanged.
  • run waits for its own session's reflection at exit for at most 60 seconds.
  • import-reviews keeps only comments by owners, members and collaborators made before merge (--any-author widens it).

Known limits (in the docs):

  • Retrieval is keyword-based.
  • A second, unrelated request in a session gets only about 60% of the lessons it needs.
  • A keyword match occasionally applies a lesson outside its scope.
  • Review import is GitHub only.
  • Team sharing per lesson needs workspace-memory backend work (backend-requirements.md).

How did you verify your code works?

  • Unit tests: 1,818 tests in test/altimate/learn/, including a frozen guard-rail corpus for redaction and lint, and regressions for every review finding (lock recovery, rejected-only reflections, malformed stores, redaction forms, publish adoption).
  • Checks on the current head: typecheck, Marker Guard and the tracker-leak check pass. The full suite was last run before the review fixes: 15,988 pass / 8 fail, the same 8 that fail on main (mcp.headers ×5, HttpApi SDK, mcp status, v2 SDK error shape); CI runs it on this head.
  • Benchmarks: 610 agent runs on a dbt project (Gemini 3.5 Flash; Claude was unavailable to the benchmark project).
    • No lessons → 4 real lessons: held-out tasks 2/9 → 9/9 (only runs whose agent turn completed count).
    • Retrieval with 50, 300 and 1,000 stored lessons: every needed lesson shown, held-out 9/9 at each size, first call constant at ~21k tokens.
    • Loading all 1,000 lessons instead costs 34% more per run.
    • Outdated lessons are retired by learning: 0/9 → 9/9.
    • Bootstrap from 8 past sessions → 9/9, for $0.20.
    • Over-application: a staging-only lesson was applied to an analysis in 1 of 6 control runs in three retrieval arms.
    • Full tables: learn-v1-results.md.
  • Smoke test: enable → correct the agent → reflect → promote → a new session receives the lesson.
  • Reviews: an independent review (15 findings) and 130 bot comments, all fixed or answered on the threads.

Screenshots / recordings

No visual UI except the one-line TUI tip.

Checklist

  • I have tested my changes locally
  • I have not included unrelated changes in this PR

🤖 Generated with Claude Code


Note

Medium Risk
New prompt-injection surface (approved lessons in every session) and model calls during reflection/bootstrap, mitigated by curator lint, redaction, and mandatory promote; delivery is fail-soft but can affect agent behavior once lessons are approved.

Overview
Introduces altimate-code learn: an opt-in pipeline that turns corrections, repeated tool failures, and review/CI feedback into human-approved lessons stored under .altimate-code/learn/, then injects relevant rules into agent sessions without a tool call.

Pipeline: capture (pattern-based correction classifier + tool-retry signals, redacted) → reflect (model proposes edits; deterministic curator lints, dedupes, conflict handling via backtick anchors) → promote gate → delivery (frozen session-start “Team rules”, BM25 retrieval, per-request and per-file hooks, token budgets). Cross-process locks/claims protect concurrent reflection; bootstrap and import-reviews seed history with confirmation/dry-run.

Integration: learn settings added to shared config schema (packages/core); instance bootstrap starts capture/scheduler when enabled; run can auto-reflect on exit (bounded wait). TUI nudge when capture is off; ALTIMATE_LEARN master switch. Docs (configure/learn.md) and GitGuardian ignores for synthetic credential tests.

Reviewed by Cursor Bugbot for commit ceec300. Bugbot is set up for automated code reviews on this repo. Configure here.

Summary by CodeRabbit

  • New Features

    • Added learn workflows to capture corrections and repeated tool failures, reflect on feedback, and manage lessons through review, approval, search, pinning, and rollback.
    • Approved lessons can be delivered at session start and in response to relevant requests or files, with configurable limits.
    • Added imports from historical sessions and GitHub reviews, and optional prompt inclusion in trajectory exports.
    • Added a TUI reminder when learning is off after repeated corrections.
    • Added skill publish --replace for eligible name conflicts.
  • Documentation

    • Added a guide to configuring and using learning workflows.

anandgupta42 and others added 30 commits September 30, 2026 21:00
…feedback, guidance

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- `validateCandidate` rejects unmanaged body lines, hand-edited frontmatter and duplicate ids; promote publishes the canonical serialization
- `promote` and `rollback` delete `candidate.md` so a rolled-back version is not resurrected
- `redactSecrets` `ASSIGNMENT` regex is linear; feedback and tool inputs are clipped before redaction
- lint: NFKC + `\p{Cf}`/`\p{Cc}` stripping, line separators, markdown links, URL schemes, `//host`, bare domains, emails, session ids, more shell rules; text is stored normalized
- CRLF playbooks parse; duplicate ids are re-id'd on parse
- edit budget: 1 HELPFUL per bullet, 3 EDITs and 3 REMOVEs per reflection
- distinct-feedback hash uses normalized content plus session or trajectory
- model call has a `--timeout` abort signal; `--feedback -` on a TTY fails fast
- `candidate.md`, `harmful.json` and `SKILL.md` are written via temp file and rename

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…retries, review/ci) for learn

- opt-in capture (`learn.capture` / `ALTIMATE_LEARN_CAPTURE=1`) via a Bus subscription started from instance bootstrap; local-only `.altimate-code/learn/signals.jsonl` with redaction, clipping and dedupe
- conservative `correctionReason` classifier; first prompt of a session is never a correction
- `tool_retry` signal when the same tool fails 3+ consecutive times (one per episode)
- `learn signals`, `learn signal add`, `learn reflect --session` without `--feedback`, `learn reflect --pending`; signals are consumed only after a successful reflection
- auto-reflect at the end of `run` (`learn.auto_reflect` / `ALTIMATE_LEARN_AUTO=1`, model `learn.model` / `ALTIMATE_LEARN_MODEL`); stages the candidate only, never promotes, never affects the exit code
- reflector prompt: user feedback is in-chat corrections (expected behaviour), prior agent behaviour is the mistake

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
How coding-agent harnesses and memory frameworks bound always-on knowledge,
evidence on instruction-count scaling and prompt caching, and a four-tier
design (core, task-matched, archive, checks) for `altimate-code learn`.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…sults

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… them

Measured: a playbook holding a stale lesson next to its replacement drops held-out
pass to 1/9 (the agent follows the stale one), and a weak reflector that only marks
stale lessons HARMFUL lets the curator auto-remove them, losing the convention
(3/9 and 4/9 vs 9/9 with a strong reflector).

- anchors: lessons overlap when they share a code identifier in backticks
  (affix-aware, generic SQL functions ignored)
- curator: an overlapping ADD/EDIT must declare `supersedes` or `coexists`;
  explicit supersede replaces in place and runs before fuzzy dedupe; an overlapping
  ADD next to a lesson marked harmful in the same reflection supersedes it implicitly
- replacement step: a lesson removed for contradiction without a replacement gets
  one narrow model call for the corrected rule (or NONE); failures are queued in
  pending-replacements.jsonl and retried (5 attempts), shown in `learn show`
- coexist links persist in the playbook; promote refuses undeclared overlaps
- reflector prompt: rewrite a contradicted lesson; HARMFUL alone loses it

End to end (Gemini agent, outdated playbook, teammate corrections, 2 iterations):
weak reflector 9/9, 9/9, 9/9 held-out vs 3/9, 4/9 on the previous code; strong
reflector 9/9 unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…etirement

- harness: --seed-playbook for the corrections loop, REVIEWER_MAX_TURNS env
- budget/: applicable/overgeneral/conflict arms, stale seed, drift scripts
- research: experiments 2 and 3 results

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
From an independent review of the full `learn` diff:
- cross-process lock (Flock) around candidate, harmful, history, pending and
  signal state; reflection re-reads and re-curates under the lock; promote
  commits only the candidate hash it showed
- redact credential CLI arguments, connection strings, emails and SSNs in every
  model input (digest, feedback, existing bullets); lint rejects them in lessons
- lint rejects verification-weakening lessons (skip tests/review, `commit -n`)
- trajectory export includes user prompts only with `--include-prompts`, redacted
- drain capture at run end and instance disposal; quarantine malformed state
- consume only signals that fit in the model request
- reject/rollback discard pending recovery work
- external user corrections reflect without a session
- accurate advice for a publish rename that collides on update
- test: bootstrap and run-end leave learn untouched when learning is disabled

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… locking

- stale EDIT/REMOVE/HARMFUL on a bullet changed during reflection is rejected and its signals retried
- redact before clipping; cover curl --user, Authorization Basic/Bearer, scheme://user:pass@ for any scheme
- verification lint catches skip/disable of tests, CI, review and hooks, respecting negation
- -p/-P treated as a password only for tools where it is one (mysql, sqlcmd, bcp...)
- learn lock verifies its lease before every write; 10-minute stale threshold
- bootstrap does not import capture when learning is disabled

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…le rejections

- redaction: attached -p for mysql/mariadb tools, line continuations, --proxy-user, key/value and YAML block credentials, NFKC + zero-width stripping
- verification lint: broad targets restored, markdown stripped, clause-level negation for lists
- a reflection with any stale-rejected delta keeps all its signals open
- cap eviction never removes a bullet changed since the snapshot
- replacement model calls run outside the lock
- document the security posture (promote is the boundary; redaction and lint are guard rails) and the lock residual risk; promote prints a review notice

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ore password flags

- reflection model: explicit -m > learn.model/ALTIMATE_LEARN_MODEL > the source session model > global default
- verification lint: negation stops at a new clause
- redact and lint mongosh/mongo -p, sshpass -p, redis-cli -a again
- model test lets unrelated requests reach the real fetch

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…il corpus

- each credential rule scans the text independently; matches are merged, so adding a rule cannot hide another
- verification lint negates a bypass only when the negator directly precedes the verb, carried across or/nor only
- guardrail-corpus.test.ts: 316 rows collected from five review rounds, asserted through redactor, lint, curate and promote validation

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… instead of parsing negation

Regex negation parsing could not be right in both directions across six review rounds.
Lessons mentioning skipped or disabled verification are now staged with a warning; show and
promote display it, and `promote --yes` refuses flagged lessons without `--allow-flagged`.
Credentials and PII are still rejected.

Redaction: linear scanning (100 KB < 10 ms), tool names only at command positions and outside
quotes, docker login passwords, Bearer terminology left intact. Corpus: 352 rows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…flags

A real-CLI smoke test stored `sqlcmd ... -P hunter2` unredacted in signals.jsonl because
tools were only recognized at command start. Tools now count at any word occurrence unless
quoted or the value of a preceding flag; Bearer in Authorization headers always redacts;
flags cover validation/verification and "no need to test", CI only as the bypass object;
warning scan is bounded and linear; numeric IDs are no longer redacted.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…enchmark matrix)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…, 140-char lessons

- approved/candidate/versions/retired JSON snapshots under .altimate-code/learn/<name>/, atomic, locked
- curator 25-bullet cap replaced by learn.max_stored (default 1000), pinned lessons never evicted
- idempotent, resumable migration from a learn-managed playbook skill; malformed files quarantined
- learn-managed skills excluded from auto-load; publish exports on demand
- new/edited lessons <=140 chars, longer ones grandfathered; no dbt defaults; learn search

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- core tier (pinned, then net helpful) + local BM25 retrieval, frozen per session for cache safety
- per-request additions appended to the user turn; file hook appends lessons for touched paths once
- limits: learn.core_lessons, retrieved_lessons, budget_tokens, session_max_lessons (config + env)
- shown.jsonl selection log; applied counters updated at flush, never in the prompt
- projects without approved lessons are untouched

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… scheduler

- enable/disable write learn.capture/auto_reflect to project config; status shows counts and limits
- threshold (3 open signals), 10-min idle debounce cancelled by activity, flush-only exit
- startup recovery after first idle, bounded by count/time with backoff
- cross-process claims so the same signals are never reflected twice; run auto-reflect uses them

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- scope preview and confirmation (--yes for non-TTY, --dry-run sends nothing)
- bounded, paged traversal of this project's root sessions; resumable, deduped extraction
- batched signal writes; claimed, budgeted reflection; summary with tokens and estimated cost

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a7b80f9fb3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread packages/core/src/v1/config/config.ts
Comment thread packages/opencode/src/session/prompt.ts Outdated
Comment thread packages/opencode/src/altimate/learn/capture.ts Outdated
Comment thread packages/opencode/src/session/system.ts Outdated
Comment thread packages/opencode/src/cli/cmd/learn.ts
Comment thread packages/opencode/src/altimate/learn/store.ts Outdated
- shared config schema accepts every `learn` key (including `enabled`), so
  the config API no longer drops them; SDK types and OpenAPI regenerated
- file lessons are delivered for tools run inside the batch tool
- a first message is never classified as a correction because of a later
  assistant reply
- published playbooks load on teammates' checkouts; the exported skill is
  suppressed only where a same-name local approved store exists
- a successful manual reflection clears the automatic backoff
- export refuses unmanaged skill directories, unexpected files and symlinks
- results doc: disclose two benchmark confounds found in review (topic-switch
  work directory name, compression-pair lesson order), tracked in #1408

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@cursor

cursor Bot commented Oct 5, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: b808a086-9922-4555-a7ea-2e0642a31b4a)

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 22 files (changes from recent commits).

Tip: Review your code locally with the cubic CLI to iterate faster.

Re-trigger cubic

Comment thread packages/opencode/src/altimate/learn/store.ts Outdated
Comment thread packages/opencode/test/tool/batch.test.ts
Comment thread packages/opencode/src/altimate/learn/store.ts Outdated
Comment thread packages/opencode/src/session/system.ts
Comment thread packages/opencode/src/cli/cmd/learn.ts Outdated
Comment thread packages/opencode/src/altimate/learn/capture.ts Outdated
Comment thread packages/core/src/v1/config/config.ts
Comment thread packages/opencode/src/altimate/learn/store.ts Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 3e714a4919

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread packages/opencode/src/altimate/learn/delivery.ts Outdated
Comment thread packages/opencode/src/altimate/learn/store.ts Outdated
Comment thread packages/opencode/src/altimate/learn/schedule.ts
Comment thread packages/opencode/src/cli/cmd/learn.ts
Comment thread packages/opencode/src/altimate/learn/capture.ts Outdated
- every learn write goes through one guard that rejects symlinks at the
  target and in any ancestor under the project root (no-follow opens); the
  prompt path skips delivery, CLI commands fail loudly
- the kill switch also stops published learn skills from auto-loading
- `learn disable` reports environment overrides that keep capture on
- export recovers its own stale temps and detects edited managed exports;
  rollback refreshes the export
- capture orders by stored message order, not id, and keeps text parts that
  arrive before their message; direct-feedback reflection clears backoff;
  shared-worktree reflections run in their originating instance

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@cursor

cursor Bot commented Oct 5, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: fdf0ff2a-92e0-47a0-8089-20905b8ea18f)

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 22 files (changes from recent commits).

Tip: Review your code locally with the cubic CLI to iterate faster.

Re-trigger cubic

Comment thread packages/opencode/src/altimate/learn/lock.ts
Comment thread packages/opencode/src/altimate/learn/delivery.ts
Comment thread packages/opencode/src/altimate/learn/safe-fs.ts
Comment thread packages/opencode/src/altimate/learn/store.ts Outdated
Comment thread packages/opencode/src/altimate/learn/store.ts Outdated
Comment thread packages/opencode/src/altimate/learn/schedule.ts
Comment thread packages/opencode/src/session/system.ts
Comment thread packages/opencode/test/altimate/learn/symlink-writes.test.ts Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 03a0415492

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread packages/opencode/src/altimate/workspace/skill-publish.ts Outdated
Comment thread packages/opencode/src/session/prompt.ts Outdated
Comment thread packages/opencode/src/altimate/learn/delivery.ts
- the kill switch also hides learn-managed skills from the skill list and
  the skill tool
- path triggers are linted at promotion and delivery; unsafe lessons dropped
- export keeps temp files it cannot prove it created; rollback restores
  approved.json even when the export was hand-edited, with a note
- deferred startup recovery is retried when its directory opens
- skill publish lookup dedupes ids repeated across pages
- over-budget request notes are ignored on resume
- symlink write test asserts with stat instead of fs.promises.exists

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@cursor

cursor Bot commented Oct 5, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: 9cd0e8f4-97f9-44c5-8ebb-f8475b5579c5)

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 17 files (changes from recent commits).

Tip: Review your code locally with the cubic CLI to iterate faster.

Re-trigger cubic

Comment thread packages/opencode/src/session/system.ts Outdated
Comment thread packages/opencode/src/session/prompt.ts Outdated
Comment thread packages/opencode/src/altimate/learn/schedule.ts
Comment thread packages/opencode/src/altimate/learn/schedule.ts Outdated
Comment thread packages/opencode/src/tool/skill.ts Outdated
Comment thread packages/opencode/src/session/prompt.ts Outdated
Comment thread packages/opencode/test/tool/skill.test.ts Outdated
- request-note cleanup trusts only server-recorded provenance, so client
  text is never dropped; notes are retired immutably, which keeps the
  plan-layer source-structure test passing
- disabled learn skills are filtered before LLM skill selection and from
  cached selections
- cancelled reflections with open signals stay retryable; deferred recovery
  lookups do not consume the time budget
- a skill test removes an originally unset OPENCODE_TEST_HOME

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@cursor

cursor Bot commented Oct 5, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: 1dfb67d2-ca7f-4c68-b380-b98963609772)

Comment thread packages/opencode/src/altimate/learn/schedule.ts Outdated
Comment thread packages/opencode/src/session/system.ts
- a failed signal-store read after a cancelled startup reflection refunds
  the recovery attempt, so open signals stay retryable
- the skill selector cache is keyed by the candidate skill set, so lesson
  skills return after learning is re-enabled in the same process

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@cursor

cursor Bot commented Oct 5, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: 2880b9b1-d985-4faa-8baa-a7c65d1415cb)

Resolve compaction.ts: keep client trace inheritance from main ahead of the
learn replay filter for server-recorded request notes. Add the MCP listMeta
and snapshot stubs that the new trace-context-loop test mock lacks.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@cursor

cursor Bot commented Oct 5, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: f3f11e80-3159-45c4-9704-219d9023afd9)

@anandgupta42
anandgupta42 merged commit cc22cda into main Oct 5, 2026
27 of 34 checks passed
@altimate-harness-bot

Copy link
Copy Markdown
Contributor

🔁 @anandgupta42 — this change was promoted to the help-docs hub: https://github.com/AltimateAI/help-docs/pull/108 — please review/merge to keep the docs in sync.

anandgupta42 added a commit that referenced this pull request Oct 5, 2026
* test: [#1406] learn benchmark harness, task project and lesson pools

Moved out of #1405 unchanged (from `feat/rsi-workspace-learning`), so the
product PR stays reviewable. Produces the numbers in
`research/rsi-workspace-learning-2026-09-30/learn-v1-results.md`.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: [#1406] harden the learn benchmark harness

Review fixes for research tooling; the benchmark definition (task project,
seed data, `to_utc` macro and verifier checks) is unchanged so the published
runs stay comparable.

- safety: validate ids and paths before any deletion; fake backend binds to
  127.0.0.1, debug routes need a token, authorization gaps closed
- reproducibility: no machine-specific paths (dbt, CLI, worktrees); env
  overrides for the CLI and models; the fake backend is the default and the
  real SaaS needs an explicit opt-in and workspace id
- scoring: runs whose agent turn did not complete no longer count; topic
  sessions must complete both turns in one session; analysis keeps unique
  directory identities
- README with prerequisites (needs the `learn` features from #1405), suites
  and env vars; self-tests for the harness, v1bench and fake backend

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: [#1406] address the second round of harness review comments

Benchmark definition (verifier checks, task project, seeds, macros) unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: [#1406] restore the published benchmark configuration; third review round

Model, arm, lesson and pool defaults match the commit that produced the
published runs (780168a); comparison checkouts are required env vars.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: [#1406] require an explicit workspace id for real-SaaS harness runs

Removes the internal workspace id default from `run_corrections.sh` and
`run_drift.sh`; SaaS mode already refuses to run without a positive
`WORKSPACE_ID`, and fake-backend runs do not use it. Self-test fixtures use a
neutral id.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: [#1406] combine topic-switch leak scans; safer workdir reset and baseline default

- topic-switch records combine first- and second-turn leak scans
- prepare_workdir deletes only a destination it created (marker + git origin)
- run_corrections evaluates a fresh baseline unless one is given explicitly

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: [#1406] fix five harness review findings

- publish_replace requires an explicit workspace ID (flag or WORKSPACE_ID)
- bootstrap keeps training sessions from iterations 10 and later
- ablation retries start from fresh learning state
- topic-switch rejects unknown session IDs before touching output
- leaked records make an arm incomplete and are excluded from analysis

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: [#1406] fix two regressions from the previous harness fix

- ablation validates --from-loop before resetting learning state
- publish_replace requires a workspace ID only for the SaaS backend

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: [#1406] reset ablation state only after preflight; scope WORKSPACE_ID to SaaS

- ablation keeps prior learning state when task, user or model setup fails
- publish_replace reads WORKSPACE_ID only for the SaaS backend

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant