Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
103 changes: 103 additions & 0 deletions PLANS.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,109 @@ plan_schema_version: 2

Use this file for active, blocked, ready-for-closure, or recently completed execution work. The canonical lifecycle is the installed `engineering-workflow` planning reference.

## Active Plan: Import GPT-6 Workflow 0.9.8 And Release Marketplace 1.0.8

Status: active
Owner: root
Last Updated: 2026-09-23

### Goal

Import the immutable engineering-workflow 0.9.8 release, publish xeonvs-engineering 1.0.8, refresh local managed marketplaces, and close the release state.

### Plan Origin

direct_execution

### Requested Scope

- Update this distribution marketplace for the released GPT-6 engineering-workflow model profiles and target migration behavior.

### Requirement Traceability

| Requirement | Complete outcome | Source | Work queue | Acceptance or validation | Status |
| --- | --- | --- | --- | --- | --- |
| REQ-001 | The catalog imports exact released 0.9.8 source bytes and provenance, preserving tgrep-search 1.0.3. | User request; upstream annotated v0.9.8 | WQ-01 | Sync report, generated diff, recorded-byte verification | done |
| REQ-002 | The catalog passes local checks and publishes a reviewed 1.0.8 release with a checksummed archive. | User request; marketplace release contract | WQ-02 | Catalog validator, tests, package validators, PR merge, annotated tag, release workflow/assets | pending |
| REQ-003 | Local Codex and Claude managed installations resolve workflow 0.9.8 from this marketplace. | Prior user preference | WQ-03 | Native marketplace/plugin update and installed-version readback | pending |
| REQ-004 | Durable source and marketplace plans close truthfully after delivery. | Workflow lifecycle contract | WQ-04 | Lifecycle checks and closure readback | pending |

### Explicit Non-Goals

- Modify tgrep-search, edit vendored bytes by hand, change marketplace identity, or alter unrelated repositories.

### Constraints

- Synchronize only stable annotated upstream tags with exact source provenance.
- Keep this repository a catalog; use repository validators and final redacted security gates before push.
- Source repository remains canonical for runtime code and model profiles.

### Inputs And Sources

- https://github.com/xeonvs/codex-engineering-workflow/releases/tag/v0.9.8
- https://github.com/xeonvs/codex-engineering-workflow/pull/15
- Current catalog source policy, synchronizer, provenance, and release workflow.

### User Decisions And Answers

- 2026-09-23: Release support for new models through the marketplace, retain Terra as an explicit fallback, keep simple commands/tests model-free, and require confirmation for agent-initiated Astra escalation.

### Completed Baseline State

- [x] Catalog main is clean at 1.0.7; source 0.9.8 is merged and published as a stable annotated tag.

### Current Work Queue

- [x] WQ-01 — Synchronize and review exact upstream bundle/provenance for REQ-001. `done`
- [ ] WQ-02 — Validate, PR/merge, tag and verify 1.0.8 release for REQ-002. `in_progress`
- [ ] WQ-03 — Refresh/read back managed local installations for REQ-003. `pending`
- [ ] WQ-04 — Reconcile and close source/marketplace plans for REQ-004. `pending`

### Locked Decisions

- Publish a marketplace patch release 1.0.8; leave tgrep-search at 1.0.3.
- Use the synchronizer's generated output; retain one canonical upstream source.

### Verification

- Run catalog validator, all marketplace tests, `--verify-recorded`, external Codex/Claude plugin checks, aggregate generated diff review, security tree/history scans, and release asset readback.

### Latest Validation Results

- 2026-09-23: Source tag v0.9.8 is annotated and peels to merged commit `9a28224af9efc329ed420ddcbd25b1f0aa354565`; upstream GitHub release is public.
- 2026-09-23: Synchronizer imported exact 0.9.8 bundle from the annotated tag with SHA-256 `df06c716221a780eb92d2570a2602ad8034e219194da52f9feb4135c24f6874c`; tgrep-search remains 1.0.3. All 16 marketplace tests, catalog validator, recorded-byte verification for both plugins, Codex plugin validation, skill validation, and strict Claude plugin/marketplace validation passed. Generated diff contains only workflow bundle/provenance/version-table changes plus this plan.

### Risks And Recovery

- If sync rejects the tag or bytes, stop before publication and diagnose the exact provenance rule; never edit bundle bytes manually.
- If remote main advances, refresh and revalidate before merge/tag.
- If local installation points to an old snapshot, use native marketplace updates and active-version readback.

### Resume Point

- WQ-02: run final redacted security scans, review and commit the import, then publish the marketplace PR/tag/release.

### Plan Fidelity Check

- [x] Source import, marketplace release, local refresh, and durable closure are mapped to ordered work.
- [x] Constraints, sources, decisions, validation, recovery, and exact resume point are recorded.

### Reconciliation Check

- [ ] Final catalog, release, local installation, and plan state agree.

### Closure Gate

- [ ] All requirements and queue items are terminal with applicable evidence.

### Post-Close Delivery

- Marketplace publication and local refresh are active work under WQ-02 and WQ-03.

### Handoff Notes

- None.

## Recently Completed

- [x] 2026-09-20: Completed Marketplace Repository Audit Fix 1.0.7.
Expand Down
8 changes: 4 additions & 4 deletions PROVENANCE.json
Original file line number Diff line number Diff line change
Expand Up @@ -4,13 +4,13 @@
"bundles": [
{
"name": "engineering-workflow",
"version": "0.9.7",
"version": "0.9.8",
"source_repository": "https://github.com/xeonvs/codex-engineering-workflow",
"source_policy": "latest-tag",
"source_ref": "v0.9.7",
"source_commit": "2ceeba7e92497040b99a3bc1d302e4e26bf2a853",
"source_ref": "v0.9.8",
"source_commit": "9a28224af9efc329ed420ddcbd25b1f0aa354565",
"source_path": "plugins/engineering-workflow",
"bundle_sha256": "09dcce2e3b3cd7ab9f017e6eedce1e5b0f76d4507c6a1a1c852566ae011c1d67"
"bundle_sha256": "df06c716221a780eb92d2570a2602ad8034e219194da52f9feb4135c24f6874c"
},
{
"name": "tgrep-search",
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ reviewed local bundle held in this repository.
<!-- BEGIN GENERATED PLUGIN VERSIONS -->
| Plugin | Version | Purpose | Canonical source |
| --- | --- | --- | --- |
| [`engineering-workflow`](plugins/engineering-workflow/) | 0.9.7 | Audit, plan, migrate, validate, and maintain repository engineering workflows. | [`xeonvs/codex-engineering-workflow`](https://github.com/xeonvs/codex-engineering-workflow) |
| [`engineering-workflow`](plugins/engineering-workflow/) | 0.9.8 | Audit, plan, migrate, validate, and maintain repository engineering workflows. | [`xeonvs/codex-engineering-workflow`](https://github.com/xeonvs/codex-engineering-workflow) |
| [`tgrep-search`](plugins/tgrep-search/) | 1.0.3 | Search local source trees efficiently with the tgrep trigram index. | [`xeonvs/tgrep-search`](https://github.com/xeonvs/tgrep-search) |
<!-- END GENERATED PLUGIN VERSIONS -->

Expand Down
2 changes: 1 addition & 1 deletion plugins/engineering-workflow/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "engineering-workflow",
"version": "0.9.7",
"version": "0.9.8",
"description": "Audit, plan, migrate, validate, and maintain repository engineering workflows.",
"author": {
"name": "xeonvs",
Expand Down
2 changes: 1 addition & 1 deletion plugins/engineering-workflow/.codex-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "engineering-workflow",
"version": "0.9.7",
"version": "0.9.8",
"description": "Audit, plan, migrate, validate, and maintain repository engineering workflows.",
"author": {
"name": "xeonvs",
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@
name: engineering-workflow
description: Set up, audit, or upgrade repository workflow instructions and planning. Use for workflow changes or explicit skill refresh/update; ordinary repository work does not invoke migration.
metadata:
version: 0.9.7
version: 0.9.8
---

# Engineering Workflow
Expand Down
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
name = "workflow-explorer"
description = "Bounded read-heavy repository evidence collection"
developer_instructions = "Use only the bounded self-contained packet and accessible path scope supplied by the root. Return compact status, distilled findings, file references, checks, blockers, and accessible artifact paths. Do not paste unbounded raw output, reconstruct missing context, write shared state, or spawn child agents; report an inaccessible required input."
model = "gpt-5.6-terra"
model = "gpt-6-sol"
model_reasoning_effort = "medium"
sandbox_mode = "read-only"
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
name = "workflow-reviewer"
description = "Evidence-first review for correctness and high-risk changes"
description = "Bounded evidence-first review for ordinary changes"
developer_instructions = "Review only the bounded self-contained change packet and accessible evidence supplied by the root without modifying shared state. Return compact status, findings with severity and confidence, evidence paths, checks, blockers, assumptions, and a stopping or escalation condition; report missing required context rather than inferring it."
model = "gpt-6-astra"
model_reasoning_effort = "high"
model = "gpt-6-sol"
model_reasoning_effort = "medium"
sandbox_mode = "read-only"
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
name = "workflow-utility"
description = "Bounded read-only interpretation with a fixed result schema"
developer_instructions = "Use only the bounded self-contained packet and accessible inputs supplied by the root. Do not expand scope, write shared state, spawn child agents, or retry more than once. Return compact status, findings, checks, blockers, accessible evidence paths, needs_escalation, and the stopping condition; report an inaccessible required input instead of guessing."
model = "gpt-5.6-terra"
model = "gpt-6-luna"
model_reasoning_effort = "low"
sandbox_mode = "read-only"
Original file line number Diff line number Diff line change
Expand Up @@ -51,6 +51,7 @@ Use a tool, script, scheduler, hook, or harness layer instead of an LLM subagent
- unambiguous JSON status reads
- sorting, filtering, joining, ranking, aggregation, or deduplication
- repeating one command
- running a known shell command or test suite and reporting its exit status
- bounded retries and backoff
- deterministic stop conditions

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -2,10 +2,13 @@

Use this file only in Codex as the single canonical owner of current concrete model mappings. Keep task-shape policy in `agent_orchestration.md`; Claude Code follows native model and effort selection in `platform_compatibility.md`.

Choose a model only after the task-shape route calls for semantic work. Deterministic command execution, test runs, polling, and status aggregation use tools or scripts without creating a model worker.

## Source Snapshot

Verified against current official guidance on 2026-09-05:
Verified against current official guidance on 2026-09-23:

- `https://developers.openai.com/api/docs/models`
- `https://developers.openai.com/api/docs/guides/latest-model`
- `https://learn.chatgpt.com/docs/models`
- `https://learn.chatgpt.com/docs/agent-configuration/subagents`
Expand All @@ -16,41 +19,45 @@ Revalidate this mapping when supported Codex models or reasoning levels change.

### `utility`

- model: `gpt-5.6-terra`
- model: `gpt-6-luna`
- `model_reasoning_effort`: `low`
- `sandbox_mode`: `read-only`
- allow `minimal` or `none` only when the selected model supports it, the task needs almost no reasoning, and regression tests or evaluation preserve quality
- allow `none` only when the selected model supports it, the task needs almost no reasoning, and regression tests or evaluation preserve quality
- forbid `high`, `xhigh`, `max`, `ultra`, and pro mode by default

### `explorer`

- model: `gpt-5.6-terra`
- model: `gpt-6-sol`
- `model_reasoning_effort`: `low` or `medium`
- `sandbox_mode`: `read-only`
- use bounded path scope and distilled evidence

### `standard`

- model: `gpt-6-astra`
- model: `gpt-6-sol`
- `model_reasoning_effort`: `medium`
- use the minimum sandbox needed by the bounded work

### `review`

- model: `gpt-6-astra`
- `model_reasoning_effort`: `high`
- model: `gpt-6-sol`
- `model_reasoning_effort`: `medium`
- normally use `sandbox_mode: read-only`
- use `xhigh` only after representative evaluation shows a material quality gain
- use `high` on Sol when review complexity warrants it; consider `gpt-6-astra` with `high` effort for high-consequence correctness or security review under the confirmation rule below; use `xhigh` only after representative evaluation shows a material quality gain

### `exceptional_quality`

- keep the selected supported model; use the standard profile when no model is selected
- use `gpt-6-astra` with `high` effort for unusually difficult, high-consequence semantic work when the user requests it or confirms a proposed escalation; keep a selected supported user-pinned model
- consider `max`, `ultra`, or API pro mode only for difficult quality-first work with measurable acceptance criteria and high error cost
- compare against the cheaper baseline instead of assuming maximum reasoning wins

Before selecting Astra for a new worker or changing a saved profile from Sol/Luna to Astra on the agent's initiative, explain the concrete task risk or quality gap and obtain the user's confirmation. Do not treat a routine shell command, test run, broad task label, or available model slot as a reason to escalate. A user-selected Astra session or an explicit request for Astra already supplies that choice; do not ask again. This rule governs optional model selection, not the host's current model or an API-wide approval mechanism.

## API And Codex Boundary

When explicitly migrating a profile to Astra, preserve its effective supported reasoning effort. Replace `none` or `minimal` with `low` as the initial evaluated baseline; Astra does not support those efforts. The standard and reviewer defaults above retain `medium` and `high` respectively.
When migrating a profile to a GPT-6 model, preserve its effective supported reasoning effort unless deliberately changing the task profile after evaluation. Astra does not support `none`; use `low` as the initial evaluated baseline. Sol and Luna support `none`. If an older profile used `minimal`, start with `low` and compare representative tasks. Standard and routine review default to `medium`; higher review effort requires a task-specific reason.

Keep `gpt-5.6-terra` as an explicit compatibility fallback for a utility or explorer profile when its GPT-6 recommendation is unavailable in the active Codex client, or when representative evaluation favors the existing profile. Retain the role's `low` or `medium` effort and read-only boundary. This is a deliberate profile choice, not automatic retry or a silent replacement of a user pin. The published API token prices do not make Terra cheaper than Luna for utility work or Sol for explorer work; Codex subscription usage should be assessed in its own environment.

In the Responses API, pro is a reasoning mode selected with `reasoning.mode: "pro"`; it is not a separate model slug. Persisted reasoning and Programmatic Tool Calling are also API features.

Expand All @@ -61,6 +68,6 @@ Do not write API-only fields into Codex custom-agent TOML unless current Codex d
- Keep concrete model slugs out of `agent_orchestration.md` and other runtime references.
- Optional custom-agent templates may repeat the concrete slug they instantiate.
- Keep user-pinned supported models unless the user requests a migration.
- These recommendations and templates apply to newly requested profiles, not the user's global model selection. Verify the chosen model is exposed by the actual client before installing optional configuration. If Astra is unavailable, retain the current supported model and report the limitation; do not silently overwrite a pin or invent a fallback.
- These recommendations and templates apply to newly requested profiles, not the user's global model selection. Verify each chosen model is exposed by the actual client before installing optional configuration. If it is unavailable, retain the current supported model and report the limitation; do not silently overwrite a pin or invent a fallback.
- Treat reasoning and model selection as evaluation decisions, not status symbols.
- Preserve an existing profile when current repository evidence shows it is intentional and supported.
Loading
Loading