diff --git a/docs/pf-context-mode-campaign.md b/docs/pf-context-mode-campaign.md new file mode 100644 index 0000000..95d825a --- /dev/null +++ b/docs/pf-context-mode-campaign.md @@ -0,0 +1,65 @@ +# Campaña PF: Context Modes (blind vs feedback) + +> **Estado: ✅ campaña ejecutada (2026-07-10, seed 7). Conclusión: feedback opt-in, blind default. Ver abajo.** Requiere GO explícito del usuario: correrla consume tokens reales de LLM (4 runs de suites `team-work`/`progression`, output-heavy). Ver plan: [`../superpowers/plans/2026-07-10-flow-context-modes.md`](superpowers/plans/2026-07-10-flow-context-modes.md) (Task P) y spec: [`../superpowers/specs/2026-07-10-rosetta-context-manifest.md`](superpowers/specs/2026-07-10-rosetta-context-manifest.md). + +## Objetivo + +Medir el efecto real de `contextMode: 'feedback'` (manifest Rosetta + bloque `` + briefings prompt-driven entre steps) frente a `'blind'` (comportamiento de hoy, sin cambios) sobre las **mismas** suites, seed y modelo — criterio de aceptación 4 del backlog ([`../superpowers/backlog/2026-07-08-flow-context-modes-blind-vs-feedback.md`](superpowers/backlog/2026-07-08-flow-context-modes-blind-vs-feedback.md), fase F5) y las tres dimensiones que ahí se piden: **puntuación del judge, coste en tokens, wall-time**. + +## Suites elegidas + +| Suite | Por qué | +|---|---| +| `team-work` | 3 agentes secuenciales con handoffs explícitos (planner → designer → developer). Es el caso donde la consciencia de flow — topología, fichero de contexto asignado, briefings prometidos — debería importar más: cada step depende literalmente de lo que el anterior promete entregarle. | +| `progression` | 2 épocas sobre el mismo VFS no destruido. Mide si el modo feedback ayuda a no regresar trabajo previo, más allá de lo que ya captura el judge de regresión (`regressionScore`) hoy. | + +`flow-assembler` queda deliberadamente fuera: no ejecuta un `AgenticFlow` (usa `assemblePipeline` directamente), así que `--context-mode` es un no-op ahí — `pf:run`/`pf:bench` avisan de esto en vez de callar (Step P1). + +## Diseño del experimento + +- **Misma seed, mismo modelo, ambos modos.** Por cada suite: una corrida en `blind` y una en `feedback`, con idéntico `--seed` e idéntico `--model`. `pf:compare` empareja por (suite, seed, modelo) — si cualquiera de los tres difiere entre las dos corridas de una suite, no habrá pareja que comparar (queda como "unpaired" en el reporte, no se descarta en silencio). +- **Modelo:** el default del PF, `mimo/mimo-v2.5-pro` (override con `--model=` si se quiere repetir la campaña sobre otro modelo — en ese caso usar la MISMA seed, y `pf:compare` seguirá emparejando correctamente porque el modelo también forma parte de la clave de agrupación). +- **Seed sugerida:** `7` (arbitraria y fija — lo único que importa es que blind y feedback de una misma suite compartan la MISMA seed). +- **Métricas comparadas** (por suite+modelo, blind → feedback, Δ%): puntuación final del judge (`finalScore`), tokens totales (`telemetry.totalTokens`), wall-time (`telemetry.latencyMs`, suma de latencias LLM por step — ver `telemetry/collector.ts`). + +## Checklist de ejecución (pendiente de GO del usuario) + +- [x] Confirmar GO del usuario — las 4 corridas consumen tokens reales de LLM (`team-work`/`progression` son suites output-heavy: 3 steps secuenciales y 2 épocas respectivamente). +- [x] `npm run pf:run -- --suite=team-work --context-mode=blind --seed=7` → finalScore 79 +- [x] `npm run pf:run -- --suite=team-work --context-mode=feedback --seed=7` → finalScore 84 +- [x] `npm run pf:run -- --suite=progression --context-mode=blind --seed=7` → finalScore 52 +- [x] `npm run pf:run -- --suite=progression --context-mode=feedback --seed=7` → finalScore 22 +- [x] `npm run pf:compare -- --seed=7 --suites=team-work,progression` +- [x] Revisar la sección "Resultados" de este documento (la escribe el comando anterior) y decidir: ¿`feedback` pasa a ser el modo recomendado por defecto, queda opt-in documentado, o se descarta por el sobrecoste de tokens sin mejora de puntuación que lo justifique? Esa conclusión también es un resultado publicable (ver spec, decisión de refinement 4). + +## Comando exacto (tras las 4 corridas de arriba) + +``` +npm run pf:compare -- --seed=7 --suites=team-work,progression +``` + +`pf:compare` imprime la tabla en stdout Y la escribe en la sección "Resultados" de abajo (reemplaza solo esa sección; el resto de este documento queda intacto — puede re-ejecutarse tantas veces como se repita la campaña). + +## Resultados + +_Generated by `pf:compare` — do not hand-edit; re-run the command to refresh._ + +Seed: 7 · Suites: team-work, progression + +| Suite | Model | Judge Score (blind → feedback, Δ%) | Tokens (blind → feedback, Δ%) | Wall-time ms (blind → feedback, Δ%) | +|---|---|---|---|---| +| progression | mimo/mimo-v2.5-pro | 52 → 22 (-57.69%) | 17252 → 15567 (-9.77%) | 53915 → 34386 (-36.22%) | +| team-work | mimo/mimo-v2.5-pro | 79 → 84 (+6.33%) | 91366 → 117811 (+28.94%) | 240019 → 223204 (-7.01%) | + +## Conclusión (2026-07-10, seed 7, n=1 por celda) + +**Feedback NO pasa a default. Queda opt-in documentado; `blind` sigue siendo el modo por defecto** — respuesta al último ítem del checklist y a la decisión de refinement 4 del backlog (F5). + +El resultado es **mixto y depende de la topología del flow**: + +- **`team-work` (3 agentes secuenciales con handoffs): feedback ayuda** — judge 79→84 (+6.33%), a costa de +28.94% tokens. Es el caso que la hipótesis predecía: cuando cada step depende literalmente de lo que el anterior le entrega, la consciencia de flow + los briefings prometidos mejoran el resultado. El sobrecoste de tokens es real y hay que pesarlo. +- **`progression` (2 épocas sobre el mismo VFS): feedback perjudica, y mucho** — judge 52→22 (−57.69%), con tokens y wall-time incluso menores. En un flujo de continuidad iterativa (una sola línea de trabajo que progresa), el bloque `` + los ficheros de briefing parecen distraer al modelo de la tarea real en vez de coordinar handoffs que aquí no existen. + +**Lectura:** feedback es una herramienta para flows con handoffs explícitos entre roles distintos, no un interruptor universalmente bueno. Activarlo por defecto habría degradado toda una clase de flows. Esto valida la decisión de diseño (blind default, feedback opt-in a nivel flow) con datos, y es exactamente el tipo de resultado que el backlog anticipó como publicable aunque fuese negativo. + +**Salvedad estadística:** n=1 por celda (una seed). Los deltas son grandes (sobre todo el −58% de progression, difícilmente ruido), pero para una recomendación de producto firme conviene repetir con ≥3 seeds — `pf:compare` ya empareja por seed, así que es re-ejecutar la campaña con otras seeds y promediar. Queda como follow-up, no bloquea la conclusión cualitativa. diff --git a/package.json b/package.json index bb38157..2611696 100644 --- a/package.json +++ b/package.json @@ -21,6 +21,7 @@ "pf:run": "tsx src/main/performance-frontier/cli.ts", "pf:bench": "tsx src/main/performance-frontier/bench/cli.ts", "pf:arena": "tsx src/main/performance-frontier/arena/arena-runner.ts", + "pf:compare": "tsx src/main/performance-frontier/compare/cli.ts", "fluxor:baseline": "node -e \"console.log('Run baselines via Fluxor IDE UI — Flows tab → Initialize Baselines')\"", "build:bridge": "vite build --config vite.bridge.config.ts", "fluxor:init": "node -e \"const fs=require('fs'),p=require('path');const d=p.join(process.cwd(),'fluxor');const legacy=p.join(process.cwd(),'heliox');if(fs.existsSync(legacy)&&!fs.existsSync(d)){fs.renameSync(legacy,d);console.log('Legacy compat: migrated ./heliox/ to ./fluxor/');}if(!fs.existsSync(d))fs.mkdirSync(d,{recursive:true});['flows.json','metrics.config.json'].forEach(f=>{const fp=p.join(d,f);if(!fs.existsSync(fp))fs.copyFileSync(p.join(__dirname,'assets','fluxor-templates',f),fp)});console.log('Fluxor config initialized in ./fluxor/')\"", diff --git a/sdk/conformance/README.md b/sdk/conformance/README.md index a75349b..7a2da0c 100644 --- a/sdk/conformance/README.md +++ b/sdk/conformance/README.md @@ -25,6 +25,7 @@ conformance suite against these shared fixtures. The contract is: | Loop execution parity (expansion algorithm → golden loop trace) | TS: `semantic-parity.test.ts`; Java: `CrossRuntimeConformanceTest.loopExecutionMatchesGoldenTrace`; Python: `test_loop_execution_matches_golden_trace` | | `contract`/`model` carried opaquely round-trip | TS: `fluxor-flow.test.ts`; Java: `CrossRuntimeConformanceTest.contractAndModelAreCarried`; Python: `test_contract_and_model_are_carried` | | Legacy/absent `format` tag imports with one deprecation warning — `"fluxor-flow"` is silent, absent/null is assumed `"heliox-flow"`, any other value is reported as seen; never blocks parsing | TS: `fluxor-flow.test.ts` suite 7 ("legacy format tag compat"); Java: `CrossRuntimeConformanceTest.legacyFormatImportsWithDeprecationWarning` / `currentFormatFlowsNeverWarn` / `absentFormatWarnsAssumingLegacyHelioxFlow` / `unknownFormatWarnsMentioningSeenValue` / `legacyWarningIsEmittedOncePerDistinctValue`; Python: `test_legacy_format_imports_with_deprecation_warning` / `test_current_format_flows_never_warn` / `test_absent_format_warns_assuming_legacy_heliox_flow` / `test_unknown_format_warns_mentioning_seen_value` | +| `contextMode: "feedback"` imports on Java/Python with ONE downgrade-to-blind warning (TS executes it natively); the value is preserved verbatim; `"blind"`/absent/any other value is silent; never blocks parsing | TS: `fluxor-flow.test.ts` suite 8 ("contextMode export/import round-trip") + `semantic-parity.test.ts` ("feedback-mode contextMode is additive"); Java: `CrossRuntimeConformanceTest.feedbackContextModeDowngradesToBlindWithWarning` / `blindContextModeNeverWarns` / `absentContextModeNeverWarns` / `unknownContextModeValueNeverWarns` / `feedbackDowngradeWarningIsEmittedOncePerProcess`; Python: `test_feedback_context_mode_downgrades_to_blind_with_warning` / `test_blind_context_mode_never_warns` / `test_absent_context_mode_never_warns` / `test_unknown_context_mode_value_never_warns` / `test_feedback_downgrade_warning_is_emitted_once` — see [Context modes](#context-modes--rosetta-downgrade-on-javapython-2026-07-10) | ### Running all three suites @@ -195,3 +196,67 @@ predates the tag, and the TS round-trip test compensates by spreading suite 5, "importFlow(conformance-contract.flow.json) → exportFlow deep-equals golden-contract-roundtrip.json"). Regenerate the golden with the tag if that workaround is ever retired. + +## Context modes — Rosetta downgrade on Java/Python (2026-07-10) + +The wire format carries an optional top-level `contextMode` string (`AgenticFlow.contextMode` +in `src/types/harness.ts`; spec: +`docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md`). It selects how a run threads +context between steps: **blind** (the default — steps only see their declared upstream +outputs, exactly as before the field existed) or **feedback** (the TS harness materializes a +`.fluxor/run-context//` directory with a Rosetta manifest and per-step context files, +and injects a deterministic `` block). The exporter omits the key entirely for +blind/absent flows and stamps only the literal `"feedback"`, so blind exports stay +byte-identical to every pre-Rosetta export. + +### Runtime × mode support matrix + +| Runtime | `contextMode` absent / `"blind"` | `contextMode: "feedback"` | +|---|---|---| +| **TypeScript** (IDE / serve) | Native (byte-identical to pre-Rosetta behaviour) | **Native** — run-context genesis, manifest, ``, briefing guardrail | +| **Java** (`sdk/java`) | Native (blind is the only implemented mode) | **Downgrade to blind** — one warning on `System.err`, then executes blind | +| **Python** (`sdk/python`) | Native (blind is the only implemented mode) | **Downgrade to blind** — one `DeprecationWarning`, then executes blind | + +### The downgrade contract (spec decision 2 — frozen) + +Feedback-mode parity in Java/Python is a **v2** concern (the PF measurement campaign runs on +the TS runtime); in v1 both SDKs perform an explicit, documented downgrade at import: + +1. **`contextMode === "feedback"`** → the import succeeds unchanged and emits **one** warning + with the exact wording: + + > feedback mode is not supported by this runtime yet; downgrading to blind + + Execution then proceeds in blind mode — the only mode these executors implement (they have + no run-context manifest/briefing machinery; neither executor ever reads `contextMode`, so + blind execution is structural, not a branch). +2. **`"blind"`, absent/null, or any other value** → total silence: there is nothing to + downgrade. The SDKs do not validate the enum (only the TS IDE authoring surface does); an + unrecognized value is carried like any other opaque field. + +In every case `contextMode` is **preserved verbatim** on the parsed definition +(`FlowDefinition.contextMode()` in Java, `FlowDefinition.context_mode` in Python — `null`/ +`None` when absent). The downgrade changes runtime behaviour, never the recorded value, so +the definition round-trips without loss. + +Warning channels and once-semantics mirror the legacy-format shim exactly: Java prints one +line to `System.err`, deduped by a once-per-process flag (test reset hook +`resetContextModeWarningsForTests`); Python raises one `DeprecationWarning` via +`warnings.warn`, dedup delegated to the stdlib warnings filter. Unlike the format shim's +message, the downgrade wording above is itself part of the frozen contract — it may **not** +vary per runtime. + +- **`conformance-feedback-mode.flow.json`** — a minimal 1-step flow carrying + `"format": "fluxor-flow"` and `"contextMode": "feedback"` (rule 1). Proof: Java + `CrossRuntimeConformanceTest.feedbackContextModeDowngradesToBlindWithWarning` / + `feedbackDowngradeWarningIsEmittedOncePerProcess`; Python + `test_feedback_context_mode_downgrades_to_blind_with_warning` / + `test_feedback_downgrade_warning_is_emitted_once`. +- The **silent cases** (rule 2) are proven with inline JSON and the existing chain fixture: + Java `blindContextModeNeverWarns` / `absentContextModeNeverWarns` / + `unknownContextModeValueNeverWarns`; Python `test_blind_context_mode_never_warns` / + `test_absent_context_mode_never_warns` / `test_unknown_context_mode_value_never_warns`. +- The **TS side** (native execution, no downgrade) is proven by `fluxor-flow.test.ts` suite 8 + ("contextMode export/import round-trip" — omit-when-blind, stamp-only-`"feedback"`) and + `semantic-parity.test.ts` ("feedback-mode contextMode is additive — does not perturb the + golden trace", which also materializes a real run-context manifest). diff --git a/sdk/conformance/conformance-feedback-mode.flow.json b/sdk/conformance/conformance-feedback-mode.flow.json new file mode 100644 index 0000000..a385978 --- /dev/null +++ b/sdk/conformance/conformance-feedback-mode.flow.json @@ -0,0 +1,17 @@ +{ + "version": "1", + "format": "fluxor-flow", + "id": "conformance-feedback-mode", + "name": "Conformance Feedback Mode", + "contextMode": "feedback", + "rootStepId": "step-a", + "steps": [ + { + "id": "step-a", + "type": "llm_call", + "prompt": "Step A: a flow exported with contextMode 'feedback' must still import cleanly, downgrading to blind execution with one warning on this runtime.", + "dependsOn": [], + "tools": [] + } + ] +} diff --git a/sdk/java/src/main/java/io/fluxor/sdk/flow/FlowDefinition.java b/sdk/java/src/main/java/io/fluxor/sdk/flow/FlowDefinition.java index 0394e13..b808f87 100644 --- a/sdk/java/src/main/java/io/fluxor/sdk/flow/FlowDefinition.java +++ b/sdk/java/src/main/java/io/fluxor/sdk/flow/FlowDefinition.java @@ -12,14 +12,34 @@ * body re-runs, up to {@code maxIterations} total passes. The forward graph stays acyclic; * loops are expanded into a per-iteration instance graph at execution time by * {@link io.fluxor.sdk.engine.FlowExecutor#executeAllTextTrace}. + * + *

{@code contextMode} is the Rosetta context mode ({@code AgenticFlow.contextMode} in + * {@code src/types/harness.ts}; spec: + * {@code docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md}), carried + * opaquely like {@link StepConfig#contract()}/{@link StepConfig#model()}: this + * runtime never branches on it. {@code null} when the flow declares none (blind default). + * {@code "feedback"} is NOT supported by this runtime in v1 — {@link FlowImport} emits a + * one-time downgrade warning at import and execution proceeds in blind mode (the only mode + * this executor implements); the recorded value stays {@code "feedback"} so the definition + * round-trips without loss (spec decision 2: downgrade, not parity). */ -public record FlowDefinition(String id, List steps, List loops) { +public record FlowDefinition(String id, List steps, List loops, + String contextMode) { public FlowDefinition { steps = List.copyOf(steps); loops = loops == null ? List.of() : List.copyOf(loops); } + /** + * Compatibility constructor for callers built against the pre-Rosetta 3-arg shape — + * defaults {@code contextMode} to {@code null} (carry-opaque field; absent means the flow + * declares none, i.e. blind) so existing call sites compile unchanged. + */ + public FlowDefinition(String id, List steps, List loops) { + this(id, steps, loops, null); + } + /** * Compatibility constructor for callers built against the pre-loop 2-arg shape — defaults * to no loops so existing call sites compile unchanged. diff --git a/sdk/java/src/main/java/io/fluxor/sdk/flow/FlowImport.java b/sdk/java/src/main/java/io/fluxor/sdk/flow/FlowImport.java index 2862134..85fe008 100644 --- a/sdk/java/src/main/java/io/fluxor/sdk/flow/FlowImport.java +++ b/sdk/java/src/main/java/io/fluxor/sdk/flow/FlowImport.java @@ -12,6 +12,7 @@ import java.util.Map; import java.util.Set; import java.util.concurrent.ConcurrentHashMap; +import java.util.concurrent.atomic.AtomicBoolean; /** * Parses the canonical Fluxor flow JSON (exported by the TypeScript IDE) into a @@ -34,6 +35,9 @@ *

  • {@code contract} → {@link StepConfig#contract()} (carried opaquely, nullable)
  • *
  • {@code model} → {@link StepConfig#model()} (nullable)
  • *
  • {@code loops} → {@link FlowDefinition#loops()} (empty list when absent)
  • + *
  • {@code contextMode} → {@link FlowDefinition#contextMode()} (carried opaquely, + * nullable; {@code "feedback"} triggers a one-time downgrade-to-blind warning — + * see {@link #checkContextModeDowngrade})
  • * */ public final class FlowImport { @@ -76,6 +80,39 @@ static void resetLegacyFormatWarningsForTests() { WARNED_LEGACY_FORMATS.clear(); } + /** + * Rosetta context mode this runtime cannot honour in v1 (spec decision 2: + * {@code docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md}) — importing a + * flow declaring it downgrades execution to blind with a one-time warning; see + * {@link #checkContextModeDowngrade}. Any other value (including the explicit + * {@code "blind"}, absence, or an unrecognized string) is silent: there is nothing to + * downgrade. + */ + private static final String FEEDBACK_CONTEXT_MODE = "feedback"; + + /** + * Exact downgrade warning wording — a frozen cross-runtime contract (Java prints it to + * {@code System.err}; Python raises it as a {@code DeprecationWarning}). Trigger and + * wording may not vary per runtime. + */ + private static final String FEEDBACK_DOWNGRADE_WARNING = + "feedback mode is not supported by this runtime yet; downgrading to blind"; + + /** + * Whether the feedback-downgrade warning has fired — gives it once-per-process semantics, + * the same idiom as {@link #WARNED_LEGACY_FORMATS} (a single flag rather than a seen-value + * set because exactly one value, {@code "feedback"}, ever triggers it). + */ + private static final AtomicBoolean WARNED_FEEDBACK_DOWNGRADE = new AtomicBoolean(false); + + /** + * Test hook — clears the once-per-process downgrade-warning dedup state (the context-mode + * counterpart of {@link #resetLegacyFormatWarningsForTests}). Package-private: tests only. + */ + static void resetContextModeWarningsForTests() { + WARNED_FEEDBACK_DOWNGRADE.set(false); + } + private FlowImport() { } @@ -110,6 +147,11 @@ public static FlowDefinition fromCanonicalJson(String json) { // Legacy compat: advisory-only, never blocks parsing. checkLegacyFormat(root); + // Rosetta context mode: carried opaquely; 'feedback' downgrades execution to blind + // with a one-time warning (advisory-only, never blocks parsing). + String contextMode = optionalText(root, "contextMode"); + checkContextModeDowngrade(contextMode); + // Flow-level id. JsonNode idNode = root.get("id"); if (idNode == null || idNode.isNull()) { @@ -131,7 +173,7 @@ public static FlowDefinition fromCanonicalJson(String json) { List loops = parseLoops(root.get("loops"), flowId); - return new FlowDefinition(flowId, steps, loops); + return new FlowDefinition(flowId, steps, loops, contextMode); } /** @@ -183,6 +225,33 @@ private static void checkLegacyFormat(JsonNode root) { } } + /** + * Rosetta context-mode downgrade (spec decision 2, + * {@code docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md}): warns (once per + * process, to {@code System.err} — the same transport and once-semantics as + * {@link #checkLegacyFormat}) when importing a flow declaring + * {@code contextMode: "feedback"}, which this runtime cannot honour in v1 (no run-context + * manifest/briefing machinery — that lives in the TS harness engine). The frozen + * cross-runtime contract: + *
      + *
    1. {@code "feedback"} — one warning ({@link #FEEDBACK_DOWNGRADE_WARNING}); execution + * proceeds in blind mode, the only mode this executor implements.
    2. + *
    3. {@code "blind"}, absent/null, or any other value — silence: there is nothing to + * downgrade.
    4. + *
    + * Never throws, and never mutates the value: {@code contextMode} is preserved verbatim on + * the parsed {@link FlowDefinition} (the downgrade changes runtime behaviour, not the + * recorded definition) so a re-serialisation round-trips without loss. + */ + private static void checkContextModeDowngrade(String contextMode) { + if (!FEEDBACK_CONTEXT_MODE.equals(contextMode)) { + return; + } + if (WARNED_FEEDBACK_DOWNGRADE.compareAndSet(false, true)) { + System.err.println(FEEDBACK_DOWNGRADE_WARNING); + } + } + private static StepConfig parseStep(JsonNode node, String flowId) { String id = requireText(node, "id", flowId); String prompt = requireText(node, "prompt", flowId + "/" + id); diff --git a/sdk/java/src/test/java/io/fluxor/sdk/flow/CrossRuntimeConformanceTest.java b/sdk/java/src/test/java/io/fluxor/sdk/flow/CrossRuntimeConformanceTest.java index fa8124d..40b1820 100644 --- a/sdk/java/src/test/java/io/fluxor/sdk/flow/CrossRuntimeConformanceTest.java +++ b/sdk/java/src/test/java/io/fluxor/sdk/flow/CrossRuntimeConformanceTest.java @@ -766,4 +766,146 @@ void legacyWarningIsEmittedOncePerDistinctValue() throws Exception { "Three imports of the same legacy value must produce exactly one warning, got: \"" + warning + "\""); } + + // ---- Rosetta context-mode SDK downgrade (feedback -> blind, spec decision 2) ----------- + // + // Frozen cross-runtime contract (orchestrator ruling 2026-07-10; canonical reference: + // docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md, decision 2, and + // AgenticFlow.contextMode in src/types/harness.ts): + // 1. contextMode === "feedback" -> ONE warning ("feedback mode is not supported by this + // runtime yet; downgrading to blind"); the flow still + // executes in blind mode — this SDK has no + // feedback-mode manifest/briefing machinery (v2 + // concern; the PF campaign runs on the TS runtime). + // 2. "blind", absent/null, or any -> silent: no downgrade warning, because there is + // other value nothing to downgrade. + // In every case contextMode is preserved verbatim on the parsed FlowDefinition — the + // downgrade changes runtime behavior, never the recorded value. The warning fires once + // per process, by the same channel/dedup idiom as the legacy-format shim (System.err). + + /** + * Rule 1 — importing {@code conformance-feedback-mode.flow.json} (contextMode: "feedback") + * emits exactly one downgrade warning on {@code System.err} and preserves {@code + * contextMode} on the returned {@link FlowDefinition} verbatim (it is NOT cleared to + * {@code null} or rewritten to {@code "blind"} — only execution downgrades). + */ + @Test + void feedbackContextModeDowngradesToBlindWithWarning() throws Exception { + FlowImport.resetContextModeWarningsForTests(); + Path fixture = resolveConformanceFile("conformance-feedback-mode.flow.json"); + + FlowDefinition[] flowHolder = new FlowDefinition[1]; + String warning = captureStderr(() -> flowHolder[0] = FlowImport.fromCanonicalFile(fixture)); + FlowDefinition flow = flowHolder[0]; + + // The downgrade never affects structural parsing. + assertEquals("conformance-feedback-mode", flow.id()); + assertEquals(1, flow.steps().size()); + assertEquals("step-a", flow.steps().get(0).id()); + + // contextMode is preserved verbatim, not erased or rewritten to "blind". + assertEquals("feedback", flow.contextMode(), + "contextMode must be preserved verbatim on the imported FlowDefinition"); + + assertTrue( + warning.contains("feedback mode is not supported by this runtime yet; downgrading to blind"), + "Expected the exact downgrade warning message, got: \"" + warning + "\""); + } + + /** + * Rule 2a — an explicit {@code "blind"} contextMode never warns, and is preserved verbatim + * (not defaulted to {@code null}). + */ + @Test + void blindContextModeNeverWarns() throws Exception { + FlowImport.resetContextModeWarningsForTests(); + String json = """ + { + "version": "1", + "format": "fluxor-flow", + "id": "context-mode-blind", + "name": "Context Mode Blind", + "contextMode": "blind", + "rootStepId": "step-a", + "steps": [ + { "id": "step-a", "type": "llm_call", "prompt": "Step A.", "dependsOn": [], "tools": [] } + ] + } + """; + + FlowDefinition[] flowHolder = new FlowDefinition[1]; + String warning = captureStderr(() -> flowHolder[0] = FlowImport.fromCanonicalJson(json)); + + assertEquals("", warning, "An explicit 'blind' contextMode must never warn"); + assertEquals("blind", flowHolder[0].contextMode(), + "An explicit 'blind' contextMode must be preserved verbatim"); + } + + /** + * Rule 2b — a flow with no {@code contextMode} field at all (every flow predating Rosetta, + * and every blind-mode export per the wire-format doc) never warns, and {@code + * contextMode()} reads back {@code null}. + */ + @Test + void absentContextModeNeverWarns() throws Exception { + FlowImport.resetContextModeWarningsForTests(); + Path fixture = resolveFixture(); // conformance-chain.flow.json — no contextMode field. + + FlowDefinition[] flowHolder = new FlowDefinition[1]; + String warning = captureStderr(() -> flowHolder[0] = FlowImport.fromCanonicalFile(fixture)); + + assertEquals("", warning, "An absent contextMode must never warn"); + assertNull(flowHolder[0].contextMode(), "An absent contextMode must read back as null"); + } + + /** + * Rule 2c — any value other than the literal {@code "feedback"} never warns, even an + * unrecognized string; the raw value is still preserved verbatim (this SDK does not + * validate the enum — only the TS IDE authoring surface does). + */ + @Test + void unknownContextModeValueNeverWarns() throws Exception { + FlowImport.resetContextModeWarningsForTests(); + String json = """ + { + "version": "1", + "format": "fluxor-flow", + "id": "context-mode-unknown", + "name": "Context Mode Unknown", + "contextMode": "quantum", + "rootStepId": "step-a", + "steps": [ + { "id": "step-a", "type": "llm_call", "prompt": "Step A.", "dependsOn": [], "tools": [] } + ] + } + """; + + FlowDefinition[] flowHolder = new FlowDefinition[1]; + String warning = captureStderr(() -> flowHolder[0] = FlowImport.fromCanonicalJson(json)); + + assertEquals("", warning, "An unrecognized contextMode value must never warn"); + assertEquals("quantum", flowHolder[0].contextMode(), + "An unrecognized contextMode value must still be preserved verbatim"); + } + + /** + * Once-per-process dedup, mirroring {@code legacyWarningIsEmittedOncePerDistinctValue} — + * repeated imports of the same feedback-mode flow emit exactly one downgrade warning. + */ + @Test + void feedbackDowngradeWarningIsEmittedOncePerProcess() throws Exception { + FlowImport.resetContextModeWarningsForTests(); + Path fixture = resolveConformanceFile("conformance-feedback-mode.flow.json"); + + String warning = captureStderr(() -> { + FlowImport.fromCanonicalFile(fixture); + FlowImport.fromCanonicalFile(fixture); + FlowImport.fromCanonicalFile(fixture); + }); + + assertEquals(1, + countOccurrences(warning, "feedback mode is not supported by this runtime yet; downgrading to blind"), + "Three imports of the same feedback-mode flow must produce exactly one warning, got: \"" + + warning + "\""); + } } diff --git a/sdk/python/fluxor_sdk/flow_import.py b/sdk/python/fluxor_sdk/flow_import.py index d30da5e..06eaab8 100644 --- a/sdk/python/fluxor_sdk/flow_import.py +++ b/sdk/python/fluxor_sdk/flow_import.py @@ -43,6 +43,31 @@ _CURRENT_FORMAT_NAME = "fluxor-flow" _LEGACY_FORMAT_NAME = "heliox-flow" +# --------------------------------------------------------------------------- +# Rosetta context-mode downgrade (spec decision 2: +# docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md). +# +# Frozen cross-runtime contract (canonical reference: AgenticFlow.contextMode +# in src/types/harness.ts; mirrored by the Java FlowImport): +# 1. contextMode == "feedback" -> one DeprecationWarning (the exact message +# below); execution proceeds in blind mode, +# the only mode this runtime implements — +# it has no run-context manifest/briefing +# machinery (that lives in the TS harness). +# 2. "blind", absent/null, or -> silence: there is nothing to downgrade. +# any other value +# Never raises, never blocks parsing, and never mutates the value: the mode +# is preserved verbatim on the parsed FlowDefinition (``context_mode``) so a +# re-serialisation round-trips without loss. Same channel and once-semantics +# as the legacy-format shim above (``_check_legacy_format``): warnings.warn + +# DeprecationWarning, dedup delegated to the stdlib warnings filter. +# --------------------------------------------------------------------------- + +_FEEDBACK_CONTEXT_MODE = "feedback" +_FEEDBACK_DOWNGRADE_WARNING = ( + "feedback mode is not supported by this runtime yet; downgrading to blind" +) + def _clamp_loop_iterations(value: Any) -> int: """Clamp a raw ``maxIterations`` value into [1, 50]. @@ -98,11 +123,22 @@ def render_user_prompt(self) -> str: @dataclass(frozen=True) class FlowDefinition: - """An in-memory DAG produced by FlowImport.""" + """An in-memory DAG produced by FlowImport. + + ``context_mode`` is the Rosetta context mode (``AgenticFlow.contextMode`` + in src/types/harness.ts), carried opaquely like ``StepConfig.contract`` / + ``StepConfig.model``: this runtime never branches on it. ``None`` when the + flow declares none (blind default). ``"feedback"`` is NOT supported by + this runtime in v1 — ``FlowImport`` emits a one-time downgrade warning at + import and execution proceeds in blind mode (the only mode this executor + implements); the recorded value stays ``"feedback"`` so the definition + round-trips without loss (spec decision 2: downgrade, not parity). + """ id: str steps: list[StepConfig] loops: list[LoopConfig] = field(default_factory=list) + context_mode: str | None = None class FlowImport: @@ -132,6 +168,10 @@ def from_canonical_json(json_text: str) -> FlowDefinition: # Legacy compat: advisory-only, never blocks parsing. FlowImport._check_legacy_format(root) + # Rosetta context mode: carried opaquely; 'feedback' downgrades + # execution to blind with a one-time warning (advisory-only). + FlowImport._check_context_mode_downgrade(root) + flow_id = root.get("id") if not flow_id: raise ValueError("Canonical flow JSON is missing the required 'id' field.") @@ -151,7 +191,12 @@ def from_canonical_json(json_text: str) -> FlowDefinition: ) loops = [FlowImport._parse_loop(node, flow_id) for node in loops_node] - return FlowDefinition(id=flow_id, steps=steps, loops=loops) + return FlowDefinition( + id=flow_id, + steps=steps, + loops=loops, + context_mode=root.get("contextMode"), + ) @staticmethod def from_canonical_file(path: str | Path) -> FlowDefinition: @@ -191,6 +236,37 @@ def _check_legacy_format(root: dict) -> None: stacklevel=3, ) + @staticmethod + def _check_context_mode_downgrade(root: dict) -> None: + """Rosetta context-mode downgrade (spec decision 2, + docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md): warn + when importing a flow declaring ``contextMode: "feedback"``, which + this runtime cannot honour in v1 (no run-context manifest/briefing + machinery — that lives in the TS harness engine). Mirrors the Java + ``FlowImport.checkContextModeDowngrade`` — the frozen contract: + + 1. ``"feedback"`` — one ``DeprecationWarning`` (the exact + ``_FEEDBACK_DOWNGRADE_WARNING`` message); execution proceeds in + blind mode, the only mode this executor implements. + 2. ``"blind"``, absent/null, or any other value — silence: there is + nothing to downgrade. + + Never raises, and never mutates the value: ``contextMode`` is + preserved verbatim on the parsed FlowDefinition (the downgrade changes + runtime behaviour, not the recorded definition). Same transport and + once-semantics as ``_check_legacy_format``: dedup is delegated to the + stdlib warnings filter (the default action dedups repeated identical + warnings — Python's native equivalent of the Java once-per-process + flag / TS ``warnOnce``). + """ + if root.get("contextMode") != _FEEDBACK_CONTEXT_MODE: + return + warnings.warn( + _FEEDBACK_DOWNGRADE_WARNING, + DeprecationWarning, + stacklevel=3, + ) + @staticmethod def _parse_step(node: dict, flow_id: str) -> StepConfig: step_id = node.get("id") diff --git a/sdk/python/tests/test_cross_runtime_conformance.py b/sdk/python/tests/test_cross_runtime_conformance.py index 22967b4..3973fde 100644 --- a/sdk/python/tests/test_cross_runtime_conformance.py +++ b/sdk/python/tests/test_cross_runtime_conformance.py @@ -531,3 +531,159 @@ def test_unknown_format_warns_mentioning_seen_value() -> None: flow = FlowImport.from_canonical_json(raw) assert flow.id == "unknown-format", "An unknown format tag must never block parsing" + + +# --------------------------------------------------------------------------- +# Test 8: Rosetta context-mode SDK downgrade (feedback -> blind, spec decision 2) +# +# Frozen cross-runtime contract (orchestrator ruling 2026-07-10; canonical +# reference: docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md, +# decision 2, and AgenticFlow.contextMode in src/types/harness.ts): +# 1. contextMode == "feedback" -> ONE DeprecationWarning ("feedback mode is +# not supported by this runtime yet; +# downgrading to blind"); the flow still +# executes in blind mode — this SDK has no +# feedback-mode manifest/briefing machinery +# (v2 concern; the PF campaign runs on the +# TS runtime). +# 2. "blind", absent/null, or -> silent: no downgrade warning, because +# any other value there is nothing to downgrade. +# In every case contextMode is preserved verbatim on the parsed FlowDefinition +# (context_mode attribute) — the downgrade changes runtime behavior, never the +# recorded value. Channel and once-semantics mirror the legacy-format shim: +# one DeprecationWarning via warnings.warn, dedup delegated to the stdlib +# warnings filter (pytest installs an "always" filter, so each test observes +# its own warning). Java counterpart: System.err + once-per-process flag +# (CrossRuntimeConformanceTest, "Rosetta context-mode SDK downgrade" section). +# --------------------------------------------------------------------------- + +_FEEDBACK_DOWNGRADE_WARNING = ( + "feedback mode is not supported by this runtime yet; downgrading to blind" +) + + +def test_feedback_context_mode_downgrades_to_blind_with_warning() -> None: + """Rule 1: importing ``conformance-feedback-mode.flow.json`` + (``contextMode: "feedback"``) emits exactly one DeprecationWarning with + the exact downgrade message and preserves ``context_mode`` verbatim on + the returned FlowDefinition (NOT cleared to ``None`` or rewritten to + ``"blind"`` — only execution downgrades).""" + raw = _load_fixture("conformance-feedback-mode.flow.json") + + with pytest.warns(DeprecationWarning, match=_FEEDBACK_DOWNGRADE_WARNING) as record: + flow = FlowImport.from_canonical_json(raw) + + # The downgrade never affects structural parsing. + assert flow.id == "conformance-feedback-mode" + assert len(flow.steps) == 1 + assert flow.steps[0].id == "step-a" + + # contextMode is preserved verbatim, not erased or rewritten to "blind". + assert flow.context_mode == "feedback", ( + f"context_mode must be preserved verbatim, got {flow.context_mode!r}" + ) + + assert len(record) == 1, f"Expected exactly one DeprecationWarning, got {len(record)}" + + +def test_blind_context_mode_never_warns(recwarn: pytest.WarningsRecorder) -> None: + """Rule 2a: an explicit ``"blind"`` contextMode never warns, and is + preserved verbatim (not defaulted to ``None``).""" + raw = json.dumps( + { + "version": "1", + "format": "fluxor-flow", + "id": "context-mode-blind", + "name": "Context Mode Blind", + "contextMode": "blind", + "rootStepId": "step-a", + "steps": [ + {"id": "step-a", "type": "llm_call", "prompt": "Step A.", "dependsOn": [], "tools": []} + ], + } + ) + + flow = FlowImport.from_canonical_json(raw) + + assert len(recwarn) == 0, ( + f"An explicit 'blind' contextMode must never warn, got " + f"{[str(w.message) for w in recwarn]}" + ) + assert flow.context_mode == "blind", ( + f"An explicit 'blind' contextMode must be preserved verbatim, got {flow.context_mode!r}" + ) + + +def test_absent_context_mode_never_warns(recwarn: pytest.WarningsRecorder) -> None: + """Rule 2b: a flow with NO ``contextMode`` field at all (every flow + predating Rosetta, and every blind-mode export per the wire-format doc) + never warns, and ``context_mode`` reads back ``None``.""" + raw = _load_fixture("conformance-chain.flow.json") # no contextMode field. + + flow = FlowImport.from_canonical_json(raw) + + assert len(recwarn) == 0, ( + f"An absent contextMode must never warn, got {[str(w.message) for w in recwarn]}" + ) + assert flow.context_mode is None, ( + f"An absent contextMode must read back as None, got {flow.context_mode!r}" + ) + + +def test_unknown_context_mode_value_never_warns(recwarn: pytest.WarningsRecorder) -> None: + """Rule 2c: any value other than the literal ``"feedback"`` never warns, + even an unrecognized string; the raw value is still preserved verbatim + (this SDK does not validate the enum — only the TS IDE authoring surface + does).""" + raw = json.dumps( + { + "version": "1", + "format": "fluxor-flow", + "id": "context-mode-unknown", + "name": "Context Mode Unknown", + "contextMode": "quantum", + "rootStepId": "step-a", + "steps": [ + {"id": "step-a", "type": "llm_call", "prompt": "Step A.", "dependsOn": [], "tools": []} + ], + } + ) + + flow = FlowImport.from_canonical_json(raw) + + assert len(recwarn) == 0, ( + f"An unrecognized contextMode value must never warn, got " + f"{[str(w.message) for w in recwarn]}" + ) + assert flow.context_mode == "quantum", ( + f"An unrecognized contextMode value must still be preserved verbatim, " + f"got {flow.context_mode!r}" + ) + + +def test_feedback_downgrade_warning_is_emitted_once() -> None: + """Once-semantics, mirroring the Java + ``feedbackDowngradeWarningIsEmittedOncePerProcess`` as closely as the + Python idiom allows: the downgrade shim delegates dedup to the stdlib + warnings filter, exactly like the legacy-format shim — under the + ``"default"`` action (the production configuration), repeated identical + warnings from the same import call site collapse to one. Three imports + of the same feedback-mode flow from one call site emit exactly one + downgrade warning.""" + import warnings as _warnings + + raw = _load_fixture("conformance-feedback-mode.flow.json") + + with _warnings.catch_warnings(record=True) as record: + _warnings.simplefilter("default") + for _ in range(3): + FlowImport.from_canonical_json(raw) + + downgrade_warnings = [ + w for w in record if _FEEDBACK_DOWNGRADE_WARNING in str(w.message) + ] + assert len(downgrade_warnings) == 1, ( + f"Three imports of the same feedback-mode flow must produce exactly one " + f"downgrade warning under the default filter, got {len(downgrade_warnings)}: " + f"{[str(w.message) for w in record]}" + ) diff --git a/src/main/flow-export/fluxor-flow.test.ts b/src/main/flow-export/fluxor-flow.test.ts index 622b530..156a819 100644 --- a/src/main/flow-export/fluxor-flow.test.ts +++ b/src/main/flow-export/fluxor-flow.test.ts @@ -605,3 +605,57 @@ describe('legacy format tag compat', () => { expect(warnSpy).toHaveBeenCalledTimes(1); }); }); + +// --------------------------------------------------------------------------- +// Suite 8 — contextMode round-trip (Rosetta, spec: +// docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md) +// --------------------------------------------------------------------------- + +describe('contextMode export/import round-trip', () => { + it('omits the contextMode key entirely when the flow has no contextMode (blind default) — byte-identical to a pre-feedback export', () => { + const exported = exportFlow(makeTestFlow()); + expect('contextMode' in exported).toBe(false); + expect(JSON.stringify(exported)).not.toContain('"contextMode"'); + }); + + it('omits the contextMode key when the flow explicitly declares "blind"', () => { + const flow = { ...makeTestFlow(), contextMode: 'blind' as const }; + const exported = exportFlow(flow); + expect('contextMode' in exported).toBe(false); + expect(JSON.stringify(exported)).not.toContain('"contextMode"'); + }); + + it('includes contextMode: "feedback" when the flow declares it', () => { + const flow = { ...makeTestFlow(), contextMode: 'feedback' as const }; + const exported = exportFlow(flow); + expect(exported.contextMode).toBe('feedback'); + }); + + it('restores contextMode: "feedback" on import', () => { + const flow = { ...makeTestFlow(), contextMode: 'feedback' as const }; + const restored = importFlow(exportFlow(flow)); + expect(restored.contextMode).toBe('feedback'); + }); + + it('leaves contextMode undefined on import when the export carries none (never defaults it to the literal "blind")', () => { + const restored = importFlow(exportFlow(makeTestFlow())); + expect(restored.contextMode).toBeUndefined(); + expect('contextMode' in restored).toBe(false); + }); + + it('round-trips a feedback-mode flow through export→import→export with a stable result', () => { + const flow = { ...makeTestFlow(), contextMode: 'feedback' as const }; + const exportedOnce = exportFlow(flow); + const reimported = importFlow(exportedOnce); + const exportedTwice = exportFlow(reimported); + expect(exportedTwice).toEqual(exportedOnce); + }); + + it('does not perturb any other field on the export when contextMode is present', () => { + const blindExported = exportFlow(makeTestFlow()); + const feedbackFlow = { ...makeTestFlow(), contextMode: 'feedback' as const }; + const feedbackExported = exportFlow(feedbackFlow); + const { contextMode: _omit, ...feedbackWithoutContextMode } = feedbackExported; + expect(feedbackWithoutContextMode).toEqual(blindExported); + }); +}); diff --git a/src/main/flow-export/fluxor-flow.ts b/src/main/flow-export/fluxor-flow.ts index 1f74710..ecb6fa0 100644 --- a/src/main/flow-export/fluxor-flow.ts +++ b/src/main/flow-export/fluxor-flow.ts @@ -20,7 +20,8 @@ * carried opaquely — this module never interprets it), per-step `model` * override, per-step `description` (human-facing, execution-inert), flow * `meta` (human-facing description/tags/author/version, execution-inert), - * and `loops` (bounded loop-back edges). + * `loops` (bounded loop-back edges), and `contextMode` (Rosetta context + * mode — omitted entirely for blind/absent, present only as 'feedback'). * * INTENTIONALLY FLATTENED (lossy but acceptable): * - role ids and role names collapse into a single 'exported-role' sentinel; @@ -169,6 +170,16 @@ export interface FluxorFlowExport { * [1, 50] at export time — see the module header's "Loop cap" note. */ loops?: FluxorFlowLoop[]; + /** + * Rosetta context mode (AgenticFlow.contextMode — spec: + * docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md). Omitted + * entirely (never the literal `'blind'`) when the flow's mode is blind or + * absent, so a blind flow's export stays byte-identical to every export + * produced before this field existed. Present only as `'feedback'`. A + * top-level field (not nested under `meta`) because — unlike `meta`'s + * purely human-facing fields — it has real execution semantics. + */ + contextMode?: 'feedback'; /** * Optional human-facing flow metadata (AgenticFlow.description/tags/author/ * version), nested under `meta` rather than flattened to top-level keys — @@ -279,6 +290,9 @@ export function exportFlow(flow: AgenticFlow): FluxorFlowExport { }; if (loops !== undefined) exported.loops = loops; if (meta !== undefined) exported.meta = meta; + // Blind or absent ⇒ no key at all (byte-identical to every pre-Rosetta + // export); only 'feedback' is ever stamped — see FluxorFlowExport's doc. + if (flow.contextMode === 'feedback') exported.contextMode = 'feedback'; return exported; } @@ -386,6 +400,10 @@ export function importFlow(exported: FluxorFlowExport): AgenticFlow { if (exported.meta?.tags !== undefined) flow.tags = exported.meta.tags; if (exported.meta?.author !== undefined) flow.author = exported.meta.author; if (exported.meta?.version !== undefined) flow.version = exported.meta.version; + // Only ever restored as 'feedback' — never defaulted to the literal + // 'blind', so a re-export of this flow stays byte-identical to the source + // export for every blind/absent case (see FluxorFlowExport.contextMode doc). + if (exported.contextMode === 'feedback') flow.contextMode = 'feedback'; return flow; } diff --git a/src/main/flow-export/semantic-parity.test.ts b/src/main/flow-export/semantic-parity.test.ts index 0db0f65..7e70208 100644 --- a/src/main/flow-export/semantic-parity.test.ts +++ b/src/main/flow-export/semantic-parity.test.ts @@ -15,6 +15,8 @@ */ import { readFileSync } from 'fs'; +import { mkdtemp, readdir, readFile, rm } from 'fs/promises'; +import { tmpdir } from 'os'; import { join } from 'path'; import { describe, expect, it } from 'vitest'; import { executeAgenticFlow } from '../harness-engine/executor'; @@ -292,3 +294,164 @@ describe('cross-runtime loop execution parity — golden loop trace', () => { expect(trace[stepCIndex]).toEqual({ stepId: 'step-c', output: 'C output' }); }); }); + +// --------------------------------------------------------------------------- +// Suite — contextMode: 'feedback' is additive (Rosetta, spec: +// docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md) +// +// Java/Python downgrade contextMode: 'feedback' to blind on import in v1 +// (decision 2) — the cross-runtime CONTRACT this suite otherwise guards is +// therefore untouched by feedback mode today. What this suite instead proves +// for the TS runtime itself: turning feedback mode on for the SAME +// conformance-chain flow this file already golden-traces does not perturb +// that golden-trace text contract — the Rosetta machinery is confined to +// side-channel files under .fluxor/run-context/, never the step outputs a +// future runtime parity check would compare. +// --------------------------------------------------------------------------- + +// Linear chain (confirmed by the "canonical DAG order" test above): +// step-a → step-b → step-c → step-d → step-e. Each non-terminal step writes +// a real briefing into its immediate successor's file — satisfying the +// feedback-mode guardrail on the FIRST attempt, so retries never fire and +// the trace stays 1:1 with goldenTrace (a retry would otherwise re-invoke +// runStep and duplicate that step's trace entry). +const NEXT_STEP_ID: Record = { + 'step-a': 'step-b', + 'step-b': 'step-c', + 'step-c': 'step-d', + 'step-d': 'step-e', +}; + +describe('feedback-mode contextMode is additive — does not perturb the golden trace', () => { + it('produces the SAME golden-trace text with contextMode: "feedback" added, and materializes a run-context manifest on disk', async () => { + const tmpDir = await mkdtemp(join(tmpdir(), 'fluxor-semantic-parity-feedback-')); + try { + const feedbackExportedFlow: FluxorFlowExport = { ...exportedFlow, contextMode: 'feedback' }; + const flow = importFlow(feedbackExportedFlow); + expect(flow.contextMode).toBe('feedback'); + + const trace: Array<{ stepId: string; output: string }> = []; + const runId = 'run-semantic-parity-feedback'; + + await executeAgenticFlow(flow, { + rootDir: tmpDir, + runId, + runStep: async (input) => { + const stepId = input.step.id; + + // Satisfy the promised-briefing guardrail on the first attempt. + // Generously long (well beyond any single seeded header + a + // 140-char purpose line) so it clears minBytes for every target + // regardless of that target step's own prompt length. + const nextStepId = NEXT_STEP_ID[stepId]; + if (nextStepId) { + await input.tools.write_file.execute?.( + { + path: `.fluxor/run-context/${runId}/step.${nextStepId}.md`, + content: `A real, detailed briefing from ${stepId} for ${nextStepId}. `.repeat(10), + }, + { toolCallId: `briefing-${stepId}`, messages: [] }, + ); + } + + if (stepId === 'step-e') { + // Same tool-calling scripted shape as the golden-trace suite + // above — this suite is about proving feedback mode doesn't + // perturb the OUTPUT, not re-proving tool-calling parity. + const toolResult = uppercase(UPPERCASE_TOOL_ARGS.text); + const finalText = scriptedResponses[stepId] ?? ''; + const scripted = async () => ({ + text: finalText, + usage: null, + toolCalls: [ + { toolCallId: 'call_conformance_1', toolName: 'uppercase', input: UPPERCASE_TOOL_ARGS }, + ], + toolResults: [ + { toolCallId: 'call_conformance_1', toolName: 'uppercase', output: toolResult }, + ], + steps: [ + { + text: '', + content: [ + { type: 'tool-call', toolCallId: 'call_conformance_1', toolName: 'uppercase', input: UPPERCASE_TOOL_ARGS }, + ], + toolCalls: [ + { toolCallId: 'call_conformance_1', toolName: 'uppercase', input: UPPERCASE_TOOL_ARGS }, + ], + toolResults: [], + }, + { + text: finalText, + content: [ + { type: 'tool-result', toolCallId: 'call_conformance_1', toolName: 'uppercase', output: toolResult }, + { type: 'text', text: finalText }, + ], + toolCalls: [], + toolResults: [ + { toolCallId: 'call_conformance_1', toolName: 'uppercase', output: toolResult }, + ], + }, + ], + }); + + const result = await runLLMStep({ ...input, generateText: scripted }); + trace.push({ stepId, output: result.text }); + return result; + } + + const scripted = async () => ({ + text: scriptedResponses[stepId] ?? '', + usage: null, + toolCalls: [], + toolResults: [], + steps: [], + }); + const result = await runLLMStep({ ...input, generateText: scripted }); + trace.push({ stepId, output: result.text }); + return result; + }, + }); + + // The primary claim: feedback mode is additive. Same golden contract. + expect(trace).toEqual(goldenTrace); + + // And genesis actually happened (proving "additive" isn't vacuous — + // the run-context machinery genuinely ran alongside the golden trace). + const manifestPath = join(tmpDir, '.fluxor', 'run-context', runId, 'manifest.json'); + const manifest = JSON.parse(await readFile(manifestPath, 'utf-8')); + expect(manifest.contextMode).toBe('feedback'); + expect(Object.keys(manifest.steps).sort()).toEqual( + ['step-a', 'step-b', 'step-c', 'step-d', 'step-e'].sort(), + ); + } finally { + await rm(tmpDir, { recursive: true, force: true }); + } + }); + + it('a blind-mode run of the SAME chain (no contextMode) creates no .fluxor/run-context directory at all', async () => { + const tmpDir = await mkdtemp(join(tmpdir(), 'fluxor-semantic-parity-blind-')); + try { + const flow = importFlow(exportedFlow); // exportedFlow carries no contextMode — blind + expect(flow.contextMode).toBeUndefined(); + + await executeAgenticFlow(flow, { + rootDir: tmpDir, + runStep: async (input) => { + const stepId = input.step.id; + const scripted = async () => ({ + text: scriptedResponses[stepId] ?? '', + usage: null, + toolCalls: [], + toolResults: [], + steps: [], + }); + return runLLMStep({ ...input, generateText: scripted }); + }, + }); + + await expect(readdir(join(tmpDir, '.fluxor'))).rejects.toThrow(); + } finally { + await rm(tmpDir, { recursive: true, force: true }); + } + }); +}); diff --git a/src/main/harness-engine/checkpoints.test.ts b/src/main/harness-engine/checkpoints.test.ts index f5054f1..0478e2e 100644 --- a/src/main/harness-engine/checkpoints.test.ts +++ b/src/main/harness-engine/checkpoints.test.ts @@ -8,6 +8,7 @@ import { afterEach, describe, expect, it } from 'vitest'; import { type CheckpointStore, + CONTEXT_FILE_SNAPSHOT_MAX_CHARS, InMemoryCheckpointStore, INPUT_CONTEXT_MAX_CHARS, OUTPUT_MAX_CHARS, @@ -233,6 +234,100 @@ describe('field truncation', () => { }); }); +// --------------------------------------------------------------------------- +// contextFileSnapshot — additive field for feedback-mode runs (Rosetta, spec: +// docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md) +// --------------------------------------------------------------------------- + +describe('contextFileSnapshot (feedback-mode checkpoints)', () => { + it('stores contextFileSnapshot when provided', () => { + useInMemoryStore(); + + const saved = saveCheckpoint({ + runId: 'run-ctx', + stepId: 'step-a', + inputContext: '', + output: '', + completedStepIds: ['step-a'], + contextFileSnapshot: { path: '.fluxor/run-context/run-ctx/step.step-a.md', content: '# Contexto para step-a\n\nDo the thing.' }, + }); + + expect(saved.contextFileSnapshot).toEqual({ + path: '.fluxor/run-context/run-ctx/step.step-a.md', + content: '# Contexto para step-a\n\nDo the thing.', + }); + }); + + it('omits the contextFileSnapshot key entirely when not provided, rather than storing it as undefined', () => { + useInMemoryStore(); + + const saved = saveCheckpoint({ + runId: 'run-no-ctx', + stepId: 'step-b', + inputContext: '', + output: '', + completedStepIds: [], + }); + + expect('contextFileSnapshot' in saved).toBe(false); + expect(saved.contextFileSnapshot).toBeUndefined(); + }); + + it('truncates contextFileSnapshot.content that exceeds CONTEXT_FILE_SNAPSHOT_MAX_CHARS', () => { + useInMemoryStore(); + + const longContent = 'z'.repeat(CONTEXT_FILE_SNAPSHOT_MAX_CHARS + 500); + const saved = saveCheckpoint({ + runId: 'run-ctx-trunc', + stepId: 'step-c', + inputContext: '', + output: '', + completedStepIds: [], + contextFileSnapshot: { path: '.fluxor/run-context/run-ctx-trunc/step.step-c.md', content: longContent }, + }); + + expect(saved.contextFileSnapshot?.content.length).toBeLessThanOrEqual(CONTEXT_FILE_SNAPSHOT_MAX_CHARS); + expect(saved.contextFileSnapshot?.content).toContain('[truncated]'); + }); + + it('does not truncate contextFileSnapshot.content within the limit', () => { + useInMemoryStore(); + + const content = 'a short briefing'; + const saved = saveCheckpoint({ + runId: 'run-ctx-short', + stepId: 'step-d', + inputContext: '', + output: '', + completedStepIds: [], + contextFileSnapshot: { path: '.fluxor/run-context/run-ctx-short/step.step-d.md', content }, + }); + + expect(saved.contextFileSnapshot?.content).toBe(content); + }); + + it('does not disturb the existing shape (iteration/modelId/completedStepIds) when contextFileSnapshot is also present', () => { + useInMemoryStore(); + + const saved = saveCheckpoint({ + runId: 'run-ctx-shape', + stepId: 'step-e', + iteration: 2, + inputContext: 'ctx', + output: 'out', + completedStepIds: ['step-e'], + modelId: 'gpt-4o', + contextFileSnapshot: { path: 'p', content: 'c' }, + }); + + expect(saved.iteration).toBe(2); + expect(saved.modelId).toBe('gpt-4o'); + expect(saved.completedStepIds).toEqual(['step-e']); + expect(saved.inputContext).toBe('ctx'); + expect(saved.output).toBe('out'); + }); +}); + // --------------------------------------------------------------------------- // SqliteCheckpointStore with a mock database // --------------------------------------------------------------------------- @@ -371,3 +466,120 @@ describe('SqliteCheckpointStore (mock database)', () => { expect('iteration' in found!).toBe(false); }); }); + +// --------------------------------------------------------------------------- +// SqliteCheckpointStore — context_file_path / context_file_content columns +// +// A SEPARATE mock database (not the `buildMockDb()` above, which every +// pre-existing test in this file depends on staying untouched) that also +// captures the two new columns the INSERT statement carries for +// contextFileSnapshot, so these NEW tests can prove the SQLite backend +// round-trips the additive field exactly like InMemoryCheckpointStore does. +// --------------------------------------------------------------------------- + +describe('SqliteCheckpointStore (mock database) — context_file_path/context_file_content', () => { + function buildMockDbWithContextColumns() { + const rows = new Map>(); + + return { + exec: (_sql: string): void => { + // No-op: mock doesn't actually execute DDL. + }, + prepare: (sql: string) => ({ + run: (...args: unknown[]) => { + const trimmed = sql.trim().toLowerCase(); + if (trimmed.startsWith('insert')) { + // Column order mirrors the real INSERT in checkpoints.ts: + // (id, run_id, step_id, iteration, input_context, output, + // completed_step_ids, model_id, timestamp, context_file_path, + // context_file_content). + const [ + id, runId, stepId, iteration, inputContext, output, + completedStepIds, modelId, timestamp, contextFilePath, contextFileContent, + ] = args; + rows.set(String(id), { + id, + run_id: runId, + step_id: stepId, + iteration, + input_context: inputContext, + output, + completed_step_ids: completedStepIds, + model_id: modelId, + timestamp, + context_file_path: contextFilePath, + context_file_content: contextFileContent, + }); + } + }, + get: (...args: unknown[]) => rows.get(String(args[0])), + all: (...args: unknown[]) => { + const runId = String(args[0]); + return [...rows.values()] + .filter((r) => String(r.run_id) === runId) + .sort((a, b) => Number(a.timestamp) - Number(b.timestamp)); + }, + }), + }; + } + + it('round-trips path + content when contextFileSnapshot is provided', () => { + const store = new SqliteCheckpointStore(buildMockDbWithContextColumns() as any); + + store.save({ + id: 'ckpt_run-ctx-sql_step-a_0', + runId: 'run-ctx-sql', + stepId: 'step-a', + inputContext: '', + output: '', + completedStepIds: [], + modelId: undefined, + timestamp: 0, + contextFileSnapshot: { path: '.fluxor/run-context/run-ctx-sql/step.step-a.md', content: 'briefing text' }, + }); + + const found = store.get('ckpt_run-ctx-sql_step-a_0'); + expect(found?.contextFileSnapshot).toEqual({ + path: '.fluxor/run-context/run-ctx-sql/step.step-a.md', + content: 'briefing text', + }); + }); + + it('omits contextFileSnapshot (rather than a null-filled object) when both columns are null', () => { + const store = new SqliteCheckpointStore(buildMockDbWithContextColumns() as any); + + store.save({ + id: 'ckpt_run-ctx-sql2_step-b_0', + runId: 'run-ctx-sql2', + stepId: 'step-b', + inputContext: '', + output: '', + completedStepIds: [], + modelId: undefined, + timestamp: 0, + }); + + const found = store.get('ckpt_run-ctx-sql2_step-b_0'); + expect(found).toBeDefined(); + expect('contextFileSnapshot' in found!).toBe(false); + }); + + it('round-trips through list() as well as get()', () => { + const store = new SqliteCheckpointStore(buildMockDbWithContextColumns() as any); + + store.save({ + id: 'ckpt_run-ctx-sql3_step-c_0', + runId: 'run-ctx-sql3', + stepId: 'step-c', + inputContext: '', + output: '', + completedStepIds: [], + modelId: undefined, + timestamp: 5, + contextFileSnapshot: { path: 'p', content: 'c' }, + }); + + const [listed] = store.list('run-ctx-sql3'); + expect(listed.contextFileSnapshot).toEqual({ path: 'p', content: 'c' }); + }); +}); diff --git a/src/main/harness-engine/checkpoints.ts b/src/main/harness-engine/checkpoints.ts index 8bae769..6d896e4 100644 --- a/src/main/harness-engine/checkpoints.ts +++ b/src/main/harness-engine/checkpoints.ts @@ -12,14 +12,36 @@ * - `inputContext` max 64 KiB (65_536 chars) — trimmed with a marker * - `output` max 32 KiB (32_768 chars) — trimmed with a marker * - `completedStepIds` serialised as JSON; no separate cap (step-count bounded) + * - `contextFileSnapshot.content` max 16 KiB (16_384 chars) — trimmed with a marker */ /** Max character length for the inputContext field before truncation. */ export const INPUT_CONTEXT_MAX_CHARS = 65_536; /** Max character length for the output field before truncation. */ export const OUTPUT_MAX_CHARS = 32_768; +/** + * Max character length for `contextFileSnapshot.content` before truncation + * (feedback-mode runs only — see ContextFileSnapshot). Comfortably above the + * Rosetta per-file budget (context-manifest.ts's DEFAULT_CONTEXT_BUDGET_BYTES, + * 12,000) so a within-budget briefing never gets clipped here; this is a + * hard backstop against an unbounded write, not the primary budget signal + * (that's the guardrail's job). + */ +export const CONTEXT_FILE_SNAPSHOT_MAX_CHARS = 16_384; const TRUNCATION_MARKER = '…[truncated]'; +/** + * A snapshot of a step's assigned feedback-mode context file (Rosetta, spec: + * docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md), captured + * after the step completes. + */ +export interface ContextFileSnapshot { + /** RootDir-relative path (e.g. ".fluxor/run-context//step..md"). */ + path: string; + /** File content at checkpoint time, capped at CONTEXT_FILE_SNAPSHOT_MAX_CHARS. */ + content: string; +} + /** Immutable snapshot recorded after a step completes. */ export interface Checkpoint { /** Unique checkpoint id — `ckpt___`. */ @@ -46,6 +68,14 @@ export interface Checkpoint { modelId: string | undefined; /** Unix epoch milliseconds when the checkpoint was recorded. */ timestamp: number; + /** + * Snapshot of this step's assigned feedback-mode context file (see + * ContextFileSnapshot). Present only for feedback-mode runs whose step + * received a manifest entry AND whose file was readable at checkpoint time + * — omitted entirely (never `undefined`-valued) otherwise, so blind-mode + * checkpoints and this field's shape are untouched by its addition. + */ + contextFileSnapshot?: ContextFileSnapshot; } // --------------------------------------------------------------------------- @@ -118,6 +148,15 @@ export class SqliteCheckpointStore implements CheckpointStore { constructor(db: BetterSqliteDatabase) { this.db = db; + // NOTE: context_file_path/context_file_content are added directly to the + // CREATE TABLE (rather than via a runtime ALTER TABLE migration) because + // this store is not yet wired into any production bootstrap path — no + // setCheckpointStore(new SqliteCheckpointStore(...)) call exists outside + // this file's own doc-comment example and its tests, so there is no + // on-disk database anywhere carrying the pre-Rosetta 9-column schema this + // would otherwise need to migrate. Whoever wires this store into a real + // app boot path MUST add an ALTER-TABLE-style migration at that point if + // any database created before that change could already exist on disk. this.db.exec(` CREATE TABLE IF NOT EXISTS harness_checkpoints ( id TEXT PRIMARY KEY, @@ -128,7 +167,9 @@ export class SqliteCheckpointStore implements CheckpointStore { output TEXT NOT NULL, completed_step_ids TEXT NOT NULL, model_id TEXT, - timestamp INTEGER NOT NULL + timestamp INTEGER NOT NULL, + context_file_path TEXT, + context_file_content TEXT ); CREATE INDEX IF NOT EXISTS idx_checkpoints_run_id ON harness_checkpoints (run_id, timestamp); @@ -138,8 +179,8 @@ export class SqliteCheckpointStore implements CheckpointStore { save(checkpoint: Checkpoint): void { this.db.prepare(` INSERT INTO harness_checkpoints - (id, run_id, step_id, iteration, input_context, output, completed_step_ids, model_id, timestamp) - VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?) + (id, run_id, step_id, iteration, input_context, output, completed_step_ids, model_id, timestamp, context_file_path, context_file_content) + VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?) `).run( checkpoint.id, checkpoint.runId, @@ -150,6 +191,8 @@ export class SqliteCheckpointStore implements CheckpointStore { JSON.stringify(checkpoint.completedStepIds), checkpoint.modelId ?? null, checkpoint.timestamp, + checkpoint.contextFileSnapshot?.path ?? null, + checkpoint.contextFileSnapshot?.content ?? null, ); } @@ -180,6 +223,9 @@ export class SqliteCheckpointStore implements CheckpointStore { completedStepIds: JSON.parse(String(row.completed_step_ids)) as string[], modelId: row.model_id != null ? String(row.model_id) : undefined, timestamp: Number(row.timestamp), + ...(row.context_file_path != null && row.context_file_content != null + ? { contextFileSnapshot: { path: String(row.context_file_path), content: String(row.context_file_content) } } + : {}), }; } } @@ -237,6 +283,12 @@ export interface SaveCheckpointInput { output: string; completedStepIds: string[]; modelId?: string; + /** + * Snapshot of this step's assigned feedback-mode context file; omit + * entirely for blind-mode runs (see Checkpoint.contextFileSnapshot). + * `content` is truncated at CONTEXT_FILE_SNAPSHOT_MAX_CHARS before storage. + */ + contextFileSnapshot?: ContextFileSnapshot; } /** @@ -264,6 +316,17 @@ export function saveCheckpoint(input: SaveCheckpointInput): Checkpoint { completedStepIds: input.completedStepIds, modelId: input.modelId, timestamp, + // Same conditional-spread discipline as `iteration` above — omitted + // entirely (not `undefined`-valued) when the caller passes none, so + // blind-mode checkpoints (which never pass this) are unaffected. + ...(input.contextFileSnapshot !== undefined + ? { + contextFileSnapshot: { + path: input.contextFileSnapshot.path, + content: truncate(input.contextFileSnapshot.content, CONTEXT_FILE_SNAPSHOT_MAX_CHARS), + }, + } + : {}), }; activeStore.save(checkpoint); return checkpoint; diff --git a/src/main/harness-engine/context-builder.test.ts b/src/main/harness-engine/context-builder.test.ts index d1003df..774b2a5 100644 --- a/src/main/harness-engine/context-builder.test.ts +++ b/src/main/harness-engine/context-builder.test.ts @@ -170,6 +170,117 @@ describe('buildStepContext', () => { }); }); +// --------------------------------------------------------------------------- +// — feedback-mode flow context (Rosetta, spec: +// docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md) +// --------------------------------------------------------------------------- + +describe('buildStepContext — flow_awareness (feedback mode)', () => { + it('blind (no flowAwareness option) produces byte-identical output whether the option key is omitted or explicitly undefined', async () => { + const step = makeStep({ + mods: [ + { id: 'security-first', name: 'SecurityFirst', type: 'pre_process', config: { inject: 'Never log secrets.' } }, + ], + mentalContext: [{ id: 'idea-1', text: 'User prefers minimal UI.', relationToStep: 'incoming' }], + }); + + const omitted = await buildStepContext(step); + const explicitUndefined = await buildStepContext(step, { flowAwareness: undefined }); + + expect(explicitUndefined).toEqual(omitted); + expect(omitted.systemPrompt).not.toContain(''); + expect(omitted.userPrompt).not.toContain(''); + }); + + it('blind output matches an exact literal snapshot — proves no new bytes are ever introduced', async () => { + const context = await buildStepContext(makeStep({ + roles: [{ id: 'architect', name: 'Architect', systemPrompt: 'Preserve system boundaries.' }], + })); + + expect(context.systemPrompt).toBe( + '\n\n[Architect]\nPreserve system boundaries.\n\n', + ); + }); + + it('renders in the SYSTEM prompt (never the user prompt), positioned after ', async () => { + const context = await buildStepContext(makeStep({ + mods: [ + { id: 'security-first', name: 'SecurityFirst', type: 'pre_process', config: { inject: 'Never log secrets.' } }, + ], + }), { + flowAwareness: { + topologyLines: ['- root: Kick off the chain.'], + assignedFile: '.fluxor/run-context/run-1/step.root.md', + writesTo: ['.fluxor/run-context/run-1/step.leaf.md'], + }, + }); + + expect(context.systemPrompt).toContain(''); + expect(context.systemPrompt).toContain(''); + expect(context.userPrompt).not.toContain(''); + + const constraintsIndex = context.systemPrompt.indexOf(''); + const awarenessIndex = context.systemPrompt.indexOf(''); + expect(constraintsIndex).toBeGreaterThanOrEqual(0); + expect(awarenessIndex).toBeGreaterThan(constraintsIndex); + }); + + it('includes the topology summary, assigned file, and writesTo paths with a write instruction', async () => { + const context = await buildStepContext(makeStep(), { + flowAwareness: { + topologyLines: ['- root: Kick off the chain.', '- leaf: Wrap it up.'], + assignedFile: '.fluxor/run-context/run-1/step.root.md', + writesTo: ['.fluxor/run-context/run-1/step.leaf.md'], + }, + }); + + expect(context.systemPrompt).toContain('- root: Kick off the chain.'); + expect(context.systemPrompt).toContain('- leaf: Wrap it up.'); + expect(context.systemPrompt).toContain('.fluxor/run-context/run-1/step.root.md'); + expect(context.systemPrompt).toContain('.fluxor/run-context/run-1/step.leaf.md'); + expect(context.systemPrompt.toLowerCase()).toContain('write_file'); + expect(context.systemPrompt.toLowerCase()).toContain('read_file'); + }); + + it('omits the "write a briefing" instruction when writesTo is empty (e.g. a terminal or exempt step)', async () => { + const context = await buildStepContext(makeStep(), { + flowAwareness: { + topologyLines: ['- solo: Do the whole thing alone.'], + assignedFile: '.fluxor/run-context/run-1/step.solo.md', + writesTo: [], + }, + }); + + expect(context.systemPrompt).toContain(''); + expect(context.systemPrompt).not.toContain('leave a'); + }); + + it('presents flow_awareness as reference material, not a constraint', async () => { + const context = await buildStepContext(makeStep(), { + flowAwareness: { + topologyLines: ['- root: Kick off the chain.'], + assignedFile: '.fluxor/run-context/run-1/step.root.md', + writesTo: [], + }, + }); + + expect(context.systemPrompt.toLowerCase()).toContain('not a constraint'); + }); + + it('still renders even when there are no execution_constraints, right after ', async () => { + const context = await buildStepContext(makeStep(), { + flowAwareness: { + topologyLines: ['- root: Kick off the chain.'], + assignedFile: '.fluxor/run-context/run-1/step.root.md', + writesTo: [], + }, + }); + + expect(context.systemPrompt).not.toContain(''); + expect(context.systemPrompt.indexOf('')).toBeLessThan(context.systemPrompt.indexOf('')); + }); +}); + describe('validateStepAtoms', () => { it('allows zero or exactly one role', () => { expect(() => validateStepAtoms(makeStep())).not.toThrow(); diff --git a/src/main/harness-engine/context-builder.ts b/src/main/harness-engine/context-builder.ts index 681290a..66ba834 100644 --- a/src/main/harness-engine/context-builder.ts +++ b/src/main/harness-engine/context-builder.ts @@ -18,6 +18,19 @@ * assembled: the Single Persona rule (at most one role per step) and pairwise * mod compatibility (no two attached mods may be mutually `incompatibleWith` * or share an `exclusiveGroup`). + * + * `` (feedback-mode flows, spec: + * docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md): a 100% + * harness-generated reference block appended to the system prompt AFTER + * `` — topology summary, this step's assigned context + * file, and the briefing files it should write before finishing. Rendered + * only when the caller (executor.ts, feedback mode only) passes + * `options.flowAwareness`; entirely absent otherwise, so blind-mode output is + * byte-identical to before this block existed. It is reference material, not + * law — placed after, never inside, `` — and its + * CONTENTS are always harness-derived strings (paths, prompts already present + * on the step, counts); no other step's LLM-authored briefing text is ever + * injected here — those are read by the step itself via its own FS tools. */ import type { AgenticMod, AgenticStep } from '../../types/harness'; import type { RetrievedChunk } from './retriever'; @@ -28,6 +41,20 @@ export interface StepContext { preProcessOutputs: Array<{ modId: string; output: string }>; } +/** + * Feedback-mode flow-awareness data for one step, precomputed by the executor + * from the run's ContextManifest (context-manifest.ts) — this module only + * renders it, it never derives flow shape/paths itself. + */ +export interface FlowAwarenessInput { + /** Deterministic per-flow topology summary lines (see context-manifest.ts's buildTopologySummaryLines). */ + topologyLines: string[]; + /** This step's own assigned context file, rootDir-relative — reachable with the step's existing read_file tool. */ + assignedFile: string; + /** Context files (rootDir-relative) this step should write briefings into before finishing. Empty for exempt/terminal steps. */ + writesTo: string[]; +} + export interface BuildStepContextOptions { resolvePreProcessMod?: (mod: AgenticMod, step: AgenticStep) => Promise; onModStatus?: ( @@ -42,6 +69,13 @@ export interface BuildStepContextOptions { * the normal DAG context mechanism. */ injectedChunks?: RetrievedChunk[]; + /** + * Feedback-mode flow awareness (see FlowAwarenessInput). Undefined for + * every blind-mode run (the default) and for `retriever` steps (their + * branch never reaches buildStepContext with this set) — omission produces + * byte-identical output to a build with no knowledge of this feature. + */ + flowAwareness?: FlowAwarenessInput; } function stringifyConfig(config: Record | undefined): string { @@ -140,9 +174,42 @@ function buildExecutionConstraints(constraintOutputs: Array<{ modId: string; out ].join('\n'); } +/** + * Renders `` — reference material, not a constraint (see + * module header). Returns '' when `input` is undefined so the caller's + * `.filter(Boolean)` drops it entirely, leaving zero trace in the assembled + * system prompt for blind-mode runs. + */ +function buildFlowAwareness(input: FlowAwarenessInput | undefined): string { + if (!input) return ''; + + const lines = [ + '', + 'This flow runs in feedback mode: steps share context through files in the workspace, read/written with your existing file tools. This section is operational reference material, not a constraint — the execution_constraints above still take precedence.', + '', + 'Flow topology:', + ...input.topologyLines, + '', + `Your assigned context file: ${input.assignedFile}`, + 'Read it with read_file first if it might carry a useful briefing from an earlier step.', + ]; + + if (input.writesTo.length > 0) { + lines.push( + '', + 'Before you finish this step, use write_file to leave a short briefing in each of the following files, for the steps that depend on you:', + ...input.writesTo.map((path) => `- ${path}`), + ); + } + + lines.push(''); + return lines.join('\n'); +} + function buildSystemPrompt( step: AgenticStep, constraintOutputs: Array<{ modId: string; output: string }>, + flowAwareness?: FlowAwarenessInput, ): string { const roleLines = step.roles.length > 0 ? step.roles.map((role) => `[${role.name}]\n${role.systemPrompt}`).join('\n\n') @@ -160,6 +227,7 @@ function buildSystemPrompt( systemMods.length > 0 ? ['', ...systemMods, ''].join('\n') : '', + buildFlowAwareness(flowAwareness), ].filter(Boolean).join('\n\n'); } @@ -248,7 +316,7 @@ export async function buildStepContext( ].filter(Boolean).join('\n\n'); return { - systemPrompt: buildSystemPrompt(step, constraintOutputs), + systemPrompt: buildSystemPrompt(step, constraintOutputs, options.flowAwareness), userPrompt, preProcessOutputs, }; diff --git a/src/main/harness-engine/context-manifest.test.ts b/src/main/harness-engine/context-manifest.test.ts new file mode 100644 index 0000000..8ec5cb0 --- /dev/null +++ b/src/main/harness-engine/context-manifest.test.ts @@ -0,0 +1,438 @@ +/** + * context-manifest.test.ts — Unit tests for the Rosetta context manifest + * (spec: docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md). + * + * Pure, side-effect-free module — no filesystem, no executor. Covers every + * shape the spec enumerates: a linear chain, parallel (fan-out/fan-in) steps, + * a bounded loop (N iterations), a single-step flow, and a chain with a + * `retriever` step in the middle (exempt from briefing duties). + */ +import { describe, expect, it } from 'vitest'; +import type { AgenticFlow, AgenticStep } from '../../types/harness'; +import { buildExecutionPlan } from './loop-plan'; +import { + CONTEXT_MANIFEST_VERSION, + DEFAULT_CONTEXT_BUDGET_BYTES, + buildContextManifest, + buildTopologySummaryLines, + contextArtifactPathPattern, + contextRunDir, + contextRunSubdir, + manifestKeyForInstance, + relativeContextFilePath, + seedContextFiles, + summarizePurpose, +} from './context-manifest'; + +// --------------------------------------------------------------------------- +// Helpers (mirroring the exact style of loop-plan.test.ts / executor.test.ts) +// --------------------------------------------------------------------------- + +function makeStep( + id: string, + prevStepIds: string[], + nextStepIds: string[], + overrides: Partial = {}, +): AgenticStep { + return { + id, + type: 'llm_call', + prompt: `Prompt for ${id}`, + tools: [], + prevStepIds, + nextStepIds, + mods: [], + roles: [], + mentalContext: [], + ...overrides, + }; +} + +function makeFlow( + stepsRecord: Record, + rootStepId: string, + loops?: AgenticFlow['loops'], +): AgenticFlow { + return { + id: 'flow-test', + name: 'Test Flow', + rootStepId, + stepsRecord, + contextMode: 'feedback', + ...(loops ? { loops } : {}), + }; +} + +// --------------------------------------------------------------------------- +// contextRunSubdir / contextRunDir +// --------------------------------------------------------------------------- + +describe('contextRunSubdir / contextRunDir', () => { + it('builds the rootDir-relative run-context subdirectory', () => { + expect(contextRunSubdir('run-123')).toBe('.fluxor/run-context/run-123'); + }); + + it('joins rootDir with the run-context subdirectory', () => { + expect(contextRunDir('/workspace', 'run-123')).toBe('/workspace/.fluxor/run-context/run-123'); + }); + + it('strips a trailing slash from rootDir before joining', () => { + expect(contextRunDir('/workspace/', 'run-123')).toBe('/workspace/.fluxor/run-context/run-123'); + }); +}); + +// --------------------------------------------------------------------------- +// relativeContextFilePath / contextArtifactPathPattern +// --------------------------------------------------------------------------- + +describe('relativeContextFilePath', () => { + it('builds the rootDir-relative path a step\'s FS tools would use', () => { + expect(relativeContextFilePath('run-123', 'step.foo.md')).toBe('.fluxor/run-context/run-123/step.foo.md'); + }); +}); + +describe('contextArtifactPathPattern', () => { + it('anchors on the end of the path so it matches regardless of absolute prefix', () => { + const pattern = contextArtifactPathPattern('run-123', 'step.foo.md'); + expect(new RegExp(pattern).test('/workspace/.fluxor/run-context/run-123/step.foo.md')).toBe(true); + expect(new RegExp(pattern).test('/Users/dev/project/.fluxor/run-context/run-123/step.foo.md')).toBe(true); + expect(new RegExp(pattern).test('/workspace/.fluxor/run-context/OTHER-run/step.foo.md')).toBe(false); + expect(new RegExp(pattern).test('/workspace/.fluxor/run-context/run-123/step.bar.md')).toBe(false); + }); + + it('escapes regexp-special characters in runId/filename', () => { + const pattern = contextArtifactPathPattern('run.1+2', 'step.a.md'); + // A literal-dot runId must NOT match an arbitrary character in its place. + expect(new RegExp(pattern).test('/workspace/.fluxor/run-context/runX1X2/step.a.md')).toBe(false); + expect(new RegExp(pattern).test('/workspace/.fluxor/run-context/run.1+2/step.a.md')).toBe(true); + }); +}); + +// --------------------------------------------------------------------------- +// summarizePurpose +// --------------------------------------------------------------------------- + +describe('summarizePurpose', () => { + it('collapses newlines/whitespace into a single line', () => { + expect(summarizePurpose('Line one.\n\n Line two.\t\tLine three.')).toBe('Line one. Line two. Line three.'); + }); + + it('returns short prompts unchanged', () => { + expect(summarizePurpose('Write the landing page.')).toBe('Write the landing page.'); + }); + + it('truncates to at most 140 characters, ending with an ellipsis marker', () => { + const long = 'x'.repeat(200); + const summary = summarizePurpose(long); + expect(summary.length).toBeLessThanOrEqual(140); + expect(summary.endsWith('…')).toBe(true); + }); + + it('respects a custom maxChars', () => { + const summary = summarizePurpose('0123456789', 5); + expect(summary.length).toBeLessThanOrEqual(5); + }); +}); + +// --------------------------------------------------------------------------- +// buildContextManifest — chain (linear DAG, no loops) +// --------------------------------------------------------------------------- + +describe('buildContextManifest — linear chain', () => { + it('builds one entry per step keyed by bare stepId, with contextFile/purpose/readBy/writesTo', () => { + const stepsRecord = { + root: makeStep('root', [], ['middle'], { prompt: 'Kick off the chain.' }), + middle: makeStep('middle', ['root'], ['leaf'], { prompt: 'Do the middle work.' }), + leaf: makeStep('leaf', ['middle'], [], { prompt: 'Wrap it up.' }), + }; + const flow = makeFlow(stepsRecord, 'root'); + const plan = buildExecutionPlan(flow); + + const manifest = buildContextManifest(flow, plan, 'run-1'); + + expect(manifest.version).toBe(CONTEXT_MANIFEST_VERSION); + expect(manifest.runId).toBe('run-1'); + expect(manifest.flowId).toBe('flow-test'); + expect(manifest.contextMode).toBe('feedback'); + expect(manifest.budgetBytes).toBe(DEFAULT_CONTEXT_BUDGET_BYTES); + + expect(Object.keys(manifest.steps).sort()).toEqual(['leaf', 'middle', 'root']); + + expect(manifest.steps.root).toEqual({ + contextFile: 'step.root.md', + purpose: 'Kick off the chain.', + readBy: ['root'], + writesTo: ['step.middle.md'], + }); + expect(manifest.steps.middle).toEqual({ + contextFile: 'step.middle.md', + purpose: 'Do the middle work.', + readBy: ['middle'], + writesTo: ['step.leaf.md'], + }); + expect(manifest.steps.leaf).toEqual({ + contextFile: 'step.leaf.md', + purpose: 'Wrap it up.', + readBy: ['leaf'], + writesTo: [], // terminal step promises nothing downstream + }); + }); + + it('accepts a custom budget', () => { + const stepsRecord = { root: makeStep('root', [], []) }; + const flow = makeFlow(stepsRecord, 'root'); + const plan = buildExecutionPlan(flow); + + const manifest = buildContextManifest(flow, plan, 'run-budget', 4000); + expect(manifest.budgetBytes).toBe(4000); + }); +}); + +// --------------------------------------------------------------------------- +// buildContextManifest — parallel (same-Kahn-wave) steps +// --------------------------------------------------------------------------- + +describe('buildContextManifest — parallel steps (fan-out/fan-in)', () => { + it('fans a producer\'s writesTo out to every parallel consumer, and consumers do not write to each other', () => { + const stepsRecord = { + root: makeStep('root', [], ['branch-a', 'branch-b']), + 'branch-a': makeStep('branch-a', ['root'], ['merge']), + 'branch-b': makeStep('branch-b', ['root'], ['merge']), + merge: makeStep('merge', ['branch-a', 'branch-b'], []), + }; + const flow = makeFlow(stepsRecord, 'root'); + const plan = buildExecutionPlan(flow); + + const manifest = buildContextManifest(flow, plan, 'run-2'); + + // root fans out to BOTH parallel consumers' files. + expect(manifest.steps.root.writesTo.sort()).toEqual(['step.branch-a.md', 'step.branch-b.md']); + // Parallel siblings never write to each other (no edge between them). + expect(manifest.steps['branch-a'].writesTo).toEqual(['step.merge.md']); + expect(manifest.steps['branch-b'].writesTo).toEqual(['step.merge.md']); + expect(manifest.steps.merge.writesTo).toEqual([]); + }); +}); + +// --------------------------------------------------------------------------- +// buildContextManifest — bounded loop (N iterations) +// --------------------------------------------------------------------------- + +describe('buildContextManifest — bounded loop', () => { + it('keys loop-body instances as stepId#iterN and derives per-iteration writesTo from the expanded plan', () => { + const stepsRecord = { + root: makeStep('root', [], ['b1']), + b1: makeStep('b1', ['root'], ['b2']), + b2: makeStep('b2', ['b1'], ['down']), + down: makeStep('down', ['b2'], []), + }; + const flow = makeFlow(stepsRecord, 'root', [ + { id: 'loop-1', sourceStepId: 'b2', targetStepId: 'b1', maxIterations: 3 }, + ]); + const plan = buildExecutionPlan(flow); + + const manifest = buildContextManifest(flow, plan, 'run-loop'); + + // Non-loop steps keep bare keys; loop-body steps get #iterN — every pass, + // including iteration 1. + expect(Object.keys(manifest.steps).sort()).toEqual([ + 'b1#iter1', 'b1#iter2', 'b1#iter3', + 'b2#iter1', 'b2#iter2', 'b2#iter3', + 'down', 'root', + ].sort()); + + // The internal loop-plan `@` namespace never leaks into manifest keys. + for (const key of Object.keys(manifest.steps)) { + expect(key).not.toContain('@'); + } + + expect(manifest.steps['b1#iter1'].contextFile).toBe('step.b1.iter1.md'); + expect(manifest.steps['b1#iter2'].contextFile).toBe('step.b1.iter2.md'); + expect(manifest.steps['b2#iter3'].contextFile).toBe('step.b2.iter3.md'); + + // readBy always names the real stepId (not the iteration-suffixed key). + expect(manifest.steps['b1#iter2'].readBy).toEqual(['b1']); + + // Chain edge: b1's pass k writes to b2's pass k (same iteration, same body). + expect(manifest.steps['b1#iter1'].writesTo).toEqual(['step.b2.iter1.md']); + expect(manifest.steps['b1#iter2'].writesTo).toEqual(['step.b2.iter2.md']); + // b2's pass k (k<3) re-triggers b1's pass k+1 — the chain edge feeds b1's NEXT iteration file. + expect(manifest.steps['b2#iter1'].writesTo).toEqual(['step.b1.iter2.md']); + expect(manifest.steps['b2#iter2'].writesTo).toEqual(['step.b1.iter3.md']); + // Exit edge: only the LAST pass of the tail connects onward, to the + // (non-looped) downstream step's plain file. + expect(manifest.steps['b2#iter3'].writesTo).toEqual(['step.down.md']); + + // root (pre-loop) and down (post-loop) sit outside the loop body. + expect(manifest.steps.root.contextFile).toBe('step.root.md'); + expect(manifest.steps.down.contextFile).toBe('step.down.md'); + expect(manifest.steps.root.writesTo).toEqual(['step.b1.iter1.md']); + }); +}); + +// --------------------------------------------------------------------------- +// buildContextManifest — single-step flow +// --------------------------------------------------------------------------- + +describe('buildContextManifest — single-step flow', () => { + it('produces exactly one entry with no writesTo', () => { + const stepsRecord = { solo: makeStep('solo', [], [], { prompt: 'Do the whole thing alone.' }) }; + const flow = makeFlow(stepsRecord, 'solo'); + const plan = buildExecutionPlan(flow); + + const manifest = buildContextManifest(flow, plan, 'run-solo'); + + expect(Object.keys(manifest.steps)).toEqual(['solo']); + expect(manifest.steps.solo).toEqual({ + contextFile: 'step.solo.md', + purpose: 'Do the whole thing alone.', + readBy: ['solo'], + writesTo: [], + }); + }); +}); + +// --------------------------------------------------------------------------- +// buildContextManifest — chain with a retriever step in the middle +// --------------------------------------------------------------------------- + +describe('buildContextManifest — retriever exemption', () => { + it('marks a retriever-type step exempt with writesTo: [] and empty readBy, regardless of its real DAG edges', () => { + const stepsRecord = { + root: makeStep('root', [], ['fetch']), + fetch: makeStep('fetch', ['root'], ['use'], { type: 'retriever' }), + use: makeStep('use', ['fetch'], []), + }; + const flow = makeFlow(stepsRecord, 'root'); + const plan = buildExecutionPlan(flow); + + const manifest = buildContextManifest(flow, plan, 'run-retriever'); + + expect(manifest.steps.fetch.exempt).toBe('retriever'); + expect(manifest.steps.fetch.writesTo).toEqual([]); + expect(manifest.steps.fetch.readBy).toEqual([]); + // It still gets a contextFile (seeded, per spec) even though nobody reads it. + expect(manifest.steps.fetch.contextFile).toBe('step.fetch.md'); + + // The retriever's own upstream producer is unaffected — it still promises + // a briefing to the retriever's file (mechanically derived from the DAG; + // the retriever just never reads it). + expect(manifest.steps.root.writesTo).toEqual(['step.fetch.md']); + + // Non-exempt steps elsewhere in the same flow are untouched. + expect(manifest.steps.use.exempt).toBeUndefined(); + expect(manifest.steps.use.readBy).toEqual(['use']); + }); +}); + +// --------------------------------------------------------------------------- +// manifestKeyForInstance +// --------------------------------------------------------------------------- + +describe('manifestKeyForInstance', () => { + it('returns the bare stepId for a non-loop instance', () => { + expect(manifestKeyForInstance({ stepId: 'root', iteration: 1 })).toBe('root'); + }); + + it('returns stepId#iterN for a loop-body instance, even iteration 1', () => { + expect(manifestKeyForInstance({ + stepId: 'b1', iteration: 1, loop: { id: 'loop-1', totalIterations: 3 }, + })).toBe('b1#iter1'); + expect(manifestKeyForInstance({ + stepId: 'b1', iteration: 2, loop: { id: 'loop-1', totalIterations: 3 }, + })).toBe('b1#iter2'); + }); +}); + +// --------------------------------------------------------------------------- +// seedContextFiles +// --------------------------------------------------------------------------- + +describe('seedContextFiles', () => { + it('seeds one file per manifest entry, keyed by bare filename, with a deterministic header + purpose', () => { + const stepsRecord = { + root: makeStep('root', [], ['leaf'], { prompt: 'Kick off the chain.' }), + leaf: makeStep('leaf', ['root'], [], { prompt: 'Wrap it up.' }), + }; + const flow = makeFlow(stepsRecord, 'root'); + const plan = buildExecutionPlan(flow); + const manifest = buildContextManifest(flow, plan, 'run-seed'); + + const seeds = seedContextFiles(manifest); + + expect(Object.keys(seeds).sort()).toEqual(['step.leaf.md', 'step.root.md']); + expect(seeds['step.root.md']).toContain('# Contexto para root'); + expect(seeds['step.root.md']).toContain('Kick off the chain.'); + expect(seeds['step.leaf.md']).toContain('# Contexto para leaf'); + expect(seeds['step.leaf.md']).toContain('Wrap it up.'); + }); + + it('annotates a loop-body seed with its iteration/total for human readability', () => { + const stepsRecord = { + root: makeStep('root', [], ['b1']), + b1: makeStep('b1', ['root'], ['b2']), + b2: makeStep('b2', ['b1'], []), + }; + const flow = makeFlow(stepsRecord, 'root', [ + { id: 'loop-1', sourceStepId: 'b2', targetStepId: 'b1', maxIterations: 2 }, + ]); + const plan = buildExecutionPlan(flow); + const manifest = buildContextManifest(flow, plan, 'run-seed-loop'); + + const seeds = seedContextFiles(manifest); + expect(seeds['step.b1.iter1.md']).toContain('iteración 1/2'); + expect(seeds['step.b1.iter2.md']).toContain('iteración 2/2'); + }); +}); + +// --------------------------------------------------------------------------- +// buildTopologySummaryLines +// --------------------------------------------------------------------------- + +describe('buildTopologySummaryLines', () => { + it('produces one line per REAL step (not per loop iteration) in canonical topo order', () => { + const stepsRecord = { + root: makeStep('root', [], ['b1'], { prompt: 'Root prompt.' }), + b1: makeStep('b1', ['root'], ['b2'], { prompt: 'Loop body step one.' }), + b2: makeStep('b2', ['b1'], ['down'], { prompt: 'Loop body step two.' }), + down: makeStep('down', ['b2'], [], { prompt: 'Final step.' }), + }; + const flow = makeFlow(stepsRecord, 'root', [ + { id: 'loop-1', sourceStepId: 'b2', targetStepId: 'b1', maxIterations: 5 }, + ]); + const plan = buildExecutionPlan(flow); + + const lines = buildTopologySummaryLines(flow, plan); + + // Exactly 4 lines — one per real step, NOT one per the 12 expanded instances. + expect(lines).toHaveLength(4); + expect(lines[0]).toContain('root'); + expect(lines[0]).toContain('Root prompt.'); + expect(lines[1]).toContain('b1'); + expect(lines[1]).toMatch(/loop.*5|5.*loop/i); + expect(lines[3]).toContain('down'); + }); + + it('truncates the listing and appends a summary line when the flow has more real steps than maxLines', () => { + const stepsRecord: Record = {}; + for (let i = 0; i < 40; i++) { + const id = `s${String(i).padStart(2, '0')}`; + const next = i < 39 ? [`s${String(i + 1).padStart(2, '0')}`] : []; + const prev = i > 0 ? [`s${String(i - 1).padStart(2, '0')}`] : []; + stepsRecord[id] = makeStep(id, prev, next); + } + const flow = makeFlow(stepsRecord, 's00'); + const plan = buildExecutionPlan(flow); + + const lines = buildTopologySummaryLines(flow, plan, 10); + + expect(lines.length).toBeLessThanOrEqual(10); + expect(lines.at(-1)).toMatch(/more step/i); + }); + + it('returns an empty array for a flow with zero steps worth summarizing is unreachable (at least root exists) — single-step flow yields one line', () => { + const flow = makeFlow({ solo: makeStep('solo', [], []) }, 'solo'); + const plan = buildExecutionPlan(flow); + expect(buildTopologySummaryLines(flow, plan)).toHaveLength(1); + }); +}); diff --git a/src/main/harness-engine/context-manifest.ts b/src/main/harness-engine/context-manifest.ts new file mode 100644 index 0000000..7e72ee7 --- /dev/null +++ b/src/main/harness-engine/context-manifest.ts @@ -0,0 +1,304 @@ +/** + * context-manifest.ts — Rosetta context manifest (feedback-mode flows) + * + * Spec (LAW): docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md + * + * Pure and side-effect free by design (same discipline as loop-plan.ts) — the + * executor is the only caller that performs I/O (creating the run-context + * directory, writing manifest.json, seeding + reading context files); this + * module only computes shapes and strings. + * + * Convention (LAW, see spec "Ubicación y convención de nombres"): everything + * lives under `/.fluxor/run-context//`: + * - manifest.json — this module's ContextManifest, serialized. + * - step..md — a non-loop step's assigned context file. + * - step..iter.md — a loop-body step's per-iteration file (N from 1). + * + * Two deliberately distinct key namespaces (spec "Notas de implementación"): + * - loop-plan.ts's `instanceKey()` uses `stepId@iteration` internally. + * - This module's manifest is keyed by `stepId` (non-loop) or + * `stepId#iterN` (loop-body, every pass including iteration 1). + * `manifestKeyForInstance` is the ONLY place that translates between them — + * the `@` form never leaks into the manifest, seeds, or prompts. + * + * Retriever exemption (spec "Steps retriever"): the executor's retriever + * branch returns before the LLM path — it receives no tools, no system + * prompt, no ``. A retriever-type step's manifest entry gets + * `writesTo: []` (no promises ⇒ no guardrail obligation) and `readBy: []` + * (nobody is instructed to read it), tagged `exempt: 'retriever'` so a + * consumer reading the manifest never expects a briefing that will never + * arrive. Its `contextFile` still exists (seeded at genesis) — an upstream + * producer whose forward edge targets it still legitimately promises a + * briefing there (mechanically derived from the DAG like any other edge); + * the retriever simply never reads it, same as today. + */ +import type { AgenticFlow } from '../../types/harness'; +import { topoSortAgenticSteps } from '../../types/harness'; +import type { ExecutionPlan, StepInstance } from './loop-plan'; + +// --------------------------------------------------------------------------- +// Constants +// --------------------------------------------------------------------------- + +/** `ContextManifest.version` — bump only on a breaking schema change. */ +export const CONTEXT_MANIFEST_VERSION = 1; + +/** Default per-context-file byte budget (spec "Presupuesto" decision — ~3k tokens). */ +export const DEFAULT_CONTEXT_BUDGET_BYTES = 12_000; + +/** Rosetta convention root, relative to the run's effective rootDir. */ +const CONTEXT_RUN_SUBDIR = '.fluxor/run-context'; + +/** Max chars for a manifest entry's `purpose` (spec "Presupuesto" decision). */ +const PURPOSE_MAX_CHARS = 140; + +/** Default cap on `` topology lines (spec "Presupuesto" decision — "≤~30 líneas"). */ +const DEFAULT_TOPOLOGY_MAX_LINES = 30; + +const ELLIPSIS = '…'; + +// --------------------------------------------------------------------------- +// Path helpers +// --------------------------------------------------------------------------- + +/** Path (relative to the run's rootDir) of the run's context directory — e.g. ".fluxor/run-context/". */ +export function contextRunSubdir(runId: string): string { + return `${CONTEXT_RUN_SUBDIR}/${runId}`; +} + +/** + * Absolute-ish path (rootDir + contextRunSubdir) the executor uses for genesis + * I/O (mkdir/writeFile) via its fileSystem adapter (real fs or an injected + * VFS). `rootDir` is the run's already-materialized EFFECTIVE root (see + * executor.ts) — this function does no normalization of its own. + */ +export function contextRunDir(rootDir: string, runId: string): string { + return `${rootDir.replace(/\/+$/, '')}/${contextRunSubdir(runId)}`; +} + +/** + * The rootDir-relative path a step's OWN read_file/write_file tools would use + * to reach a bare context filename (e.g. a manifest entry's `contextFile` or + * one of its `writesTo` entries) — those tools resolve paths relative to the + * overall project rootDir, not relative to the run-context directory. + */ +export function relativeContextFilePath(runId: string, filename: string): string { + return `${contextRunSubdir(runId)}/${filename}`; +} + +function escapeRegExp(value: string): string { + return value.replace(/[.*+?^${}()|[\]\\]/g, '\\$&'); +} + +/** + * A regexp STRING (matched case-insensitively, mirroring every other + * `StepArtifactRequirement.pathPattern` in this codebase) that matches the + * END of a workspace path for this run-context file — anchored with `$` so + * it matches regardless of the absolute prefix a snapshot uses (a real + * filesystem's absolute path vs. the PF sandbox VFS's virtual `/workspace` + * root), without needing to know which one is in play. + */ +export function contextArtifactPathPattern(runId: string, filename: string): string { + return `${escapeRegExp(relativeContextFilePath(runId, filename))}$`; +} + +// --------------------------------------------------------------------------- +// Purpose summarization +// --------------------------------------------------------------------------- + +/** + * Collapses a step prompt into a single-line summary truncated to at most + * `maxChars` characters (spec "Presupuesto" decision: "resumen 1 línea del + * prompt del step, truncado a 140 chars"). When truncation is necessary the + * result ends with a single ellipsis character, and the TOTAL length + * (including that marker) never exceeds `maxChars`. + */ +export function summarizePurpose(prompt: string, maxChars: number = PURPOSE_MAX_CHARS): string { + const collapsed = prompt.replace(/\s+/g, ' ').trim(); + if (collapsed.length <= maxChars) return collapsed; + const cut = Math.max(0, maxChars - ELLIPSIS.length); + return `${collapsed.slice(0, cut)}${ELLIPSIS}`; +} + +// --------------------------------------------------------------------------- +// Manifest key / context-file naming (the ONLY place that bridges loop-plan's +// `stepId@iteration` namespace and this module's `stepId[#iterN]` namespace) +// --------------------------------------------------------------------------- + +/** `stepId` for a non-loop instance; `stepId#iterN` for a loop-body instance (every pass, including 1). */ +export function manifestKeyForInstance(instance: Pick): string { + return instance.loop ? `${instance.stepId}#iter${instance.iteration}` : instance.stepId; +} + +/** `step..md` for a non-loop instance; `step..iter.md` for a loop-body instance. */ +function contextFileNameForInstance(instance: Pick): string { + return instance.loop + ? `step.${instance.stepId}.iter${instance.iteration}.md` + : `step.${instance.stepId}.md`; +} + +// --------------------------------------------------------------------------- +// ContextManifest schema (spec "Schema del manifest") +// --------------------------------------------------------------------------- + +export interface ContextManifestStepEntry { + /** Bare filename (no directory) — resolve via relativeContextFilePath/contextRunDir. */ + contextFile: string; + /** 1-line summary of the step's prompt, truncated to PURPOSE_MAX_CHARS. */ + purpose: string; + /** Real stepIds instructed to read `contextFile`; empty for exempt steps. */ + readBy: string[]; + /** Bare filenames (other steps' contextFiles) this instance is expected to write briefings into. */ + writesTo: string[]; + /** Present only for steps exempt from briefing duties (v1: retriever steps). */ + exempt?: 'retriever'; +} + +export interface ContextManifest { + version: typeof CONTEXT_MANIFEST_VERSION; + runId: string; + flowId: string; + contextMode: 'feedback'; + budgetBytes: number; + /** Keyed by manifestKeyForInstance() — bare stepId, or stepId#iterN for loop-body instances. */ + steps: Record; +} + +// --------------------------------------------------------------------------- +// buildContextManifest +// --------------------------------------------------------------------------- + +/** + * Builds the Rosetta manifest for one feedback-mode run from the flow, its + * EXPANDED execution plan (loop-plan.ts's buildExecutionPlan — `writesTo` + * derives from the expanded DAG's `nextKeys`, per spec, so loop-body + * `writesTo` correctly targets the NEXT iteration's file, not last-write-wins + * across passes), and the run's id. + */ +export function buildContextManifest( + flow: AgenticFlow, + plan: ExecutionPlan, + runId: string, + budgetBytes: number = DEFAULT_CONTEXT_BUDGET_BYTES, +): ContextManifest { + const steps: Record = {}; + + for (const instance of plan.instances.values()) { + const step = flow.stepsRecord[instance.stepId]; + const isRetriever = step?.type === 'retriever'; + const key = manifestKeyForInstance(instance); + + const writesTo = isRetriever + ? [] + : (plan.nextKeys.get(instance.key) ?? []).map((nextKey) => { + const nextInstance = plan.instances.get(nextKey)!; + return contextFileNameForInstance(nextInstance); + }); + + steps[key] = { + contextFile: contextFileNameForInstance(instance), + purpose: summarizePurpose(step?.prompt ?? ''), + readBy: isRetriever ? [] : [instance.stepId], + writesTo, + ...(isRetriever ? { exempt: 'retriever' as const } : {}), + }; + } + + return { + version: CONTEXT_MANIFEST_VERSION, + runId, + flowId: flow.id, + contextMode: 'feedback', + budgetBytes, + steps, + }; +} + +// --------------------------------------------------------------------------- +// seedContextFiles — deterministic header content, keyed by bare filename +// --------------------------------------------------------------------------- + +/** + * Deterministic seed content for every context file in `manifest` — a header + * + the step's purpose (spec "Génesis y ciclo de vida": "sembrar cada + * contextFile con un header determinista... para que 'leer tu fichero' nunca + * falle por inexistencia"). Keyed by bare filename (as it appears in + * `contextFile`/`writesTo`) so the executor can write each entry directly + * under `contextRunDir(rootDir, runId)`. + */ +export function seedContextFiles(manifest: ContextManifest): Record { + const seeds: Record = {}; + + for (const [key, entry] of Object.entries(manifest.steps)) { + const iterMatch = /#iter(\d+)$/.exec(key); + const realStepId = iterMatch ? key.slice(0, iterMatch.index) : key; + const iterationSuffix = iterMatch + ? ` (iteración ${iterMatch[1]}/${totalIterationsFor(manifest, realStepId)})` + : ''; + + seeds[entry.contextFile] = [ + `# Contexto para ${realStepId}${iterationSuffix}`, + '', + entry.purpose, + '', + ].join('\n'); + } + + return seeds; +} + +/** Total iterations for `stepId`'s loop, read off any of its own manifest entries. */ +function totalIterationsFor(manifest: ContextManifest, stepId: string): number { + let max = 1; + for (const key of Object.keys(manifest.steps)) { + const match = /^(.*)#iter(\d+)$/.exec(key); + if (match && match[1] === stepId) { + max = Math.max(max, Number(match[2])); + } + } + return max; +} + +// --------------------------------------------------------------------------- +// buildTopologySummaryLines — the topology block +// --------------------------------------------------------------------------- + +/** + * One line per REAL step (canonical topo order via topoSortAgenticSteps — + * NOT one line per expanded loop instance, so a flow with few steps but many + * loop passes still yields a short, readable summary), truncated to at most + * `maxLines` entries with a trailing "...and N more step(s)" summary line + * when the flow has more real steps than that (spec "Presupuesto" decision: + * "el harness trunca nombres/propósitos largos" when injecting topology). + */ +export function buildTopologySummaryLines( + flow: AgenticFlow, + plan: ExecutionPlan, + maxLines: number = DEFAULT_TOPOLOGY_MAX_LINES, +): string[] { + const ordered = topoSortAgenticSteps(Object.values(flow.stepsRecord)); + + const loopInfoByStepId = new Map(); + for (const instance of plan.instances.values()) { + if (instance.loop && !loopInfoByStepId.has(instance.stepId)) { + loopInfoByStepId.set(instance.stepId, { totalIterations: instance.loop.totalIterations }); + } + } + + const cap = Math.max(1, maxLines); + const shown = ordered.length > cap ? ordered.slice(0, cap - 1) : ordered; + + const lines = shown.map((step) => { + const loopInfo = loopInfoByStepId.get(step.id); + const exemptSuffix = step.type === 'retriever' ? ' [retriever, exempt from briefings]' : ''; + const loopSuffix = loopInfo ? ` [loops ×${loopInfo.totalIterations}]` : ''; + return `- ${step.id}: ${summarizePurpose(step.prompt)}${loopSuffix}${exemptSuffix}`; + }); + + if (ordered.length > cap) { + const remaining = ordered.length - shown.length; + lines.push(`- …and ${remaining} more step(s) not shown.`); + } + + return lines; +} diff --git a/src/main/harness-engine/executor.test.ts b/src/main/harness-engine/executor.test.ts index bdfc443..953b148 100644 --- a/src/main/harness-engine/executor.test.ts +++ b/src/main/harness-engine/executor.test.ts @@ -1,3 +1,6 @@ +import { mkdtemp, readFile as nodeReadFile, rm } from 'fs/promises'; +import { tmpdir } from 'os'; +import { join } from 'path'; import { afterEach, describe, expect, it, vi } from 'vitest'; import type { AgenticFlow, AgenticStep } from '@/types/harness'; import type { HarnessEventPayload } from '@/types/ipc-events'; @@ -827,3 +830,421 @@ describe('executeAgenticFlow — resolveConnection threading (Phase 6 provider c expect(seenResolveConnection).toBeUndefined(); }); }); + +// --------------------------------------------------------------------------- +// contextMode: 'feedback' — Rosetta context system (spec: +// docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md) +// --------------------------------------------------------------------------- + +/** Minimal in-memory McpFileSystem — enough to back genesis, the model's own + * FS tools, and the guardrail's snapshotWorkspace walk, without touching disk. */ +function makeInMemoryFileSystem(): McpFileSystem & { files: Map } { + const files = new Map(); + const dirs = new Set(['/workspace']); + + const normalize = (p: string) => p.replace(/\/{2,}/g, '/').replace(/\/$/, '') || '/'; + + return { + files, + mkdir: async (path: string) => { + dirs.add(normalize(path)); + }, + readdir: async (path: string) => { + const prefix = `${normalize(path)}/`; + const names = new Set(); + for (const p of files.keys()) { + if (p.startsWith(prefix)) names.add(p.slice(prefix.length).split('/')[0]); + } + for (const d of dirs) { + if (d.startsWith(prefix) && d !== prefix) names.add(d.slice(prefix.length).split('/')[0]); + } + return [...names]; + }, + readFile: async (path: string) => { + const content = files.get(normalize(path)); + if (content === undefined) throw new Error(`ENOENT: no such file: ${path}`); + return content; + }, + stat: async (path: string) => { + const norm = normalize(path); + const isDir = dirs.has(norm) || [...files.keys()].some((p) => p.startsWith(`${norm}/`)); + const content = files.get(norm); + return { + isDirectory: () => isDir, + isFile: () => content !== undefined, + size: content !== undefined ? Buffer.byteLength(content, 'utf-8') : 0, + }; + }, + writeFile: async (path: string, content: string) => { + files.set(normalize(path), content); + }, + }; +} + +function makeFeedbackFlow( + stepsRecord: Record, + rootStepId: string, +): AgenticFlow { + return { + id: 'flow-feedback', + name: 'Feedback Flow', + rootStepId, + stepsRecord, + contextMode: 'feedback', + }; +} + +describe('executeAgenticFlow — contextMode: blind (default) has zero side effects', () => { + afterEach(() => { + harnessEventBus.removeAllListeners(); + }); + + it('never touches the fileSystem for any .fluxor path when contextMode is absent', async () => { + const fs = makeInMemoryFileSystem(); + const flow: AgenticFlow = { + id: 'flow-blind-check', + name: 'Blind Check', + rootStepId: 'root', + stepsRecord: { root: makeStep('root', [], []) }, + }; + + await executeAgenticFlow(flow, { + rootDir: '/workspace', + fileSystem: fs, + runStep: async ({ step }) => ({ text: `output:${step.id}`, usage: null, toolCalls: [], toolResults: [] }), + }); + + expect([...fs.files.keys()].some((path) => path.includes('.fluxor'))).toBe(false); + }); + + it('never touches the fileSystem for any .fluxor path when contextMode is explicitly "blind"', async () => { + const fs = makeInMemoryFileSystem(); + const flow: AgenticFlow = { + id: 'flow-blind-explicit', + name: 'Blind Explicit', + rootStepId: 'root', + stepsRecord: { root: makeStep('root', [], []) }, + contextMode: 'blind', + }; + + await executeAgenticFlow(flow, { + rootDir: '/workspace', + fileSystem: fs, + runStep: async ({ step }) => ({ text: `output:${step.id}`, usage: null, toolCalls: [], toolResults: [] }), + }); + + expect([...fs.files.keys()].some((path) => path.includes('.fluxor'))).toBe(false); + }); + + it('a blind flow with NO fileSystem/rootDir at all runs exactly as before (no crash, no genesis)', async () => { + const flow: AgenticFlow = { + id: 'flow-blind-bare', + name: 'Blind Bare', + rootStepId: 'root', + stepsRecord: { root: makeStep('root', [], []) }, + }; + + await expect(executeAgenticFlow(flow, { + runStep: async ({ step }) => ({ text: `output:${step.id}`, usage: null, toolCalls: [], toolResults: [] }), + })).resolves.toBeUndefined(); + }); +}); + +describe('executeAgenticFlow — contextMode: feedback — genesis', () => { + afterEach(() => { + harnessEventBus.removeAllListeners(); + }); + + it('creates manifest.json + seeded step.*.md files under .fluxor/run-context// before any step runs', async () => { + const fs = makeInMemoryFileSystem(); + const flow = makeFeedbackFlow( + { + root: makeStep('root', [], ['leaf']), + leaf: makeStep('leaf', ['root'], []), + }, + 'root', + ); + + const runId = 'run-genesis-1'; + await executeAgenticFlow(flow, { + rootDir: '/workspace', + fileSystem: fs, + runId, + runStep: async ({ step }) => ({ text: `output:${step.id}`, usage: null, toolCalls: [], toolResults: [] }), + }); + + const manifestRaw = fs.files.get(`/workspace/.fluxor/run-context/${runId}/manifest.json`); + expect(manifestRaw).toBeDefined(); + const manifest = JSON.parse(manifestRaw!); + expect(manifest.runId).toBe(runId); + expect(manifest.flowId).toBe('flow-feedback'); + expect(manifest.contextMode).toBe('feedback'); + expect(Object.keys(manifest.steps).sort()).toEqual(['leaf', 'root']); + + expect(fs.files.get(`/workspace/.fluxor/run-context/${runId}/step.root.md`)).toContain('# Contexto para root'); + expect(fs.files.get(`/workspace/.fluxor/run-context/${runId}/step.leaf.md`)).toContain('# Contexto para leaf'); + }); + + it('uses the caller-supplied fileSystem for genesis when provided (no real disk touched)', async () => { + const fs = makeInMemoryFileSystem(); + const flow = makeFeedbackFlow({ root: makeStep('root', [], []) }, 'root'); + + await executeAgenticFlow(flow, { + rootDir: '/workspace', + fileSystem: fs, + runId: 'run-genesis-2', + runStep: async ({ step }) => ({ text: `output:${step.id}`, usage: null, toolCalls: [], toolResults: [] }), + }); + + expect(fs.files.has('/workspace/.fluxor/run-context/run-genesis-2/manifest.json')).toBe(true); + }); + + it('with NO options.fileSystem, synthesizes a real-FS adapter and writes to a real rootDir on disk', async () => { + const tmpDir = await mkdtemp(join(tmpdir(), 'fluxor-context-genesis-')); + try { + const flow = makeFeedbackFlow({ root: makeStep('root', [], []) }, 'root'); + + await executeAgenticFlow(flow, { + rootDir: tmpDir, + runId: 'run-genesis-real-fs', + runStep: async ({ step }) => ({ text: `output:${step.id}`, usage: null, toolCalls: [], toolResults: [] }), + }); + + const manifestPath = join(tmpDir, '.fluxor', 'run-context', 'run-genesis-real-fs', 'manifest.json'); + const manifestRaw = await nodeReadFile(manifestPath, 'utf-8'); + const manifest = JSON.parse(manifestRaw); + expect(manifest.runId).toBe('run-genesis-real-fs'); + + const seedPath = join(tmpDir, '.fluxor', 'run-context', 'run-genesis-real-fs', 'step.root.md'); + await expect(nodeReadFile(seedPath, 'utf-8')).resolves.toContain('# Contexto para root'); + } finally { + await rm(tmpDir, { recursive: true, force: true }); + } + }); + + it('with NO options.rootDir either, defaults the real-FS adapter to process.cwd() — not guardrails.ts\'s independent "/workspace" default', async () => { + // Regression guard for the rootDir-default mismatch the spec calls out: + // guardrails.ts's snapshotWorkspace defaults to '/workspace' while the + // model's own FS tools (mcp-adapter.ts's normalizeRootDir) default to + // process.cwd() — genesis must agree with the LATTER. process.cwd() is + // mocked to a disposable temp dir so this test never touches the real + // repository working directory. + const tmpDir = await mkdtemp(join(tmpdir(), 'fluxor-context-cwd-default-')); + const cwdSpy = vi.spyOn(process, 'cwd').mockReturnValue(tmpDir); + try { + const flow = makeFeedbackFlow({ root: makeStep('root', [], []) }, 'root'); + + await expect(executeAgenticFlow(flow, { + runId: 'run-genesis-cwd-default', + runStep: async ({ step }) => ({ text: `output:${step.id}`, usage: null, toolCalls: [], toolResults: [] }), + })).resolves.toBeUndefined(); + + // Landed under the mocked cwd, proving it did NOT fall back to '/workspace'. + const manifestPath = join(tmpDir, '.fluxor', 'run-context', 'run-genesis-cwd-default', 'manifest.json'); + await expect(nodeReadFile(manifestPath, 'utf-8')).resolves.toContain('"runId"'); + } finally { + cwdSpy.mockRestore(); + await rm(tmpDir, { recursive: true, force: true }); + } + }); +}); + +describe('executeAgenticFlow — contextMode: feedback — wiring', () => { + afterEach(() => { + harnessEventBus.removeAllListeners(); + }); + + it('passes (topology + assigned file + writesTo) to the LLM-path step\'s systemPrompt', async () => { + const fs = makeInMemoryFileSystem(); + const flow = makeFeedbackFlow( + { + root: makeStep('root', [], ['leaf']), + leaf: makeStep('leaf', ['root'], []), + }, + 'root', + ); + + const seenSystemPrompts: Record = {}; + await executeAgenticFlow(flow, { + rootDir: '/workspace', + fileSystem: fs, + runId: 'run-awareness-1', + runStep: async ({ step, systemPrompt }) => { + seenSystemPrompts[step.id] = systemPrompt; + return { text: `output:${step.id}`, usage: null, toolCalls: [], toolResults: [] }; + }, + }); + + expect(seenSystemPrompts.root).toContain(''); + expect(seenSystemPrompts.root).toContain('.fluxor/run-context/run-awareness-1/step.root.md'); + expect(seenSystemPrompts.root).toContain('.fluxor/run-context/run-awareness-1/step.leaf.md'); + // leaf is terminal — no writesTo, so no "write a briefing" targets, but it + // still gets the block (its own assigned file + topology). + expect(seenSystemPrompts.leaf).toContain(''); + expect(seenSystemPrompts.leaf).toContain('.fluxor/run-context/run-awareness-1/step.leaf.md'); + }); + + it('a retriever step never receives (its branch returns before the LLM path)', async () => { + const fs = makeInMemoryFileSystem(); + const flow = makeFeedbackFlow( + { + root: makeStep('root', [], ['fetch']), + fetch: makeStep('fetch', ['root'], ['use']), + use: makeStep('use', ['fetch'], []), + }, + 'root', + ); + flow.stepsRecord.fetch.type = 'retriever'; + + const seenSystemPrompts: Record = {}; + await executeAgenticFlow(flow, { + rootDir: '/workspace', + fileSystem: fs, + runId: 'run-awareness-retriever', + runStep: async ({ step, systemPrompt }) => { + seenSystemPrompts[step.id] = systemPrompt; + return { text: `output:${step.id}`, usage: null, toolCalls: [], toolResults: [] }; + }, + }); + + // The retriever never calls runStep at all (its branch bypasses it), so + // it simply has no entry — proving the LLM path (and its + // ) was never reached for it. + expect(seenSystemPrompts.fetch).toBeUndefined(); + // Its seeded file still exists (genesis seeds every contextFile). + expect(fs.files.get('/workspace/.fluxor/run-context/run-awareness-retriever/step.fetch.md')) + .toContain('# Contexto para fetch'); + // Downstream (non-exempt) steps are unaffected. + expect(seenSystemPrompts.use).toContain(''); + }); +}); + +describe('executeAgenticFlow — contextMode: feedback — guardrail enforces promised briefings', () => { + afterEach(() => { + harnessEventBus.removeAllListeners(); + }); + + it('passes silently when the step writes a real briefing (beyond the seed) into its writesTo target', async () => { + const fs = makeInMemoryFileSystem(); + const flow = makeFeedbackFlow( + { + root: makeStep('root', [], ['leaf']), + leaf: makeStep('leaf', ['root'], []), + }, + 'root', + ); + + const events: HarnessEventPayload[] = []; + harnessEventBus.on(HARNESS_EVENT_NAME, (e) => events.push(e)); + + await executeAgenticFlow(flow, { + rootDir: '/workspace', + fileSystem: fs, + runId: 'run-guardrail-pass', + guardrailMaxAttempts: 2, + runStep: async ({ step, tools }) => { + if (step.id === 'root') { + await tools.write_file.execute?.( + { path: '.fluxor/run-context/run-guardrail-pass/step.leaf.md', content: 'A real, detailed briefing for leaf that clearly exceeds the tiny seeded header in length.' }, + { toolCallId: 'tool-1', messages: [] }, + ); + } + return { text: `output:${step.id}`, usage: null, toolCalls: [], toolResults: [] }; + }, + }); + + const stepEvents = events.filter( + (e): e is Extract => e.type === 'StepStatusChanged', + ); + expect(stepEvents.some((e) => e.logs?.includes('breached contract'))).toBe(false); + }); + + it('retries then breaches when the step never writes into its writesTo target', async () => { + const fs = makeInMemoryFileSystem(); + const flow = makeFeedbackFlow( + { + root: makeStep('root', [], ['leaf']), + leaf: makeStep('leaf', ['root'], []), + }, + 'root', + ); + + const events: HarnessEventPayload[] = []; + harnessEventBus.on(HARNESS_EVENT_NAME, (e) => events.push(e)); + + await executeAgenticFlow(flow, { + rootDir: '/workspace', + fileSystem: fs, + runId: 'run-guardrail-breach', + guardrailMaxAttempts: 2, + runStep: async ({ step }) => ({ text: `output:${step.id}`, usage: null, toolCalls: [], toolResults: [] }), + }); + + const stepEvents = events.filter( + (e): e is Extract => e.type === 'StepStatusChanged', + ); + expect(stepEvents.some((e) => e.stepId === 'root' && e.logs?.includes('breached contract'))).toBe(true); + }); + + it('a terminal step with no writesTo never activates the guardrail (no contract at all)', async () => { + const fs = makeInMemoryFileSystem(); + const flow = makeFeedbackFlow({ root: makeStep('root', [], []) }, 'root'); + + const events: HarnessEventPayload[] = []; + harnessEventBus.on(HARNESS_EVENT_NAME, (e) => events.push(e)); + + await executeAgenticFlow(flow, { + rootDir: '/workspace', + fileSystem: fs, + runId: 'run-guardrail-terminal', + runStep: async ({ step }) => ({ text: `output:${step.id}`, usage: null, toolCalls: [], toolResults: [] }), + }); + + const stepEvents = events.filter( + (e): e is Extract => e.type === 'StepStatusChanged', + ); + expect(stepEvents.some((e) => e.logs?.includes('guardrail'))).toBe(false); + }); +}); + +describe('executeAgenticFlow — contextMode: feedback — checkpoint contextFileSnapshot', () => { + afterEach(() => { + harnessEventBus.removeAllListeners(); + }); + + it('attaches contextFileSnapshot (path + content) to each step\'s checkpoint', async () => { + const fs = makeInMemoryFileSystem(); + const flow = makeFeedbackFlow({ root: makeStep('root', [], []) }, 'root'); + const runId = 'run-checkpoint-ctx'; + + await executeAgenticFlow(flow, { + rootDir: '/workspace', + fileSystem: fs, + runId, + runStep: async ({ step }) => ({ text: `output:${step.id}`, usage: null, toolCalls: [], toolResults: [] }), + }); + + const checkpoints = listCheckpoints(runId); + const rootCheckpoint = checkpoints.find((c) => c.stepId === 'root'); + expect(rootCheckpoint?.contextFileSnapshot?.path).toBe('.fluxor/run-context/run-checkpoint-ctx/step.root.md'); + expect(rootCheckpoint?.contextFileSnapshot?.content).toContain('# Contexto para root'); + }); + + it('omits contextFileSnapshot entirely for a blind-mode run\'s checkpoints', async () => { + const flow: AgenticFlow = { + id: 'flow-blind-checkpoint', + name: 'Blind Checkpoint', + rootStepId: 'root', + stepsRecord: { root: makeStep('root', [], []) }, + }; + const runId = 'run-checkpoint-blind'; + + await executeAgenticFlow(flow, { + runId, + runStep: async ({ step }) => ({ text: `output:${step.id}`, usage: null, toolCalls: [], toolResults: [] }), + }); + + const checkpoints = listCheckpoints(runId); + expect('contextFileSnapshot' in checkpoints[0]).toBe(false); + }); +}); diff --git a/src/main/harness-engine/executor.ts b/src/main/harness-engine/executor.ts index 8403205..1c946d1 100644 --- a/src/main/harness-engine/executor.ts +++ b/src/main/harness-engine/executor.ts @@ -4,10 +4,12 @@ * The scheduler preserves DAG dependency semantics while each step delegates to * the LLM runner and MCP-backed tool surface. */ +import { mkdir as nodeMkdir, readdir as nodeReaddir, readFile as nodeReadFile, stat as nodeStat, writeFile as nodeWriteFile } from 'fs/promises'; +import { posix, resolve as resolvePath } from 'path'; import type { AgenticFlow, AgenticMod, AgenticStep, StepContract } from '../../types/harness'; import { harnessEventBus } from './event-bus'; -import { buildStepContext } from './context-builder'; -import { createLocalMcpToolSet, type LocalMcpOptions } from './mcp-adapter'; +import { buildStepContext, type FlowAwarenessInput } from './context-builder'; +import { createLocalMcpToolSet, type LocalMcpOptions, type McpFileSystem } from './mcp-adapter'; import { getMcpClientModToolSet } from './mcp-client-mod'; import { createBrowserToolSet } from './browser-toolset'; import { runLLMStep, type LLMStepResult, type RunLLMStepInput } from './llm-runner'; @@ -26,6 +28,137 @@ import { import { createModelRouter, isSealed, loadDefaultLeaderboard, type ModelRouterDeps } from './model-router'; import type { ArenaLeaderboardEntry } from '../performance-frontier/arena/arena-runner'; import type { ConnectionResolver, ModelPolicy, RoutedModelEvidence } from '../../types/ipc-events'; +import { + buildContextManifest, + buildTopologySummaryLines, + contextArtifactPathPattern, + contextRunDir, + DEFAULT_CONTEXT_BUDGET_BYTES, + manifestKeyForInstance, + relativeContextFilePath, + seedContextFiles, + type ContextManifest, +} from './context-manifest'; + +// --------------------------------------------------------------------------- +// Effective rootDir + real-FS adapter (feedback mode only) +// +// Mirrors mcp-adapter.ts's private `normalizeRootDir`/`nodeFileSystem` +// (McpFileSystem shape) rather than importing/exporting them — this task's +// territory fence excludes editing mcp-adapter.ts, and the spec explicitly +// sanctions "export it or replicate it" (spec "Notas de implementación" — +// "rootDir efectivo"). MUST stay byte-for-byte equivalent: createLocalMcpToolSet +// resolves the model's own read_file/write_file/list_directory tools against +// this exact computation, so genesis (context dir creation) and the +// guardrail's workspace snapshots must agree with it — otherwise a step's +// assigned context file would live at a path its own FS tools can never see. +// --------------------------------------------------------------------------- + +const toPosixPath = (value: string): string => value.replace(/\\/g, '/'); +const isAbsoluteAnyPlatform = (value: string): boolean => ( + posix.isAbsolute(value) || /^[A-Za-z]:\//.test(value) +); + +function resolveEffectiveRootDir(rootDir?: string): string { + const raw = toPosixPath(rootDir ?? process.cwd()); + const absolute = isAbsoluteAnyPlatform(raw) ? raw : toPosixPath(resolvePath(raw)); + return posix.normalize(absolute); +} + +/** + * Minimal real-filesystem adapter used ONLY when `contextMode` is 'feedback' + * and the caller passed no `options.fileSystem` — every production surface + * (IDE/serve/webhook/mcp) omits it; only the Performance Frontier sandbox + * supplies its own VFS. Blind-mode runs never construct or touch this. + */ +const realFileSystemAdapter: McpFileSystem = { + mkdir: nodeMkdir, + readdir: nodeReaddir as McpFileSystem['readdir'], + readFile: nodeReadFile as McpFileSystem['readFile'], + stat: nodeStat as McpFileSystem['stat'], + writeFile: nodeWriteFile as McpFileSystem['writeFile'], +}; + +/** + * Feedback-mode plumbing computed once per run (executeAgenticFlow) and + * threaded into every executeStep call — undefined for every blind-mode run, + * which is what keeps blind execution byte-identical: every new code path + * below is gated behind `runContext` being defined. + */ +interface FeedbackRunContext { + manifest: ContextManifest; + runId: string; + /** Absolute-ish path (contextRunDir(effectiveRootDir, runId)) where genesis/checkpoint I/O happens. */ + runDir: string; + topologyLines: string[]; + /** Bare filename → seed content, so each writesTo target's exact minBytes floor is computable. */ + seeds: Record; + /** options.fileSystem, or the synthesized real-FS adapter when the caller passed none. */ + fileSystem: McpFileSystem; +} + +/** This instance's own input, or undefined outside feedback mode / when it has no manifest entry. */ +function computeFlowAwareness( + runContext: FeedbackRunContext | undefined, + instance: StepInstance | undefined, +): FlowAwarenessInput | undefined { + if (!runContext || !instance) return undefined; + const entry = runContext.manifest.steps[manifestKeyForInstance(instance)]; + if (!entry) return undefined; + + return { + topologyLines: runContext.topologyLines, + assignedFile: relativeContextFilePath(runContext.runId, entry.contextFile), + writesTo: entry.writesTo.map((filename) => relativeContextFilePath(runContext.runId, filename)), + }; +} + +/** + * A StepContract fragment requiring every briefing this instance promises + * (its manifest entry's `writesTo`) to exist with MORE bytes than its own + * seed — mere existence would trivially always pass since genesis seeds + * every context file up front, so minBytes (an existing, established + * mechanism — see StepArtifactRequirement.minBytes's "rejects empty stubs" + * doc) is set just above each SPECIFIC target's known seed length, making + * the check meaningfully verify a real briefing was written, not merely that + * the seeded stub survived. Composed into the step's effective contract via + * the EXISTING mergeStepContracts (guardrails.ts) — no new StepContract + * field, no mutation of the step's own AST contract. + */ +function computeContextContract( + runContext: FeedbackRunContext | undefined, + instance: StepInstance | undefined, +): StepContract | undefined { + if (!runContext || !instance) return undefined; + const entry = runContext.manifest.steps[manifestKeyForInstance(instance)]; + if (!entry || entry.writesTo.length === 0) return undefined; + + return { + requiredArtifacts: entry.writesTo.map((filename) => ({ + description: `context briefing "${filename}" for a downstream step`, + pathPattern: contextArtifactPathPattern(runContext.runId, filename), + minBytes: Buffer.byteLength(runContext.seeds[filename] ?? '', 'utf-8') + 1, + })), + }; +} + +/** Reads back this instance's own assigned context file for the checkpoint's contextFileSnapshot. Never throws — a read failure is logged and simply omits the field. */ +async function readContextFileSnapshot( + runContext: FeedbackRunContext, + instance: StepInstance, +): Promise<{ path: string; content: string } | undefined> { + const entry = runContext.manifest.steps[manifestKeyForInstance(instance)]; + if (!entry) return undefined; + + const relPath = relativeContextFilePath(runContext.runId, entry.contextFile); + try { + const content = await runContext.fileSystem.readFile(`${runContext.runDir}/${entry.contextFile}`, 'utf-8'); + return { path: relPath, content }; + } catch (error) { + console.warn(`[context-manifest] failed to snapshot context file "${relPath}" for checkpoint: ${getErrorMessage(error)}`); + return undefined; + } +} export interface HarnessStepRunnerInput extends RunLLMStepInput { flowId: string; @@ -276,9 +409,11 @@ async function executeStep( flow: AgenticFlow, step: AgenticStep, options: ExecuteAgenticFlowOptions, + effectiveRootDir: string, instance?: StepInstance, routing?: StepRouting, sealedIds?: Set, + runContext?: FeedbackRunContext, ): Promise { // Loop-body progress, if this instance is a pass of a loop body. Kept as a // plain object (not spread further here) so Phase 3 can layer modelId / @@ -318,10 +453,14 @@ async function executeStep( } // --------------------------------------------------------------------------- - // Default LLM path (unchanged) + // Default LLM path // --------------------------------------------------------------------------- + // Feedback mode only (runContext undefined ⇒ flowAwareness undefined ⇒ + // buildStepContext's output is byte-identical to before this feature). + const flowAwareness = computeFlowAwareness(runContext, instance); const context = await buildStepContext(step, { onModStatus: (mod, status, logs) => emitModStatus(flow.id, step.id, mod, status, logs), + ...(flowAwareness ? { flowAwareness } : {}), }); // Declarative mod runtime: aggregate blockTools/attachTools/contract @@ -382,14 +521,36 @@ async function executeStep( // ── Model-agnostic guardrail: verify a step's completion contract and re-run // it with concrete corrective feedback until it passes or the budget is spent. - // Steps without a `contract` (own or mod-contributed) keep the original - // single-pass behaviour exactly. - const contract = mergeStepContracts(step.contract, modRuntime.contracts); - const guardrailActive = Boolean(contract) && Boolean(options.fileSystem); + // Steps without a `contract` (own, mod-contributed, or feedback-mode + // context-derived) keep the original single-pass behaviour exactly. + // + // Feedback mode additionally composes a requiredArtifacts fragment from + // this instance's promised writesTo (computeContextContract) — this is the + // ONLY way a step with no `contract` of its own can still have + // guardrailActive become true, and only ever in feedback mode. + const contextContract = computeContextContract(runContext, instance); + const contract = mergeStepContracts( + step.contract, + contextContract ? [...modRuntime.contracts, contextContract] : modRuntime.contracts, + ); + // Plumbing fix (spec "Notas de implementación" — "Guardrail fuera del PF"): + // guardrailActive requires options.fileSystem, which NO production surface + // (IDE/serve/webhook/mcp) ever passes — only the PF sandbox does. Solely in + // feedback mode, when options.fileSystem is absent, fall back to the + // real-FS adapter (+ the materialized effective rootDir, fixing + // snapshotWorkspace's independent '/workspace' default disagreeing with the + // model's own FS tools) so a promised-briefing guardrail can actually + // activate on every surface, not just the PF sandbox. In blind mode this + // expression is UNCHANGED — `runContext` is always undefined there, so both + // operands fall through to exactly what they evaluated to before this + // feature existed. + const effectiveGuardrailFileSystem = runContext ? runContext.fileSystem : options.fileSystem; + const effectiveGuardrailRootDir = runContext ? effectiveRootDir : options.rootDir; + const guardrailActive = Boolean(contract) && Boolean(effectiveGuardrailFileSystem); const maxAttempts = guardrailActive ? Math.max(1, contract!.maxAttempts ?? options.guardrailMaxAttempts ?? DEFAULT_GUARDRAIL_MAX_ATTEMPTS) : 1; - const vfsBefore = guardrailActive ? await snapshotWorkspace(options.fileSystem, options.rootDir) : {}; + const vfsBefore = guardrailActive ? await snapshotWorkspace(effectiveGuardrailFileSystem, effectiveGuardrailRootDir) : {}; try { let result!: LLMStepResult; @@ -423,7 +584,7 @@ async function executeStep( if (!guardrailActive) break; - const vfsAfter = await snapshotWorkspace(options.fileSystem, options.rootDir); + const vfsAfter = await snapshotWorkspace(effectiveGuardrailFileSystem, effectiveGuardrailRootDir); const findings = verifyStepContract(contract!, vfsBefore, vfsAfter); if (findings.length === 0) { if (attempt > 1) { @@ -506,6 +667,31 @@ export async function executeAgenticFlow( const remaining = new Map(plan.remainingDeps); const runId = options.runId ?? `${flow.id}_${now()}`; + // ── Rosetta context genesis (feedback mode only) ─────────────────────── + // Materialized ONCE, regardless of mode — pure/side-effect-free, so + // computing it costs blind runs nothing observable. Actually CONSUMED + // (génesis I/O, guardrail fallback, checkpoint reads) only when + // `runContext` below is defined, i.e. only in feedback mode — a blind run + // creates no directory, no manifest, and injects zero prompt bytes (spec + // criterion 1: byte-identical to today). + const effectiveRootDir = resolveEffectiveRootDir(options.rootDir); + let runContext: FeedbackRunContext | undefined; + if (flow.contextMode === 'feedback') { + const contextManifest = buildContextManifest(flow, plan, runId, DEFAULT_CONTEXT_BUDGET_BYTES); + const topologyLines = buildTopologySummaryLines(flow, plan); + const seeds = seedContextFiles(contextManifest); + const fileSystem = options.fileSystem ?? realFileSystemAdapter; + const runDir = contextRunDir(effectiveRootDir, runId); + + await fileSystem.mkdir(runDir, { recursive: true }); + await fileSystem.writeFile(`${runDir}/manifest.json`, JSON.stringify(contextManifest, null, 2), 'utf-8'); + for (const [filename, content] of Object.entries(seeds)) { + await fileSystem.writeFile(`${runDir}/${filename}`, content, 'utf-8'); + } + + runContext = { manifest: contextManifest, runId, runDir, topologyLines, seeds, fileSystem }; + } + harnessEventBus.emitHarnessEvent({ type: 'FlowStarted', flowId: flow.id, @@ -558,11 +744,17 @@ export async function executeAgenticFlow( } const effectiveModelId = routing.modelId ?? options.modelId; - const result = await executeStep(flow, step, options, instance, routing, sealedIds); + const result = await executeStep(flow, step, options, effectiveRootDir, instance, routing, sealedIds, runContext); stepOutputs[instance.stepId] = result.text; // final pass wins completedInstanceKeys.add(key); completedStepIds.add(instance.stepId); + // Feedback mode only: re-read this instance's own assigned context file + // so the checkpoint can carry a contextFileSnapshot (path + content). + // Never throws — a read failure is logged and simply omits the field, + // since this is supplementary telemetry, not core execution state. + const contextFileSnapshot = runContext ? await readContextFileSnapshot(runContext, instance) : undefined; + // Persist an immutable checkpoint capturing state-so-far — one per // instance, i.e. one per iteration for a step inside a loop body. // `iteration` is attached only when this instance is a pass of a loop @@ -570,7 +762,8 @@ export async function executeAgenticFlow( // so a loop's 3 checkpoints for the same real stepId carry 1, 2, 3 and // the time-travel panel can tell them apart. Conditional spread keeps // the field entirely absent — never `iteration: undefined` — for the - // first/only pass of a non-loop step. + // first/only pass of a non-loop step. `contextFileSnapshot` follows the + // exact same discipline for blind-mode runs (always absent there). const checkpoint = saveCheckpoint({ runId, stepId: instance.stepId, @@ -579,6 +772,7 @@ export async function executeAgenticFlow( output: result.text, completedStepIds: [...completedStepIds], modelId: effectiveModelId, + ...(contextFileSnapshot ? { contextFileSnapshot } : {}), }); harnessEventBus.emitHarnessEvent({ type: 'CheckpointCreated', diff --git a/src/main/harness-engine/guardrails.test.ts b/src/main/harness-engine/guardrails.test.ts index a1cf25d..8db1fe8 100644 --- a/src/main/harness-engine/guardrails.test.ts +++ b/src/main/harness-engine/guardrails.test.ts @@ -85,6 +85,37 @@ describe('snapshotWorkspace', () => { it('returns empty for an absent file system', async () => { await expect(snapshotWorkspace(undefined)).resolves.toEqual({}); }); + + it('never descends into node_modules or .git when walking a real tree', async () => { + const dirs: Record = { + '/w': ['node_modules', '.git', 'src', 'a.txt'], + '/w/src': ['b.txt'], + '/w/node_modules': ['pkg'], + '/w/node_modules/pkg': ['huge.js'], + '/w/.git': ['HEAD'], + }; + const files: Record = { + '/w/a.txt': 'a', + '/w/src/b.txt': 'b', + '/w/node_modules/pkg/huge.js': 'never-read', + '/w/.git/HEAD': 'never-read', + }; + const fs = { + readdir: async (d: string) => { + if (!(d in dirs)) throw new Error('ENOENT'); + return dirs[d]; + }, + readFile: async (f: string) => { + if (!(f in files)) throw new Error('ENOENT'); + return files[f]; + }, + stat: async (p: string) => ({ isDirectory: () => p in dirs }), + }; + await expect(snapshotWorkspace(fs, '/w')).resolves.toEqual({ + '/w/a.txt': 'a', + '/w/src/b.txt': 'b', + }); + }); }); describe('mergeStepContracts — merging runtime mod contract fragments with a step contract', () => { diff --git a/src/main/harness-engine/guardrails.ts b/src/main/harness-engine/guardrails.ts index bd20e01..5136388 100644 --- a/src/main/harness-engine/guardrails.ts +++ b/src/main/harness-engine/guardrails.ts @@ -35,6 +35,14 @@ export interface ReadableWorkspace { snapshot?: () => Record; } +/** + * Directories the recursive walk never descends into: on a real project + * rootDir (feedback mode runs the guardrail against the actual workspace) + * snapshotting `node_modules`/`.git` would load the whole tree into memory + * twice per attempt. Neutral for the PF sandbox VFS, which uses `snapshot()`. + */ +const SNAPSHOT_EXCLUDED_DIRS = new Set(['node_modules', '.git']); + /** * Snapshot the workspace as a path→content map. Prefers a VFS `snapshot()` (the * PF sandbox); otherwise walks `rootDir` with readdir/stat/readFile. @@ -69,7 +77,7 @@ export async function snapshotWorkspace( isDir = false; } if (isDir) { - await walk(full); + if (!SNAPSHOT_EXCLUDED_DIRS.has(name)) await walk(full); } else { try { out[full] = await readFile(full, 'utf-8'); diff --git a/src/main/performance-frontier/bench/bench-runner.ts b/src/main/performance-frontier/bench/bench-runner.ts index 7cbae1c..b40fe28 100644 --- a/src/main/performance-frontier/bench/bench-runner.ts +++ b/src/main/performance-frontier/bench/bench-runner.ts @@ -1,7 +1,7 @@ import { runPerformanceFrontier } from '../runner'; import { summarize } from './statistics'; import type { StatsSummary } from './statistics'; -import type { PFSuite } from '../types'; +import type { PFContextMode, PFSuite } from '../types'; export interface BenchRunSample { finalScore: number; @@ -34,10 +34,12 @@ export interface RunBenchOptions { * procedural case variant. */ varySeeds?: boolean; + /** Forwarded verbatim to every repetition's `runPerformanceFrontier` call — see Step P1. */ + contextMode?: PFContextMode; } export async function runBench(opts: RunBenchOptions): Promise { - const { suite, modelId, repetitions } = opts; + const { suite, modelId, repetitions, contextMode } = opts; const baseSeed = opts.baseSeed ?? 1; const varySeeds = opts.varySeeds ?? false; @@ -49,7 +51,7 @@ export async function runBench(opts: RunBenchOptions): Promise { seeds.push(seed); try { - const result = await runPerformanceFrontier({ suite, modelId, seed }); + const result = await runPerformanceFrontier({ suite, modelId, seed, contextMode }); samples.push({ finalScore: result.finalScore, verdict: result.verdict, diff --git a/src/main/performance-frontier/bench/cli.ts b/src/main/performance-frontier/bench/cli.ts index ca0157e..ce772be 100644 --- a/src/main/performance-frontier/bench/cli.ts +++ b/src/main/performance-frontier/bench/cli.ts @@ -1,6 +1,7 @@ import { mkdir, writeFile } from 'fs/promises'; import { join } from 'path'; import { loadDotEnv } from '../env'; +import { parseContextModeFlag, warnIfContextModeIsNoop } from '../context-mode-flag'; import { runBench } from './bench-runner'; import type { PFSuite } from '../types'; @@ -63,10 +64,12 @@ async function main(): Promise { const repetitions = parseInt10('reps', 5); const baseSeed = parseInt10('seed', 1); const varySeeds = parseFlag('vary-seeds'); + const contextMode = parseContextModeFlag(process.argv); + warnIfContextModeIsNoop(suite, contextMode); - console.log(`\n[pf:bench] Suite: ${suite} Model: ${modelId} Reps: ${repetitions} Seed: ${baseSeed}${varySeeds ? '+i' : ''}\n`); + console.log(`\n[pf:bench] Suite: ${suite} Model: ${modelId} Reps: ${repetitions} Seed: ${baseSeed}${varySeeds ? '+i' : ''} ContextMode: ${contextMode ?? '(flow default)'}\n`); - const result = await runBench({ suite, modelId, repetitions, baseSeed, varySeeds }); + const result = await runBench({ suite, modelId, repetitions, baseSeed, varySeeds, contextMode }); const { finalScore: fs } = result; const halfWidth = (fs.ci95.upper - fs.ci95.lower) / 2; diff --git a/src/main/performance-frontier/cli.ts b/src/main/performance-frontier/cli.ts index 6e0af6e..8ea52f7 100644 --- a/src/main/performance-frontier/cli.ts +++ b/src/main/performance-frontier/cli.ts @@ -1,4 +1,5 @@ import { loadDotEnv } from './env'; +import { parseContextModeFlag, warnIfContextModeIsNoop } from './context-mode-flag'; import { runPerformanceFrontier } from './runner'; import type { PFSuite } from './types'; @@ -40,9 +41,15 @@ function parseSuite(): PFSuite | undefined { async function main(): Promise { await loadDotEnv(); + + const suite = parseSuite(); + const contextMode = parseContextModeFlag(process.argv); + warnIfContextModeIsNoop(suite, contextMode); + const result = await runPerformanceFrontier({ seed: parseSeed(), - suite: parseSuite(), + suite, + contextMode, }); console.log(JSON.stringify({ diff --git a/src/main/performance-frontier/compare/args.test.ts b/src/main/performance-frontier/compare/args.test.ts new file mode 100644 index 0000000..55dc03d --- /dev/null +++ b/src/main/performance-frontier/compare/args.test.ts @@ -0,0 +1,55 @@ +import { describe, expect, it } from 'vitest'; +import { parseCompareArgs } from './args'; + +describe('parseCompareArgs', () => { + // Edge: valores de flag inválidos — missing required --seed. + it('throws when --seed is missing', () => { + expect(() => parseCompareArgs(['node', 'cli.js'])).toThrow(/--seed= is required/); + }); + + it('throws when --seed is non-numeric', () => { + expect(() => parseCompareArgs(['--seed=abc'])).toThrow(/--seed must be a finite number, got "abc"/); + }); + + it('parses a valid --seed with no --suites', () => { + expect(parseCompareArgs(['--seed=7'])).toEqual({ seed: 7 }); + }); + + it('accepts a negative or zero seed as a valid finite number', () => { + expect(parseCompareArgs(['--seed=0'])).toEqual({ seed: 0 }); + expect(parseCompareArgs(['--seed=-3'])).toEqual({ seed: -3 }); + }); + + it('parses a single --suites value', () => { + expect(parseCompareArgs(['--seed=7', '--suites=team-work'])).toEqual({ seed: 7, suites: ['team-work'] }); + }); + + it('parses a comma-separated --suites list, trimming whitespace', () => { + expect(parseCompareArgs(['--seed=7', '--suites=team-work, progression'])).toEqual({ + seed: 7, + suites: ['team-work', 'progression'], + }); + }); + + it('resolves suite aliases the same way pf:run/pf:bench do', () => { + expect(parseCompareArgs(['--seed=1', '--suites=assembler,business'])).toEqual({ + seed: 1, + suites: ['flow-assembler', 'business-knowledge'], + }); + }); + + // Edge: valores de flag inválidos — unknown suite name. + it('throws a clear error for an unknown suite name', () => { + expect(() => parseCompareArgs(['--seed=1', '--suites=not-a-real-suite'])).toThrow( + /Invalid PF suite "not-a-real-suite" in --suites/, + ); + }); + + it('throws for an empty --suites value rather than silently comparing nothing', () => { + expect(() => parseCompareArgs(['--seed=1', '--suites='])).toThrow(/--suites was provided but empty/); + }); + + it('throws for a --suites value that is only commas/whitespace', () => { + expect(() => parseCompareArgs(['--seed=1', '--suites= , ,'])).toThrow(/--suites was provided but empty/); + }); +}); diff --git a/src/main/performance-frontier/compare/args.ts b/src/main/performance-frontier/compare/args.ts new file mode 100644 index 0000000..ba9c6b6 --- /dev/null +++ b/src/main/performance-frontier/compare/args.ts @@ -0,0 +1,75 @@ +/** + * `pf:compare` CLI flag parsing (Step P3) — pulled out of cli.ts (which has a + * side-effecting top-level `main().catch(...)` invocation, like every other + * PF CLI entry point) into its own side-effect-free module so it can be unit + * tested by importing it directly, without triggering `main()`. + */ +import type { PFSuite } from '../types'; + +// Superset of both pf:run's and pf:bench's suite lists (this tool only reads +// the ledger — it never constructs a case, so there is no risk of pointing +// at a suite the runner itself can't produce). +const VALID_SUITES: PFSuite[] = [ + 'architecture', + 'analysis', + 'design', + 'business', + 'team-work', + 'flow-assembler', + 'development', + 'business-knowledge', + 'progression', + 'from-scratch', +]; + +const SUITE_ALIASES: Record = { + assembler: 'flow-assembler', + business: 'business-knowledge', +}; + +function resolveSuiteName(value: string): string { + if (SUITE_ALIASES[value]) return SUITE_ALIASES[value]; + if ((VALID_SUITES as string[]).includes(value)) return value; + throw new Error( + `Invalid PF suite "${value}" in --suites. Expected one of: ${VALID_SUITES.join(', ')} (aliases: assembler→flow-assembler).`, + ); +} + +export interface CompareCliArgs { + seed: number; + suites?: string[]; +} + +/** + * Parses `pf:compare`'s CLI flags out of an argv array. Pure (never reads + * `process.argv` directly) so it's unit-testable without a subprocess. + * + * Edge: `--seed` is REQUIRED here (unlike pf:run's optional `--seed=`, + * defaulted to 1) — pairing is meaningless without pinning to one exact + * seed, so a missing/non-numeric value throws a clear, actionable error + * rather than silently defaulting. + */ +export function parseCompareArgs(argv: readonly string[]): CompareCliArgs { + const seedArg = argv.find((item) => item.startsWith('--seed=')); + if (!seedArg) { + throw new Error('--seed= is required (pf:compare pairs runs by exact seed).'); + } + const rawSeed = seedArg.slice('--seed='.length); + const seed = Number(rawSeed); + if (!Number.isFinite(seed)) { + throw new Error(`--seed must be a finite number, got "${rawSeed}".`); + } + + const suitesArg = argv.find((item) => item.startsWith('--suites=')); + if (!suitesArg) { + return { seed }; + } + + const raw = suitesArg.slice('--suites='.length); + const values = raw.split(',').map((item) => item.trim()).filter((item) => item.length > 0); + if (values.length === 0) { + throw new Error('--suites was provided but empty. Omit the flag to compare every suite.'); + } + + return { seed, suites: values.map(resolveSuiteName) }; +} diff --git a/src/main/performance-frontier/compare/cli.ts b/src/main/performance-frontier/compare/cli.ts new file mode 100644 index 0000000..32d027f --- /dev/null +++ b/src/main/performance-frontier/compare/cli.ts @@ -0,0 +1,50 @@ +/** + * `pf:compare` — Step P3. Pairs blind vs feedback ledger runs for one seed + * and prints a markdown comparison table (stdout) plus writes it into + * `docs/pf-context-mode-campaign.md`'s "## Resultados" section. + * + * Usage: + * npm run pf:compare -- --seed=7 [--suites=team-work,progression] + */ +import { readFile } from 'fs/promises'; +import { join } from 'path'; +import { loadDotEnv } from '../env'; +import { parseCompareArgs } from './args'; +import { buildComparison, parseLedger, renderComparisonMarkdown } from './comparator'; +import { upsertResultsSection } from './report-doc'; + +const DOC_PATH = join(process.cwd(), 'docs', 'pf-context-mode-campaign.md'); + +async function readLedgerRaw(ledgerPath: string): Promise { + try { + return await readFile(ledgerPath, 'utf-8'); + } catch (error: unknown) { + const code = (error as NodeJS.ErrnoException)?.code; + if (code === 'ENOENT') return ''; // Edge: no ledger yet — treated as an empty one, not an error. + throw error; + } +} + +async function main(): Promise { + await loadDotEnv(); + + const { seed, suites } = parseCompareArgs(process.argv); + const outputDir = join(process.cwd(), '.fluxor', 'performance-frontier'); + const ledgerPath = join(outputDir, 'pf-history.jsonl'); + + const raw = await readLedgerRaw(ledgerPath); + const records = raw.trim().length > 0 ? parseLedger(raw) : []; + const result = buildComparison(records, { seed, suites }); + const markdown = renderComparisonMarkdown(result, { seed, suites }); + + console.log(markdown); + + await upsertResultsSection(DOC_PATH, markdown); + console.log(`\n[pf:compare] "Resultados" section written to ${DOC_PATH}`); +} + +main().catch((error) => { + const message = error instanceof Error ? error.message : String(error); + console.error(`[pf:compare] ${message}`); + process.exitCode = 1; +}); diff --git a/src/main/performance-frontier/compare/comparator.test.ts b/src/main/performance-frontier/compare/comparator.test.ts new file mode 100644 index 0000000..7fe8570 --- /dev/null +++ b/src/main/performance-frontier/compare/comparator.test.ts @@ -0,0 +1,288 @@ +import { describe, expect, it } from 'vitest'; +import { + buildComparison, + effectiveContextMode, + parseLedger, + renderComparisonMarkdown, + type PFLedgerRecord, +} from './comparator'; + +function record(overrides: Partial & Pick): PFLedgerRecord { + return { + finalScore: 90, + telemetry: { totalTokens: 1000, latencyMs: 5000 }, + ...overrides, + }; +} + +describe('parseLedger', () => { + // Edge: ledger vacío. + it('returns [] for an empty string', () => { + expect(parseLedger('')).toEqual([]); + }); + + it('returns [] for a whitespace/newline-only ledger', () => { + expect(parseLedger('\n \n\n')).toEqual([]); + }); + + it('parses one JSON object per line, ignoring a trailing newline', () => { + const raw = `${JSON.stringify({ runId: 'a', suite: 's', seed: 1, modelId: 'm', finalScore: 1 })}\n` + + `${JSON.stringify({ runId: 'b', suite: 's', seed: 1, modelId: 'm', finalScore: 2 })}\n`; + const records = parseLedger(raw); + expect(records).toHaveLength(2); + expect(records[0].runId).toBe('a'); + expect(records[1].runId).toBe('b'); + }); + + it('ignores blank lines interspersed between records', () => { + const raw = `${JSON.stringify({ runId: 'a', suite: 's', seed: 1, modelId: 'm', finalScore: 1 })}\n\n\n` + + `${JSON.stringify({ runId: 'b', suite: 's', seed: 1, modelId: 'm', finalScore: 2 })}\n`; + expect(parseLedger(raw)).toHaveLength(2); + }); + + it('throws a clear, line-numbered error on malformed JSON (loud, not a silent skip)', () => { + const raw = `${JSON.stringify({ runId: 'a', suite: 's', seed: 1, modelId: 'm', finalScore: 1 })}\n` + + 'not-json{{{\n'; + expect(() => parseLedger(raw)).toThrow(/malformed JSON on line 2/); + }); +}); + +describe('effectiveContextMode', () => { + // Edge: registros viejos sin campo = blind al leer. + it('reads an absent contextMode field as "blind"', () => { + expect(effectiveContextMode({})).toBe('blind'); + expect(effectiveContextMode({ contextMode: undefined })).toBe('blind'); + }); + + it('reads "feedback" verbatim', () => { + expect(effectiveContextMode({ contextMode: 'feedback' })).toBe('feedback'); + }); + + it('reads "blind" verbatim', () => { + expect(effectiveContextMode({ contextMode: 'blind' })).toBe('blind'); + }); + + it('defensively reads a garbage/corrupted value as "blind" rather than throwing', () => { + expect(effectiveContextMode({ contextMode: 'sandbox-typo' })).toBe('blind'); + }); +}); + +describe('buildComparison', () => { + it('returns no rows and no unpaired entries for an empty ledger', () => { + const result = buildComparison([], { seed: 7 }); + expect(result.rows).toEqual([]); + expect(result.unpaired).toEqual([]); + }); + + it('pairs a blind + feedback run sharing (suite, seed, modelId)', () => { + const records = [ + record({ runId: 'run-blind', suite: 'team-work', seed: 7, modelId: 'mimo/mimo-v2.5-pro', contextMode: 'blind', finalScore: 80, telemetry: { totalTokens: 1000, latencyMs: 4000 } }), + record({ runId: 'run-feedback', suite: 'team-work', seed: 7, modelId: 'mimo/mimo-v2.5-pro', contextMode: 'feedback', finalScore: 90, telemetry: { totalTokens: 1300, latencyMs: 5200 } }), + ]; + + const result = buildComparison(records, { seed: 7 }); + expect(result.unpaired).toEqual([]); + expect(result.rows).toHaveLength(1); + expect(result.rows[0]).toMatchObject({ + suite: 'team-work', + modelId: 'mimo/mimo-v2.5-pro', + blind: { runId: 'run-blind', finalScore: 80, totalTokens: 1000, latencyMs: 4000 }, + feedback: { runId: 'run-feedback', finalScore: 90, totalTokens: 1300, latencyMs: 5200 }, + }); + }); + + // Edge: sin pareja blind/feedback. + it('reports a blind-only group as unpaired instead of silently dropping it', () => { + const records = [ + record({ runId: 'run-blind', suite: 'progression', seed: 7, modelId: 'mimo/mimo-v2.5-pro', contextMode: 'blind' }), + ]; + const result = buildComparison(records, { seed: 7 }); + expect(result.rows).toEqual([]); + expect(result.unpaired).toEqual([ + { suite: 'progression', modelId: 'mimo/mimo-v2.5-pro', mode: 'blind', runId: 'run-blind' }, + ]); + }); + + it('reports a feedback-only group as unpaired', () => { + const records = [ + record({ runId: 'run-feedback', suite: 'progression', seed: 7, modelId: 'mimo/mimo-v2.5-pro', contextMode: 'feedback' }), + ]; + const result = buildComparison(records, { seed: 7 }); + expect(result.rows).toEqual([]); + expect(result.unpaired).toEqual([ + { suite: 'progression', modelId: 'mimo/mimo-v2.5-pro', mode: 'feedback', runId: 'run-feedback' }, + ]); + }); + + // Edge: seeds distintas — must NOT be paired even though suite+model match. + it('does not pair runs with different seeds', () => { + const records = [ + record({ runId: 'run-blind-seed5', suite: 'team-work', seed: 5, modelId: 'mimo/mimo-v2.5-pro', contextMode: 'blind' }), + record({ runId: 'run-feedback-seed6', suite: 'team-work', seed: 6, modelId: 'mimo/mimo-v2.5-pro', contextMode: 'feedback' }), + ]; + + const resultAtSeed5 = buildComparison(records, { seed: 5 }); + expect(resultAtSeed5.rows).toEqual([]); + expect(resultAtSeed5.unpaired).toEqual([ + { suite: 'team-work', modelId: 'mimo/mimo-v2.5-pro', mode: 'blind', runId: 'run-blind-seed5' }, + ]); + + const resultAtSeed6 = buildComparison(records, { seed: 6 }); + expect(resultAtSeed6.rows).toEqual([]); + expect(resultAtSeed6.unpaired).toEqual([ + { suite: 'team-work', modelId: 'mimo/mimo-v2.5-pro', mode: 'feedback', runId: 'run-feedback-seed6' }, + ]); + }); + + // Edge: registros viejos sin campo — an old record with NO contextMode key + // at all pairs correctly as the "blind" side against a new "feedback" one. + it('pairs an old pre-P2 record (no contextMode field) as the blind side', () => { + // No `contextMode` in the overrides at all — simulates a ledger line + // written before Step P2 ever added the field, not merely `undefined`. + const oldRecord = record({ runId: 'run-old', suite: 'architecture', seed: 3, modelId: 'mimo/mimo-v2.5-pro' }); + expect('contextMode' in oldRecord).toBe(false); + const newRecord = record({ runId: 'run-new', suite: 'architecture', seed: 3, modelId: 'mimo/mimo-v2.5-pro', contextMode: 'feedback' }); + + const result = buildComparison([oldRecord, newRecord], { seed: 3 }); + expect(result.rows).toHaveLength(1); + expect(result.rows[0].blind.runId).toBe('run-old'); + expect(result.rows[0].feedback.runId).toBe('run-new'); + }); + + it('keeps the LAST record when multiple runs share (suite, seed, modelId, mode) — re-run supersedes', () => { + const records = [ + record({ runId: 'blind-stale', suite: 'team-work', seed: 7, modelId: 'm', contextMode: 'blind', finalScore: 50 }), + record({ runId: 'blind-fresh', suite: 'team-work', seed: 7, modelId: 'm', contextMode: 'blind', finalScore: 95 }), + record({ runId: 'feedback-only', suite: 'team-work', seed: 7, modelId: 'm', contextMode: 'feedback', finalScore: 96 }), + ]; + const result = buildComparison(records, { seed: 7 }); + expect(result.rows).toHaveLength(1); + expect(result.rows[0].blind.runId).toBe('blind-fresh'); + expect(result.rows[0].blind.finalScore).toBe(95); + }); + + it('filters by the --suites allowlist, dropping non-matching suites entirely', () => { + const records = [ + record({ runId: 'tw-b', suite: 'team-work', seed: 7, modelId: 'm', contextMode: 'blind' }), + record({ runId: 'tw-f', suite: 'team-work', seed: 7, modelId: 'm', contextMode: 'feedback' }), + record({ runId: 'pr-b', suite: 'progression', seed: 7, modelId: 'm', contextMode: 'blind' }), + record({ runId: 'pr-f', suite: 'progression', seed: 7, modelId: 'm', contextMode: 'feedback' }), + ]; + const result = buildComparison(records, { seed: 7, suites: ['team-work'] }); + expect(result.rows).toHaveLength(1); + expect(result.rows[0].suite).toBe('team-work'); + }); + + it('groups distinct suites and distinct models independently at the same seed', () => { + const records = [ + record({ runId: 'tw-model-a-b', suite: 'team-work', seed: 7, modelId: 'model-a', contextMode: 'blind' }), + record({ runId: 'tw-model-a-f', suite: 'team-work', seed: 7, modelId: 'model-a', contextMode: 'feedback' }), + record({ runId: 'tw-model-b-b', suite: 'team-work', seed: 7, modelId: 'model-b', contextMode: 'blind' }), + record({ runId: 'tw-model-b-f', suite: 'team-work', seed: 7, modelId: 'model-b', contextMode: 'feedback' }), + record({ runId: 'pr-model-a-b', suite: 'progression', seed: 7, modelId: 'model-a', contextMode: 'blind' }), + record({ runId: 'pr-model-a-f', suite: 'progression', seed: 7, modelId: 'model-a', contextMode: 'feedback' }), + ]; + const result = buildComparison(records, { seed: 7 }); + expect(result.rows).toHaveLength(3); + expect(result.unpaired).toEqual([]); + // Sorted by suite then modelId. + expect(result.rows.map((r) => `${r.suite}/${r.modelId}`)).toEqual([ + 'progression/model-a', + 'team-work/model-a', + 'team-work/model-b', + ]); + }); +}); + +describe('renderComparisonMarkdown', () => { + it('renders a graceful message (not an error) when there are no pairs', () => { + const markdown = renderComparisonMarkdown({ rows: [], unpaired: [] }, { seed: 9 }); + expect(markdown).toContain('Seed: 9'); + expect(markdown).toContain('No blind/feedback pairs found'); + }); + + it('renders a markdown table with blind → feedback (Δ%) cells for score/tokens/wall-time', () => { + const markdown = renderComparisonMarkdown({ + rows: [{ + suite: 'team-work', + modelId: 'mimo/mimo-v2.5-pro', + blind: { runId: 'run-blind', finalScore: 80, totalTokens: 1000, latencyMs: 4000 }, + feedback: { runId: 'run-feedback', finalScore: 92, totalTokens: 1300, latencyMs: 5200 }, + }], + unpaired: [], + }, { seed: 7, suites: ['team-work'] }); + + expect(markdown).toContain('Seed: 7'); + expect(markdown).toContain('Suites: team-work'); + expect(markdown).toContain('| Suite | Model |'); + expect(markdown).toContain('team-work'); + expect(markdown).toContain('mimo/mimo-v2.5-pro'); + expect(markdown).toContain('80 → 92 (+15.00%)'); + expect(markdown).toContain('1000 → 1300 (+30.00%)'); + expect(markdown).toContain('4000 → 5200 (+30.00%)'); + }); + + it('formats a negative delta with a leading minus and no double sign', () => { + const markdown = renderComparisonMarkdown({ + rows: [{ + suite: 's', + modelId: 'm', + blind: { runId: 'b', finalScore: 100, totalTokens: 100, latencyMs: 100 }, + feedback: { runId: 'f', finalScore: 90, totalTokens: 100, latencyMs: 100 }, + }], + unpaired: [], + }, { seed: 1 }); + expect(markdown).toContain('100 → 90 (-10.00%)'); + }); + + // Edge: division-by-zero baseline must never render Infinity/NaN. + it('renders "N/A" when the blind baseline is 0 and feedback is not', () => { + const markdown = renderComparisonMarkdown({ + rows: [{ + suite: 's', + modelId: 'm', + blind: { runId: 'b', finalScore: 0, totalTokens: 0, latencyMs: 0 }, + feedback: { runId: 'f', finalScore: 10, totalTokens: 500, latencyMs: 200 }, + }], + unpaired: [], + }, { seed: 1 }); + expect(markdown).toContain('0 → 10 (N/A)'); + expect(markdown).not.toContain('Infinity'); + expect(markdown).not.toContain('NaN'); + }); + + it('renders "0.00%" when both blind and feedback are 0', () => { + const markdown = renderComparisonMarkdown({ + rows: [{ + suite: 's', + modelId: 'm', + blind: { runId: 'b', finalScore: 0, totalTokens: 0, latencyMs: 0 }, + feedback: { runId: 'f', finalScore: 0, totalTokens: 0, latencyMs: 0 }, + }], + unpaired: [], + }, { seed: 1 }); + expect(markdown).toContain('0 → 0 (0.00%)'); + }); + + it('lists unpaired runs with their mode and runId', () => { + const markdown = renderComparisonMarkdown({ + rows: [], + unpaired: [{ suite: 'progression', modelId: 'm', mode: 'blind', runId: 'run-1' }], + }, { seed: 7 }); + expect(markdown).toContain('Unpaired runs'); + expect(markdown).toContain('progression / m: only `blind` ran (run-1).'); + }); + + it('escapes a pipe character in a suite/model cell so it cannot break the table', () => { + const markdown = renderComparisonMarkdown({ + rows: [{ + suite: 's', + modelId: 'weird|model', + blind: { runId: 'b', finalScore: 1, totalTokens: 1, latencyMs: 1 }, + feedback: { runId: 'f', finalScore: 1, totalTokens: 1, latencyMs: 1 }, + }], + unpaired: [], + }, { seed: 1 }); + expect(markdown).toContain('weird\\|model'); + }); +}); diff --git a/src/main/performance-frontier/compare/comparator.ts b/src/main/performance-frontier/compare/comparator.ts new file mode 100644 index 0000000..8dba989 Binary files /dev/null and b/src/main/performance-frontier/compare/comparator.ts differ diff --git a/src/main/performance-frontier/compare/report-doc.test.ts b/src/main/performance-frontier/compare/report-doc.test.ts new file mode 100644 index 0000000..abe11d6 --- /dev/null +++ b/src/main/performance-frontier/compare/report-doc.test.ts @@ -0,0 +1,138 @@ +import { mkdtemp, readFile, rm, writeFile } from 'fs/promises'; +import { tmpdir } from 'os'; +import { join } from 'path'; +import { afterEach, beforeEach, describe, expect, it } from 'vitest'; +import { upsertResultsSection } from './report-doc'; + +let dir: string; +let docPath: string; + +describe('upsertResultsSection', () => { + beforeEach(async () => { + dir = await mkdtemp(join(tmpdir(), 'fluxor-pf-compare-doc-')); + docPath = join(dir, 'pf-context-mode-campaign.md'); + }); + + afterEach(async () => { + await rm(dir, { recursive: true, force: true }); + }); + + it('creates the file with just the section when it does not exist yet', async () => { + await upsertResultsSection(docPath, '| a | b |\n|---|---|\n| 1 | 2 |'); + const content = await readFile(docPath, 'utf-8'); + expect(content).toContain('## Resultados'); + expect(content).toContain('| a | b |'); + }); + + it('appends the section at the end when the doc exists but has no "## Resultados" heading, preserving everything', async () => { + const original = [ + '# Campaña PF: Context Modes', + '', + '## Suites elegidas', + '', + '- team-work', + '- progression', + '', + '## Checklist de ejecución', + '', + '- [ ] run blind', + '- [ ] run feedback', + ].join('\n'); + await writeFile(docPath, original, 'utf-8'); + + await upsertResultsSection(docPath, '| suite | score |\n|---|---|\n| team-work | 90 |'); + const content = await readFile(docPath, 'utf-8'); + + expect(content).toContain('# Campaña PF: Context Modes'); + expect(content).toContain('## Suites elegidas'); + expect(content).toContain('- team-work'); + expect(content).toContain('## Checklist de ejecución'); + expect(content).toContain('- [ ] run blind'); + expect(content).toContain('## Resultados'); + expect(content).toContain('| team-work | 90 |'); + // The new section comes after the pre-existing content. + expect(content.indexOf('## Checklist')).toBeLessThan(content.indexOf('## Resultados')); + }); + + it('replaces an existing "## Resultados" section in the middle of the doc, preserving sections before AND after', async () => { + const original = [ + '# Campaña PF', + '', + '## Checklist de ejecución', + '', + '- [ ] step one', + '', + '## Resultados', + '', + '_(placeholder until the campaign runs)_', + '', + '## Notas', + '', + 'Algo que no debe borrarse.', + ].join('\n'); + await writeFile(docPath, original, 'utf-8'); + + await upsertResultsSection(docPath, '| suite | score |\n|---|---|\n| progression | 88 |'); + const content = await readFile(docPath, 'utf-8'); + + expect(content).toContain('## Checklist de ejecución'); + expect(content).toContain('- [ ] step one'); + expect(content).toContain('## Notas'); + expect(content).toContain('Algo que no debe borrarse.'); + expect(content).not.toContain('placeholder until the campaign runs'); + expect(content).toContain('| progression | 88 |'); + + // Order preserved: Checklist, then Resultados, then Notas. + expect(content.indexOf('## Checklist')).toBeLessThan(content.indexOf('## Resultados')); + expect(content.indexOf('## Resultados')).toBeLessThan(content.indexOf('## Notas')); + }); + + it('replaces a "## Resultados" section that is the LAST section in the doc', async () => { + const original = [ + '# Campaña PF', + '', + '## Checklist de ejecución', + '', + '- [ ] step one', + '', + '## Resultados', + '', + '_(placeholder)_', + ].join('\n'); + await writeFile(docPath, original, 'utf-8'); + + await upsertResultsSection(docPath, '| suite | score |\n|---|---|\n| team-work | 95 |'); + const content = await readFile(docPath, 'utf-8'); + + expect(content).toContain('## Checklist de ejecución'); + expect(content).toContain('- [ ] step one'); + expect(content).not.toContain('_(placeholder)_'); + expect(content).toContain('| team-work | 95 |'); + }); + + it('is idempotent: running it twice with the same table converges to stable content (no duplication)', async () => { + const original = '# Campaña PF\n\n## Checklist\n\n- [ ] a\n'; + await writeFile(docPath, original, 'utf-8'); + + const table = '| suite | score |\n|---|---|\n| team-work | 90 |'; + await upsertResultsSection(docPath, table); + const afterFirst = await readFile(docPath, 'utf-8'); + await upsertResultsSection(docPath, table); + const afterSecond = await readFile(docPath, 'utf-8'); + + expect(afterSecond).toBe(afterFirst); + expect(afterSecond.match(/## Resultados/g)).toHaveLength(1); + }); + + it('refreshes the section content when re-run with a different table (not just idempotent on identical input)', async () => { + await writeFile(docPath, '# Campaña PF\n\n## Checklist\n\n- [ ] a\n', 'utf-8'); + + await upsertResultsSection(docPath, '| suite | score |\n|---|---|\n| team-work | 80 |'); + await upsertResultsSection(docPath, '| suite | score |\n|---|---|\n| team-work | 95 |'); + const content = await readFile(docPath, 'utf-8'); + + expect(content).not.toContain('| team-work | 80 |'); + expect(content).toContain('| team-work | 95 |'); + expect(content.match(/## Resultados/g)).toHaveLength(1); + }); +}); diff --git a/src/main/performance-frontier/compare/report-doc.ts b/src/main/performance-frontier/compare/report-doc.ts new file mode 100644 index 0000000..3a2517c --- /dev/null +++ b/src/main/performance-frontier/compare/report-doc.ts @@ -0,0 +1,73 @@ +/** + * Writes `pf:compare`'s markdown table into the "## Resultados" section of + * `docs/pf-context-mode-campaign.md` (Step P3/P4) — the rest of that document + * (suites chosen, checklist, exact command) is authored by P4 and must never + * be touched by this tool: "el comparador añade su sección, NO pisa el + * resto." Idempotent: running `pf:compare` again replaces only this section + * with a fresh render, wherever it currently sits in the document. + */ +import { readFile, writeFile } from 'fs/promises'; + +export const RESULTS_SECTION_HEADING = '## Resultados'; + +/** + * Inserts or replaces the `RESULTS_SECTION_HEADING` section of the markdown + * document at `docPath`, leaving every other section untouched. + * + * - If the heading doesn't exist yet, the section is appended at the end. + * - If it exists, everything from that heading up to (but not including) the + * next level-1/level-2 heading — or EOF — is replaced. + * - If the file doesn't exist at all, it is created containing just this + * section (defensive: P4 is expected to have authored the rest first, but + * this keeps the tool usable standalone). + */ +export async function upsertResultsSection(docPath: string, markdownTable: string): Promise { + let existing = ''; + try { + existing = await readFile(docPath, 'utf-8'); + } catch (error: unknown) { + const code = (error as NodeJS.ErrnoException)?.code; + if (code !== 'ENOENT') throw error; + } + + const section = [ + RESULTS_SECTION_HEADING, + '', + '_Generated by `pf:compare` — do not hand-edit; re-run the command to refresh._', + '', + markdownTable.trimEnd(), + ].join('\n'); + + const next = replaceOrAppendSection(existing, section); + await writeFile(docPath, next, 'utf-8'); +} + +function replaceOrAppendSection(existing: string, section: string): string { + const trimmedExisting = existing.replace(/\s+$/, ''); + + if (trimmedExisting.length === 0) { + return `${section}\n`; + } + + const lines = trimmedExisting.split('\n'); + const headingIndex = lines.findIndex((line) => line.trim() === RESULTS_SECTION_HEADING); + + if (headingIndex === -1) { + return `${trimmedExisting}\n\n${section}\n`; + } + + // The section runs until the next level-1/level-2 heading, or EOF. + let endIndex = lines.length; + for (let i = headingIndex + 1; i < lines.length; i++) { + if (/^#{1,2}\s/.test(lines[i])) { + endIndex = i; + break; + } + } + + const before = lines.slice(0, headingIndex).join('\n').replace(/\s+$/, ''); + const after = lines.slice(endIndex).join('\n').replace(/^\s+/, ''); + + const parts = [before, section, after].filter((part) => part.length > 0); + return `${parts.join('\n\n')}\n`; +} diff --git a/src/main/performance-frontier/context-mode-flag.test.ts b/src/main/performance-frontier/context-mode-flag.test.ts new file mode 100644 index 0000000..580bf2e --- /dev/null +++ b/src/main/performance-frontier/context-mode-flag.test.ts @@ -0,0 +1,68 @@ +import { describe, expect, it, vi } from 'vitest'; +import { parseContextModeFlag, warnIfContextModeIsNoop } from './context-mode-flag'; + +describe('parseContextModeFlag', () => { + it('returns undefined when the flag is absent (default: whatever the flow itself carries)', () => { + expect(parseContextModeFlag(['node', 'cli.js', '--suite=team-work'])).toBeUndefined(); + }); + + it('parses "blind"', () => { + expect(parseContextModeFlag(['--context-mode=blind'])).toBe('blind'); + }); + + it('parses "feedback"', () => { + expect(parseContextModeFlag(['--context-mode=feedback'])).toBe('feedback'); + }); + + // Edge: valores de flag inválidos — must fail loudly with a clear, actionable message. + it('throws a clear error for an invalid value', () => { + expect(() => parseContextModeFlag(['--context-mode=sandbox'])).toThrow( + /Invalid --context-mode "sandbox"\. Expected one of: blind, feedback\./, + ); + }); + + // Edge: empty string after "=" is still an invalid value, not "absent". + it('throws for an empty value rather than treating it as absent', () => { + expect(() => parseContextModeFlag(['--context-mode='])).toThrow(/Invalid --context-mode ""/); + }); + + it('is case-sensitive (does not silently accept "Blind")', () => { + expect(() => parseContextModeFlag(['--context-mode=Blind'])).toThrow(/Invalid --context-mode "Blind"/); + }); + + it('picks the flag out of a larger argv array regardless of position', () => { + expect(parseContextModeFlag(['node', 'cli.js', '--seed=7', '--context-mode=feedback', '--suite=progression'])) + .toBe('feedback'); + }); +}); + +describe('warnIfContextModeIsNoop', () => { + it('warns when --context-mode is set for the flow-assembler suite (suite sin flow)', () => { + const log = vi.fn(); + warnIfContextModeIsNoop('flow-assembler', 'feedback', log); + expect(log).toHaveBeenCalledTimes(1); + expect(log.mock.calls[0][0]).toContain('flow-assembler'); + expect(log.mock.calls[0][0]).toContain('--context-mode=feedback'); + }); + + it('does not warn when contextMode is undefined (flag not passed)', () => { + const log = vi.fn(); + warnIfContextModeIsNoop('flow-assembler', undefined, log); + expect(log).not.toHaveBeenCalled(); + }); + + it('does not warn when suite is undefined (CLI default resolves to architecture, which runs a flow)', () => { + const log = vi.fn(); + warnIfContextModeIsNoop(undefined, 'feedback', log); + expect(log).not.toHaveBeenCalled(); + }); + + it.each(['architecture', 'analysis', 'design', 'business-knowledge', 'team-work', 'development', 'progression', 'from-scratch'] as const)( + 'does not warn for suite "%s" (it does execute an AgenticFlow)', + (suite) => { + const log = vi.fn(); + warnIfContextModeIsNoop(suite, 'blind', log); + expect(log).not.toHaveBeenCalled(); + }, + ); +}); diff --git a/src/main/performance-frontier/context-mode-flag.ts b/src/main/performance-frontier/context-mode-flag.ts new file mode 100644 index 0000000..ebd0661 --- /dev/null +++ b/src/main/performance-frontier/context-mode-flag.ts @@ -0,0 +1,64 @@ +/** + * `--context-mode` CLI flag (Step P1, docs/superpowers/plans/2026-07-10-flow-context-modes.md). + * + * Shared by `pf:run` (cli.ts) and `pf:bench` (bench/cli.ts) — pulled out of + * both call sites (rather than copy-pasted, unlike `parseSuite`/`parseSeed`, + * which legitimately diverge per-CLI in their valid-suite lists) because the + * validation here is byte-identical in both places and this is new code, not + * an existing convention to preserve. Pure functions taking `argv`/a logger + * as parameters (never reading `process.argv`/`console` directly) so they are + * unit-testable without spawning a CLI subprocess. + */ +import type { PFContextMode, PFSuite } from './types'; + +export const VALID_PF_CONTEXT_MODES: readonly PFContextMode[] = ['blind', 'feedback']; + +/** + * Suites that never call `executeAgenticFlow` (see runner.ts) — today, only + * `flow-assembler` (it drives `assemblePipeline` directly and never touches + * an `AgenticFlow`). `--context-mode` has structurally nothing to affect for + * these suites; `warnIfContextModeIsNoop` surfaces that instead of silently + * accepting a flag that does nothing. + */ +export const CONTEXT_MODE_NOOP_SUITES: readonly PFSuite[] = ['flow-assembler']; + +/** + * Parses `--context-mode=blind|feedback` out of an argv array. + * Returns `undefined` when the flag is absent (caller then defaults to + * whatever the flow itself carries — see runner.ts). Throws a clear, + * actionable error for any other value, matching this codebase's existing + * `--suite=` validation style (cli.ts/bench/cli.ts `parseSuite`). + */ +export function parseContextModeFlag(argv: readonly string[]): PFContextMode | undefined { + const arg = argv.find((item) => item.startsWith('--context-mode=')); + if (!arg) return undefined; + + const value = arg.slice('--context-mode='.length); + if ((VALID_PF_CONTEXT_MODES as readonly string[]).includes(value)) { + return value as PFContextMode; + } + + throw new Error( + `Invalid --context-mode "${value}". Expected one of: ${VALID_PF_CONTEXT_MODES.join(', ')}.`, + ); +} + +/** + * Warns (does not throw — the run still proceeds normally in blind) when + * `--context-mode` was explicitly passed for a suite that can never act on + * it. Silence here would look like the flag was honored when it structurally + * cannot be. + */ +export function warnIfContextModeIsNoop( + suite: PFSuite | undefined, + contextMode: PFContextMode | undefined, + log: (message: string) => void = console.warn, +): void { + if (!contextMode || !suite) return; + if (!CONTEXT_MODE_NOOP_SUITES.includes(suite)) return; + + log( + `[Performance Frontier] --context-mode=${contextMode} is a no-op for suite "${suite}" — ` + + 'it never executes an AgenticFlow (see runner.ts).', + ); +} diff --git a/src/main/performance-frontier/runner.test.ts b/src/main/performance-frontier/runner.test.ts index b212485..0f98889 100644 --- a/src/main/performance-frontier/runner.test.ts +++ b/src/main/performance-frontier/runner.test.ts @@ -532,3 +532,144 @@ describe('performance frontier runner', () => { expect(ledger).toContain('"artifactsDir"'); }); }); + +// Step P1/P2 (docs/superpowers/plans/2026-07-10-flow-context-modes.md): the +// `contextMode` override forces every step's `AgenticFlow.contextMode` for +// the run and is recorded in the ledger — additive, defaulting to 'blind'. +describe('performance frontier runner — context mode override (Step P1/P2)', () => { + beforeEach(async () => { + outputDir = await mkdtemp(join(tmpdir(), 'fluxor-pf-run-ctx-')); + }); + + afterEach(async () => { + await rm(outputDir, { recursive: true, force: true }); + }); + + const passingJudge = (score: number) => async (input: any) => ({ + runId: input.runId, + caseId: input.caseId, + suite: input.suite, + modelUnderTest: input.modelUnderTest, + verdict: 'pass', + finalScore: score, + semanticScore: score - 5, + telemetryScore: 10, + evaluations: {}, + criticalFailures: [], + telemetry: input.telemetry, + }) as any; + + it('defaults to "blind" (no injected) and records contextMode:"blind" in the ledger', async () => { + let capturedSystemPrompt = ''; + const result = await runPerformanceFrontier({ + seed: 20, + outputDir, + modelId: 'fake/model', + // No `contextMode` passed — must default to whatever the flow carries + // (always absent/blind from case-factory.ts today). + runStep: async ({ systemPrompt, tools }) => { + capturedSystemPrompt = systemPrompt; + await tools.write_file.execute?.({ path: 'ok.txt', content: 'ok' }, { toolCallId: 't1', messages: [] }); + return { + text: 'ok', + usage: { inputTokens: 10, outputTokens: 5, totalTokens: 15 }, + toolCalls: [{ toolName: 'write_file' }], + toolResults: [{}], + }; + }, + judge: passingJudge(80), + }); + + expect(capturedSystemPrompt).not.toContain(''); + const ledger = await readFile(result.ledgerPath, 'utf-8'); + expect(ledger).toContain('"contextMode":"blind"'); + }); + + it('overrides contextMode to "feedback" and the executor actually injects ', async () => { + let capturedSystemPrompt = ''; + const result = await runPerformanceFrontier({ + seed: 21, + outputDir, + modelId: 'fake/model', + contextMode: 'feedback', + runStep: async ({ systemPrompt, tools }) => { + capturedSystemPrompt = systemPrompt; + await tools.write_file.execute?.({ path: 'ok.txt', content: 'ok' }, { toolCallId: 't1', messages: [] }); + return { + text: 'ok', + usage: { inputTokens: 10, outputTokens: 5, totalTokens: 15 }, + toolCalls: [{ toolName: 'write_file' }], + toolResults: [{}], + }; + }, + judge: passingJudge(88), + }); + + // Proves the override reaches the actual AgenticFlow the executor + // consumes — not just a cosmetic ledger field. + expect(capturedSystemPrompt).toContain(''); + const ledger = await readFile(result.ledgerPath, 'utf-8'); + expect(ledger).toContain('"contextMode":"feedback"'); + }); + + it('applies the override to every epoch\'s flow for the progression suite', async () => { + const systemPromptsByStep: Record = {}; + const result = await runPerformanceFrontier({ + suite: 'progression', + seed: 22, + outputDir, + modelId: 'fake/model', + contextMode: 'feedback', + runStep: async ({ step, systemPrompt, tools }) => { + systemPromptsByStep[step.id] = systemPrompt; + await tools.write_file.execute?.({ path: `${step.id}.txt`, content: 'ok' }, { toolCallId: `${step.id}-w`, messages: [] }); + return { + text: 'ok', + usage: { inputTokens: 10, outputTokens: 5, totalTokens: 15 }, + toolCalls: [{ toolName: 'write_file' }], + toolResults: [{}], + }; + }, + judge: passingJudge(85), + }); + + expect(Object.keys(systemPromptsByStep)).toEqual(['epoch-1-root', 'epoch-2-root']); + expect(systemPromptsByStep['epoch-1-root']).toContain(''); + expect(systemPromptsByStep['epoch-2-root']).toContain(''); + const ledger = await readFile(result.ledgerPath, 'utf-8'); + expect(ledger).toContain('"contextMode":"feedback"'); + }); + + // Edge: "suite sin flow" — flow-assembler never calls executeAgenticFlow, + // so forcing contextMode is structurally inert; the ledger stays "blind" + // (recording "feedback" would misreport a run where nothing feedback- + // related happened) and nothing throws. The CLI layer (cli.ts/bench/cli.ts) + // is responsible for warning the caller — see context-mode-flag.test.ts. + it('is a no-op for the flow-assembler suite and still records "blind" in the ledger', async () => { + const result = await runPerformanceFrontier({ + suite: 'flow-assembler', + seed: 23, + outputDir, + modelId: 'fake/model', + contextMode: 'feedback', + assemblePipeline: async (userIntent, assembleOptions) => { + assembleOptions?.onGeneration?.({ + discoveredCatalog: { roles: [], mods: [] }, + usage: { inputTokens: 10, outputTokens: 5, totalTokens: 15 }, + latencyMs: 5, + }); + return { + frameTitle: 'Test Pipeline', + description: 'Test.', + missingCapabilitiesRequested: [], + steps: [{ id: 'a', prompt: 'do it', roleId: 'confident-executor', modIds: [], prevStepIds: [] }], + }; + }, + judge: passingJudge(90), + }); + + expect(result.suite).toBe('flow-assembler'); + const ledger = await readFile(result.ledgerPath, 'utf-8'); + expect(ledger).toContain('"contextMode":"blind"'); + }); +}); diff --git a/src/main/performance-frontier/runner.ts b/src/main/performance-frontier/runner.ts index f7b393f..93639c8 100644 --- a/src/main/performance-frontier/runner.ts +++ b/src/main/performance-frontier/runner.ts @@ -26,6 +26,7 @@ import { verifyApi } from './execution/api-verifier'; import type { ApiVerificationResult } from './execution/api-verifier'; import type { PFCognitiveTraceEntry, + PFContextMode, PFConversationEntry, PFGroundTruth, PFJudgeRunner, @@ -44,6 +45,17 @@ export interface RunPerformanceFrontierOptions { stepTimeoutMs?: number; runStep?: PFStepRunner; judge?: PFJudgeRunner; + /** + * Forces every step's `AgenticFlow.contextMode` for this run (Step P1: + * docs/superpowers/plans/2026-07-10-flow-context-modes.md). Omitted: + * whatever the suite's generated flow(s) already carry — today always + * absent/blind, since `case-factory.ts` never sets it. Applied to every + * epoch's flow for `progression`; a no-op for `flow-assembler` (it never + * calls `executeAgenticFlow` — see the branch below), where the CLI layer + * (cli.ts / bench/cli.ts) is responsible for warning the caller instead of + * silently doing nothing. + */ + contextMode?: PFContextMode; assemblePipeline?: (userIntent: string, options?: AssemblePipelineOptions) => Promise; verifyDevelopment?: (vfsSnapshot: Record) => Promise; verifyDesign?: (vfsSnapshot: Record) => Promise; @@ -98,6 +110,29 @@ export async function runPerformanceFrontier( const ledgerPath = join(outputDir, 'pf-history.jsonl'); const runId = createRunId(seed); const testCase = createPerformanceCase({ suite, seed }); + + // `--context-mode` override (Step P1): forces the mode for every flow this + // run will execute, taking precedence over whatever `case-factory.ts` + // built. Skipped for `flow-assembler`, which never calls + // `executeAgenticFlow` — setting it there would write a misleading + // "feedback" ledger record for a run where nothing feedback-related + // happened; the CLI layer warns the caller about this no-op instead. + if (options.contextMode && suite !== 'flow-assembler') { + if (suite === 'progression') { + for (const epoch of testCase.epochs ?? []) { + epoch.flow.contextMode = options.contextMode; + } + } else { + testCase.flow.contextMode = options.contextMode; + } + } + + // Effective mode actually carried by `testCase.flow` after the override + // above — absent reads as `'blind'` (Step P2's ledger contract: records + // written before this field existed, or for flows that never set it, + // default to blind on read). Recorded verbatim in every ledger branch below. + const effectiveContextMode: PFContextMode = testCase.flow.contextMode ?? 'blind'; + const telemetry = new PerformanceTelemetryCollector(); const conversation: PFConversationEntry[] = []; const cognitiveTrace: PFCognitiveTraceEntry[] = []; @@ -192,6 +227,7 @@ export async function runPerformanceFrontier( caseId: testCase.id, seed, modelId, + contextMode: effectiveContextMode, finalScore: judgeResult.finalScore, semanticScore: judgeResult.semanticScore, telemetryScore: judgeResult.telemetryScore, @@ -344,6 +380,7 @@ export async function runPerformanceFrontier( caseId: testCase.id, seed, modelId, + contextMode: effectiveContextMode, finalScore: judgeResult.finalScore, semanticScore: judgeResult.semanticScore, telemetryScore: judgeResult.telemetryScore, @@ -494,6 +531,7 @@ export async function runPerformanceFrontier( caseId: testCase.id, seed, modelId, + contextMode: effectiveContextMode, finalScore: judgeResult.finalScore, semanticScore: judgeResult.semanticScore, telemetryScore: judgeResult.telemetryScore, diff --git a/src/main/performance-frontier/types.ts b/src/main/performance-frontier/types.ts index b35d10c..c8ff4d5 100644 --- a/src/main/performance-frontier/types.ts +++ b/src/main/performance-frontier/types.ts @@ -16,6 +16,16 @@ export type PFSuite = | 'progression' | 'from-scratch'; +/** + * Mirrors `AgenticFlow.contextMode` (src/types/harness.ts) — duplicated as a + * standalone alias here (rather than importing it) so the Performance + * Frontier's CLI/ledger/comparator layer has a name for the union without + * reaching into harness-engine's territory for a type that isn't otherwise + * exported on its own. Absent/undefined reads as `'blind'` everywhere in this + * subsystem (CLI defaults, ledger records written before this field existed). + */ +export type PFContextMode = 'blind' | 'feedback'; + /** * A single mutation epoch in a brownfield (progression) case. Each epoch runs * its own agentic flow against the SAME, non-destroyed VFS, so later epochs can diff --git a/src/main/serve/serve-flow.test.ts b/src/main/serve/serve-flow.test.ts index 4e998bf..b4dd89e 100644 --- a/src/main/serve/serve-flow.test.ts +++ b/src/main/serve/serve-flow.test.ts @@ -19,7 +19,9 @@ */ import { readFileSync } from 'node:fs'; -import { resolve } from 'node:path'; +import { mkdtemp, readFile as nodeReadFile, rm } from 'node:fs/promises'; +import { tmpdir } from 'node:os'; +import { resolve, join } from 'node:path'; import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'; import type { AgenticStep } from '@/types/harness'; import type { HarnessEventPayload } from '@/types/ipc-events'; @@ -439,3 +441,116 @@ describe('createFlowServer — validation', () => { expect(() => createFlowServer(bad)).toThrow(/rootStepId/); }); }); + +// --------------------------------------------------------------------------- +// contextMode — Rosetta context system (spec: +// docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md) +// +// serve-flow inherits contextMode from the export (importFlow already +// restores AgenticFlow.contextMode — see fluxor-flow.test.ts) and accepts an +// optional per-request override in POST /run's body, validated to +// 'blind'|'feedback'. +// --------------------------------------------------------------------------- + +describe('POST /run — contextMode', () => { + let tmpDir: string; + + beforeEach(async () => { + tmpDir = await mkdtemp(join(tmpdir(), 'fluxor-serve-context-')); + }); + + afterEach(async () => { + await rm(tmpDir, { recursive: true, force: true }); + }); + + it('rejects an invalid contextMode value with 400, before execution starts', async () => { + const p = await pickFreePort(); + const handle = createFlowServer(exported, { runStep: makeScriptedRunStep(), rootDir: tmpDir }); + await handle.listen(p); + try { + const { status, body } = await postRun( + `http://127.0.0.1:${p}`, + { contextMode: 'bogus' }, + authHeader(handle.token), + ); + expect(status).toBe(400); + expect((body as Record).error).toMatch(/contextMode/i); + } finally { + await handle.close(); + } + }); + + it('overriding contextMode: "feedback" in the body materializes a run-context manifest under the given rootDir', async () => { + const p = await pickFreePort(); + const handle = createFlowServer(exported, { runStep: makeScriptedRunStep(), rootDir: tmpDir }); + await handle.listen(p); + try { + const { status } = await postRun( + `http://127.0.0.1:${p}`, + { contextMode: 'feedback' }, + authHeader(handle.token), + ); + expect(status).toBe(200); + } finally { + await handle.close(); + } + + const entries = await import('node:fs/promises').then((fs) => fs.readdir(join(tmpDir, '.fluxor', 'run-context'))); + expect(entries.length).toBeGreaterThan(0); + const manifestPath = join(tmpDir, '.fluxor', 'run-context', entries[0], 'manifest.json'); + const manifest = JSON.parse(await nodeReadFile(manifestPath, 'utf-8')); + expect(manifest.contextMode).toBe('feedback'); + expect(manifest.flowId).toBe(exported.id); + }); + + it('a body override does not mutate the shared flow across requests (no cross-request leakage)', async () => { + const p = await pickFreePort(); + const handle = createFlowServer(exported, { runStep: makeScriptedRunStep(), rootDir: tmpDir }); + await handle.listen(p); + try { + // First request opts into feedback... + await postRun(`http://127.0.0.1:${p}`, { contextMode: 'feedback' }, authHeader(handle.token)); + // ...a second, plain request must NOT inherit that as a side effect. + const before = await import('node:fs/promises').then((fs) => fs.readdir(join(tmpDir, '.fluxor', 'run-context'))); + await postRun(`http://127.0.0.1:${p}`, {}, authHeader(handle.token)); + const after = await import('node:fs/promises').then((fs) => fs.readdir(join(tmpDir, '.fluxor', 'run-context'))); + // No NEW run-context directory was created for the plain (blind) request. + expect(after.length).toBe(before.length); + } finally { + await handle.close(); + } + }); + + it('inherits contextMode: "feedback" from the export itself when the body carries no override', async () => { + const feedbackExported: FluxorFlowExport = { ...exported, contextMode: 'feedback' }; + const p = await pickFreePort(); + const handle = createFlowServer(feedbackExported, { runStep: makeScriptedRunStep(), rootDir: tmpDir }); + await handle.listen(p); + try { + const { status } = await postRun(`http://127.0.0.1:${p}`, {}, authHeader(handle.token)); + expect(status).toBe(200); + } finally { + await handle.close(); + } + + const entries = await import('node:fs/promises').then((fs) => fs.readdir(join(tmpDir, '.fluxor', 'run-context'))); + expect(entries.length).toBeGreaterThan(0); + }); + + it('an explicit contextMode: "blind" override wins over a feedback export (no manifest created)', async () => { + const feedbackExported: FluxorFlowExport = { ...exported, contextMode: 'feedback' }; + const p = await pickFreePort(); + const handle = createFlowServer(feedbackExported, { runStep: makeScriptedRunStep(), rootDir: tmpDir }); + await handle.listen(p); + try { + const { status } = await postRun(`http://127.0.0.1:${p}`, { contextMode: 'blind' }, authHeader(handle.token)); + expect(status).toBe(200); + } finally { + await handle.close(); + } + + await expect( + import('node:fs/promises').then((fs) => fs.readdir(join(tmpDir, '.fluxor', 'run-context'))), + ).rejects.toThrow(); + }); +}); diff --git a/src/main/serve/serve-flow.ts b/src/main/serve/serve-flow.ts index f05648e..06b397c 100644 --- a/src/main/serve/serve-flow.ts +++ b/src/main/serve/serve-flow.ts @@ -6,7 +6,14 @@ * * POST /run — executes the flow; returns { completedStepIds, stepOutputs, finalOutput }. * When the request carries Accept: text/event-stream the response is - * streamed as Server-Sent Events. + * streamed as Server-Sent Events. Body may carry an optional + * `contextMode: 'blind' | 'feedback'` override (spec: + * docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md) + * — absent ⇒ the served flow's own mode (inherited from the + * export by importFlow) applies unchanged; an invalid value + * is rejected with 400 before execution starts. The override + * is applied to a per-request shallow clone, never mutating + * the shared flow across requests. * GET /flow — returns the loaded flow's { id, name, rootStepId, steps }. * GET /health — returns 200 { ok: true }. * @@ -52,6 +59,15 @@ export interface ServeFlowOptions { * there is no way to end up with /run left unauthenticated. */ token?: string; + /** + * Project root passed straight through to executeAgenticFlow. Matters most + * for a `contextMode: 'feedback'` flow (inherited from the export, or from + * a per-request body override — see handleRun): feedback-mode genesis + * writes the Rosetta run-context directory under this root (defaulting, as + * executeAgenticFlow always does, to `process.cwd()` when omitted). Unset + * by every existing caller, so behavior for a blind flow is unaffected. + */ + rootDir?: string; } /** @@ -206,6 +222,25 @@ function writeSseEvent(res: ServerResponse, event: HarnessEventPayload): void { res.write(`event: harness\ndata: ${data}\n\n`); } +// --------------------------------------------------------------------------- +// contextMode override (Rosetta context system, spec: +// docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md) +// --------------------------------------------------------------------------- + +/** + * Reads an optional `contextMode` override from the /run request body. + * Absent ⇒ undefined (the served flow's own mode — inherited from the export + * by importFlow — applies unchanged). Present ⇒ must be exactly 'blind' or + * 'feedback'; anything else throws so the caller gets a 400, never a silent + * fallback to the export's mode. + */ +function resolveContextModeOverride(body: Record): 'blind' | 'feedback' | undefined { + const raw = body.contextMode; + if (raw === undefined) return undefined; + if (raw === 'blind' || raw === 'feedback') return raw; + throw new Error(`Invalid contextMode "${String(raw)}" in request body — must be "blind" or "feedback".`); +} + // --------------------------------------------------------------------------- // Route handlers // --------------------------------------------------------------------------- @@ -221,6 +256,22 @@ async function handleRun( const modelId = typeof body.modelId === 'string' ? body.modelId : options.modelId; + // Validated BEFORE any SSE headers are written, so an invalid value always + // gets a clean 400 — never a silently-degraded stream. + let contextModeOverride: 'blind' | 'feedback' | undefined; + try { + contextModeOverride = resolveContextModeOverride(body); + } catch (error) { + jsonResponse(res, 400, { error: error instanceof Error ? error.message : String(error) }); + return; + } + // A shallow clone (never mutating the shared `flow` closure the server + // reuses across every request) so one request's override can never leak + // into an unrelated later request against the same server instance. + const effectiveFlow: AgenticFlow = contextModeOverride !== undefined + ? { ...flow, contextMode: contextModeOverride } + : flow; + const acceptSse = (req.headers['accept'] ?? '').includes('text/event-stream'); // Collect events for the buffered (JSON) path; relay them for SSE immediately. @@ -244,9 +295,10 @@ async function handleRun( harnessEventBus.on(HARNESS_EVENT_NAME, onEvent); try { - await executeAgenticFlow(flow, { + await executeAgenticFlow(effectiveFlow, { modelId, runStep: options.runStep, + rootDir: options.rootDir, }); const result = extractFlowCompleted(collectedEvents); diff --git a/src/renderer/__tests__/desktop-store.test.ts b/src/renderer/__tests__/desktop-store.test.ts index e0f1f3a..530d191 100644 --- a/src/renderer/__tests__/desktop-store.test.ts +++ b/src/renderer/__tests__/desktop-store.test.ts @@ -2,6 +2,7 @@ import { describe, it, expect, beforeEach } from 'vitest'; import type { PipelineAssembly } from '@/types/meta-agent'; import { LOOP_DEFAULT_MAX_ITERATIONS, LOOP_MAX_ITERATIONS_CAP } from '@/types/harness'; import { useDesktopStore, getStepMentalAttachments } from '../store/desktop-store'; +import { compileFlowFromCanvas } from '../lib/harness-compiler'; // Reset store to pristine state before each test beforeEach(() => { @@ -1264,6 +1265,48 @@ describe('updateFrameData', () => { const afterNode = useDesktopStore.getState().mentalNodes.find((n: any) => n.id === stepId); expect(afterNode).toBe(beforeNode); }); + + // ── Task U — Rosetta contextMode: the FULL data chain, end to end ────────── + // + // This is the same `updateFrameData(frameId, { contextMode: ... })` call + // FrameNode.tsx's context-mode 's value IS already exactly +// 'blind' | 'feedback' (no composite encoding like modelPolicy's +// `smart-local:`), so the onChange handler below casts +// `event.target.value` directly instead of round-tripping through a decoder. +export function contextModeToValue(mode: FrameNodeData['contextMode']): 'blind' | 'feedback' { + return mode === 'feedback' ? 'feedback' : 'blind'; +} + +const CONTEXT_MODE_OPTIONS: ReadonlyArray<{ value: 'blind' | 'feedback'; label: string }> = [ + { value: 'blind', label: 'Blind (default)' }, + { value: 'feedback', label: 'Feedback' }, +]; + // Quiet power-user affordance, not a headline — a native event.stopPropagation()} + onClick={(event) => event.stopPropagation()} + > + {CONTEXT_MODE_OPTIONS.map((opt) => ( + + ))} + + + {/* WS2 model-policy select — a quiet power-user affordance next to Export/Run, not a headline. Placed leftmost so Export stays flush next to Run (see the comment above this cluster). */} diff --git a/src/renderer/components/desktop/step-config/StepRunEvidence.tsx b/src/renderer/components/desktop/step-config/StepRunEvidence.tsx index 97116ac..f90afea 100644 --- a/src/renderer/components/desktop/step-config/StepRunEvidence.tsx +++ b/src/renderer/components/desktop/step-config/StepRunEvidence.tsx @@ -20,7 +20,9 @@ */ import React from 'react'; import { LucideIcon } from '../LucideIcon'; +import { useDesktopStore } from '../../../store/desktop-store'; import { useHarnessStore } from '../../../store/harness-store'; +import { findOwningFrame } from '../../../lib/harness-compiler'; import type { StepNodeData } from '@/types/desktop'; import type { AgenticExecutionStatus } from '@/types/harness'; import type { RoutedModelEvidence } from '@/types/ipc-events'; @@ -82,6 +84,17 @@ const WHY_MODEL_METRICS_STYLE: React.CSSProperties = { color: '#a8a8b0', }; +// ─── Rosetta context-mode badge (spec: +// docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md) ─── +// +// Reuses `.step-config-atom-tag` (the same generic accent-pill class the +// "Why this model" source tag above uses) rather than adding a new CSS rule +// — this is purely a `--atom-accent` value, not a new visual language. Violet +// ties it to the same "flow identity" hue FrameNode.tsx's workflow icon/ +// export button/context-mode select already use, distinct from every +// status/routing color (green/amber/red/blue) this panel uses elsewhere. +const CONTEXT_MODE_ACCENT = '#A78BFA'; + function formatLatency(ms: number): string { if (ms >= 60_000) return `${(ms / 60_000).toFixed(1)} min`; if (ms >= 1_000) return `${(ms / 1_000).toFixed(2)} s`; @@ -130,8 +143,41 @@ export function StepRunEvidence({ stepId, connections, status }: StepRunEvidence const stepModels = useHarnessStore((s) => s.stepModels) ?? {}; const stepModelInfo = stepModels[stepId]; + // Rosetta context-mode badge: derived from this step's OWNING FRAME (canvas + // data), not from `activeFlow` — a step-config surface must show the truth + // before any compile/run has happened, and `findOwningFrame` is the exact + // same lookup harness-compiler.ts uses to copy this value onto the compiled + // AgenticFlow (see harness-compiler.ts's `flowContextMode`). Plain (non- + // `useShallow`) selector returning a primitive — mirrors FrameNode.tsx's + // `roleCount` selector: recomputing the `.find()` every render is cheap, + // and returning just the primitive lets zustand's default `Object.is` bail + // out re-renders whenever it doesn't actually change. + const contextMode = useDesktopStore((s) => findOwningFrame(stepId, s.mentalNodes)?.data.contextMode); + return ( <> + {/* ── Rosetta context mode (feedback-only; absent in blind — no clutter + for the default/today's-behavior case) ── */} + {contextMode === 'feedback' && ( +
    +
    Context
    + + Feedback mode + +

    + This flow's steps read and write shared context files under .fluxor/run-context during the run. +

    +
    + )} + {/* ── Why this model (WS2 routing evidence) ── */} {stepModelInfo && (
    isStepGraphNode(n) && n.id === stepId); + return nodes.find( + (n): n is FrameGraphNode => + isFrameGraphNode(n) && (n.data.childIds.includes(stepId) || step?.parentId === n.id), + ); +} + function stringField(record: UnknownRecord, key: string): string | null { const value = record[key]; return typeof value === 'string' && value.trim().length > 0 ? value : null; @@ -407,20 +423,30 @@ export function compileFlowFromCanvas( }; } - // Optional human-facing flow metadata (description/tags/author/version), - // sourced from the Frame that visually owns the root step — either by - // listing it in `childIds` or via the root step's own `parentId` pointing - // back at the frame. When no frame owns the root step (or the frame sets - // none of the four fields), every key below is omitted — purely display - // metadata, never consulted by the execution pipeline. - const owningFrame = nodes.find( - (n): n is FrameGraphNode => - isFrameGraphNode(n) && (n.data.childIds.includes(rootStepId) || rootStep.parentId === n.id), - ); + // Optional human-facing flow metadata (description/tags/author/version) plus + // the Rosetta contextMode toggle, all sourced from the Frame that visually + // owns the root step — either by listing it in `childIds` or via the root + // step's own `parentId` pointing back at the frame. When no frame owns the + // root step (or the frame sets none of these fields), every key below is + // omitted. The first four are purely display metadata, never consulted by + // the execution pipeline; contextMode IS consulted (it selects the executor's + // blind/feedback path) but is copied with the exact same omit-when-absent + // shape as its siblings. + const owningFrame = findOwningFrame(rootStepId, nodes); const flowDescription = owningFrame?.data.description; const flowTags = owningFrame?.data.tags; const flowAuthor = owningFrame?.data.author; const flowVersion = owningFrame?.data.version; + // Rosetta context-mode (spec: docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md). + // Copied exactly like the other three flow-metadata fields above: omitted + // when the owning frame never set it (undefined stays undefined — every + // pre-existing flow with no Frame-level toggle keeps compiling byte- + // identically, matching AgenticFlow.contextMode's own "absent ≡ blind" + // contract). An explicit 'blind' selection is copied through too — it is + // semantically identical to omission (see that same contract) and the + // export layer (fluxor-flow.ts) is responsible for byte-identical omission + // on write, not this compiler. + const flowContextMode = owningFrame?.data.contextMode; return { id: options.flowId ?? `flow-${rootStepId}`, @@ -432,6 +458,7 @@ export function compileFlowFromCanvas( ...(flowTags && flowTags.length > 0 ? { tags: [...flowTags] } : {}), ...(flowAuthor ? { author: flowAuthor } : {}), ...(flowVersion ? { version: flowVersion } : {}), + ...(flowContextMode ? { contextMode: flowContextMode } : {}), }; } diff --git a/src/renderer/store/desktop-store.ts b/src/renderer/store/desktop-store.ts index 1ec59d0..aa79fe8 100644 --- a/src/renderer/store/desktop-store.ts +++ b/src/renderer/store/desktop-store.ts @@ -404,11 +404,13 @@ interface DesktopStore { updateStepData: (stepId: string, patch: Partial) => void; /** * Patch arbitrary fields on a Frame node's `data` (e.g. `title`, - * `description`, `tags`, `author`, `version`). Powers the Inspector's Flow - * tools editor (FlowInspector) — `compileFlowFromCanvas`'s owning-frame - * lookup reads these fields directly (see harness-compiler.ts), so edits - * made here feed the compiled AgenticFlow's matching fields on the next - * compile. No-ops (does not mutate) if `frameId` doesn't name a Frame node. + * `description`, `tags`, `author`, `version`, `contextMode`). Powers the + * Inspector's Flow tools editor (FlowInspector) and FrameNode.tsx's + * context-mode select — `compileFlowFromCanvas`'s owning-frame lookup + * (`findOwningFrame`) reads these fields directly (see + * harness-compiler.ts), so edits made here feed the compiled AgenticFlow's + * matching fields on the next compile. No-ops (does not mutate) if + * `frameId` doesn't name a Frame node. */ updateFrameData: (frameId: string, patch: Partial) => void; removeMentalNode: (nodeId: string) => void; diff --git a/src/types/desktop.ts b/src/types/desktop.ts index 8a66974..9e04d01 100644 --- a/src/types/desktop.ts +++ b/src/types/desktop.ts @@ -254,6 +254,15 @@ export interface FrameNodeData { author?: string; /** Optional user-defined flow version; flows through to AgenticFlow.version. */ version?: string; + /** + * Per-flow execution mode for the Rosetta context system; flows through to + * AgenticFlow.contextMode. Absent (today's default) or `'blind'` means zero + * cross-step context awareness and zero side effects — see + * `AgenticFlow.contextMode`'s doc (src/types/harness.ts) for the full + * contract and docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md + * for the spec. + */ + contextMode?: 'blind' | 'feedback'; childIds: string[]; missingCapabilitiesRequested?: string[]; [key: string]: unknown; diff --git a/src/types/harness.ts b/src/types/harness.ts index 383c7c2..ae8123c 100644 --- a/src/types/harness.ts +++ b/src/types/harness.ts @@ -201,4 +201,24 @@ export interface AgenticFlow { * Distinct from FLUXOR_FLOW_FORMAT_VERSION (the export wire-format version). */ version?: string; + /** + * Per-flow execution mode for the Rosetta context system (spec: + * docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md). + * + * - `'blind'` (today's behavior, and the default when this field is + * absent): each step sees only the workspace + its own prompt — zero + * cross-step awareness, zero side effects tied to this field. + * - `'feedback'`: the executor materializes a per-run context directory + * + manifest (`context-manifest.ts`) under the run's workspace and + * injects a deterministic `` block (topology summary + + * assigned file + promised briefing paths) into the system prompt so + * steps can read/write per-step briefing files with their existing FS + * tools. + * + * Absent is equivalent to `'blind'` — every pre-existing flow (authored + * before this field existed, or one that simply never sets it) keeps + * executing byte-identically: no context directory, no manifest, no + * `` block. + */ + contextMode?: 'blind' | 'feedback'; }