Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
65 changes: 65 additions & 0 deletions docs/pf-context-mode-campaign.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,65 @@
# Campaña PF: Context Modes (blind vs feedback)

> **Estado: ✅ campaña ejecutada (2026-07-10, seed 7). Conclusión: feedback opt-in, blind default. Ver abajo.** Requiere GO explícito del usuario: correrla consume tokens reales de LLM (4 runs de suites `team-work`/`progression`, output-heavy). Ver plan: [`../superpowers/plans/2026-07-10-flow-context-modes.md`](superpowers/plans/2026-07-10-flow-context-modes.md) (Task P) y spec: [`../superpowers/specs/2026-07-10-rosetta-context-manifest.md`](superpowers/specs/2026-07-10-rosetta-context-manifest.md).

## Objetivo

Medir el efecto real de `contextMode: 'feedback'` (manifest Rosetta + bloque `<flow_awareness>` + briefings prompt-driven entre steps) frente a `'blind'` (comportamiento de hoy, sin cambios) sobre las **mismas** suites, seed y modelo — criterio de aceptación 4 del backlog ([`../superpowers/backlog/2026-07-08-flow-context-modes-blind-vs-feedback.md`](superpowers/backlog/2026-07-08-flow-context-modes-blind-vs-feedback.md), fase F5) y las tres dimensiones que ahí se piden: **puntuación del judge, coste en tokens, wall-time**.

## Suites elegidas

| Suite | Por qué |
|---|---|
| `team-work` | 3 agentes secuenciales con handoffs explícitos (planner → designer → developer). Es el caso donde la consciencia de flow — topología, fichero de contexto asignado, briefings prometidos — debería importar más: cada step depende literalmente de lo que el anterior promete entregarle. |
| `progression` | 2 épocas sobre el mismo VFS no destruido. Mide si el modo feedback ayuda a no regresar trabajo previo, más allá de lo que ya captura el judge de regresión (`regressionScore`) hoy. |

`flow-assembler` queda deliberadamente fuera: no ejecuta un `AgenticFlow` (usa `assemblePipeline` directamente), así que `--context-mode` es un no-op ahí — `pf:run`/`pf:bench` avisan de esto en vez de callar (Step P1).

## Diseño del experimento

- **Misma seed, mismo modelo, ambos modos.** Por cada suite: una corrida en `blind` y una en `feedback`, con idéntico `--seed` e idéntico `--model`. `pf:compare` empareja por (suite, seed, modelo) — si cualquiera de los tres difiere entre las dos corridas de una suite, no habrá pareja que comparar (queda como "unpaired" en el reporte, no se descarta en silencio).
- **Modelo:** el default del PF, `mimo/mimo-v2.5-pro` (override con `--model=` si se quiere repetir la campaña sobre otro modelo — en ese caso usar la MISMA seed, y `pf:compare` seguirá emparejando correctamente porque el modelo también forma parte de la clave de agrupación).
- **Seed sugerida:** `7` (arbitraria y fija — lo único que importa es que blind y feedback de una misma suite compartan la MISMA seed).
- **Métricas comparadas** (por suite+modelo, blind → feedback, Δ%): puntuación final del judge (`finalScore`), tokens totales (`telemetry.totalTokens`), wall-time (`telemetry.latencyMs`, suma de latencias LLM por step — ver `telemetry/collector.ts`).

## Checklist de ejecución (pendiente de GO del usuario)

- [x] Confirmar GO del usuario — las 4 corridas consumen tokens reales de LLM (`team-work`/`progression` son suites output-heavy: 3 steps secuenciales y 2 épocas respectivamente).
- [x] `npm run pf:run -- --suite=team-work --context-mode=blind --seed=7` → finalScore 79
- [x] `npm run pf:run -- --suite=team-work --context-mode=feedback --seed=7` → finalScore 84
- [x] `npm run pf:run -- --suite=progression --context-mode=blind --seed=7` → finalScore 52
- [x] `npm run pf:run -- --suite=progression --context-mode=feedback --seed=7` → finalScore 22
- [x] `npm run pf:compare -- --seed=7 --suites=team-work,progression`
- [x] Revisar la sección "Resultados" de este documento (la escribe el comando anterior) y decidir: ¿`feedback` pasa a ser el modo recomendado por defecto, queda opt-in documentado, o se descarta por el sobrecoste de tokens sin mejora de puntuación que lo justifique? Esa conclusión también es un resultado publicable (ver spec, decisión de refinement 4).

## Comando exacto (tras las 4 corridas de arriba)

```
npm run pf:compare -- --seed=7 --suites=team-work,progression
```

`pf:compare` imprime la tabla en stdout Y la escribe en la sección "Resultados" de abajo (reemplaza solo esa sección; el resto de este documento queda intacto — puede re-ejecutarse tantas veces como se repita la campaña).

## Resultados

_Generated by `pf:compare` — do not hand-edit; re-run the command to refresh._

Seed: 7 · Suites: team-work, progression

| Suite | Model | Judge Score (blind → feedback, Δ%) | Tokens (blind → feedback, Δ%) | Wall-time ms (blind → feedback, Δ%) |
|---|---|---|---|---|
| progression | mimo/mimo-v2.5-pro | 52 → 22 (-57.69%) | 17252 → 15567 (-9.77%) | 53915 → 34386 (-36.22%) |
| team-work | mimo/mimo-v2.5-pro | 79 → 84 (+6.33%) | 91366 → 117811 (+28.94%) | 240019 → 223204 (-7.01%) |

## Conclusión (2026-07-10, seed 7, n=1 por celda)

**Feedback NO pasa a default. Queda opt-in documentado; `blind` sigue siendo el modo por defecto** — respuesta al último ítem del checklist y a la decisión de refinement 4 del backlog (F5).

El resultado es **mixto y depende de la topología del flow**:

- **`team-work` (3 agentes secuenciales con handoffs): feedback ayuda** — judge 79→84 (+6.33%), a costa de +28.94% tokens. Es el caso que la hipótesis predecía: cuando cada step depende literalmente de lo que el anterior le entrega, la consciencia de flow + los briefings prometidos mejoran el resultado. El sobrecoste de tokens es real y hay que pesarlo.
- **`progression` (2 épocas sobre el mismo VFS): feedback perjudica, y mucho** — judge 52→22 (−57.69%), con tokens y wall-time incluso menores. En un flujo de continuidad iterativa (una sola línea de trabajo que progresa), el bloque `<flow_awareness>` + los ficheros de briefing parecen distraer al modelo de la tarea real en vez de coordinar handoffs que aquí no existen.

**Lectura:** feedback es una herramienta para flows con handoffs explícitos entre roles distintos, no un interruptor universalmente bueno. Activarlo por defecto habría degradado toda una clase de flows. Esto valida la decisión de diseño (blind default, feedback opt-in a nivel flow) con datos, y es exactamente el tipo de resultado que el backlog anticipó como publicable aunque fuese negativo.

**Salvedad estadística:** n=1 por celda (una seed). Los deltas son grandes (sobre todo el −58% de progression, difícilmente ruido), pero para una recomendación de producto firme conviene repetir con ≥3 seeds — `pf:compare` ya empareja por seed, así que es re-ejecutar la campaña con otras seeds y promediar. Queda como follow-up, no bloquea la conclusión cualitativa.
1 change: 1 addition & 0 deletions package.json
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,7 @@
"pf:run": "tsx src/main/performance-frontier/cli.ts",
"pf:bench": "tsx src/main/performance-frontier/bench/cli.ts",
"pf:arena": "tsx src/main/performance-frontier/arena/arena-runner.ts",
"pf:compare": "tsx src/main/performance-frontier/compare/cli.ts",
"fluxor:baseline": "node -e \"console.log('Run baselines via Fluxor IDE UI — Flows tab → Initialize Baselines')\"",
"build:bridge": "vite build --config vite.bridge.config.ts",
"fluxor:init": "node -e \"const fs=require('fs'),p=require('path');const d=p.join(process.cwd(),'fluxor');const legacy=p.join(process.cwd(),'heliox');if(fs.existsSync(legacy)&&!fs.existsSync(d)){fs.renameSync(legacy,d);console.log('Legacy compat: migrated ./heliox/ to ./fluxor/');}if(!fs.existsSync(d))fs.mkdirSync(d,{recursive:true});['flows.json','metrics.config.json'].forEach(f=>{const fp=p.join(d,f);if(!fs.existsSync(fp))fs.copyFileSync(p.join(__dirname,'assets','fluxor-templates',f),fp)});console.log('Fluxor config initialized in ./fluxor/')\"",
Expand Down
65 changes: 65 additions & 0 deletions sdk/conformance/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,7 @@ conformance suite against these shared fixtures. The contract is:
| Loop execution parity (expansion algorithm → golden loop trace) | TS: `semantic-parity.test.ts`; Java: `CrossRuntimeConformanceTest.loopExecutionMatchesGoldenTrace`; Python: `test_loop_execution_matches_golden_trace` |
| `contract`/`model` carried opaquely round-trip | TS: `fluxor-flow.test.ts`; Java: `CrossRuntimeConformanceTest.contractAndModelAreCarried`; Python: `test_contract_and_model_are_carried` |
| Legacy/absent `format` tag imports with one deprecation warning — `"fluxor-flow"` is silent, absent/null is assumed `"heliox-flow"`, any other value is reported as seen; never blocks parsing | TS: `fluxor-flow.test.ts` suite 7 ("legacy format tag compat"); Java: `CrossRuntimeConformanceTest.legacyFormatImportsWithDeprecationWarning` / `currentFormatFlowsNeverWarn` / `absentFormatWarnsAssumingLegacyHelioxFlow` / `unknownFormatWarnsMentioningSeenValue` / `legacyWarningIsEmittedOncePerDistinctValue`; Python: `test_legacy_format_imports_with_deprecation_warning` / `test_current_format_flows_never_warn` / `test_absent_format_warns_assuming_legacy_heliox_flow` / `test_unknown_format_warns_mentioning_seen_value` |
| `contextMode: "feedback"` imports on Java/Python with ONE downgrade-to-blind warning (TS executes it natively); the value is preserved verbatim; `"blind"`/absent/any other value is silent; never blocks parsing | TS: `fluxor-flow.test.ts` suite 8 ("contextMode export/import round-trip") + `semantic-parity.test.ts` ("feedback-mode contextMode is additive"); Java: `CrossRuntimeConformanceTest.feedbackContextModeDowngradesToBlindWithWarning` / `blindContextModeNeverWarns` / `absentContextModeNeverWarns` / `unknownContextModeValueNeverWarns` / `feedbackDowngradeWarningIsEmittedOncePerProcess`; Python: `test_feedback_context_mode_downgrades_to_blind_with_warning` / `test_blind_context_mode_never_warns` / `test_absent_context_mode_never_warns` / `test_unknown_context_mode_value_never_warns` / `test_feedback_downgrade_warning_is_emitted_once` — see [Context modes](#context-modes--rosetta-downgrade-on-javapython-2026-07-10) |

### Running all three suites

Expand Down Expand Up @@ -195,3 +196,67 @@ predates the tag, and the TS round-trip test compensates by spreading
suite 5, "importFlow(conformance-contract.flow.json) → exportFlow deep-equals
golden-contract-roundtrip.json"). Regenerate the golden with the tag if that workaround is
ever retired.

## Context modes — Rosetta downgrade on Java/Python (2026-07-10)

The wire format carries an optional top-level `contextMode` string (`AgenticFlow.contextMode`
in `src/types/harness.ts`; spec:
`docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md`). It selects how a run threads
context between steps: **blind** (the default — steps only see their declared upstream
outputs, exactly as before the field existed) or **feedback** (the TS harness materializes a
`.fluxor/run-context/<runId>/` directory with a Rosetta manifest and per-step context files,
and injects a deterministic `<flow_awareness>` block). The exporter omits the key entirely for
blind/absent flows and stamps only the literal `"feedback"`, so blind exports stay
byte-identical to every pre-Rosetta export.

### Runtime × mode support matrix

| Runtime | `contextMode` absent / `"blind"` | `contextMode: "feedback"` |
|---|---|---|
| **TypeScript** (IDE / serve) | Native (byte-identical to pre-Rosetta behaviour) | **Native** — run-context genesis, manifest, `<flow_awareness>`, briefing guardrail |
| **Java** (`sdk/java`) | Native (blind is the only implemented mode) | **Downgrade to blind** — one warning on `System.err`, then executes blind |
| **Python** (`sdk/python`) | Native (blind is the only implemented mode) | **Downgrade to blind** — one `DeprecationWarning`, then executes blind |

### The downgrade contract (spec decision 2 — frozen)

Feedback-mode parity in Java/Python is a **v2** concern (the PF measurement campaign runs on
the TS runtime); in v1 both SDKs perform an explicit, documented downgrade at import:

1. **`contextMode === "feedback"`** → the import succeeds unchanged and emits **one** warning
with the exact wording:

> feedback mode is not supported by this runtime yet; downgrading to blind

Execution then proceeds in blind mode — the only mode these executors implement (they have
no run-context manifest/briefing machinery; neither executor ever reads `contextMode`, so
blind execution is structural, not a branch).
2. **`"blind"`, absent/null, or any other value** → total silence: there is nothing to
downgrade. The SDKs do not validate the enum (only the TS IDE authoring surface does); an
unrecognized value is carried like any other opaque field.

In every case `contextMode` is **preserved verbatim** on the parsed definition
(`FlowDefinition.contextMode()` in Java, `FlowDefinition.context_mode` in Python — `null`/
`None` when absent). The downgrade changes runtime behaviour, never the recorded value, so
the definition round-trips without loss.

Warning channels and once-semantics mirror the legacy-format shim exactly: Java prints one
line to `System.err`, deduped by a once-per-process flag (test reset hook
`resetContextModeWarningsForTests`); Python raises one `DeprecationWarning` via
`warnings.warn`, dedup delegated to the stdlib warnings filter. Unlike the format shim's
message, the downgrade wording above is itself part of the frozen contract — it may **not**
vary per runtime.

- **`conformance-feedback-mode.flow.json`** — a minimal 1-step flow carrying
`"format": "fluxor-flow"` and `"contextMode": "feedback"` (rule 1). Proof: Java
`CrossRuntimeConformanceTest.feedbackContextModeDowngradesToBlindWithWarning` /
`feedbackDowngradeWarningIsEmittedOncePerProcess`; Python
`test_feedback_context_mode_downgrades_to_blind_with_warning` /
`test_feedback_downgrade_warning_is_emitted_once`.
- The **silent cases** (rule 2) are proven with inline JSON and the existing chain fixture:
Java `blindContextModeNeverWarns` / `absentContextModeNeverWarns` /
`unknownContextModeValueNeverWarns`; Python `test_blind_context_mode_never_warns` /
`test_absent_context_mode_never_warns` / `test_unknown_context_mode_value_never_warns`.
- The **TS side** (native execution, no downgrade) is proven by `fluxor-flow.test.ts` suite 8
("contextMode export/import round-trip" — omit-when-blind, stamp-only-`"feedback"`) and
`semantic-parity.test.ts` ("feedback-mode contextMode is additive — does not perturb the
golden trace", which also materializes a real run-context manifest).
17 changes: 17 additions & 0 deletions sdk/conformance/conformance-feedback-mode.flow.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
{
"version": "1",
"format": "fluxor-flow",
"id": "conformance-feedback-mode",
"name": "Conformance Feedback Mode",
"contextMode": "feedback",
"rootStepId": "step-a",
"steps": [
{
"id": "step-a",
"type": "llm_call",
"prompt": "Step A: a flow exported with contextMode 'feedback' must still import cleanly, downgrading to blind execution with one warning on this runtime.",
"dependsOn": [],
"tools": []
}
]
}
22 changes: 21 additions & 1 deletion sdk/java/src/main/java/io/fluxor/sdk/flow/FlowDefinition.java
Original file line number Diff line number Diff line change
Expand Up @@ -12,14 +12,34 @@
* body re-runs, up to {@code maxIterations} total passes. The forward graph stays acyclic;
* loops are expanded into a per-iteration instance graph at execution time by
* {@link io.fluxor.sdk.engine.FlowExecutor#executeAllTextTrace}.
*
* <p>{@code contextMode} is the Rosetta context mode ({@code AgenticFlow.contextMode} in
* {@code src/types/harness.ts}; spec:
* {@code docs/superpowers/specs/2026-07-10-rosetta-context-manifest.md}), carried
* <em>opaquely</em> like {@link StepConfig#contract()}/{@link StepConfig#model()}: this
* runtime never branches on it. {@code null} when the flow declares none (blind default).
* {@code "feedback"} is NOT supported by this runtime in v1 — {@link FlowImport} emits a
* one-time downgrade warning at import and execution proceeds in blind mode (the only mode
* this executor implements); the recorded value stays {@code "feedback"} so the definition
* round-trips without loss (spec decision 2: downgrade, not parity).
*/
public record FlowDefinition(String id, List<StepConfig> steps, List<LoopConfig> loops) {
public record FlowDefinition(String id, List<StepConfig> steps, List<LoopConfig> loops,
String contextMode) {

public FlowDefinition {
steps = List.copyOf(steps);
loops = loops == null ? List.of() : List.copyOf(loops);
}

/**
* Compatibility constructor for callers built against the pre-Rosetta 3-arg shape —
* defaults {@code contextMode} to {@code null} (carry-opaque field; absent means the flow
* declares none, i.e. blind) so existing call sites compile unchanged.
*/
public FlowDefinition(String id, List<StepConfig> steps, List<LoopConfig> loops) {
this(id, steps, loops, null);
}

/**
* Compatibility constructor for callers built against the pre-loop 2-arg shape — defaults
* to no loops so existing call sites compile unchanged.
Expand Down
Loading
Loading