Conversation
- Migrate AI execution to the OpenCode adapter and add provider/model credential management through IPC and settings. - Add the agentic harness event bus, DAG executor, context builder, local MCP tools, LLM runner, renderer compiler, and harness store. - Move mental cards into the xyflow graph, add Step nodes, mental-to-chat attachments, flow quick rail, and supporting desktop UI/state. - Add Playwright E2E launch isolation helpers, smoke scripts, the brainstorm-cards market flow, and unit coverage for harness and mental digest behavior. - Harden dependencies to resolve npm audit findings and stop tracking generated dist/test/report artifacts via gitignore.
…ce-of-truth & cross-runtime export Turns the Performance Frontier from an LLM-judge-that-opines into a benchmark that verifies, and makes the /market the single source of truth for prebuilt agentic resources. Large coherent landing on feat/dev. Market as single source of truth (Atoms architecture) - Engine now consumes mods/roles from /market via src/main/market/market-loader.ts (.md + inventory.json), instead of hard-coding them. Code-only resources (AntiVerificationInterceptor, ConfidentExecutor) relocated to src/main/market to fix a main->renderer/store layer violation. - New prebuilt mods authored as market resources: output-budget, regression-sentinel, spec-adherence, systematic-debug, self-review, edge-case-coverage. New market `steps` category (scaffold/write-failing-tests/implement-to-green/review-diff/ refactor-safely). Meta-Agent discovery catalog (roles+mods) now sourced from the market. - Design systems are no longer a market category: modeled as mods with a mutually- exclusive `exclusiveGroup`. The whole legacy design-system renderer feature was removed at the root to stop contaminating the codebase. - Guardrail tests (src/main/market/market-integrity.test.ts): inventory<->.md sync, catalog-no-drift, PF-steps-use-only-registered-mods, no layer violation, design- system-stays-removed. Performance Frontier: new suites + richer judge - New procedural suites: development (TDD), business-knowledge (domain rules), design (atomic a11y component), progression (two-epoch brownfield on a shared VFS). - New judge dimensions: algorithmicAccuracy, domainLogicAdherence, uxUiFidelity, accessibilityScore, regressionScore; per-suite schemas + scoring + CLI + HTML report. Execution-based ground truth (the cutting-edge shift) — src/main/performance-frontier/execution - sandbox-runner: materialize a VFS snapshot to a temp dir and execute (shell:false, timeout, cleanup). Untrusted-code security note documented. - development-verifier: runs the injected vitest for real -> algorithmicAccuracy anchored. - design-verifier: renders the component in jsdom + axe-core -> accessibilityScore anchored. - api-verifier: boots the agent's Express app + real HTTP checks -> regressionScore anchored. - Hybrid scoring: objective dims verified by execution, subjective dims by the LLM judge. Judge credibility & statistics - Judge hardened: retry + graceful no-throw degradation (judgeError flag), and a configurable model via HELIOX_JUDGE_MODEL (judge != agent). Fixes live crashes from intermittent NoObjectGeneratedError. - Calibration meta-eval (calibration/): measures judge vs ground-truth (MAE, overrules); stress-calibration feeds known-broken implementations. Finding: 0 overrules (safe), but the judge weights edge-case failures harder than a linear pass-rate. - pf:bench (bench/): N-repetition runs with mean +/- 95% confidence interval (Student's t). Exportable flows: canonical format + cross-runtime conformance - src/main/flow-export/heliox-flow.ts: canonical interchange format (steps[]+dependsOn[]+ flattened systemPrompt) + bidirectional exportFlow/importFlow round-trip. - sdk/conformance/ shared, byte-identical fixtures consumed by BOTH runtimes. - sdk/java FlowImport + conformance tests prove the TS and Java runtimes traverse the same DAG order AND reproduce the same golden execution trace (vitest + mvn test). Honest divergence documented: typed-output/tool-calling parity still pending. Also: Heliox Arena (OpenRouter model discovery + per-model cost/leaderboard); token/cost accounting fix (use AI SDK totalUsage, bill reasoning tokens); team-work/design/progression step timeouts. Roadmap + audit in docs/. Verification: TS 394 tests + tsc clean; Java 31 tests (mvn) incl. conformance. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
docs/superpowers holds agent-generated planning/spec scratch files, not product source. Add it to .gitignore and untrack the files that were inadvertently committed (they stay on disk, just no longer versioned). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…tive CDP/AOM) Adds "Web Browser 2.0" as two halves sharing one Chromium surface: - Embedded preview: a 'web-preview' DesktopWindow hosting an Electron <webview> with an address bar. A main-process dev-server watcher probes candidate ports (excluding the IDE's own Vite port) and auto-opens a preview when a frontend dev server goes live. - Browser Mod: a main-process BrowserController drives the same webContents via Electron's native debugger (no remote-debugging-port). observePage() returns a compact Accessibility Object Model with numeric element ids; act() handles click/fill/select/press; extractSeo() returns meta tags + Core Web Vitals. Exposed to agents as browser_goto / browser_act / browser_extract_seo via a new WebBrowserMod (discoverable by the Meta-Agent), injected by the harness executor only when a step declares it. A "link to agent" toggle binds a preview as the active agent surface; a sandboxed headless window is the fallback. Security: guest webviews deny popups and restrict navigation to http(s); the headless agent window is sandboxed + partition-isolated; all CDP sessions and the headless window are torn down on quit. 35 new unit tests (446 total green, tsc clean). Runtime-verified: auto-open, AOM snapshot, and SEO/Core-Web-Vitals extraction against a live dev server. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…rboard A new 'arena' DesktopWindow app that parses .heliox/performance-frontier/heliox-leaderboard.json and renders a sortable, filterable table of model battle results: per-suite scores (architecture/teamWork/assembler), final arena score, tokens, execution cost, derived $/1k-tokens and free/paid tier, status, and a latency column (rendered when present). Free/Paid filter, empty + error + api_error states, refresh, and a fully responsive layout (Wrapper Principle). Read-only via a new arena:read-leaderboard IPC (ENOENT → empty list, not an error). Opens from the Dock (Trophy). Pure sort/derive helpers are unit-tested. 464 tests green, tsc clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
… telemetry Three post-v1 backlog items for the Heliox Java SDK: - Dedicated tool executor: ToolRegistry offloads argument binding + the reflective @HelioxTool invocation onto its own daemon-threaded ExecutorService (heliox-tool-N), so blocking/CPU-heavy tools no longer steal the provider's HttpClient I/O threads. HelioxRuntime owns/closes the default pool (now AutoCloseable; builder.toolExecutor(...) to inject and retain ownership of your own). - Multiple sink nodes: FlowExecutor.executeMultiSink runs a flow ending in several independent sinks, returning one typed/schema-enforced Record per sink (Map<sinkId,Object>). Single-sink execute() is unchanged. Fluent: withSinkOutputs(...).executeMultiSinkAsync(). - DAG telemetry: an optional, thread-safe DagTelemetry collector captures per-node inference time, schema-validation latency, attempts/retries and status (plus tokens when the provider reports them), serializable to JSON for the Performance Frontier report. NOOP default; opt in via FlowExecution.withTelemetry(...). 48 tests pass (was 30), BUILD SUCCESS. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…eNode
Fase 6 (Text-to-Pipeline visual layer) verification fixes:
- Camera kinematics: MetaChat's `await useReactFlow().setCenter(...)` never
resolved because MentalGraphCanvas drives a CONTROLLED React Flow viewport
(canvasPan/canvasZoom + onViewportChange), which overrides the imperative call.
Symptom found at runtime: after assembly the handler hung before
setStatus('idle') — the camera never moved, the intent never cleared, and the
submit button stayed disabled. Now centers by moving the store-controlled
viewport (setCanvasPan/setCanvasZoom), so the handler completes and the camera
frames the new pipeline.
- FrameNode polish: the frame node now shows a step-count chip ("N steps") and an
optional description subtitle on the existing translucent Figma aesthetic
(data-testid / aria-label / dragHandle unchanged).
Runtime-verified end to end: MetaChat -> real LLM assembly -> frame + child
StepNodes + dependency edges + camera centering. tsc clean, 464 tests green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…al test runner The PF Development verifier already ran the agent's code physically (real `vitest run` in a temp sandbox), but was hardcoded to `calculator.test.ts`. It now discovers the test file(s) from the VFS snapshot (any `*.test.ts | *.test.tsx | *.spec.ts` under /workspace), runs them all in one vitest invocation (sorted for determinism), and returns ran:false with a clear message when none are present — so new development cases work without touching the verifier. sandbox-runner and case-factory are unchanged. 469 tests green (+5), tsc clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
… transport - Lint: ESLint 10 dropped eslintrc + `--ext`, so `npm run lint` was broken (no config at all). Add eslint.config.mjs (typescript-eslint + @eslint/js + eslint-plugin-react-hooks, non-type-checked since tsc already enforces types) and fix the script. Pre-existing debt (no-explicit-any, no-unused-vars, etc.) is surfaced as warnings so the gate is green (0 errors, 216 warnings) and ratchetable; new violations of the rest of `recommended` still fail. - Java SDK: simplify ToolRegistry.invoke's error transport — replace the custom Binding/Invocation wrapper exceptions + sneakyThrow + exceptionally unwrap with the idiomatic `throw new CompletionException(cause)`. Behavior-preserving. tsc clean, 469 TS tests + 48 Java tests green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
evaluateArenaModel now accumulates telemetry.latencyMs across successful suites and emits avgLatencyMs (rounded mean) on each entry; api_error entries omit it. Backward-compatible optional field — the Arena dashboard already renders it (shows "—" when absent). Populates on the next `pf:arena` run. 469 tests green, tsc clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…riggers, time-travel, 3-runtime conformance Implements the competitive-positioning backlog (ARCH-063..081), turning Heliox into "the IDE for portable, verifiable agents": - serve: `heliox serve` (REST+SSE) and MCP-server expose, reusing executeAgenticFlow; Arena-informed model selection (`--select best-value`) - mcp: external MCP tool-provider mod + demand-ranked connector directory - rag: `retriever` step type, local SQLite vector store, document ingestion (clean delete via VectorStore.deleteChunksByDocId) - triggers: webhook + cron scheduler over the serve runtime - time-travel: per-step checkpoints + replayFrom/fork, canvas panel + fork indicator - cross-runtime moat: typed-output + tool-calling parity (TS<->JVM) and a new Python runtime — the same portable flow passes byte-identical conformance on three runtimes - verification-as-product: Performance Frontier scorecard + one-click Arena panels, wired into the app IPC/preload bridge Verified green: tsc; vitest (317 touched-module tests); mvn test (Java SDK, 50); pytest (Python conformance, 4). Note: also captures pre-existing in-flight working-tree changes (harness/mental UI, branding) that this work builds on, intermingled in shared files. Backlog cards, plans and competitive-analysis docs are gitignored per repo convention. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ests for reworked UI The E2E suite could not collect at all (0 tests): applied-preview.spec.ts and micro-preview.spec.ts referenced the removed inventory.designSystems at module load, aborting Playwright collection. Removed both dead specs (the design-system feature was deleted earlier). The suite now runs again (366 tests). - Mounted the 3 harness panels (Scorecard/Arena/TimeTravel) on the mental canvas behind a toggle toolbar (they were unmounted dead code) + new e2e/harness-panels.spec.ts (9 tests, green in isolation) - Fixed chat empty-state + grid-cell-empty overlays intercepting clicks (pointer-events) - Refreshed stale tests to match the reworked UI: mental cards always-visible, notifications present, marketplace total = 40 (incl. builtin tools), flows decoupled from windows (connectFlow), removed design-system Settings + marketplace -tab tests, settings adapter count Status: 348 passed / 14 failed / 4 flaky (baseline was 0 runnable). The remaining failures are concentrated in the in-flight window/attachment rework (overlapping windows + attachment-overlay z-index intercept clicks), grid confirmation-modal vs dock z-index, and a few order-dependent tests — documented for follow-up; they need the intended window/overlay design decided to resolve. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add 10 web-development mods (i18n-ready, structured-data, seo-meta, web-vitals, auth-guarded, responsive-design, dark-mode, form-validation, plus the ds-tailwind/ds-shadcn design-system exclusive group) and 2 steps (landing-page, auth-pages) — each an inventory entry + .md system injection. Fix steps never appearing in the marketplace: PluginCategory/AttachableType lacked 'steps'/'step', loadInventoryPlugins skipped inventory.steps, and there was no Steps tab/badge/deploy path. Steps now load, render, and deploy as flow-like canvas attachables (not chat-attachable). Adds regression tests. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…rails Add a 'from-scratch' PF suite: a 6-step DAG (scaffold -> landing -> auth -> write-tests -> implement -> review) that builds an advanced React 19 + Tailwind 4 + shadcn project, wiring the full web-mod catalog across 5 roles. Judged via the base rubric's optional uxUiFidelity/accessibilityScore/ algorithmicAccuracy dimensions (no judge-schema change). Add a model-agnostic guardrail engine: a declarative StepContract (mustWriteFiles, forbidStubMarkers, requiredArtifacts) is verified deterministically against the workspace after each step; on failure the executor re-runs the step with concrete corrective feedback until it passes, the attempt budget is spent, or it stalls (early-exit on no progress). Steps without a contract keep the original single-pass behaviour exactly. Fix the reasoning-model "empty output" stall: set an explicit maxOutputTokens (default 16k, HELIOX_HARNESS_MAX_OUTPUT_TOKENS) so models like Mimo stop spending the cap on reasoning and emit tool calls; make the ReAct step budget env-tunable (HELIOX_HARNESS_MAX_STEPS). On Mimo the gate lifted from-scratch 21 -> 37, forcing real SEO/JSON-LD/auth-schema/route-guard output that the ungated run skipped. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…writing stall Diagnosed from the judge's own post-mortem: on the meta-steps (write-failing- tests, implement-to-green, review-diff) the agent looped on list_directory / read_file and never called write_file — producing zero artifacts across runs. Prompt-only coaching (no model or mechanism change): an EXECUTION_DISCIPLINE block (don't explore, no list_directory, read_file is for files not dirs, the only completion is write_file, reply text is discarded), concrete target paths and content scaffolds, an explicit "DONE WHEN" mirror of each step's contract, and a loop-aware guardrail corrective. write-failing-tests and review-diff now converge (real *.test.ts with 138/70 assertions, a 13KB REVIEW.md) where they previously produced nothing; from-scratch on Mimo rose 37 -> 43 (semantic 58 -> 68/150). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ep incoherence Each meta-step produced its artifact but forked divergent systems (dual i18n, mismatched imports, an unwired ProtectedRoute), stalling runs at fail / 43. Inject a canonical PROJECT_LAYOUT (exact paths + export names) into every step; give scaffold sole ownership of the i18n and an auth-schema stub with fixed exports; have auth-pages import (not re-fork) it. Add a forbiddenArtifacts contract field (deterministic gate) banning parallel systems, plus requiredArtifacts gates that force ProtectedRoute to be wired into App.tsx and Login to import the canonical schema. On Mimo this lifted from-scratch 43 -> 68 and flipped fail -> PASS (semantic 68 -> 107/150): single i18n, ProtectedRoute wired (5 refs), coherent imports. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…for real Two model-agnostic gates that turn the judge's inferred pass into a verified one, closing the run-5 critical failures (undeclared deps, tests never run). A — requireDeclaredDependencies StepContract field + an import scanner (normalizes scoped/subpath specifiers; skips relative/@-alias/node builtins) flags any package imported but absent from package.json. Wired on auth-pages; the agent now declares react-hook-form + @hookform/resolvers. B — reuse verifyDevelopment for the from-scratch suite (one runner branch): it materializes the VFS, symlinks the host node_modules and runs vitest, so the judge anchors algorithmicAccuracy on real results. The run now executes 49 passing tests (the judge cites the verified count) instead of inferring. Score holds at PASS (66, within run-to-run variance of run-5's 68) but is now backed by executed tests and a buildable manifest rather than inference. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…link, honest review Close the run-6 polish failures: auth forms now build on the shadcn Button/ Input primitives (contract + prompt, no raw inputs), the app shell ships a skip-to-main-content link (WCAG 2.4.1), and review-diff must read the test files and report real coverage (no more "zero tests" while 49 pass). All three land (verified in the artifacts). The score holds in the PASS band (65, vs 68/66 on the prior two runs) — now variance-bound: each polish round fixes its cited failures but the judge surfaces a deeper one (here: zod messages bypass i18n; a hreflang useEffect race), so the number no longer moves. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…id, widgets, tests 53-file dead-code removal (center/ viewers, legacy panels, unused e2e suites) plus untracked feature work: HUD grid logic, session/notification/text-to-flow widgets, mental-graph extraction, and colocated unit tests. Tree validated before commit: vitest 873/873 green, eslint 0 errors. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EgGYF6UVRZ3G1KdxhVS3MW
…lease→web + Nextra) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EgGYF6UVRZ3G1KdxhVS3MW
… status README Audit actions 0.4-0.6: lint+test on every push/PR across ubuntu/macos/windows, soft-gated e2e smoke under xvfb (bridge + context-map specs, forge package step because .vite/build is gitignored), CONTRIBUTING/SECURITY/CoC/templates, README badges + truthful July 2026 status + clone URL owner fix. Implemented-by: Sonnet subagent (devops-engineer role + spec-adherence mod) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EgGYF6UVRZ3G1KdxhVS3MW
…et signing root-of-trust Audit actions 1.3-1.5: - Fonts: only 4 of 8 CDN families were actually rendered; installed those as @fontsource(-variable), removed Google Fonts links. Monaco AMD loader was silently CDN-served — now self-hosted via loader.config({ monaco }) (would have broken under CSP otherwise); vitest stub added for monaco resolution. - CSP via onHeadersReceived, app.isPackaged only; WebPreview webview traffic is exempt by architecture (separate persist:heliox-preview session). - MCP stdio spawn gate (mcp-command-policy.ts): curated-prefix match, exact persisted user approvals, HELIOX_MCP_ALLOW_ALL escape hatch, typed MCPCommandBlockedError surfaced through step status; approve/revoke IPC. Consent modal left IPC-ready (no existing dialog pattern to reuse). - market-trust.ts: ed25519 detached signature over sorted sha256 manifest; fail-closed when packaged with trusted keys, bootstrap warn while empty; scripts/market-sign.ts for key generation + signing. market/ untouched. Validated: 895/895 tests (22 new), lint 0 errors (242-warning baseline), tsc clean. Implemented-by: Sonnet subagent (security-researcher role + security-hardened mod) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EgGYF6UVRZ3G1KdxhVS3MW
…pipeline Bridge hardening (audit 1.6, Sonnet subagent: security-researcher + edge-case-coverage): - QR carries single-use 120s pairing token in URL fragment (#pt=, never hits server logs); exchange via POST body; manual PIN fallback kept - WS session token moved from query string to Sec-WebSocket-Protocol - Removed unauthenticated /bridge/qr endpoint that leaked the live PIN to LAN - Per-IP+global lockout (5 fails/60s), timingSafeEqual over SHA-256 digests, per-request session expiry, revoke-all on re-init; threat model in SECURITY.md Distribution (audit 1.1/1.2/1.7/1.8/1.10, Sonnet subagent: devops-engineer + regression-sentinel): - Env-gated macOS notarization (APPLE_ID/APPLE_ID_PASSWORD/APPLE_TEAM_ID) and Windows signing (WINDOWS_SIGN_PARAMS | WINDOWS_CERTIFICATE_FILE); unsigned builds remain the graceful default; GitHub releases now draft-first - Playwright browsers on demand: removed 350MB extraResource; installer module targets userData/pw-browsers via ELECTRON_RUN_AS_NODE; fixed load-order bug (playwright-core caches PLAYWRIGHT_BROWSERS_PATH at module load → dynamic import in SnapshotRunner.init) - update-electron-app (packaged-only, autoUpdateEnabled setting, default on) - crashReporter local-only + opt-in anonymous telemetry ping (default OFF, endpoint via env/settings, getOptIn/setOptIn IPC) - release.yml: 3-OS matrix npm ci + publish with signing secrets pass-through, checksums job uploads SHA256SUMS to the draft; CHANGELOG.md + RELEASE_CHECKLIST.md Gate (combined tree): tsc clean, 914/914 tests, lint 0 errors/241 warnings, build:bridge green, bridge e2e 8/8, forge config loads (unsigned mode logged). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EgGYF6UVRZ3G1KdxhVS3MW
…ave 1-2 implementation Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EgGYF6UVRZ3G1KdxhVS3MW
… new mods - Constraints compile into the step system prompt; validateStepAtoms gate - MarketModRuntime + MarketDomain types; domain hints in MarketplaceApp - New roles (ai/design/full-stack/mobile engineer, incident-responder) and mods (api-contract-first, conventional-commits, db-migrations-safe, docs-sync, error-ux, explain-to-me, observability-ready, privacy-guard) - StepInfoModal shows attached atoms; demo assets for README Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EgGYF6UVRZ3G1KdxhVS3MW
… cleanup)
Three root causes behind the red CI run on all OSes:
- Workflows pinned Node 20, but undici@8.5.0 (sandbox/network.ts) requires
Node >= 22.19 — importing it threw markAsUncloneable TypeError and killed
3 performance-frontier suites on every OS. Bump CI + release to Node 24.
- mcp-adapter did path arithmetic with platform resolve(), mangling virtual
posix roots ('/workspace' -> 'D:\workspace') on Windows and breaking
injected-filesystem lookups. All root/path math is now posix-form.
- api-verifier removed its temp dir before the killed child released its
cwd handles -> EBUSY on Windows, thrown from finally in violation of the
never-throws contract. Await child exit, retry rm, swallow cleanup errors.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EgGYF6UVRZ3G1KdxhVS3MW
ubuntu-latest restricts unprivileged user namespaces, so electron.launch
died instantly under xvfb ("Process failed to launch!"). Relax the
AppArmor sysctl in the e2e job, per Electron's CI guidance, instead of
baking --no-sandbox into the app under test.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EgGYF6UVRZ3G1KdxhVS3MW
The Electron binary download intermittently fails to materialize during npm ci on macos-latest (2 of 3 runs), so every suite importing 'electron' dies at collection. Verify the binary resolves after install and re-run electron's install.js when it doesn't; a genuine download failure now fails this step loudly instead. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EgGYF6UVRZ3G1KdxhVS3MW
Adds a Python SDK runtime targeting the conformance contract (DAG + scripted trace), a LoopConfig type and loop-execution test on the Java runtime, and shared conformance fixtures (contract + loop goldens, scripted loop responses) plus a dedicated sdk-conformance CI workflow. Widens the cross-runtime portable-flow moat toward text+typed+tools parity across the TS/Java/Python runtimes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013nbxSeU5FDHmFY3A1akRgE
…harness WIP Bundles two sessions of previously-uncommitted work. They are intermingled at file level (this session's edits share files with the prior WIP), so they are committed together rather than split artificially. This session — plan 2026-07-05 (canvas-inspector-widgets-providers-ci): - Edges: direction follows drag order (orient-connection), invertMentalEdge + invert-order menu on flow/loop edges, persisted-edge normalization, double-click to edit loop iterations. - Visual: step connector handles render as full circles (.step-node-clip); window selection ring drawn on the outer frame, uncut by attachments. - Inspector perf: narrowed store subscriptions + React.memo on canvas node/edge components; collapsed-expand button docked top-right. - HUD widgets: collision-free grid placement with a top-right safe zone (72x120) sized to reserve the collapsed-inspector button. - Sidebar: Flows/Steps section; Mental Cards lists only mental cards. - Providers: DBeaver-style connection profiles (safeStorage-encrypted tokens that never cross IPC), protocol-aware model listing, harness execution against conn:<id>/<model>; replaces the opencode section. Prior sessions (canvas overhaul + loops/routing/export): multi-board switcher with debounced storage, loop-back edges, dual-source model router, canvas->flow export (markdown/.flow.json), step quick-add/thinking popovers, Monaco workers, right-hand inspector, shared EdgeChrome, checkpoint iteration ids. Full suite green: 1387 tests / 90 files, tsc 0, lint 0 errors. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013nbxSeU5FDHmFY3A1akRgE
Adds a build job (ubuntu/macos/windows matrix, push-only) that runs `npm run make` and uploads the out/make artifacts, so each merge yields downloadable installers. Reuses the test job's Electron-binary self-heal step. Signing stays a release.yml concern; forge degrades to unsigned. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013nbxSeU5FDHmFY3A1akRgE
…data Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013nbxSeU5FDHmFY3A1akRgE
…ds on all 3 OS
The new build job failed on every runner:
- Linux deb/rpm: "could not find the Electron app binary at .../heliox-ide" —
Packager names the binary after `name` ("Heliox IDE") but the makers expect
the lowercase package name ("heliox-ide").
- macOS maker-dmg: "Cannot find module 'appdmg'" — its darwin-only native dep
is absent from the cross-platform lockfile on the hosted runner.
- Windows Squirrel: exited 1.
Fix:
- packagerConfig.executableName: 'heliox-ide' — one consistent binary name every
maker resolves (fixes deb/rpm; the packaged .app now ships MacOS/heliox-ide).
- Add @electron-forge/maker-zip and emit a portable .zip of the packaged app on
darwin/linux/win32 — pure-JS, no native/optional toolchain, reliable on every
runner.
- Gate the polished installers (dmg/squirrel/deb/rpm) behind HELIOX_MAKE_ZIP_ONLY;
ci.yml's build job sets it so per-push builds are fast and green. Tagged
releases (release.yml) and local `make` still build the full installer set.
Validated locally: `HELIOX_MAKE_ZIP_ONLY=1 npm run make` produces
out/make/zip/darwin/arm64/Heliox IDE-darwin-arm64-0.1.0.zip.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nbxSeU5FDHmFY3A1akRgE
…n every push to main Adds a publish-latest job (needs: build, main pushes only) that aggregates the three OS zip builds, writes SHA256SUMS, and recreates the `latest` prerelease pointing at the current commit — a stable, always-current public download the marketing site serves via the GitHub Releases API. Feature branches keep only the ephemeral per-run build artifacts. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013nbxSeU5FDHmFY3A1akRgE
CMolG
added a commit
that referenced
this pull request
Jul 10, 2026
Campaña blind vs feedback ejecutada (mimo/mimo-v2.5-pro, seed 7): - team-work (handoffs): 79→84 (+6.33%), +28.94% tokens → feedback ayuda - progression (iterativa): 52→22 (-57.69%) → feedback perjudica Conclusión (F5): feedback queda OPT-IN, blind sigue siendo default — depende de la topología del flow, activarlo por defecto degradaría los flows de continuidad iterativa. Resultado y salvedad estadística (n=1/celda, follow-up ≥3 seeds) documentados en docs/pf-context-mode-campaign.md. Además: redondeo de valores en la tabla de pf:compare (latencyMs traía ruido float; Δ% se sigue calculando sobre los valores crudos). 41 tests verdes. Backlog #1 F0–F5 COMPLETO. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
CMolG
added a commit
that referenced
this pull request
Jul 10, 2026
…) + tooling PF (#6) * feat(harness): flow context modes — blind vs feedback (piedra Rosetta) + tooling PF Ejecuta la tarea de backlog 2026-07-08-flow-context-modes (F0–F4 + tooling F5; la campaña PF comparativa queda a un comando, pendiente de GO por coste): - AgenticFlow.contextMode ('blind' default | 'feedback'); blind byte-idéntico probado empíricamente (captura de prompts/checkpoints pre/post: 0 diffs). - Rosetta: context-manifest.ts (manifest determinista por run, ficheros step.<id>[.iterN].md bajo .fluxor/run-context/<runId>/, retriever exento); génesis solo en feedback; <flow_awareness> 100% determinista en system prompt; briefings prompt-driven leídos via FS tools (nunca inyectados). - Guardrail de briefings activo en TODAS las superficies: adaptador real-FS cuando falta options.fileSystem (solo feedback), rootDir efectivo materializado, walk de snapshotWorkspace con exclusión node_modules/.git. - Checkpoints: campo aditivo contextFileSnapshot (shape existente intacto). - Formato: export estampa contextMode (omitido en blind); round-trip export→import→serve verde; serve acepta override validado por request. - UI: toggle de modo en FrameNode (pipeline-frame-actions) → FrameNodeData → harness-compiler → AgenticFlow (integración sin mocks); badge en StepRunEvidence; WCAG AAA auditado. - SDKs: downgrade explícito a blind con warning once en java/python (contextMode preservado verbatim); conformance + README matriz runtime×modo. - PF: --context-mode en pf:run/pf:bench, contextMode en ledger (viejos=blind), nuevo pf:compare (markdown, empareja seed+suite+modelo, upsert idempotente en docs/pf-context-mode-campaign.md). Gates: vitest 1566/1566 · tsc limpio · eslint 0 err · java 66/66 (1 skip gateado) · python 15/15 · round-trip manual OK. E2e desktop local pendiente de puerto 5173 (ocupado por otro dev server del usuario) — gate pre-merge. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * pf(context-modes): run F5 campaign (seed 7) + round wall-time display Campaña blind vs feedback ejecutada (mimo/mimo-v2.5-pro, seed 7): - team-work (handoffs): 79→84 (+6.33%), +28.94% tokens → feedback ayuda - progression (iterativa): 52→22 (-57.69%) → feedback perjudica Conclusión (F5): feedback queda OPT-IN, blind sigue siendo default — depende de la topología del flow, activarlo por defecto degradaría los flows de continuidad iterativa. Resultado y salvedad estadística (n=1/celda, follow-up ≥3 seeds) documentados en docs/pf-context-mode-campaign.md. Además: redondeo de valores en la tabla de pf:compare (latencyMs traía ruido float; Δ% se sigue calculando sobre los valores crudos). 41 tests verdes. Backlog #1 F0–F5 COMPLETO. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
CMolG
added a commit
that referenced
this pull request
Jul 10, 2026
…e on code (#7) Ejecuta la tarea de backlog 2026-07-08-chats-to-flow-steps (F0–F5): - Window type 'chat' RETIRADO (desktop.ts); migración perezosa v19 (tombstone, no crash de boards guardados) + v20 (rename HUD widget). El auto-chat es un panel fijo del HUD (HudAutoChatPanel), caja de intent one-shot vía meta-agent assemblePipeline → insertPipelineAssembly — sin agent-manager, sin edición de código por construcción. - Gesto mono-step: doble-click en canvas vacío crea un step con input enfocado (pendingStepFocusId) + streaming; Cmd+N y "New Step" del context-menu igual. - 6 lanzadores de flows (BacklogKanban executeCard/auto-architect, FlowDeck, BacklogCardModal, AgentSessions, App Cmd+N) re-cableados: materializan al board + run por harness, ya NO abren chat (criterio #1 cumplido). - Helms en steps: rol único (Single Persona, listbox/aria-selected), mods attached; input directo + adjuntos @file (FileContextBuilder reciclado); nota de autoridad de ejecución (validateStepAtoms veto visible). - StepRunEvidence gana transcript real (react-markdown + remark-gfm, tool-call cards, file-change chips → diff viewer); degrada con gracia donde el harness no emite datos estructurados (follow-up: evento FileChanged). - Market: MarketFlow.steps[] estructurado (superset de PipelineAssemblyStep), "Add to board" construye el assembly directo → insertPipelineAssembly; AttachableFlow/RightFlowAttachment retirados. Re-firma diferida. - Barrido de muerto: AgenticChatApp, MessageRenderer, SlashCommandHandler, TextToFlowWidget borrados (agent-manager CONSERVADO: FlowsEditor aún lo usa). Gates: vitest 1615/1615 · tsc 0 · eslint 0 err · e2e desktop.spec.ts sin regresiones (19 fallos pre-existentes idénticos a origin/main) · context-menus verde · new-features recuperado 25→9 (los 9 = tests e2e del panel auto-chat, non-CI, feature cubierta por unit tests — follow-up de persist/rehidratación). Backlog #2 completo. Las 3 tareas del backlog (rebranding, context-modes, chats→steps) entregadas. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Lands the previously-uncommitted
feat/devwork as four thematic commits. This bundles two sessions of work — they are intermingled at file level, so they are committed by area rather than split artificially.Highlights
Canvas / edges
invertMentalEdge+ "Invert order" on flow/loop edge menus; persisted-edge normalization heals old boards; double-click to edit loop iterations.Visual fixes
.step-node-clip); window selection ring drawn on the outer frame, uncut by role/mod attachments.Inspector performance
React.memoon canvas node/edge components → selection re-renders only the affected nodes; collapsed-expand button docked top-right.HUD widgets
Sidebar
Provider Connections (DBeaver-style) — replaces the opencode providers section
safeStorage-encrypted tokens that never cross IPC (onlyhasToken); protocol-aware model listing (openai/anthropic); the harness executes steps againstconn:<id>/<model>.CI
buildjob (ubuntu/macos/windows, push-only) emits a portable.zipof the packaged app on every OS (executableName: heliox-ide+maker-zip, zip-only viaHELIOX_MAKE_ZIP_ONLY). Rich installers (dmg/squirrel/deb/rpm) stay arelease.ymlconcern.Cross-runtime SDK
LoopConfig+ loop test; shared conformance goldens; dedicatedsdk-conformanceworkflow.Verification
tsc --noEmitclean,eslint0 errors.test+buildpass on all 3 OS; downloadable per-OS executables uploaded (windows ~147 MB, linux ~122 MB, macOS ~119 MB).e2e-smokeis a pre-existing soft-gate (Electron-under-xvfb flake,continue-on-error).Follow-ups (not blocking)
npm start) of the touched surfaces.npx playwright test) not run locally (heavy; soft-gate in CI).🤖 Generated with Claude Code
https://claude.ai/code/session_013nbxSeU5FDHmFY3A1akRgE