Skip to content

Canvas/inspector/widgets/Provider Connections + CI builds + cross-runtime SDK - #1

Merged
CMolG merged 34 commits into
mainfrom
feat/dev
Jul 5, 2026
Merged

CMolG merged 34 commits into
mainfrom
feat/dev

Conversation

@CMolG

@CMolG CMolG commented Jul 5, 2026 •

Copy link
Copy Markdown
Owner

Summary

Lands the previously-uncommitted feat/dev work as four thematic commits. This bundles two sessions of work — they are intermingled at file level, so they are committed by area rather than split artificially.

Highlights

Canvas / edges

  • Edge direction follows drag order (no more phantom loop badges on forward connectors); invertMentalEdge + "Invert order" on flow/loop edge menus; persisted-edge normalization heals old boards; double-click to edit loop iterations.

Visual fixes

  • Step connector handles render as full circles (inner .step-node-clip); window selection ring drawn on the outer frame, uncut by role/mod attachments.

Inspector performance

  • Narrowed store subscriptions + React.memo on canvas node/edge components → selection re-renders only the affected nodes; collapsed-expand button docked top-right.

HUD widgets

  • Collision-free grid placement with a top-right safe zone (72×120) reserving the collapsed-inspector button.

Sidebar

  • Dedicated Flows/Steps section; Mental Cards lists only mental cards (orphan steps get a home).

Provider Connections (DBeaver-style) — replaces the opencode providers section

  • Connection profiles with safeStorage-encrypted tokens that never cross IPC (only hasToken); protocol-aware model listing (openai/anthropic); the harness executes steps against conn:<id>/<model>.

CI

  • New build job (ubuntu/macos/windows, push-only) emits a portable .zip of the packaged app on every OS (executableName: heliox-ide + maker-zip, zip-only via HELIOX_MAKE_ZIP_ONLY). Rich installers (dmg/squirrel/deb/rpm) stay a release.yml concern.

Cross-runtime SDK

  • Python runtime targeting the conformance contract; Java LoopConfig + loop test; shared conformance goldens; dedicated sdk-conformance workflow.

Verification

  • 1387 tests / 90 files green, tsc --noEmit clean, eslint 0 errors.
  • CI green (run 28746070841): test + build pass on all 3 OS; downloadable per-OS executables uploaded (windows ~147 MB, linux ~122 MB, macOS ~119 MB). e2e-smoke is a pre-existing soft-gate (Electron-under-xvfb flake, continue-on-error).

Follow-ups (not blocking)

  • Manual visual QA sweep (npm start) of the touched surfaces.
  • E2E (npx playwright test) not run locally (heavy; soft-gate in CI).

🤖 Generated with Claude Code

https://claude.ai/code/session_013nbxSeU5FDHmFY3A1akRgE

CMolG and others added 30 commits June 22, 2026 15:44
- Migrate AI execution to the OpenCode adapter and add provider/model credential management through IPC and settings.

- Add the agentic harness event bus, DAG executor, context builder, local MCP tools, LLM runner, renderer compiler, and harness store.

- Move mental cards into the xyflow graph, add Step nodes, mental-to-chat attachments, flow quick rail, and supporting desktop UI/state.

- Add Playwright E2E launch isolation helpers, smoke scripts, the brainstorm-cards market flow, and unit coverage for harness and mental digest behavior.

- Harden dependencies to resolve npm audit findings and stop tracking generated dist/test/report artifacts via gitignore.
…ce-of-truth & cross-runtime export

Turns the Performance Frontier from an LLM-judge-that-opines into a benchmark
that verifies, and makes the /market the single source of truth for prebuilt
agentic resources. Large coherent landing on feat/dev.

Market as single source of truth (Atoms architecture)
- Engine now consumes mods/roles from /market via src/main/market/market-loader.ts
  (.md + inventory.json), instead of hard-coding them. Code-only resources
  (AntiVerificationInterceptor, ConfidentExecutor) relocated to src/main/market
  to fix a main->renderer/store layer violation.
- New prebuilt mods authored as market resources: output-budget, regression-sentinel,
  spec-adherence, systematic-debug, self-review, edge-case-coverage. New market
  `steps` category (scaffold/write-failing-tests/implement-to-green/review-diff/
  refactor-safely). Meta-Agent discovery catalog (roles+mods) now sourced from the market.
- Design systems are no longer a market category: modeled as mods with a mutually-
  exclusive `exclusiveGroup`. The whole legacy design-system renderer feature was
  removed at the root to stop contaminating the codebase.
- Guardrail tests (src/main/market/market-integrity.test.ts): inventory<->.md sync,
  catalog-no-drift, PF-steps-use-only-registered-mods, no layer violation, design-
  system-stays-removed.

Performance Frontier: new suites + richer judge
- New procedural suites: development (TDD), business-knowledge (domain rules),
  design (atomic a11y component), progression (two-epoch brownfield on a shared VFS).
- New judge dimensions: algorithmicAccuracy, domainLogicAdherence, uxUiFidelity,
  accessibilityScore, regressionScore; per-suite schemas + scoring + CLI + HTML report.

Execution-based ground truth (the cutting-edge shift) — src/main/performance-frontier/execution
- sandbox-runner: materialize a VFS snapshot to a temp dir and execute (shell:false,
  timeout, cleanup). Untrusted-code security note documented.
- development-verifier: runs the injected vitest for real -> algorithmicAccuracy anchored.
- design-verifier: renders the component in jsdom + axe-core -> accessibilityScore anchored.
- api-verifier: boots the agent's Express app + real HTTP checks -> regressionScore anchored.
- Hybrid scoring: objective dims verified by execution, subjective dims by the LLM judge.

Judge credibility & statistics
- Judge hardened: retry + graceful no-throw degradation (judgeError flag), and a
  configurable model via HELIOX_JUDGE_MODEL (judge != agent). Fixes live crashes
  from intermittent NoObjectGeneratedError.
- Calibration meta-eval (calibration/): measures judge vs ground-truth (MAE, overrules);
  stress-calibration feeds known-broken implementations. Finding: 0 overrules (safe),
  but the judge weights edge-case failures harder than a linear pass-rate.
- pf:bench (bench/): N-repetition runs with mean +/- 95% confidence interval (Student's t).

Exportable flows: canonical format + cross-runtime conformance
- src/main/flow-export/heliox-flow.ts: canonical interchange format (steps[]+dependsOn[]+
  flattened systemPrompt) + bidirectional exportFlow/importFlow round-trip.
- sdk/conformance/ shared, byte-identical fixtures consumed by BOTH runtimes.
- sdk/java FlowImport + conformance tests prove the TS and Java runtimes traverse the
  same DAG order AND reproduce the same golden execution trace (vitest + mvn test).
  Honest divergence documented: typed-output/tool-calling parity still pending.

Also: Heliox Arena (OpenRouter model discovery + per-model cost/leaderboard); token/cost
accounting fix (use AI SDK totalUsage, bill reasoning tokens); team-work/design/progression
step timeouts. Roadmap + audit in docs/.

Verification: TS 394 tests + tsc clean; Java 31 tests (mvn) incl. conformance.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
docs/superpowers holds agent-generated planning/spec scratch files, not product
source. Add it to .gitignore and untrack the files that were inadvertently
committed (they stay on disk, just no longer versioned).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…tive CDP/AOM)

Adds "Web Browser 2.0" as two halves sharing one Chromium surface:

- Embedded preview: a 'web-preview' DesktopWindow hosting an Electron <webview>
  with an address bar. A main-process dev-server watcher probes candidate ports
  (excluding the IDE's own Vite port) and auto-opens a preview when a frontend
  dev server goes live.
- Browser Mod: a main-process BrowserController drives the same webContents via
  Electron's native debugger (no remote-debugging-port). observePage() returns a
  compact Accessibility Object Model with numeric element ids; act() handles
  click/fill/select/press; extractSeo() returns meta tags + Core Web Vitals.
  Exposed to agents as browser_goto / browser_act / browser_extract_seo via a new
  WebBrowserMod (discoverable by the Meta-Agent), injected by the harness executor
  only when a step declares it. A "link to agent" toggle binds a preview as the
  active agent surface; a sandboxed headless window is the fallback.

Security: guest webviews deny popups and restrict navigation to http(s); the
headless agent window is sandboxed + partition-isolated; all CDP sessions and the
headless window are torn down on quit.

35 new unit tests (446 total green, tsc clean). Runtime-verified: auto-open,
AOM snapshot, and SEO/Core-Web-Vitals extraction against a live dev server.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…rboard

A new 'arena' DesktopWindow app that parses
.heliox/performance-frontier/heliox-leaderboard.json and renders a sortable,
filterable table of model battle results: per-suite scores
(architecture/teamWork/assembler), final arena score, tokens, execution cost,
derived $/1k-tokens and free/paid tier, status, and a latency column (rendered
when present). Free/Paid filter, empty + error + api_error states, refresh, and
a fully responsive layout (Wrapper Principle).

Read-only via a new arena:read-leaderboard IPC (ENOENT → empty list, not an
error). Opens from the Dock (Trophy). Pure sort/derive helpers are unit-tested.

464 tests green, tsc clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
… telemetry

Three post-v1 backlog items for the Heliox Java SDK:

- Dedicated tool executor: ToolRegistry offloads argument binding + the reflective
  @HelioxTool invocation onto its own daemon-threaded ExecutorService (heliox-tool-N),
  so blocking/CPU-heavy tools no longer steal the provider's HttpClient I/O threads.
  HelioxRuntime owns/closes the default pool (now AutoCloseable; builder.toolExecutor(...)
  to inject and retain ownership of your own).
- Multiple sink nodes: FlowExecutor.executeMultiSink runs a flow ending in several
  independent sinks, returning one typed/schema-enforced Record per sink
  (Map<sinkId,Object>). Single-sink execute() is unchanged. Fluent:
  withSinkOutputs(...).executeMultiSinkAsync().
- DAG telemetry: an optional, thread-safe DagTelemetry collector captures per-node
  inference time, schema-validation latency, attempts/retries and status (plus tokens
  when the provider reports them), serializable to JSON for the Performance Frontier
  report. NOOP default; opt in via FlowExecution.withTelemetry(...).

48 tests pass (was 30), BUILD SUCCESS.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…eNode

Fase 6 (Text-to-Pipeline visual layer) verification fixes:

- Camera kinematics: MetaChat's `await useReactFlow().setCenter(...)` never
  resolved because MentalGraphCanvas drives a CONTROLLED React Flow viewport
  (canvasPan/canvasZoom + onViewportChange), which overrides the imperative call.
  Symptom found at runtime: after assembly the handler hung before
  setStatus('idle') — the camera never moved, the intent never cleared, and the
  submit button stayed disabled. Now centers by moving the store-controlled
  viewport (setCanvasPan/setCanvasZoom), so the handler completes and the camera
  frames the new pipeline.
- FrameNode polish: the frame node now shows a step-count chip ("N steps") and an
  optional description subtitle on the existing translucent Figma aesthetic
  (data-testid / aria-label / dragHandle unchanged).

Runtime-verified end to end: MetaChat -> real LLM assembly -> frame + child
StepNodes + dependency edges + camera centering. tsc clean, 464 tests green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…al test runner

The PF Development verifier already ran the agent's code physically (real
`vitest run` in a temp sandbox), but was hardcoded to `calculator.test.ts`. It
now discovers the test file(s) from the VFS snapshot (any
`*.test.ts | *.test.tsx | *.spec.ts` under /workspace), runs them all in one
vitest invocation (sorted for determinism), and returns ran:false with a clear
message when none are present — so new development cases work without touching
the verifier. sandbox-runner and case-factory are unchanged.

469 tests green (+5), tsc clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
… transport

- Lint: ESLint 10 dropped eslintrc + `--ext`, so `npm run lint` was broken (no
  config at all). Add eslint.config.mjs (typescript-eslint + @eslint/js +
  eslint-plugin-react-hooks, non-type-checked since tsc already enforces types)
  and fix the script. Pre-existing debt (no-explicit-any, no-unused-vars, etc.)
  is surfaced as warnings so the gate is green (0 errors, 216 warnings) and
  ratchetable; new violations of the rest of `recommended` still fail.
- Java SDK: simplify ToolRegistry.invoke's error transport — replace the custom
  Binding/Invocation wrapper exceptions + sneakyThrow + exceptionally unwrap with
  the idiomatic `throw new CompletionException(cause)`. Behavior-preserving.

tsc clean, 469 TS tests + 48 Java tests green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
evaluateArenaModel now accumulates telemetry.latencyMs across successful suites
and emits avgLatencyMs (rounded mean) on each entry; api_error entries omit it.
Backward-compatible optional field — the Arena dashboard already renders it
(shows "—" when absent). Populates on the next `pf:arena` run.

469 tests green, tsc clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…riggers, time-travel, 3-runtime conformance

Implements the competitive-positioning backlog (ARCH-063..081), turning Heliox into
"the IDE for portable, verifiable agents":

- serve: `heliox serve` (REST+SSE) and MCP-server expose, reusing executeAgenticFlow;
  Arena-informed model selection (`--select best-value`)
- mcp: external MCP tool-provider mod + demand-ranked connector directory
- rag: `retriever` step type, local SQLite vector store, document ingestion
  (clean delete via VectorStore.deleteChunksByDocId)
- triggers: webhook + cron scheduler over the serve runtime
- time-travel: per-step checkpoints + replayFrom/fork, canvas panel + fork indicator
- cross-runtime moat: typed-output + tool-calling parity (TS<->JVM) and a new Python
  runtime — the same portable flow passes byte-identical conformance on three runtimes
- verification-as-product: Performance Frontier scorecard + one-click Arena panels,
  wired into the app IPC/preload bridge

Verified green: tsc; vitest (317 touched-module tests); mvn test (Java SDK, 50);
pytest (Python conformance, 4).

Note: also captures pre-existing in-flight working-tree changes (harness/mental UI,
branding) that this work builds on, intermingled in shared files. Backlog cards,
plans and competitive-analysis docs are gitignored per repo convention.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ests for reworked UI

The E2E suite could not collect at all (0 tests): applied-preview.spec.ts and
micro-preview.spec.ts referenced the removed inventory.designSystems at module
load, aborting Playwright collection. Removed both dead specs (the design-system
feature was deleted earlier). The suite now runs again (366 tests).

- Mounted the 3 harness panels (Scorecard/Arena/TimeTravel) on the mental canvas
  behind a toggle toolbar (they were unmounted dead code) + new
  e2e/harness-panels.spec.ts (9 tests, green in isolation)
- Fixed chat empty-state + grid-cell-empty overlays intercepting clicks (pointer-events)
- Refreshed stale tests to match the reworked UI: mental cards always-visible,
  notifications present, marketplace total = 40 (incl. builtin tools), flows
  decoupled from windows (connectFlow), removed design-system Settings + marketplace
  -tab tests, settings adapter count

Status: 348 passed / 14 failed / 4 flaky (baseline was 0 runnable). The remaining
failures are concentrated in the in-flight window/attachment rework (overlapping
windows + attachment-overlay z-index intercept clicks), grid confirmation-modal vs
dock z-index, and a few order-dependent tests — documented for follow-up; they need
the intended window/overlay design decided to resolve.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add 10 web-development mods (i18n-ready, structured-data, seo-meta,
web-vitals, auth-guarded, responsive-design, dark-mode, form-validation,
plus the ds-tailwind/ds-shadcn design-system exclusive group) and 2 steps
(landing-page, auth-pages) — each an inventory entry + .md system injection.

Fix steps never appearing in the marketplace: PluginCategory/AttachableType
lacked 'steps'/'step', loadInventoryPlugins skipped inventory.steps, and
there was no Steps tab/badge/deploy path. Steps now load, render, and deploy
as flow-like canvas attachables (not chat-attachable). Adds regression tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…rails

Add a 'from-scratch' PF suite: a 6-step DAG (scaffold -> landing -> auth ->
write-tests -> implement -> review) that builds an advanced React 19 +
Tailwind 4 + shadcn project, wiring the full web-mod catalog across 5 roles.
Judged via the base rubric's optional uxUiFidelity/accessibilityScore/
algorithmicAccuracy dimensions (no judge-schema change).

Add a model-agnostic guardrail engine: a declarative StepContract
(mustWriteFiles, forbidStubMarkers, requiredArtifacts) is verified
deterministically against the workspace after each step; on failure the
executor re-runs the step with concrete corrective feedback until it passes,
the attempt budget is spent, or it stalls (early-exit on no progress). Steps
without a contract keep the original single-pass behaviour exactly.

Fix the reasoning-model "empty output" stall: set an explicit maxOutputTokens
(default 16k, HELIOX_HARNESS_MAX_OUTPUT_TOKENS) so models like Mimo stop
spending the cap on reasoning and emit tool calls; make the ReAct step budget
env-tunable (HELIOX_HARNESS_MAX_STEPS). On Mimo the gate lifted from-scratch
21 -> 37, forcing real SEO/JSON-LD/auth-schema/route-guard output that the
ungated run skipped.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…writing stall

Diagnosed from the judge's own post-mortem: on the meta-steps (write-failing-
tests, implement-to-green, review-diff) the agent looped on list_directory /
read_file and never called write_file — producing zero artifacts across runs.

Prompt-only coaching (no model or mechanism change): an EXECUTION_DISCIPLINE
block (don't explore, no list_directory, read_file is for files not dirs, the
only completion is write_file, reply text is discarded), concrete target paths
and content scaffolds, an explicit "DONE WHEN" mirror of each step's contract,
and a loop-aware guardrail corrective.

write-failing-tests and review-diff now converge (real *.test.ts with 138/70
assertions, a 13KB REVIEW.md) where they previously produced nothing;
from-scratch on Mimo rose 37 -> 43 (semantic 58 -> 68/150).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ep incoherence

Each meta-step produced its artifact but forked divergent systems (dual i18n,
mismatched imports, an unwired ProtectedRoute), stalling runs at fail / 43.

Inject a canonical PROJECT_LAYOUT (exact paths + export names) into every step;
give scaffold sole ownership of the i18n and an auth-schema stub with fixed
exports; have auth-pages import (not re-fork) it. Add a forbiddenArtifacts
contract field (deterministic gate) banning parallel systems, plus
requiredArtifacts gates that force ProtectedRoute to be wired into App.tsx and
Login to import the canonical schema.

On Mimo this lifted from-scratch 43 -> 68 and flipped fail -> PASS (semantic
68 -> 107/150): single i18n, ProtectedRoute wired (5 refs), coherent imports.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…for real

Two model-agnostic gates that turn the judge's inferred pass into a verified
one, closing the run-5 critical failures (undeclared deps, tests never run).

A — requireDeclaredDependencies StepContract field + an import scanner
(normalizes scoped/subpath specifiers; skips relative/@-alias/node builtins)
flags any package imported but absent from package.json. Wired on auth-pages;
the agent now declares react-hook-form + @hookform/resolvers.

B — reuse verifyDevelopment for the from-scratch suite (one runner branch):
it materializes the VFS, symlinks the host node_modules and runs vitest, so
the judge anchors algorithmicAccuracy on real results. The run now executes
49 passing tests (the judge cites the verified count) instead of inferring.

Score holds at PASS (66, within run-to-run variance of run-5's 68) but is now
backed by executed tests and a buildable manifest rather than inference.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…link, honest review

Close the run-6 polish failures: auth forms now build on the shadcn Button/
Input primitives (contract + prompt, no raw inputs), the app shell ships a
skip-to-main-content link (WCAG 2.4.1), and review-diff must read the test
files and report real coverage (no more "zero tests" while 49 pass).

All three land (verified in the artifacts). The score holds in the PASS band
(65, vs 68/66 on the prior two runs) — now variance-bound: each polish round
fixes its cited failures but the judge surfaces a deeper one (here: zod
messages bypass i18n; a hreflang useEffect race), so the number no longer moves.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…id, widgets, tests

53-file dead-code removal (center/ viewers, legacy panels, unused e2e suites)
plus untracked feature work: HUD grid logic, session/notification/text-to-flow
widgets, mental-graph extraction, and colocated unit tests. Tree validated
before commit: vitest 873/873 green, eslint 0 errors.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EgGYF6UVRZ3G1KdxhVS3MW
…lease→web + Nextra)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EgGYF6UVRZ3G1KdxhVS3MW
… status README

Audit actions 0.4-0.6: lint+test on every push/PR across ubuntu/macos/windows,
soft-gated e2e smoke under xvfb (bridge + context-map specs, forge package
step because .vite/build is gitignored), CONTRIBUTING/SECURITY/CoC/templates,
README badges + truthful July 2026 status + clone URL owner fix.

Implemented-by: Sonnet subagent (devops-engineer role + spec-adherence mod)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EgGYF6UVRZ3G1KdxhVS3MW
…et signing root-of-trust

Audit actions 1.3-1.5:
- Fonts: only 4 of 8 CDN families were actually rendered; installed those as
  @fontsource(-variable), removed Google Fonts links. Monaco AMD loader was
  silently CDN-served — now self-hosted via loader.config({ monaco }) (would
  have broken under CSP otherwise); vitest stub added for monaco resolution.
- CSP via onHeadersReceived, app.isPackaged only; WebPreview webview traffic
  is exempt by architecture (separate persist:heliox-preview session).
- MCP stdio spawn gate (mcp-command-policy.ts): curated-prefix match, exact
  persisted user approvals, HELIOX_MCP_ALLOW_ALL escape hatch, typed
  MCPCommandBlockedError surfaced through step status; approve/revoke IPC.
  Consent modal left IPC-ready (no existing dialog pattern to reuse).
- market-trust.ts: ed25519 detached signature over sorted sha256 manifest;
  fail-closed when packaged with trusted keys, bootstrap warn while empty;
  scripts/market-sign.ts for key generation + signing. market/ untouched.

Validated: 895/895 tests (22 new), lint 0 errors (242-warning baseline), tsc clean.

Implemented-by: Sonnet subagent (security-researcher role + security-hardened mod)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EgGYF6UVRZ3G1KdxhVS3MW
…pipeline

Bridge hardening (audit 1.6, Sonnet subagent: security-researcher + edge-case-coverage):
- QR carries single-use 120s pairing token in URL fragment (#pt=, never hits
  server logs); exchange via POST body; manual PIN fallback kept
- WS session token moved from query string to Sec-WebSocket-Protocol
- Removed unauthenticated /bridge/qr endpoint that leaked the live PIN to LAN
- Per-IP+global lockout (5 fails/60s), timingSafeEqual over SHA-256 digests,
  per-request session expiry, revoke-all on re-init; threat model in SECURITY.md

Distribution (audit 1.1/1.2/1.7/1.8/1.10, Sonnet subagent: devops-engineer + regression-sentinel):
- Env-gated macOS notarization (APPLE_ID/APPLE_ID_PASSWORD/APPLE_TEAM_ID) and
  Windows signing (WINDOWS_SIGN_PARAMS | WINDOWS_CERTIFICATE_FILE); unsigned
  builds remain the graceful default; GitHub releases now draft-first
- Playwright browsers on demand: removed 350MB extraResource; installer module
  targets userData/pw-browsers via ELECTRON_RUN_AS_NODE; fixed load-order bug
  (playwright-core caches PLAYWRIGHT_BROWSERS_PATH at module load → dynamic
  import in SnapshotRunner.init)
- update-electron-app (packaged-only, autoUpdateEnabled setting, default on)
- crashReporter local-only + opt-in anonymous telemetry ping (default OFF,
  endpoint via env/settings, getOptIn/setOptIn IPC)
- release.yml: 3-OS matrix npm ci + publish with signing secrets pass-through,
  checksums job uploads SHA256SUMS to the draft; CHANGELOG.md + RELEASE_CHECKLIST.md

Gate (combined tree): tsc clean, 914/914 tests, lint 0 errors/241 warnings,
build:bridge green, bridge e2e 8/8, forge config loads (unsigned mode logged).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EgGYF6UVRZ3G1KdxhVS3MW
…ave 1-2 implementation

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EgGYF6UVRZ3G1KdxhVS3MW
… new mods

- Constraints compile into the step system prompt; validateStepAtoms gate
- MarketModRuntime + MarketDomain types; domain hints in MarketplaceApp
- New roles (ai/design/full-stack/mobile engineer, incident-responder) and
  mods (api-contract-first, conventional-commits, db-migrations-safe,
  docs-sync, error-ux, explain-to-me, observability-ready, privacy-guard)
- StepInfoModal shows attached atoms; demo assets for README

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EgGYF6UVRZ3G1KdxhVS3MW
… cleanup)

Three root causes behind the red CI run on all OSes:

- Workflows pinned Node 20, but undici@8.5.0 (sandbox/network.ts) requires
  Node >= 22.19 — importing it threw markAsUncloneable TypeError and killed
  3 performance-frontier suites on every OS. Bump CI + release to Node 24.
- mcp-adapter did path arithmetic with platform resolve(), mangling virtual
  posix roots ('/workspace' -> 'D:\workspace') on Windows and breaking
  injected-filesystem lookups. All root/path math is now posix-form.
- api-verifier removed its temp dir before the killed child released its
  cwd handles -> EBUSY on Windows, thrown from finally in violation of the
  never-throws contract. Await child exit, retry rm, swallow cleanup errors.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EgGYF6UVRZ3G1KdxhVS3MW
ubuntu-latest restricts unprivileged user namespaces, so electron.launch
died instantly under xvfb ("Process failed to launch!"). Relax the
AppArmor sysctl in the e2e job, per Electron's CI guidance, instead of
baking --no-sandbox into the app under test.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EgGYF6UVRZ3G1KdxhVS3MW
The Electron binary download intermittently fails to materialize during
npm ci on macos-latest (2 of 3 runs), so every suite importing 'electron'
dies at collection. Verify the binary resolves after install and re-run
electron's install.js when it doesn't; a genuine download failure now
fails this step loudly instead.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EgGYF6UVRZ3G1KdxhVS3MW
Adds a Python SDK runtime targeting the conformance contract (DAG +
scripted trace), a LoopConfig type and loop-execution test on the Java
runtime, and shared conformance fixtures (contract + loop goldens,
scripted loop responses) plus a dedicated sdk-conformance CI workflow.
Widens the cross-runtime portable-flow moat toward text+typed+tools
parity across the TS/Java/Python runtimes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nbxSeU5FDHmFY3A1akRgE
…harness WIP

Bundles two sessions of previously-uncommitted work. They are intermingled
at file level (this session's edits share files with the prior WIP), so
they are committed together rather than split artificially.

This session — plan 2026-07-05 (canvas-inspector-widgets-providers-ci):
- Edges: direction follows drag order (orient-connection), invertMentalEdge
  + invert-order menu on flow/loop edges, persisted-edge normalization,
  double-click to edit loop iterations.
- Visual: step connector handles render as full circles (.step-node-clip);
  window selection ring drawn on the outer frame, uncut by attachments.
- Inspector perf: narrowed store subscriptions + React.memo on canvas
  node/edge components; collapsed-expand button docked top-right.
- HUD widgets: collision-free grid placement with a top-right safe zone
  (72x120) sized to reserve the collapsed-inspector button.
- Sidebar: Flows/Steps section; Mental Cards lists only mental cards.
- Providers: DBeaver-style connection profiles (safeStorage-encrypted
  tokens that never cross IPC), protocol-aware model listing, harness
  execution against conn:<id>/<model>; replaces the opencode section.

Prior sessions (canvas overhaul + loops/routing/export): multi-board
switcher with debounced storage, loop-back edges, dual-source model
router, canvas->flow export (markdown/.flow.json), step quick-add/thinking
popovers, Monaco workers, right-hand inspector, shared EdgeChrome,
checkpoint iteration ids.

Full suite green: 1387 tests / 90 files, tsc 0, lint 0 errors.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nbxSeU5FDHmFY3A1akRgE
CMolG and others added 4 commits July 5, 2026 17:28
Adds a build job (ubuntu/macos/windows matrix, push-only) that runs
`npm run make` and uploads the out/make artifacts, so each merge yields
downloadable installers. Reuses the test job's Electron-binary self-heal
step. Signing stays a release.yml concern; forge degrades to unsigned.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nbxSeU5FDHmFY3A1akRgE
…ds on all 3 OS

The new build job failed on every runner:
- Linux deb/rpm: "could not find the Electron app binary at .../heliox-ide" —
  Packager names the binary after `name` ("Heliox IDE") but the makers expect
  the lowercase package name ("heliox-ide").
- macOS maker-dmg: "Cannot find module 'appdmg'" — its darwin-only native dep
  is absent from the cross-platform lockfile on the hosted runner.
- Windows Squirrel: exited 1.

Fix:
- packagerConfig.executableName: 'heliox-ide' — one consistent binary name every
  maker resolves (fixes deb/rpm; the packaged .app now ships MacOS/heliox-ide).
- Add @electron-forge/maker-zip and emit a portable .zip of the packaged app on
  darwin/linux/win32 — pure-JS, no native/optional toolchain, reliable on every
  runner.
- Gate the polished installers (dmg/squirrel/deb/rpm) behind HELIOX_MAKE_ZIP_ONLY;
  ci.yml's build job sets it so per-push builds are fast and green. Tagged
  releases (release.yml) and local `make` still build the full installer set.

Validated locally: `HELIOX_MAKE_ZIP_ONLY=1 npm run make` produces
out/make/zip/darwin/arm64/Heliox IDE-darwin-arm64-0.1.0.zip.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nbxSeU5FDHmFY3A1akRgE
…n every push to main

Adds a publish-latest job (needs: build, main pushes only) that aggregates the
three OS zip builds, writes SHA256SUMS, and recreates the `latest` prerelease
pointing at the current commit — a stable, always-current public download the
marketing site serves via the GitHub Releases API. Feature branches keep only
the ephemeral per-run build artifacts.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nbxSeU5FDHmFY3A1akRgE
@CMolG
CMolG merged commit 8bd6603 into main Jul 5, 2026
15 of 17 checks passed
CMolG added a commit that referenced this pull request Jul 10, 2026
Campaña blind vs feedback ejecutada (mimo/mimo-v2.5-pro, seed 7):
- team-work (handoffs): 79→84 (+6.33%), +28.94% tokens → feedback ayuda
- progression (iterativa): 52→22 (-57.69%) → feedback perjudica

Conclusión (F5): feedback queda OPT-IN, blind sigue siendo default —
depende de la topología del flow, activarlo por defecto degradaría los
flows de continuidad iterativa. Resultado y salvedad estadística (n=1/celda,
follow-up ≥3 seeds) documentados en docs/pf-context-mode-campaign.md.

Además: redondeo de valores en la tabla de pf:compare (latencyMs traía ruido
float; Δ% se sigue calculando sobre los valores crudos). 41 tests verdes.

Backlog #1 F0–F5 COMPLETO.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
CMolG added a commit that referenced this pull request Jul 10, 2026
…) + tooling PF (#6)

* feat(harness): flow context modes — blind vs feedback (piedra Rosetta) + tooling PF

Ejecuta la tarea de backlog 2026-07-08-flow-context-modes (F0–F4 + tooling F5;
la campaña PF comparativa queda a un comando, pendiente de GO por coste):

- AgenticFlow.contextMode ('blind' default | 'feedback'); blind byte-idéntico
  probado empíricamente (captura de prompts/checkpoints pre/post: 0 diffs).
- Rosetta: context-manifest.ts (manifest determinista por run, ficheros
  step.<id>[.iterN].md bajo .fluxor/run-context/<runId>/, retriever exento);
  génesis solo en feedback; <flow_awareness> 100% determinista en system
  prompt; briefings prompt-driven leídos via FS tools (nunca inyectados).
- Guardrail de briefings activo en TODAS las superficies: adaptador real-FS
  cuando falta options.fileSystem (solo feedback), rootDir efectivo
  materializado, walk de snapshotWorkspace con exclusión node_modules/.git.
- Checkpoints: campo aditivo contextFileSnapshot (shape existente intacto).
- Formato: export estampa contextMode (omitido en blind); round-trip
  export→import→serve verde; serve acepta override validado por request.
- UI: toggle de modo en FrameNode (pipeline-frame-actions) → FrameNodeData →
  harness-compiler → AgenticFlow (integración sin mocks); badge en
  StepRunEvidence; WCAG AAA auditado.
- SDKs: downgrade explícito a blind con warning once en java/python
  (contextMode preservado verbatim); conformance + README matriz runtime×modo.
- PF: --context-mode en pf:run/pf:bench, contextMode en ledger (viejos=blind),
  nuevo pf:compare (markdown, empareja seed+suite+modelo, upsert idempotente
  en docs/pf-context-mode-campaign.md).

Gates: vitest 1566/1566 · tsc limpio · eslint 0 err · java 66/66 (1 skip
gateado) · python 15/15 · round-trip manual OK. E2e desktop local pendiente
de puerto 5173 (ocupado por otro dev server del usuario) — gate pre-merge.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* pf(context-modes): run F5 campaign (seed 7) + round wall-time display

Campaña blind vs feedback ejecutada (mimo/mimo-v2.5-pro, seed 7):
- team-work (handoffs): 79→84 (+6.33%), +28.94% tokens → feedback ayuda
- progression (iterativa): 52→22 (-57.69%) → feedback perjudica

Conclusión (F5): feedback queda OPT-IN, blind sigue siendo default —
depende de la topología del flow, activarlo por defecto degradaría los
flows de continuidad iterativa. Resultado y salvedad estadística (n=1/celda,
follow-up ≥3 seeds) documentados en docs/pf-context-mode-campaign.md.

Además: redondeo de valores en la tabla de pf:compare (latencyMs traía ruido
float; Δ% se sigue calculando sobre los valores crudos). 41 tests verdes.

Backlog #1 F0–F5 COMPLETO.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
CMolG added a commit that referenced this pull request Jul 10, 2026
…e on code (#7)

Ejecuta la tarea de backlog 2026-07-08-chats-to-flow-steps (F0–F5):

- Window type 'chat' RETIRADO (desktop.ts); migración perezosa v19 (tombstone,
  no crash de boards guardados) + v20 (rename HUD widget). El auto-chat es un
  panel fijo del HUD (HudAutoChatPanel), caja de intent one-shot vía meta-agent
  assemblePipeline → insertPipelineAssembly — sin agent-manager, sin edición de
  código por construcción.
- Gesto mono-step: doble-click en canvas vacío crea un step con input enfocado
  (pendingStepFocusId) + streaming; Cmd+N y "New Step" del context-menu igual.
- 6 lanzadores de flows (BacklogKanban executeCard/auto-architect, FlowDeck,
  BacklogCardModal, AgentSessions, App Cmd+N) re-cableados: materializan al
  board + run por harness, ya NO abren chat (criterio #1 cumplido).
- Helms en steps: rol único (Single Persona, listbox/aria-selected), mods
  attached; input directo + adjuntos @file (FileContextBuilder reciclado);
  nota de autoridad de ejecución (validateStepAtoms veto visible).
- StepRunEvidence gana transcript real (react-markdown + remark-gfm, tool-call
  cards, file-change chips → diff viewer); degrada con gracia donde el harness
  no emite datos estructurados (follow-up: evento FileChanged).
- Market: MarketFlow.steps[] estructurado (superset de PipelineAssemblyStep),
  "Add to board" construye el assembly directo → insertPipelineAssembly;
  AttachableFlow/RightFlowAttachment retirados. Re-firma diferida.
- Barrido de muerto: AgenticChatApp, MessageRenderer, SlashCommandHandler,
  TextToFlowWidget borrados (agent-manager CONSERVADO: FlowsEditor aún lo usa).

Gates: vitest 1615/1615 · tsc 0 · eslint 0 err · e2e desktop.spec.ts sin
regresiones (19 fallos pre-existentes idénticos a origin/main) · context-menus
verde · new-features recuperado 25→9 (los 9 = tests e2e del panel auto-chat,
non-CI, feature cubierta por unit tests — follow-up de persist/rehidratación).

Backlog #2 completo. Las 3 tareas del backlog (rebranding, context-modes,
chats→steps) entregadas.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant