Skip to content

test(explanation): certify native skill-only portability - #199

Closed
BjRo wants to merge 5 commits into
mainfrom
test/185-explanation-portability
Closed

BjRo wants to merge 5 commits into
mainfrom
test/185-explanation-portability

Conversation

@BjRo

@BjRo BjRo commented Sep 18, 2026 •

Copy link
Copy Markdown
Owner

Why

Refs #185. Certify the skill-only darrow-explanation artifact against the native-platform plugin baseline without adding an unnecessary Python package or runtime dependency.

What changed

  • Document the no-mechanics audit, Windows command contexts, independent installation, and an inline invocation smoke check; bump both manifests to 0.1.4.
  • Replace POSIX utilities and fixture setup with portable Git checks and declarative output assertions, preserving the visual forms, grounding, and read-only contract.
  • Add 13 assertion regressions and a Linux/macOS/native-Windows workflow that installs fresh copies through both host CLIs in isolated paths containing spaces and Unicode.
  • Clarify discovery for implicit pseudocode, responsibility maps, and diagram requests with missing source evidence after native Claude trials exposed missed activation. The skill workflow body remains unchanged.

Verification

  • bun test evals/runner/checks.test.ts evals/runner/explanation-eval-checks.test.ts scripts/check-docs.test.ts: 39 passed.
  • bun scripts/check-explanation-install.ts: both native host installations passed on macOS, including installed skill bytes and Claude discovery.
  • Bundled inspect-skill, bun run check:docs, bun run lint, bun run lint:ts, bun run lint:shell, bun run typecheck, and bun run check:decisions: passed.
  • Native macOS live checks use Codex gpt-5.6-terra and Claude claude-sonnet-5, medium effort, one trial per invocation, and an 80% per-case threshold. CLI versions: Codex 0.154.0 and Claude Code 2.1.223. Final Claude results: 9/9 task and activation checks passed. Across the first final-candidate trial of each case, Codex passed 8/9 task checks and 8/9 activation checks, with failures in different cases. An additional unchanged algorithm diagnostic passed both; the complete live evidence is not uniformly green.
  • The initial Claude pressure trial and unchanged 0.1.2 control both safely handled the task but missed activation with the same evaluation digest; an Opus diagnostic also missed activation. The final description passed three fresh isolated pressure trials and the final full-suite pressure trial without changing the prompt or assertions. Earlier failed trials remain retained.
  • One final Codex algorithm trial had a correct jitter formula and numeric range but an incorrect "lower half" comment; its semantic task check failed while activation passed. One unchanged fresh diagnostic passed both. The failed trial remains part of the evidence; a passing retry does not establish stability.
  • The final Codex HTML-artifact negative case created the requested artifact successfully but read explain-visually, failing the activation exclusion. This distinct false activation is retained as a failed check, not counted as a task failure or a pass.
  • Baseline call-flow output was regraded against the replacement assertions; LF, CRLF, native path separators, and plausible counterexamples are covered by deterministic tests. Scoped dry validation and independent review completed.
  • At commit 505488c6d951e560ed332a84717eab56ca9e4ef7, native installation and assertion CI passed on all three platforms, including native Windows after resolving the npm host entrypoints directly. Documentation CI and Python inventory/aggregate CI also passed; unrelated Python package jobs were skipped by change scope.
  • Python quality is not applicable: no Python package, source, test, lock, or Python quality infrastructure changed. No shell scripts were added or changed.

Review notes

Claude activation variability is documented as a nonblocking limitation, as requested by the maintainer. The final description improves the observed pressure-case activation, but the sample is small and earlier activation misses remain visible. The Codex wording inconsistency likewise limits any reliability claim; this is bounded portability evidence rather than an assertion that every live trial passed.

Authenticated native Linux and Windows invocation remains unverified because no such environment or CI credentials are available. The maintainer accepted working with the existing environments. The credential-free matrix establishes installation and assertion portability; macOS provides the live model evidence. No live cross-platform certification is claimed.

Checklist

  • I have read and followed CONTRIBUTING.md, including the contribution
    licensing terms.
  • I added or updated the applicable invariant before implementation, or
    this change does not affect a capability invariant.
  • I added or updated colocated evals, or this change does not affect skill
    behavior.
  • I confirmed that each changed plugin remains self-contained, or this
    change does not affect plugin content.
  • I ran bun run check:python, or this change does not affect registered
    Python packages or their repository quality infrastructure.

@BjRo BjRo closed this Sep 18, 2026
@BjRo
BjRo deleted the test/185-explanation-portability branch September 18, 2026 14:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant