Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions .gitleaks.toml
Original file line number Diff line number Diff line change
Expand Up @@ -50,3 +50,11 @@ regexes = [
paths = [
'''(^|/)contracts/fixtures/random-stream-vectors/.*''',
]

[[allowlists]]
description = "Issue #1350: exact synthetic DBOS workflow ID in the source-hashed experiment; not a credential. Other values, paths and scanner rules remain checked."
targetRules = ["generic-api-key"]
condition = "AND"
regexTarget = "secret"
regexes = ['''^dbos-r2-cancel$''']
paths = ['''(^|/)docs/research/execution-architecture/experiments/probe_dbos\.py$''']
9 changes: 9 additions & 0 deletions docs/decisions/adrs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -145,10 +145,18 @@ adr-098-portable-artifact-requirement-satisfaction
adr-099-participant-relative-predicate-opacity
adr-100-participant-crossing-bisimulation
adr-101-adversarial-participant-flow-control
adr-102-mixed-cross-backend-participant-control
adr-103-branch-aware-python-coverage-policy
adr-104-runtime-control-plane-architecture
adr-105-recursive-partial-description-semantics
adr-106-developer-package-and-artifact-management
adr-107-artifact-promotion-and-release-admission
adr-108-modular-participant-control-and-governed-effects
adr-109-participant-identity-and-objective-assignment
adr-110-reusable-mixed-control-policies-and-occurrences
adr-111-control-applicability-and-effect-decisions
adr-112-external-inject-triggering-and-execution
adr-113-reusable-execution-machinery
```

| ADR | Title | Status | Date |
Expand Down Expand Up @@ -265,3 +273,4 @@ adr-108-modular-participant-control-and-governed-effects
| [110](adr-110-reusable-mixed-control-policies-and-occurrences.md) | Reusable Mixed-Control Policies and Occurrences | accepted | 2026-09-22 |
| [111](adr-111-control-applicability-and-effect-decisions.md) | Control Applicability and Effect Decisions | accepted | 2026-09-22 |
| [112](adr-112-external-inject-triggering-and-execution.md) | External Inject Triggering and Execution | accepted | 2026-09-22 |
| [113](adr-113-reusable-execution-machinery.md) | Reusable Execution Machinery Under RAE Authority | accepted | 2026-09-23 |
325 changes: 325 additions & 0 deletions docs/decisions/adrs/adr-113-reusable-execution-machinery.md

Large diffs are not rendered by default.

3 changes: 3 additions & 0 deletions docs/decisions/adrs/adr-index.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -643,3 +643,6 @@ adrs:
- id: ADR-112
path: docs/decisions/adrs/adr-112-external-inject-triggering-and-execution.md
pin: 65e54addfb3a2dba74941f3c8030dec8c7a39661cc5aebb456e59c03af120ea3
- id: ADR-113
path: docs/decisions/adrs/adr-113-reusable-execution-machinery.md
pin: 6d904befc0d37d44b07ed083d297c7a0e3ba0d520d79f9ddce1999477c78e91d
296 changes: 296 additions & 0 deletions docs/decisions/issue-1350-execution-architecture-preflight.md

Large diffs are not rendered by default.

9 changes: 8 additions & 1 deletion docs/requirements/API-404/requirement.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ type: FUNCTIONAL
priority: MUST
wave: 1
created_at: 2026-04-03T05:55:58.825305Z
updated_at: 2026-09-22T00:00:00.000000Z
updated_at: 2026-09-23T00:00:00.000000Z
---

# API-404 — Secure, Durable, And Idempotent Control-Plane Semantics
Expand Down Expand Up @@ -104,6 +104,13 @@ identifies those implementation gaps and the retained canonical requirements.

## Traceability

- DOCUMENTS → GITHUB_ISSUE `1350` (Execution architecture selection; no new executable profile claim)
- DOCUMENTS → ADR `docs/decisions/adrs/adr-113-reusable-execution-machinery.md` (Reusable machinery, retained RAE authority and deployment boundaries)
- DOCUMENTS → DOCUMENTATION `docs/research/execution-architecture/execution-protocol.md` (Ledger/engine reconciliation, scoped worker authorization and conservative ownership recovery design)
- DOCUMENTS → DOCUMENTATION `docs/research/execution-architecture/authored-retry-policy.md` (Contextual author failure/retry policy, scoped defaults and fresh-trial distinctions; public support requires implementation)
- DOCUMENTS → DOCUMENTATION `docs/research/execution-architecture/experiment-report.md` (Bounded mechanism observations, not distributed conformance evidence)
- TESTS → TEST `implementations/python/tests/test_issue_1350_fixture_secret_scan.py` (Design-evidence publication checks: retained experiment source hashes match their record, and the exact synthetic fixture exception preserves detection of other values, paths and rules through the real-scanner integration lane; no runtime conformance claim)

- DOCUMENTS → GITHUB_ISSUE `1348` (Operation supervision decision; no new executable profile claim)
- DOCUMENTS → DOCUMENTATION `docs/decisions/issue-1348-operation-lifecycle-preflight.md` (Supervision architecture guardrails)
- DOCUMENTS → DOCUMENTATION `docs/decisions/issue-1348-operation-lifecycle.md` (Decision, requirement dispositions and implementation boundaries)
Expand Down
53 changes: 53 additions & 0 deletions docs/research/execution-architecture/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,53 @@
# Execution architecture — issue #1350

The accepted decision is [ADR-113](../../decisions/adrs/adr-113-reusable-execution-machinery.md):
self-hosted Temporal and PostgreSQL for distributed execution, existing library
and SQLite compositions for local profiles, with RAE retaining semantic authority.
No paid software is required. This delivery selects and documents the design;
it does not deploy a service, add runtime dependencies or enable P3.

The scope includes CTFs, OT twins, AI security research/sandboxes, product testing
and mixed IT/OT disaster recovery. Their failure responses and validity rules
can differ, including within one scenario. Authors select contextual policies,
scenario defaults, nested defaults and local overrides; engine retries cannot
supply those decisions.

## Reading order

1. [ADR-113](../../decisions/adrs/adr-113-reusable-execution-machinery.md): selection,
component/interface diagram, responsibilities, deployment and consequences.
2. [Authored failure/retry policy](authored-retry-policy.md): contextual responses,
scoped defaults, attempts, fresh trials and the implementation boundary.
3. [Execution protocol](execution-protocol.md): authorization, ledger/engine
reconciliation, stale ownership, partitions and restore.
4. [Candidate comparison](candidate-comparison.md): mechanisms, free-software
observations, retained duties and selection rationale.
5. [Experiment report](experiment-report.md), [sources](experiments/README.md),
[original evidence](evidence.json), and [local recheck](local-recheck.json).
6. [Requirement disposition](decision-discussion.md) and
[preflight](../../decisions/issue-1350-execution-architecture-preflight.md).

## Evidence provenance

The workspace contained the seven-candidate experiment bundle when this delivery
began. Its thirteen source hashes match the retained files. The original report
and [cleanup record](cloud-cleanup.json) describe that earlier synthetic cloud
run; this delivery did not provision cloud resources. A new local run rechecked
AnyIO, SimPy, DBOS, Temporal and SQLite publication with the same source bytes,
and ran the three witness tests. A further Temporal run checked terminal status
and result/error types through a verifier with four negative-control tests,
while keeping the original thirteen source files unchanged.
ROS/Ray/BehaviorTree.CPP results remain from the
original retained run. The two evidence records must not be conflated.

The [issue](https://github.com/OpenRAE/rae/issues/1350),
[#1348 decision](../../decisions/issue-1348-operation-lifecycle.md) and
[Hub #3](https://github.com/OpenRAE/hub/issues/3) establish scope and responsibility.
RAE determines admitted execution and validates outcomes; backends realize
resources, report effects and enforce contextual limits. A cancellation receipt,
engine success, checkpoint or process death cannot establish physical cessation.

Production persistence, tenant isolation, distributed ownership, co-simulation
fidelity and backend containment are not demonstrated by these bounded probes.
The design states how to handle their limits and what executable delivery must
verify. It does not use the local quickstart to limit either backend product.
169 changes: 169 additions & 0 deletions docs/research/execution-architecture/authored-retry-policy.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,169 @@
# Author-controlled retries and scoped defaults

Design adopted by [ADR-113](../../decisions/adrs/adr-113-reusable-execution-machinery.md)
for #1350. This is a semantic integration design, not published SDL syntax or a
new executable retry contract. Extend the existing owning contracts before
claiming support. No engine-specific retry fields enter the authored language.
Retry is one part of a contextual failure-response policy, not a universal
failure handler shared indiscriminately by experiments, twins, OT and IT.

## Context determines the permitted response

The common operation lifecycle records facts; it does not impose the same
recovery action or validity meaning on every world. Resolve policy against the
authored purpose, effect contract, operating mode, time/fidelity requirements,
and actual backend guarantees. The following are examples of selectable
policies, not automatic rules inferred from an `IT` or `OT` label:

| Context | Example response and required evidence |
| --- | --- |
| Controlled experiment | Mark this trial failed/invalid under the experiment contract, retain every attempt and observation, then request a separately admitted fresh trial if allowed. An idempotent retry may still bias results and is not automatically permitted. |
| Discrete or continuous digital twin | Hold advancement at a supported boundary, invalidate or qualify the affected state/time interval, and reconcile model/plant state before continuation. Re-reading a sensor, rewinding time or restoring a model checkpoint may change the experiment or break coupling; none is an ordinary transport retry. |
| Live OT or hardware-in-the-loop | Request the backend's admitted safe-hold/controlled-stop procedure, or a specified continuing mode where stopping is unsafe. Require process-state, interlock and operator authorization evidence where the authored contract demands it. Reissuing an actuator command, rebooting a controller or restoring an IT snapshot is not a generic recovery strategy. |
| Disposable IT test/CTF | An author may allow rebuilding a VM, restoring a known snapshot, then retrying after verified clean-state admission. That policy is inappropriate for resources the scenario is required to preserve. |
| Stateful IT/product or disaster recovery | Reconcile transactions, external services, data consistency and recovery objectives before selective retry/restore. IT is not inherently reversible or idempotent; externally visible effects may be as non-repeatable as OT actions. |
| Mixed IT/OT scenario | Apply each scope's effect and failure policy while preserving cross-scope dependencies, shared resources and time coupling. Restarting an IT gateway must not silently authorize another plant action or invalidate a twin's admitted state. |

Failure handling must distinguish infrastructure/delivery failure, backend
refusal, operation failure, deliberate in-world faults, loss of fidelity/time
continuity, and invalid experimental evidence. Do not let an infrastructure
reconciler repair an intentionally injected outage, or let a cleanup failure
disappear behind a primary success. RAE selects only the permitted response
using the canonical workflow/time/trial owners; the backend establishes whether
the concrete response is possible and what it actually did.

Convenience defaults can select reusable context policies and narrower scopes
can choose different ones. At admission, validate their interactions: one
scope's reset can affect another scope's equipment, data or clock. An
incompatible composition is refused, not resolved by whichever scope retries
first. No backend or engine supplies an unrecorded domain policy.

## Choices and identities

Idempotency describes an effect's repeat behavior; permission describes whether
the author wants repetition. Both must hold where required. An idempotent call
can still invalidate an experiment through repeated observations, elapsed time,
cost or participant exposure. RAE must support these distinct authored choices:

| Author choice | Meaning |
| --- | --- |
| Never repeat | At most one effect invocation for this admitted occurrence. A failed or uncertain invocation cannot cause another effect merely through retry/recovery. |
| Bounded retry in this trial | Retry only for selected failure/effect classes, within attempt and time budgets, after the declared safety and cleanup conditions hold. This can require verified absence, scoped backend idempotency, reset or compensation. |
| End this trial; request a fresh trial | Preserve the failed/invalid trial and its evidence. Admit a new trial through the experiment authority, with a new run identity, declared initialization/variation and verified isolation/cleanup. Do not increment a workflow retry counter and relabel the old trial as successful. |

Fresh-trial policy applies only where a trial context exists and its owning
experiment plan permits allocation; reject it on an ordinary operation lacking
that authority. Authors can allow bounded retries followed by a fresh trial on
exhaustion, or choose a fresh trial immediately. The trial allocation budget is
independent of within-trial attempts; neither can be unbounded by omission.

Keep four identities separate: engine delivery/task attempt; one permitted
backend invocation; authored workflow/execution attempt; experimental trial.
Redelivering a reference to the same invocation never allocates another one.
An explicitly admitted retry gets its own invocation/attempt identity and
provenance; a fresh trial gets its own run identity. Engine reset, restart and
Continue-As-New are not any of those author decisions.

## Scope resolution

Use the existing [lexical scope principle](../../../specs/sdl/recursive-realization-constraints.md#2-records-scopes-and-closure):
scenario default, enclosing scope defaults, then the most-specific explicit
operation/occurrence policy. Reuse canonical semantic addresses and bounded
reference resolution. This reuses scope mechanics, not the open/closed value
vocabulary: closure does not imply a retry policy.

1. A missing local policy inherits. An explicit never-repeat policy is a value,
not omission; a locally stricter or more permissive default affects only that
subtree. A concrete child choice can override a convenience default, subject
to binding requirements and admission. Siblings keep their inherited policy.
2. Resolve a **complete policy** at the nearest defining scope. Do not combine
unrelated fields from different levels into an accidental policy, such as a
child's reset mode with a parent's idempotency assumptions. Reusable named
policies provide concise authoring; any later partial-overlay syntax must
normalize to the same complete, validated value with field provenance.
3. Duplicate definitions at the same semantic scope, ambiguous targets, cycles,
unsupported versions and conflicting binding constraints are admission
errors. Import order, list order and engine defaults never break a tie.
Reusable definitions retain lexical binding; execution at a call site cannot
silently change them. An explicit application-site policy is validated as a
new local choice before the compiled plan is admitted.
4. Defaults are convenience, not a way to weaken a required guarantee, bypass
backend policy or grant authority. Resolve the author's desired policy, then
validate the entire composition. Refuse incompatibility instead of silently
clamping to fewer retries or substituting a new trial.
5. If no scope supplies permission to repeat, permit no second effect. This is
a specified conservative fallback, not an implementation library's default.
The author remains free to select a different scenario-wide default.
6. Materialize the resolved policy and defining scope/reference/version into the
admitted plan and operation commitment. Record changes as new admissions;
editing a scenario default cannot retroactively alter an active invocation,
a restored operation or historical evidence.

Conceptual examples, **not YAML syntax**:

| Scenario default | Nearer scope or local choice | Effective result |
| --- | --- | --- |
| Never repeat | Setup scope: at most three total attempts, admitted idempotency required | Setup may retry only under that condition; other operations remain one-shot. |
| Setup policy above | One actuator operation: never repeat | That operation remains one-shot, including after worker loss. |
| Bounded idempotent retry | Measurement scope: end this trial and request a fresh trial | Measurements never repeat within the same trial, even if technically idempotent. |
| Never repeat | No local choice | Inherit never repeat; no annotations are needed on every operation. |

## Effective policy and runtime admission

The normalized contract must identify its context/profile and revision,
failure classification, validity consequences, permitted response (for example
abort, hold, reconcile, continue at a verified boundary, retry or fresh trial),
and required operator/backend evidence. It must also identify the retry unit,
permitted effect-knowledge classes, total-attempt limit including the first attempt,
delay/backoff and clock basis, total budget, required safety/cleanup evidence,
and exhaustion disposition. Fresh-trial selection additionally binds trial
allocation authority and limits. Reuse `ExecutionRetryPolicyModel`'s existing
`max_attempts` and `after_effect_policy` meanings (`disallow`, `idempotent`,
`reset`, `compensate`) where applicable. Do not infer never-repeat solely from
`after_effect_policy=disallow`: that field addresses after-effect repetition,
while verified absence can have different admitted semantics.

Current SCE-007 cleanup policies do not provide general lexical inheritance,
failure-class/backoff selection or fresh-trial policy. These need governed
extensions under DSL-113/SEM-203 and SCE-002/006/007/EXP-706, not a free-form
metadata dictionary or a second workflow interpreter. Existing explicit workflow
retry bounds remain binding; nested policies must account for every actual
effect and must not multiply hidden retries at transport, activity and workflow
layers. Exhausted budgets do not reset on process restart or scope inheritance.

Before each permitted invocation, RAE rechecks current scope authorization,
policy identity, remaining budgets, backend capability/willingness, reservation
conflicts and required evidence. Backend idempotency needs a defined key, scope,
retention interval and behavior under duplicate/concurrent requests; an author
label is insufficient. Known absence is not cessation if an old caller can
still act later. Concurrent repetition requires explicit proven-safe semantics;
sequential attempts remain the default. Unknown effect state never becomes
safe merely because a worker disappeared.

Reset, compensation and fresh-trial initialization are effects with their own
admission, failure, observation and cleanup obligations. A requested reset does
not establish clean state. An unreconciled old trial cannot share resources with
a new trial unless admitted isolation actually separates their effects. Preserve
the old failure, cleanup result, lineage and any experiment validity decision.

## Engine mapping and verification obligations

Use Temporal's bounded scheduling and retry facilities when they can enforce
the effective policy through current RAE admission on every dispatch. Otherwise
schedule separately admitted single-attempt activities. Never turn engine retry
numbers into the author-visible attempt history. Bookkeeping retry and workflow
replay must not invoke an effect again. Observation can itself affect a system,
so classify it by its admitted effect contract rather than assuming all reads
are harmless.

The public-contract and implementation deliveries must test scenario defaults,
nested overrides, sibling isolation, explicit never-repeat, duplicate scopes,
definition/call-site binding, bounded normalization, incompatible guarantees,
scope reordering, missing policy, policy edits during recovery, budget exhaustion,
and no retry multiplication. Exercise different failure responses for the same
technical exception under experiment, twin, OT and IT policies, plus conflicting
responses across mixed scopes and preservation of intentional faults.
Exercise the same idempotent backend under both
retry-allowed and retry-forbidden policies, plus fresh-trial allocation and
cleanup failure. Those are behavioral gates for the later executable surface;
this design does not claim they pass today.
Loading
Loading