Skip to content

feat(examples): add a participant behavior pattern catalog with validated example links - #1468

Open
doublewhy wants to merge 4 commits into
devfrom
379-participant-behavior-patterns
Open

doublewhy wants to merge 4 commits into
devfrom
379-participant-behavior-patterns

Conversation

@doublewhy

@doublewhy doublewhy commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Builds on #1458 (and through it #1445 and #1438); review those first. Until they merge, this PR also shows their commits, 6fa7bdab, f397d0a4 and 88975279. This PR's own change is commit 7f4599ea.

Issue #379 asks for a short catalog of participant behavior patterns tied to current SDL support. Each entry states its intended use, required fields, validation status and known limits, and links an executable example only when that example passes the normal validation path.

Requirement-level verification of the behavior surfaces is in two open PRs: #1426 (#297, DSL-116 behavior specifications, reusable and scenario-local) and #1440 (#298, DSL-117 tool affordances). On dev, DSL-116 is DRAFT and DSL-117 is ACTIVE. This PR does not depend on either PR, and no committed file references them. The catalog's claims are reproduced on this branch: the new tests cover the import composition and the four limits that describe behavior, and probes run on this branch cover the required_fields lists and the realization_profile_ref limit (Test Plan). Merge order does not matter: I merged #1440's head (600080a9), #1426's head (6fb57a4b), and then both, into this head without committing, and each time the new tests, that PR's own tests and the gate passed.

Three of the limits restate gaps that #1426 and #1440 recorded. The fourth, that raes semantic validate rejects a document that declares imports, is a designed limit: docs/explain/sdl/parser.md says in-memory parsing rejects imports by design. The reusable-unit pattern checks an importing scenario with raes sdl resolve followed by raes sdl verify-imports, or with parse_sdl_file.

The participant behavior surface now has four patterns, three of them new:

  • participant-behavior-contract-binding (existing): binds agent actions and views to action contracts and observation boundaries. It gains required_fields and validation commands that run.
  • participant-behavior-specification (new): one scenario-local behavior specification over a participant's declarations.
  • participant-behavior-reusable-unit (new): a behavior specification exported from a module unit and imported under a namespace.
  • participant-behavior-tool-affordance (new): a tool content item bound to the action contracts and observation boundaries it serves.

Two new templates are the executable examples. tool-affordance is the first library example that authors a tool affordance. reusable-behavior-unit is a module unit that the gate validates as a standalone scenario. The tool-affordance pattern names the tool by its bare content name. #1440 rewrites a bare name exactly as dev does, so the pattern composes the same way before and after #1440 merges.

How this meets the issue's points:

  • Intended use and required fields. Each pattern file has intent, use_when and a new required_fields list. Each list is checked in both directions (Test Plan). Removing or breaking a listed field in a copy of a template body fails validation, and a scenario that holds only name and the listed fields validates with no advisories.
  • Validation status. Each pattern entry stays guidance, which feat(examples): add machine-checked metadata to example library entries #1438 requires for patterns. A new catalog field, example_refs, links its examples. The gate requires each ID to name a worked example or template whose validation_status is validated, so a pattern links only to examples that the gate itself parses and validates.
  • Known limits. Each entry's limits say what it does not show. The reusable-unit entry states the three import limits: role refs are not scoped to one import, raes semantic validate rejects an importing scenario, and an invalid imported declaration raises a raw pydantic ValidationError. The tool-affordance entry states that an importing scenario cannot add an affordance for an imported participant.
  • No implied semantics and no capability language. Every limit is a non-claim. No entry says that a participant runs or that a backend installs, exposes or admits anything.
  • Fewer entries. Four patterns, each tied to one surface that current validation checks. Interactive access, mixed control, inject deliveries and autonomous execution get no pattern here.

Requirement UIDs

Related Issues

Closes #379

ADR Impact

  • No ADR required, and no ADR text changes.

Changes

  • examples/library/patterns/participant-behavior-specification.yaml, participant-behavior-reusable-unit.yaml and participant-behavior-tool-affordance.yaml (new): each has intent, use_when, required_fields, authoring_steps and validation. The tool-affordance required_fields also name the fields of the view_rules entry that classifies the affordance. The reusable-unit validation lists the gate, a parse_sdl_file one-liner, raes sdl resolve and raes sdl verify-imports; the other two list the gate and raes semantic validate.
  • examples/library/patterns/participant-behavior-contract-binding.yaml: adds required_fields, including that an observation boundary needs at least one observable_refs or evidence_refs entry. Its one validation entry, command: parse_sdl_file, was a function name rather than a command. It becomes the gate command and a raes semantic validate command. One authoring step now names the observation_boundaries section.
  • examples/library/templates/participant_behavior/tool-affordance.yaml (new): a standalone scenario with a tool content item, a scan action contract, an observation boundary that lists and classifies the affordance, and a behavior specification whose affordance binds the three.
  • examples/library/templates/participant_behavior/reusable-behavior-unit.yaml (new): a module unit (example/scan-behavior 1.0.0) that exports its participant, action contract, observation boundary and behavior specification. The behavior specification names its participant with participant_refs.
  • examples/library/catalog.yaml: catalogs the two templates and three patterns, adds example_refs to all four participant behavior patterns, and gives each new entry its sdl_sections and limits. Other surfaces are unchanged.
  • tools/example_library_checks.py: new check_example_refs with rule id example-library-entry-examples. Only pattern entries may record example_refs. The value must be a non-empty list of IDs of validated worked examples or templates, on any surface. An entry whose id is not a string is left to the existing example-library-entry-id rule.
  • tools/check_example_library.py: calls check_example_refs once per catalog and checks that a pattern's optional required_fields is a list (480 lines, wc -l).
  • examples/README.md: the library table lists the new files under a "Patterns" column. The metadata table gains an example_refs row. A new "Participant behavior patterns" subsection lists the four patterns, their validated templates and what they do not claim, and names raes sdl resolve followed by raes sdl verify-imports as the command-line check for an importing scenario.
  • implementations/python/tests/test_issue_379_participant_behavior_patterns.py (new), 24 cases in six functions:
    • 14 gate cases. Two links pass: validated templates on two surfaces, and a validated worked example. Seven links fail with example-library-entry-examples: a guidance worked example, a pattern, an unknown ID, an empty list, a string, and example_refs on a template or on a worked example. A string required_fields fails with example-library-pattern-field-type. Four malformed catalogs report only their existing failure: a surface, a patterns field and an entry that have the wrong type, and a validated template whose id is a list (example-library-entry-id).
    • raes semantic validate, the documented command, succeeds on the three participant behavior template bodies and rejects a scenario that imports the unit.
    • raes sdl resolve then raes sdl verify-imports on that importing scenario: resolve writes raes.lock.json and prints its path, and verify-imports prints imports verified. With a bare-name ref (red-agent for alpha.red-agent), verify-imports exits 1 with SDLValidationError.
    • The unit template, imported as alpha (1.0.0) and bravo (>=1,<2), compiles with each import's specification covering only its own participant. With participant_role_refs instead, each covers both imports' participants, as the limit states.
    • An importing scenario that adds an affordance for alpha.red-agent fails because alpha.red-view does not classify it.
    • An invalid lifecycle_state in the unit raises SDLParseError when the unit is parsed alone and a pydantic ValidationError when it is imported.
    • Four cases pin the documented limits: role refs across imports, raes semantic validate rejecting an importing scenario, the importing-scenario affordance and the raw ValidationError. A change to any of these behaviors fails its case, so the catalog limit has to change with it.

Test Plan

  • Unit tests pass
  • Integration tests pass if applicable
  • Full completion suite required in CI before merge
  • No coverage regression

All commands ran from the worktree root on head 7f4599ea, under CPython 3.14.4 from the project environment, unless a paragraph names another head. The base is origin/dev at 35122105 plus the #1438, #1445 and #1458 commits.

I ran each documented command:

  • uv run --project implementations/python --frozen python tools/check_example_library.py exited 0 and printed nothing.
  • I saved the tool-affordance body, then the action-contract-observation-boundary body, as my-scenario.sdl.yaml. Each time, uv run --project implementations/python --frozen raes semantic validate my-scenario.sdl.yaml printed validate: success and exited 0.
  • I saved the reusable-behavior-unit body as scan-behavior.sdl.yaml, next to a my-scenario.sdl.yaml with one import (namespace alpha, version 1.0.0). The documented parse_sdl_file command printed ['alpha.red-scan-behavior'] and exited 0. raes semantic validate printed validate: success on the unit, and on the importing scenario it printed validate: invalid and error [sdl.parse] SDL input was rejected at the parse stage., with exit 1. uv run --project implementations/python --frozen raes sdl resolve my-scenario.sdl.yaml printed raes.lock.json and exited 0, and raes sdl verify-imports my-scenario.sdl.yaml then printed imports verified and exited 0.
  • I removed each file, including raes.lock.json, after its run.

The required_fields lists rest on two kinds of probe, each run by a scratch script (since deleted) in the same environment.

Necessity: removing or breaking a listed field in an edited copy of a template body fails validation. These probes ran on the first head, 60ba5216, except the boundary and view-rule probes, which ran on this head. The template bodies are the same on both heads.

  • Contract binding: removing any of the eight action-contract fields the pattern lists failed validation, as did removing any of the three precondition fields, the three effect fields or the three named boundary fields. A boundary with only projection_basis, redaction_policy and latency_profile failed with "participant observation boundary requires observable_refs or evidence_refs", both as the reusable-behavior-unit red-view without its evidence_refs and in a minimal scenario; adding one evidence_refs or one observable_refs entry passed. An observation_effect without target_refs or evidence_refs failed; a no_effect without them passed. An agent action that is not an action_contracts key and an undeclared agent boundary each failed.
  • Tool affordance: emptying the specification's action_contract_refs failed with "widens its owning behavior specification". Removing the boundary's view rule or its observable_refs entry failed the classification check. An affordance action that the participant does not declare failed with "is outside participant". A tool_ref that names a node failed. A second, classified affordance with the same relation failed as a duplicate. Omitting tool_ref passed.
  • Tool affordance view rule: removing information_ref, boundary_class, disposition or visibility_basis failed with "Field required". An information_ref that names another affordance failed with "must be explicitly classified by observation boundary 'red-view'". An evidence_only rule failed without evidence_refs and passed with them. A disclosed rule without disclosure_rule failed for each of the seven classes tried. A discovered, inferred or deceptive rule without disclosure_rule failed for each of the five classes the pattern names and passed for observable_resource and tool_output. Each of these passed with a disclosure_rule.
  • Behavior specification: a realization_profile_ref of no-such-manifest passed and a blank one failed. An ungoverned behavior_mode, an unknown feature, an unpublished evidence contract and an undeclared authority scope ref each failed.
  • Imports: omitting version passed, because it defaults to any version. version: '>=2' failed with "requested version '>=2' but module declares '1.0.0'". References to unexported declarations failed with "does not reference a declared agent" (and action_contract), and so did a bare name. A unit without a module block failed with "Imported SDL units require an explicit module descriptor".

Completeness, on this head: a scenario with only name and the fields that a list names validated with parse_sdl and gave no advisories. That held for contract binding alone, for a behavior specification added to it, and for a tool affordance added to that, with its boundary entry and view_rules entry. A unit with a module block (id, version, and exports naming the behavior specification) over the same declarations, imported by a scenario with only name and one import with source and namespace, parsed with parse_sdl_file and gave ['alpha.spec']. On the previous head, the contract-binding list did not name the boundary's observable_refs or evidence_refs requirement, and the tool-affordance list did not name the fields of the view_rules entry it requires. This head lists both.

uv run --project implementations/python --frozen --all-extras python -m pytest on test_issue_379_participant_behavior_patterns.py, test_issue_380_experiment_templates.py, test_issue_378_scenario_templates.py, test_issue_381_example_catalog_metadata.py and test_example_library_policy.py with -q -p no:cacheprovider passed: 63 passed (24 new). None of the modules has integration-marked cases, and -m integration deselects all 63.

Regression checks:

  • I restored tools/check_example_library.py and tools/example_library_checks.py from 88975279 and kept the new tests, templates and catalog. 8 of the 24 new cases failed: the seven example_refs violations and the string required_fields. The other 16 pass on both versions. They are the two allowed links, the four malformed-catalog cases, and the ten template, CLI and import cases, which do not depend on the checker. With the change restored, all 24 pass.
  • With only the string-id filter removed from check_example_refs, the template-id-not-text case failed with TypeError: cannot use 'list' as a set element (unhashable type: 'list'). With the filter, it passes.

A coverage.py branch report over the same five modules covers every line and branch of tools/example_library_checks.py (125 statements, 44 branches).

Merge order with #1426 and #1440: in this worktree I ran git merge --no-commit --no-ff with #1440's head 600080a9, then with #1426's head 6fb57a4b, then with both, and ran git merge --abort after each. Each merge applied without conflicts. With 600080a9, the 24 new cases and the 27 in test_issue_298_tool_authoring.py passed (51 passed) and the gate exited 0. With 6fb57a4b, the 24 new cases and the 18 in test_issue_297_behavior_authoring.py passed (42 passed) and the gate exited 0. With both, 69 passed and the gate exited 0.

nox -s verify-fast-feedback -- --base-rev origin/dev, nox -s lint and make policy passed, with every non-skipped stage green. The runner skipped policy / requirement governance on its own, because the branch resolves no requirement UID and no requirement-scope file exists for #379.

CI on head 7f4599ea: run 37960114996 (CI) passed, as did Docs (37960114513), Bootstrap qualification (37960114792) and CodeQL (37960108972). All 32 checks pass, with deploy skipped. The SonarCloud quality gate is OK on 7f4599ea: 0 new issues, 99.5% coverage on new code and 0.0% duplication.

Earlier heads: CI run 37946004145 on the first head, 60ba5216, passed every check except sonar. SonarCloud reported python:S3776, cognitive complexity 16 where 15 is allowed, in _catalog_entries; 7c0a1a64 moved the per-field loop into a new helper, _field_entries. CI run 37949167549 on 7c0a1a64 passed all 32 checks, with deploy skipped, and the SonarCloud gate was OK.

Changes since 7c0a1a64, after a pre-push review: the two required_fields additions above, action_contracts in the tool-affordance pattern's catalog sdl_sections, the string-id filter in check_example_refs with its test case, raes sdl resolve and raes sdl verify-imports in the reusable-unit pattern and the README with their test, and this description, which no longer says that the patterns rely on #1426 or #1440.

Not run locally: the docs lanes, because no file under docs/ or a Vale entry point changed; examples/README.md is outside the Sphinx source and the Vale file set. No package source, pyproject.toml or uv.lock changed, and no pinned scenario is touched, so the research evidence captures need no republish. No test needs Docker.

Ground Control Checks

  • Repository policy command passes: make policy exited 0, and the only skipped stage was requirement governance.
  • Regression check fails without the checker change and passes with it, as described in the Test Plan.

Pre-push code review: a review of 7c0a1a64 returned findings, and this head addresses each one ("Changes since 7c0a1a64" in the Test Plan). No separate test-quality review was run.

Traceability

  • IMPLEMENTS: examples/library/patterns/participant-behavior-specification.yaml, examples/library/patterns/participant-behavior-reusable-unit.yaml, examples/library/patterns/participant-behavior-tool-affordance.yaml, examples/library/patterns/participant-behavior-contract-binding.yaml, examples/library/templates/participant_behavior/tool-affordance.yaml, examples/library/templates/participant_behavior/reusable-behavior-unit.yaml, examples/library/catalog.yaml, tools/example_library_checks.py, tools/check_example_library.py, examples/README.md
  • TESTS: implementations/python/tests/test_issue_379_participant_behavior_patterns.py

Checklist

  • Code follows the project coding standards: the new check sits in the support module from feat(examples): add machine-checked metadata to example library entries #1438, with at most two returns per function, and the main checker gains two lines.
  • FM: not applicable. There is no SDL, contract, schema or runtime semantic change; the new rule is repository tooling over the example catalog, and the templates only use existing SDL. No executable or runtime adoption is claimed.
  • Published contract schemas regenerated if models changed: no model or schema changed.
  • PR title is a Conventional Commit: feat(examples), because the new library entries are user-visible.
  • Architectural docs updated if applicable: there is no architecture change. examples/README.md describes the new entries.

Documentation

Updated: examples/README.md (library table, metadata table, new "Participant behavior patterns" subsection). Verified unchanged: docs/explain/getting-started.md line 63 still describes the library accurately, and line 66 already names raes sdl resolve and raes sdl verify-imports for imports.

Every worked example, template, and pattern entry in
examples/library/catalog.yaml now records validation_status (validated
or guidance), sdl_sections, intended_user, and limits. The gate rejects
an entry whose status is outside the vocabulary, a template that is not
validated, a pattern that claims validation, an unknown intended user,
empty limits, or a section name that is not a current SDL section.
Validated worked examples are now parsed through parse_sdl_file and
fail on errors or advisories, and a validated entry must list only
sections present in its SDL. The catalog description no longer calls
every entry validated.

The entry checks live in a new support module,
tools/example_library_checks.py, because tools/check_example_library.py
was already 488 lines against SonarCloud's 500-line file limit. The
template-body SDL check moves there with its rule ids and messages. The
module imports raes at module level, so the
example-library-template-import rule for a failed raes import is
removed. The top-level fields that are not sections come from
tools/sdl_catalog_parity/_paths.py, which the SDL catalog parity gate
keeps in step with specs/sdl/sections.md.

examples/README.md documents the fields and what the gate checks.
…laim notes

Add two scenario templates beside minimal-validated-scenario:
segmented-network-scenario (two switched segments, linked hosts with
fixed addresses, HTTPS-only access rules, a service feature) and
parameterized-scenario (host operating system, size, instance count, and
service port from declared variables). Both bodies pass the SDL parser
and semantic validator with no advisories, and both are cataloged with
the entry metadata.

Every template now carries a validation list of commands with their
expected results. The gate requires the field, and
tools/example_library_checks.py checks that it is a non-empty list of
mappings with non-empty command and expected strings. The gate does not
run the commands.

examples/README.md lists what each scenario template shows and does not
claim, how to validate an adapted copy with raes semantic validate, and
that a ${name} placeholder in a template body must name a declared
variable.
… studies

Add two templates whose bodies are experiment-authoring-input-v1
documents: seeded-run-plan (run count, seed, and episode controls for one
task) and two-condition-study-design (a treatment factor, two compared
conditions, and the runs per condition). Both pass the experiment
authoring-input loader, parse_experiment_spec.

A template's catalog entry may now declare contract:
experiment-authoring-input-v1. Such an entry names experiment-author as
its intended user and lists no sdl_sections, and the gate validates its
body with the experiment loader instead of the SDL parser. Worked
examples and patterns stay SDL.

examples/README.md separates the SDL, experiment-input, guidance, and
unsupported parts of the task, run, and study surfaces, including why
there is no experiment-task-v1 template, and gives a command to validate
an adapted experiment template.
@doublewhy
doublewhy force-pushed the 379-participant-behavior-patterns branch from 60ba521 to 7c0a1a6 Compare October 9, 2026 15:05
…ated example links

Add three participant behavior patterns (scenario-local behavior
specification, reusable module unit, tool affordance) and give the existing
contract-binding pattern the same shape: required_fields, authoring steps,
and validation commands that run. Two new validated templates,
tool-affordance and reusable-behavior-unit, are the executable examples.
The reusable-unit pattern checks an importing scenario with raes sdl resolve
and raes sdl verify-imports, or with parse_sdl_file.

Pattern catalog entries may now record example_refs. The example-library
gate requires each one to name a worked example or template whose
validation_status is validated, so a pattern links only to examples that the
gate itself parses and validates. Each entry's limits state what it does not
show, including the import limits recorded by the DSL-116 and DSL-117
verification.
@doublewhy
doublewhy force-pushed the 379-participant-behavior-patterns branch from 7c0a1a6 to 7f4599e Compare October 9, 2026 16:34
@doublewhy
doublewhy marked this pull request as ready for review October 9, 2026 17:02

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant