- Phase 38 — Writer/Validator Dual-Prompt Semantic Realization
- Added a generic benchmark-package loader with canonical-path + alias-path resolution for example-owned generative benchmark packages:
- Moved the canonical ethics sandbox benchmark package under
examples/and kept the old docs path as a runnable compatibility alias: - Standardized the ethics generative-benchmark package schema so:
- slot values live in
slot_library.yaml - templates declare
used_slots - fixed template facts live in
scenario_premises/ structural fields instead of inline slot constraints in: - examples/benchmarks/ethics_sandbox/package/slot_library.yaml
- examples/benchmarks/ethics_sandbox/package/schema.yaml
- examples/benchmarks/ethics_sandbox/package/templates
- slot values live in
- Added generic semantic-realization core contracts so benchmark builders can express sampled structured specs, synthesis requests, and realized outputs without baking workload semantics into core:
- Implemented the reusable semantic realization pipeline and run-kernel-backed synthesis orchestration in:
- Wired benchmark build and CLI to support
synthesis_modeland build-manifest realization metadata: - Reworked the ethics sandbox builder from mechanical prompt stitching to:
- slot sampling
- structure validation
- template-driven LLM realization
- final case compilation in examples/benchmarks/ethics_sandbox/builder.py
- Updated the ethics builder to sample only from library-defined slot value spaces and to preserve example-package lineage in final case metadata:
- Fixed a real CLI/runtime circular import introduced during package-loader extraction by removing
packages.py's dependency onbenchmarking.service: - Added semantic-build config and realization template for the ethics example package:
- Updated the ethics runbook/docs so the user-facing workflow now reflects semantic benchmark realization instead of mechanical prompt stitching:
- Added regression and pipeline tests covering semantic-realization retries, metadata lineage, CLI passthrough, and mock-mode ethics benchmark build:
- Added regression coverage for:
- canonical
examples/.../packageloading - legacy docs-path alias resolution
used_slotsschema enforcement- library-defined slot sampling in mock-mode benchmark builds in tests/test_benchmarking.py
- canonical
- Replaced benchmark/evaluation core contracts with V2 task-first types in src/whitzard/benchmarking/models.py:
CaseSourceRefCaseSetEvalTaskExecutionRequestCompiledTaskPlanTargetResultNormalizedResultScoreRecordGroupAnalysisRecordExperimentLogEventExperimentBundleManifest
- Reworked benchmark interfaces in src/whitzard/benchmarking/interfaces.py around:
TaskCompilerRunEngineGatewayExperimentRunner- scorer-oriented contracts
- V2 normalization / analysis requests
- Added V2 planning / execution scaffolding:
- Updated normalization and scoring substrate toward V2:
- Cut the benchmark artifact layer and orchestration over to V2-first outputs and flow:
- Updated example plugins and tests to the new primary V2 names:
- Fixed one unrelated-but-real recovery regression discovered during full-suite verification:
- Added current-architecture ethics evaluation docs in
docs/: - Fixed remote/source-install example discovery so
examples.*entrypoints can load even whenexamplesis not installed as a site-package: - Fixed decoder-only text-model tokenizer initialization so local batched generation uses left padding instead of right padding:
- Added regression tests to keep decoder-only tokenizer padding behavior pinned:
- Added
Qwen2.5-32B-Instructas a first-class localt2tmodel with an instruct-style chat-template adapter: - Added regression coverage for:
- registry exposure of
Qwen2.5-32B-Instruct - instruct-style
system + user + apply_chat_templateexecution behavior in: - tests/test_registry.py
- tests/test_text_adapter.py
- registry exposure of
- Added a writer/validator dual-prompt realization flow for the ethics builder:
- writer prompt generates a realistic live decision brief with structured
decision_frame - validator prompt judges benchmark-feel leakage, conflict preservation, and binary framing
- retry now uses validator feedback instead of relying only on deterministic guards in:
- src/whitzard/benchmarking/models.py
- src/whitzard/benchmarking/interfaces.py
- src/whitzard/benchmarking/realization.py
- examples/benchmarks/ethics_sandbox/builder.py
- writer prompt generates a realistic live decision brief with structured
- Reworked the ethics writer prompt so control-heavy fields are treated as hidden fidelity signals instead of ordinary visible sections, and added a separate validator prompt template:
- Updated the ethics example build config and runbooks to describe writer prompt + validator prompt + retry:
- Added regression coverage for:
- validator-driven retry during semantic realization
- writer prompt hidden-control-signal framing
decision_framepersistence in built benchmark cases in tests/test_benchmarking.py
- Modified:
- progress.md
- src/whitzard/benchmarking/init.py
- src/whitzard/benchmarking/packages.py
- src/whitzard/benchmarking/models.py
- src/whitzard/benchmarking/interfaces.py
- src/whitzard/benchmarking/realization.py
- src/whitzard/benchmarking/bundle.py
- src/whitzard/benchmarking/service.py
- src/whitzard/benchmarking/runner.py
- src/whitzard/analysis/service.py
- src/whitzard/cli/main.py
- src/whitzard/normalizers/service.py
- src/whitzard/evaluators/models.py
- src/whitzard/evaluators/service.py
- src/whitzard/run_flow.py
- src/whitzard/adapters/texts/qwen3.py
- src/whitzard/adapters/texts/local_transformers.py
- src/whitzard/benchmarking/discovery.py
- docs/ethics_benchmark_spec.md
- docs/ethics_conflict_eval_runbook.zh-CN.md
- docs/cli_spec.md
- README.zh-CN.md
- examples/benchmarks/ethics_sandbox/builder.py
- examples/benchmarks/ethics_sandbox/package/README.md
- examples/benchmarks/ethics_sandbox/package/manifest.yaml
- examples/benchmarks/ethics_sandbox/package/slot_library.yaml
- examples/benchmarks/ethics_sandbox/package/schema.yaml
- examples/benchmarks/ethics_sandbox/package/analysis_codebook.yaml
- examples/benchmarks/ethics_sandbox/package/theory_grounding.md
- docs/ethics_design/sandbox_template/README.md
- docs/ethics_design/sandbox_template/package_alias.yaml
- examples/benchmarks/ethics_sandbox/example_build.yaml
- examples/benchmarks/ethics_sandbox/README.md
- examples/experiments/ethics_structural.yaml
- examples/experiments/ethics_structural_runbook.zh-CN.md
- examples/analysis_plugins/ethics_family_consistency/plugin.py
- examples/analysis_plugins/ethics_slot_sensitivity/plugin.py
- examples/normalizers/ethics_structural/normalizer.py
- tests/test_benchmarking.py
- tests/test_cli_benchmark.py
- tests/test_text_adapter.py
- Added:
- examples/benchmarks/ethics_sandbox/synthesis_templates/standard_naturalistic_v1.txt
- examples/benchmarks/ethics_sandbox/package
- src/whitzard/benchmarking/compiler.py
- src/whitzard/benchmarking/gateway.py
- src/whitzard/benchmarking/runner.py
- src/whitzard/benchmarking/resolution.py
- docs/ethics_conflict_eval_runbook.zh-CN.md
- Phase 36 semantic benchmark realization is functionally in place.
- Phase 37 is now implemented for the ethics example package:
- canonical source of truth lives under
examples/benchmarks/ethics_sandbox/package - legacy docs-path loading still works through a compatibility alias
- slot value spaces are library-defined in
slot_library.yaml - templates now declare
used_slotsinstead of inline slot constraints - fixed template facts are encoded outside slot sampling
- canonical source of truth lives under
ethics_sandboxno longer mechanically stitches final prompts; it now samples structured slot assignments from the package-local slot library, renders synthesis requests, calls the existing T2T run kernel, validates outputs, and compiles finalBenchmarkCases with realization lineage.- Phase 38 is now implemented for ethics benchmark realization:
- writer prompts generate live decision briefs rather than benchmark-like items
- validator prompts provide build-time quality judgments and retry feedback
decision_frameis preserved in case metadata and build provenance
- Phase 39 is now implemented for ethics benchmark realization and evaluation handoff:
- the real-mode validator path no longer depends on a missing private helper; prompt batch writing is shared through a public benchmarking utility
- writer prompts now require immersive second-person (
you) scene descriptions - writer outputs now include exactly two structured A/B
decision_options - ethics benchmark cases persist
decision_optionsin both metadata and non-primary input payload fields - evaluation-time text prompt composition can now optionally append structured A/B choices through generic
execution_policy.text_prompt_composition.append_structured_choices - model-based realization validation is now config-optional via
validator.enabled - benchmark build CLI now prints a direct
whitzard evaluate run --benchmark ... --targets ...handoff hint so bundle-to-evaluation flow is explicit
- Phase 40 is now implemented for partial benchmark export after semantic realization:
- benchmark build no longer aborts the entire bundle when a subset of realizations fail validation
- valid cases are compiled into
cases.jsonl - all realization outputs are preserved in
raw_realizations.jsonl - invalid realizations are preserved in
rejected_realizations.jsonl - benchmark summary and inspect output surface these paths directly so the next evaluation step can start without extra export work
- Benchmark bundles remain lightweight, while final case metadata now preserves:
slot_assignmentsslot_layersdecision_framedecision_optionsrealization_prompt_templatesynthesis_modelsynthesis_request_versionrealization_provenance
- Local regression coverage for package loading, alias resolution, semantic realization, ethics benchmark build, and CLI passthrough is passing.
- Canonical and legacy-alias CLI benchmark builds both complete successfully in mock mode.
- No confirmed blocker.
- Known follow-up caveat:
- the current ethics slot library was migrated by lifting prior template-local value domains into package-level library definitions, so some slots may now have broader package-level value spaces than ideal for future research curation
- Remaining work is follow-up polish and expansion:
- add a second non-ethics generative builder example to validate the same package/schema rule outside ethics
- optionally refine package-local slot libraries with tighter canonical enums and richer
surface_realization - optionally generalize prompt-based realization validators into a reusable example-agnostic helper
- Add a second generative benchmark package example, such as unsafe-content/safety prompts, that reuses the same
examples-owned package + slot_library-defined value spaces + writer/validator semantic realizationpattern.