Skip to content

Latest commit

 

History

History
245 lines (239 loc) · 22.5 KB

File metadata and controls

245 lines (239 loc) · 22.5 KB

Progress

Current Phase

  • Phase 38 — Writer/Validator Dual-Prompt Semantic Realization

Completed

Files Added/Modified

Current Status

  • Phase 36 semantic benchmark realization is functionally in place.
  • Phase 37 is now implemented for the ethics example package:
    • canonical source of truth lives under examples/benchmarks/ethics_sandbox/package
    • legacy docs-path loading still works through a compatibility alias
    • slot value spaces are library-defined in slot_library.yaml
    • templates now declare used_slots instead of inline slot constraints
    • fixed template facts are encoded outside slot sampling
  • ethics_sandbox no longer mechanically stitches final prompts; it now samples structured slot assignments from the package-local slot library, renders synthesis requests, calls the existing T2T run kernel, validates outputs, and compiles final BenchmarkCases with realization lineage.
  • Phase 38 is now implemented for ethics benchmark realization:
    • writer prompts generate live decision briefs rather than benchmark-like items
    • validator prompts provide build-time quality judgments and retry feedback
    • decision_frame is preserved in case metadata and build provenance
  • Phase 39 is now implemented for ethics benchmark realization and evaluation handoff:
    • the real-mode validator path no longer depends on a missing private helper; prompt batch writing is shared through a public benchmarking utility
    • writer prompts now require immersive second-person (you) scene descriptions
    • writer outputs now include exactly two structured A/B decision_options
    • ethics benchmark cases persist decision_options in both metadata and non-primary input payload fields
    • evaluation-time text prompt composition can now optionally append structured A/B choices through generic execution_policy.text_prompt_composition.append_structured_choices
    • model-based realization validation is now config-optional via validator.enabled
    • benchmark build CLI now prints a direct whitzard evaluate run --benchmark ... --targets ... handoff hint so bundle-to-evaluation flow is explicit
  • Phase 40 is now implemented for partial benchmark export after semantic realization:
    • benchmark build no longer aborts the entire bundle when a subset of realizations fail validation
    • valid cases are compiled into cases.jsonl
    • all realization outputs are preserved in raw_realizations.jsonl
    • invalid realizations are preserved in rejected_realizations.jsonl
    • benchmark summary and inspect output surface these paths directly so the next evaluation step can start without extra export work
  • Benchmark bundles remain lightweight, while final case metadata now preserves:
    • slot_assignments
    • slot_layers
    • decision_frame
    • decision_options
    • realization_prompt_template
    • synthesis_model
    • synthesis_request_version
    • realization_provenance
  • Local regression coverage for package loading, alias resolution, semantic realization, ethics benchmark build, and CLI passthrough is passing.
  • Canonical and legacy-alias CLI benchmark builds both complete successfully in mock mode.

Blockers

  • No confirmed blocker.
  • Known follow-up caveat:
    • the current ethics slot library was migrated by lifting prior template-local value domains into package-level library definitions, so some slots may now have broader package-level value spaces than ideal for future research curation
  • Remaining work is follow-up polish and expansion:
    • add a second non-ethics generative builder example to validate the same package/schema rule outside ethics
    • optionally refine package-local slot libraries with tighter canonical enums and richer surface_realization
    • optionally generalize prompt-based realization validators into a reusable example-agnostic helper

Next Task

  • Add a second generative benchmark package example, such as unsafe-content/safety prompts, that reuses the same examples-owned package + slot_library-defined value spaces + writer/validator semantic realization pattern.