Skip to content

Latest commit

 

History

History
95 lines (75 loc) · 5.14 KB

File metadata and controls

95 lines (75 loc) · 5.14 KB

Validation

Automated tests

.venv/bin/pytest -q                          # everything, engine + pre-existing app tests
.venv/bin/pytest tests/ -q -k "engine"       # engine-only
.venv/bin/ruff check divesync tests          # lint

345 tests pass as of this pass (242 pre-existing app tests, unmodified; 103 new engine tests across 7 test modules). Coverage by module:

Test file Slice Covers
test_engine_annotations.py 1 Schema loading, validation, ontology, legacy adapter, real fixtures
test_engine_inference.py 4-5 Drive sanity check, sparse regularity table, Cm/Ca split, somersault/twist separation, float window, D_e_required
test_engine_editorial.py 6 Style-profile divergence, relative-delta drive target, budget/floor formulas, subtractive buildup, motif recurrence
test_engine_placement.py 3, 7 Monotone alignment, hard-sync tolerance, Appendix B rotation snap, speed bounds, baseline determinism, hierarchical bounds, planner comparison
test_engine_treatment.py 8-11 Role/footprint independence, drive-split cut count, secondary mask density, smear constraint, constraint convergence, choke variety, subtractive buildup at treatment level
test_engine_render.py 3 Filter chain, forbidden-filter absence, rasterisation drift, PWL expansion, a real synthetic render, atomic-write-on-failure
test_engine_invariants.py all Static checks: label/threshold coupling, layering, decision-log opacity rejection, no scattered tunables

Manual / real-media validation

Run against testdata/engine/real/ (KAS:ST "Astral Talk", 174 BPM, 5 annotated real dive clips; see testdata/engine/real/README.md):

.venv/bin/python -m divesync.engine render \
  --project testdata/engine/real/project.json \
  --out output/sample --profile peak_time --seed 7

Produces output/sample/{skeleton.mp4, timeline.json, decision-log.jsonl, summary.json, state.json, targets.json, placement.json, rhythm.json, effects.json, planner-comparison.json}.

Grid accuracy

The music fixture's grid derivation (174 BPM, bar-8 origin) is verified two ways: testdata/engine/real/README.md documents the arithmetic (all four hard-sync events land within 0.72 ms of an exact beat), and divesync.engine.timebase.anchor_residuals recomputes it from the built MetricGrid — max_abs_residual_ms in output/sample/summary.json's grid block is the live check. python -m divesync.engine grid --music testdata/engine/real/astral-talk.music.json --click /tmp/click.wav renders an audible click for a by-ear check (spec 26 "no audible flam").

Phrase survival through a sparse section

validation.py::_phrase_survival (gate check 3) locates the sparsest 8-bar window by mean Ca, then confirms the bar index still advances at the expected rate across it and Cm stays above 0.5. On the real track's 224.8-233.1s silence, Ca measured ~0.0 while Cm stayed above 0.83 in manual inspection (see the Slice 2 smoke test in the implementation log).

Drive split (high-D_v gets fewer cuts)

Directly unit-tested at three levels — test_engine_inference.py::test_high_visual_drive_requires_less_articulation_than_low_visual_drive, test_engine_treatment.py::test_high_d_v_candidate_requires_fewer_cuts_than_low_d_v_at_same_target, and the pipeline-level gate check 4. The gate check compares placed slots grouped by d_v, restricted to comparable D_target; on short real-media excerpts (60-150s) the two groups' targets were not comparable enough for the check to reach a verdict (see docs/known-limitations.md) — it reported "inconclusive", not "failed", which is the correct behaviour when the premise (comparable targets) does not hold on a small sample.

Baseline vs. hierarchical

planner-comparison.json reports both. On the 150s real-media excerpt: baseline cost 41.28 (0.32s to plan) vs. hierarchical cost 45.33 (5.70s to plan); hierarchical was still selected because its cost is within the configured 20% regression tolerance and it produces a more even clip-reuse distribution across the four annotated clips. The regression guard was tightened from a 100% to a 20% tolerance during this pass — the original default never actually fired in practice (see docs/known-limitations.md).

Readability and complexity constraints

test_engine_treatment.py::test_readability_floor_never_violated_where_importance_high and the gate's readability_violations count (0 in both the synthetic and real runs) confirm the hard floor holds. Complexity-budget suppressions are real and logged: the 150s real-media render suppressed effects at essentially every busy bar (summary.json's log.by_kind.suppressed count), all attributable to complexity_budget, all with a matching DecisionKind.SUPPRESSED log record — test_nothing_is_suppressed_silently asserts the schedule and the log never disagree on this.

Decision log completeness

Every render writes decision-log.jsonl. On the 150s real render this was several MB and tens of thousands of records; summarise_log()'s by_stage/by_kind/by_action breakdown is embedded in summary.json so the log's shape is inspectable without opening the raw file.