Skip to content

perf: replace public benchmark evidence pipeline - #10

Merged
goblinmode2700 merged 7 commits into
mainfrom
infra/r-owned-benchmark-evidence-public
Aug 14, 2026
Merged

goblinmode2700 merged 7 commits into
mainfrom
infra/r-owned-benchmark-evidence-public

Conversation

@goblinmode2700

Copy link
Copy Markdown
Owner

Intent

Replace the public performance benchmark harness so Python only collects raw timings, R owns all statistical inference with family-wise multiplicity control, release benchmarking is reduced by roughly seventeen minutes, and every release GitHub Actions workflow benchmarks the exact verified ABI3 wheel and publishes both raw and inferred evidence; validate, push, open the pull request, and take it through CI for merge.

What Changed

  • Replace Python-side benchmark inference with manifest-driven raw timing collection and R analyzers using simultaneous Bonferroni intervals and Holm family-wise adjustment.
  • Reduce release gating from the 100-cell exploratory grid to a preregistered 22-endpoint guard, and batch absolute-report measurements into complete process panels.
  • Verify and benchmark the exact Linux x86-64 CPython 3.13 ABI3 release wheel, then publish its identity record plus raw and R-inferred evidence in workflow artifacts and tagged GitHub releases.

Risk Assessment

🚨 High: The change should not merge without resolving the exact-wheel isolation gap and clarifying or correcting the claimed family-wise inference contract.

Testing

No pre-run baseline commands were supplied. The 63 targeted automated contracts passed, the exact verified ABI3 wheel completed the full 253-endpoint raw-to-R design in 185.86 seconds, semantic workflow checks confirmed publication of both evidence layers, and representative historical measurements support a 16.76-minute reduction. A full historical run exceeded the local execution-session ceiling, so the runtime baseline uses completed category samples weighted across the old execution geometry. The emitted R evidence reports five PASS families and G4 FAIL; this demonstrates inference output and is not a harness failure.

Evidence: Consolidated acceptance evidence
{
  "candidate": {
    "source_revision": "49b8121b4a1c59336a22f6dbd95774dc5d3fe2bc",
    "wheel": "msgspec_toon-0.3.0b3-cp313-abi3-macosx_11_0_arm64.whl",
    "wheel_sha256": "f3f1ae88cd556e97008e9d3bddcb3a31ba5438e34469609fcbbd6c543fd15083",
    "python_abi": "cp313-abi3",
    "installed_verification": "passed",
    "native_sha256": "bada2a2050ac7748d228c055955acd12659e95f8f5530f3f5a6659567c0316de",
    "raw_workers_used_that_native_sha256": true
  },
  "python_raw_evidence": {
    "kind": "absolute_report_raw",
    "analysis_contract": "Python collects raw timings; R owns inference",
    "endpoints": 253,
    "vectorized_rows": 37,
    "workers": 10,
    "observations": 7590,
    "contains_inferential_fields": false,
    "sha256": "968779e9fc67899bb2014e52a8d471bf60271c39d63febd83b8fddc30551d81d"
  },
  "r_inference_evidence": {
    "engine": "R stats",
    "adjustment": "holm within each declared gate family",
    "interval": "simultaneous Bonferroni t intervals",
    "raw_sha256_matches": true,
    "analyzer_sha256": "25daac600ebebf2c85a8c8b8bd7a59b3541890b121b0fb023541e3730c394d0f",
    "family_decisions": {
      "G3_typed_decode_beats_wrapper": "PASS",
      "G4_whole_encode_beats_to_builtins_alone": "FAIL",
      "typed_beats_incumbent_pipeline_decode": "PASS",
      "typed_beats_incumbent_pipeline_encode": "PASS",
      "G5_encode_not_slower_than_toons": "PASS",
      "G5_decode_not_slower_than_toons": "PASS"
    }
  },
  "release_workflow": {
    "release_workflows_found": [
      "wheels.yml"
    ],
    "evidence_job_needs": [
      "collect",
      "validate"
    ],
    "benchmark_command": "make release-performance",
    "published_evidence_paths": [
      "benches/ab-guard-r.json",
      "benches/ab-guard-raw.json",
      "benches/report-performance-raw.json",
      "benches/report-performance.json",
      "release/benchmark-wheel.verified.json"
    ]
  },
  "runtime": {
    "base_representative_37_row_estimate_seconds": 1191.547,
    "base_representative_37_row_estimate_minutes": 19.86,
    "verified_abi3_current_observed_seconds": 185.86,
    "verified_abi3_current_observed_minutes": 3.1,
    "estimated_reduction_seconds": 1005.687,
    "estimated_reduction_minutes": 16.76,
    "base_worker_processes": 407,
    "current_worker_processes": 11
  }
}
Evidence: Verified ABI3 benchmark-wheel identity
{
  "artifact": {
    "filename": "msgspec_toon-0.3.0b3-cp313-abi3-macosx_11_0_arm64.whl",
    "kind": "wheel",
    "sha256": "f3f1ae88cd556e97008e9d3bddcb3a31ba5438e34469609fcbbd6c543fd15083",
    "size": 367277
  },
  "benchmark_verification": {
    "distribution_version": "0.3.0b3",
    "gil_enabled": true,
    "machine": "arm64",
    "native_path": "/private/var/folders/pt/19wwr80d5bs38zmww5rm2vsc0000gn/T/no-mistakes-evidence/01KZZ1DNBPEZYZM9D6GQF5CPF0/abi3-venv/lib/python3.13/site-packages/msgspec_toon/_native.abi3.so",
    "package_path": "/private/var/folders/pt/19wwr80d5bs38zmww5rm2vsc0000gn/T/no-mistakes-evidence/01KZZ1DNBPEZYZM9D6GQF5CPF0/abi3-venv/lib/python3.13/site-packages/msgspec_toon/__init__.py",
    "platform": "macOS-15.1-arm64-arm-64bit-Mach-O",
    "python": "3.13.12",
    "python_implementation": "CPython",
    "status": "passed"
  },
  "schema_version": 1,
  "source_revision": "49b8121b4a1c59336a22f6dbd95774dc5d3fe2bc",
  "source_verification": {
    "distribution_version": "0.3.0b3",
    "gil_enabled": true,
    "machine": "arm64",
    "native_path": "/private/var/folders/pt/19wwr80d5bs38zmww5rm2vsc0000gn/T/no-mistakes-evidence/01KZZ1DNBPEZYZM9D6GQF5CPF0/abi3-venv/lib/python3.13/site-packages/msgspec_toon/_native.abi3.so",
    "package_path": "/private/var/folders/pt/19wwr80d5bs38zmww5rm2vsc0000gn/T/no-mistakes-evidence/01KZZ1DNBPEZYZM9D6GQF5CPF0/abi3-venv/lib/python3.13/site-packages/msgspec_toon/__init__.py",
    "platform": "macOS-15.1-arm64-arm-64bit-Mach-O",
    "python": "3.13.12",
    "python_implementation": "CPython",
    "status": "passed"
  },
  "target": {
    "architecture": "aarch64",
    "operating_system": "macos",
    "python_abi": "cp313-abi3"
  }
}
Evidence: Verified ABI3 benchmark transcript

Full 253-endpoint, 10-worker ABI3 panel completed in 185.86 seconds; R emitted five PASS decisions and one G4 FAIL decision.

calibrating 253 endpoints in 37 vectorized rows (5 ms target)
worker  1/10
worker  2/10
worker  3/10
worker  4/10
worker  5/10
worker  6/10
worker  7/10
worker  8/10
worker  9/10
worker 10/10
R gate decisions: {"G3_typed_decode_beats_wrapper": "PASS", "G4_whole_encode_beats_to_builtins_alone": "FAIL", "G5_decode_not_slower_than_toons": "PASS", "G5_encode_not_slower_than_toons": "PASS", "typed_beats_incumbent_pipeline_decode": "PASS", "typed_beats_incumbent_pipeline_encode": "PASS"}
raw observations: /private/var/folders/pt/19wwr80d5bs38zmww5rm2vsc0000gn/T/no-mistakes-evidence/01KZZ1DNBPEZYZM9D6GQF5CPF0/abi3-report-default-raw.json
R analysis:       /private/var/folders/pt/19wwr80d5bs38zmww5rm2vsc0000gn/T/no-mistakes-evidence/01KZZ1DNBPEZYZM9D6GQF5CPF0/abi3-report-default-r.json
real 185.86
user 160.59
sys 22.20
Evidence: Python raw timing evidence
{
  "schema_version": 1,
  "kind": "absolute_report_raw",
  "run_id": "3d07cb21-6ece-4598-b574-475cdf7f7fd0",
  "created_at": "2026-08-14T03:06:29.202103+00:00",
  "source_revision": "49b8121b4a1c59336a22f6dbd95774dc5d3fe2bc",
  "dependency_lock_sha256": "e67c9caeb32b29082dd1bb5839d8be3fbf3a3f20ef12705393cca9d14168374c",
  "manifest_sha256": "b814fb0812a10cb1bdf3454c3a3b36c849fcae556aad131391dd5f51779fbbdf",
  "analysis_contract": "Python collects raw timings; R owns inference",
  "seed": 20260813,
  "qualification_override": false,
  "design": {
    "workers": 10,
    "samples_per_process": 3,
    "target_milliseconds": 5.0,
    "alpha": 0.05
  },
  "endpoints": [
    {
      "id": "absolute:typed@16:decoder-construction",
      "panel": "typed",
      "row_id": "typed@16",
      "metric_slug": "decoder-construction",
      "module": "bench_typed",
      "metadata_function": "metadata_run",
      "args": [
        16
      ],
      "section": "plan_us",
      "metric": "decoder_construction_cached",
      "result_path": [
        "plan_us",
        "decoder_construction_cached"
      ],
      "sampler_metric": "decode.decoder_construction"
    },
    {
      "id": "absolute:typed@16:typed-decode",
      "panel": "typed",
      "row_id": "typed@16",
      "metric_slug": "typed-decode",
      "module": "bench_typed",
      "metadata_function": "metadata_run",
      "args": [
        16
      ],
      "section": "decode_us",
      "metric": "typed_direct",
      "result_path": [
        "decode_us",
        "typed_direct"
      ],
      "sampler_metric": "decode.typed_direct"
    },
    {
      "id": "absolute:typed@16:functional-decode",
      "panel": "typed",
      "row_id": "typed@16",
      "metric_slug": "functional-decode",
      "module": "bench_typed",
      "metadata_function": "metadata_run",
      "args": [
        16
      ],
      "section": "decode_us",
      "metric": "functional",
      "result_path": [
        "decode_us",
        "functional"
      ],
      "sampler_metric": "decode.functional"
    },
    {
      "id": "absolute:typed@16:untyped-decode",
      "panel": "typed",
      "row_id": "typed@16",
      "metric_slug": "untyped-decode",
      "module": "bench_typed",
      "metadata_function": "metadata_run",
      "args": [
        16
      ],
      "section": "decode_us",
      "metric": "untyped_tree",
      "result_path": [
        "decode_us",
        "untyped_tree"
      ],
      "sampler_metric": "decode.untyped_tree"
    },
    {
      "id": "absolute:typed@16:keyed-decode",
      "panel": "typed",
      "row_id": "typed@16",
      "metric_slug": "keyed-decode",
      "module": "bench_typed",
      "metadata_function": "metadata_run",
      "args": [
        16
      ],
      "section": "decode_us",
      "metric": "keyed_document",
      "result_path": [
        "decode_us",
        "keyed_document"
      ],
      "sampler_metric": "decode.keyed_document"
    },
    {
      "id": "absolute:typed@16:entry-decode",
      "panel": "typed",
      "row_id": "typed@16",
      "metric_slug": "entry-decode",
      "module": "bench_typed",
      "metadata_function": "metadata_run",
      "args": [
        16
      ],
      "section": "decode_us",
      "metric": "entry_document",
      "result_path": [
        "decode_us",
        "entry_document"
      ],
      "sampler_metric": "decode.entry_document"
    },
    {
      "id": "absolute:typed@16:convert-only",
      "panel": "typed",
      "row_id": "typed@16",
      "metric_slug": "convert-only",
      "module": "bench_typed",
      "metadata_function": "metadata_run",
      "args": [
        16
      ],
      "section": "decode_us",
      "metric": "convert_only",
      "result_path": [
        "decode_us",
        "convert_only"
      ],
      "sampler_metric": "decode.convert_only"
    },
    {
      "id": "absolute:typed@16:wrapper-decode",
      "panel": "typed",
      "row_id": "typed@16",
      "metric_slug": "wrapper-decode",
      "module": "bench_typed",
      "metadata_function": "metadata_run",
      "args": [
        16
      ],
      "section": "decode_us",
      "metric": "wrapper_tree_plus_convert",
      "result_path": [
        "decode_us",
        "wrapper_tree_plus_convert"
      ],
      "sampler_metric": "decode.wrapper"
    },
    {
      "id": "absolute:typed@16:incumbent-decode",
      "panel": "typed",
      "row_id": "typed@16",
      "metric_slug": "incumbent-decode",
      "module": "bench_typed",
      "metadata_function": "metadata_run",
      "args": [
        16
      ],
      "section": "decode_us",
      "metric": "incumbent_pipeline_python_toon_plus_convert",
      "result_path": [
        "decode_us",
        "incumbent_pipeline_python_toon_plus_convert"
      ],
      "sampler_metric": "decode.incumbent"
    },
    {
      "id": "absolute:typed@16:json-decode",
      "panel": "typed",
      "row_id": "typed@16",
      "metric_slug": "json-decode",
      "module": "bench_typed",
      "metadata_function": "metadata_run",
      "args": [
        16
      ],
      "section": "decode_us",
      "metric": "msgspec_json_native",
      "result_path": [
        "decode_us",
        "msgspec_json_native"
      ],
      "sampler_metric": "decode.json_native"
    },
    {
      "id": "absolute:typed@16:typed-encode",
      "panel": "typed",
      "row_id": "typed@16",
      "metric_slug": "typed-encode",
      "module": "bench_typed",
      "metadata_function": "metadata_run",
      "args": [
        16
      ],
      "section": "encode_us",
      "metric": "typed_direct_whole",
      "result_path": [
        "encode_us",
        "typed_direct_whole"
      ],
      "sampler_metric": "encode.typed_direct"
    },
    {
      "id": "absolute:typed@16:functional-encode",
      "panel": "typed",
      "row_id": "typed@16",
      "metric_slug": "functional-encode",
      "module": "bench_typed",
      "metadata_function": "metadata_run",
      "args": [
        16
      ],
      "section": "encode_us",
      "metric": "functional",
      "result_path": [
        "encode_us",
        "functional"
      ],
      "sampler_metric": "encode.functional"
    },
    {
      "id": "absolute:typed@16:entry-encode",
      "panel": "typed",
      "row_id": "typed@16",
      "metric_slug": "entry-encode",
      "module": "bench_typed",
      "metadata_function": "metadata_run",
      "args": [
        16
      ],
      "section": "encode_us",
      "metric": "entry_document",
      "result_path": [
        "encode_us",
        "entry_document"
      ],
      "sampler_metric": "encode.entry_document"
    },
    {
      "id": "absolute:typed@16:wide-dict-encode",
      "panel": "typed",
      "row_id": "typed@16",
      "metric_slug": "wide-dict-encode",
      "module": "bench_typed",
      "metadata_function": "metadata_run",
      "args": [
        16
      ],
      "section": "encode_us",
      "metric": "wide_dict_document",
      "result_path": [
        "encode_us",
        "wide_dict_document"
      ],
      "sampler_metric": "encode.wide_dict_document"
    },
    {
      "id": "absolute:typed@16:to-builtins",
      "panel": "typed",
      "row_id": "typed@16",
      "metric_slug": "to-builtins",
      "module": "bench_typed",
      "metadata_function": "metadata_run",
      "args": [
        16
      ],
      "section": "encode_us",
      "metric": "to_builtins_alone",
      "result_path": [
        "encode_us",
        "to_builtins_alone"
      ],
      "sampler_metric": "encode.to_builtins"
    },
    {
      "id": "absolute:typed@16:incumbent-encode",
      "panel": "typed",
      "row_id": "typed@16",
      "metric_slug": "incumbent-encode",
      "module": "bench_typed",
      "metadata_function": "metadata_run",
      "args": [
        16
      ],
      "section": "encode_us",
      "metric": "incumbent_pipeline_to_builtins_plus_python_toon",
      "result_path": [
        "encode_us",
        "incumbent_pipeline_to_builtins_plus_python_toon"
      ],
      "sampler_metric": "encode.incumbent"
    },
    {
      "id": "absolute:typed@16:json-encode",
      "panel": "typed",
      "row_id": "typed@16",
      "metric_slug": "json-encode",
    

... [2102692 bytes truncated] ...

r": 9,
      "order_index": 33,
      "sample": 2,
      "loops": 8,
      "elapsed_ns": 7895458
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:ours-encode",
      "worker": 9,
      "order_index": 34,
      "sample": 0,
      "loops": 512,
      "elapsed_ns": 10060084
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:ours-encode",
      "worker": 9,
      "order_index": 34,
      "sample": 1,
      "loops": 512,
      "elapsed_ns": 9814333
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:ours-encode",
      "worker": 9,
      "order_index": 34,
      "sample": 2,
      "loops": 512,
      "elapsed_ns": 9797292
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:toons-encode",
      "worker": 9,
      "order_index": 34,
      "sample": 0,
      "loops": 128,
      "elapsed_ns": 9648250
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:toons-encode",
      "worker": 9,
      "order_index": 34,
      "sample": 1,
      "loops": 128,
      "elapsed_ns": 9637541
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:toons-encode",
      "worker": 9,
      "order_index": 34,
      "sample": 2,
      "loops": 128,
      "elapsed_ns": 9670042
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:python-toon-encode",
      "worker": 9,
      "order_index": 34,
      "sample": 0,
      "loops": 32,
      "elapsed_ns": 7355125
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:python-toon-encode",
      "worker": 9,
      "order_index": 34,
      "sample": 1,
      "loops": 32,
      "elapsed_ns": 7348042
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:python-toon-encode",
      "worker": 9,
      "order_index": 34,
      "sample": 2,
      "loops": 32,
      "elapsed_ns": 7391000
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:json-encode",
      "worker": 9,
      "order_index": 34,
      "sample": 0,
      "loops": 1024,
      "elapsed_ns": 5498041
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:json-encode",
      "worker": 9,
      "order_index": 34,
      "sample": 1,
      "loops": 1024,
      "elapsed_ns": 5549000
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:json-encode",
      "worker": 9,
      "order_index": 34,
      "sample": 2,
      "loops": 1024,
      "elapsed_ns": 5493667
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:ours-decode",
      "worker": 9,
      "order_index": 34,
      "sample": 0,
      "loops": 512,
      "elapsed_ns": 8647791
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:ours-decode",
      "worker": 9,
      "order_index": 34,
      "sample": 1,
      "loops": 512,
      "elapsed_ns": 8618209
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:ours-decode",
      "worker": 9,
      "order_index": 34,
      "sample": 2,
      "loops": 512,
      "elapsed_ns": 8592375
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:toons-decode",
      "worker": 9,
      "order_index": 34,
      "sample": 0,
      "loops": 256,
      "elapsed_ns": 7150000
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:toons-decode",
      "worker": 9,
      "order_index": 34,
      "sample": 1,
      "loops": 256,
      "elapsed_ns": 7143125
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:toons-decode",
      "worker": 9,
      "order_index": 34,
      "sample": 2,
      "loops": 256,
      "elapsed_ns": 7118708
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:python-toon-decode",
      "worker": 9,
      "order_index": 34,
      "sample": 0,
      "loops": 16,
      "elapsed_ns": 7260292
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:python-toon-decode",
      "worker": 9,
      "order_index": 34,
      "sample": 1,
      "loops": 16,
      "elapsed_ns": 7249792
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:python-toon-decode",
      "worker": 9,
      "order_index": 34,
      "sample": 2,
      "loops": 16,
      "elapsed_ns": 7237833
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:json-decode",
      "worker": 9,
      "order_index": 34,
      "sample": 0,
      "loops": 512,
      "elapsed_ns": 6125875
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:json-decode",
      "worker": 9,
      "order_index": 34,
      "sample": 1,
      "loops": 512,
      "elapsed_ns": 6135291
    },
    {
      "cell_id": "absolute:codecs:numeric-heavy@64:json-decode",
      "worker": 9,
      "order_index": 34,
      "sample": 2,
      "loops": 512,
      "elapsed_ns": 6132125
    },
    {
      "cell_id": "absolute:integration:uniform-records@4096:ours",
      "worker": 9,
      "order_index": 35,
      "sample": 0,
      "loops": 2,
      "elapsed_ns": 5611666
    },
    {
      "cell_id": "absolute:integration:uniform-records@4096:ours",
      "worker": 9,
      "order_index": 35,
      "sample": 1,
      "loops": 2,
      "elapsed_ns": 5622000
    },
    {
      "cell_id": "absolute:integration:uniform-records@4096:ours",
      "worker": 9,
      "order_index": 35,
      "sample": 2,
      "loops": 2,
      "elapsed_ns": 5614458
    },
    {
      "cell_id": "absolute:integration:uniform-records@4096:python-toon",
      "worker": 9,
      "order_index": 35,
      "sample": 0,
      "loops": 1,
      "elapsed_ns": 62823583
    },
    {
      "cell_id": "absolute:integration:uniform-records@4096:python-toon",
      "worker": 9,
      "order_index": 35,
      "sample": 1,
      "loops": 1,
      "elapsed_ns": 62349375
    },
    {
      "cell_id": "absolute:integration:uniform-records@4096:python-toon",
      "worker": 9,
      "order_index": 35,
      "sample": 2,
      "loops": 1,
      "elapsed_ns": 62687291
    },
    {
      "cell_id": "absolute:integration:uniform-records@4096:python-toon-cli",
      "worker": 9,
      "order_index": 35,
      "sample": 0,
      "loops": 1,
      "elapsed_ns": 136795750
    },
    {
      "cell_id": "absolute:integration:uniform-records@4096:python-toon-cli",
      "worker": 9,
      "order_index": 35,
      "sample": 1,
      "loops": 1,
      "elapsed_ns": 133910208
    },
    {
      "cell_id": "absolute:integration:uniform-records@4096:python-toon-cli",
      "worker": 9,
      "order_index": 35,
      "sample": 2,
      "loops": 1,
      "elapsed_ns": 134377500
    },
    {
      "cell_id": "absolute:integration:string-heavy@4096:ours",
      "worker": 9,
      "order_index": 36,
      "sample": 0,
      "loops": 2,
      "elapsed_ns": 7432125
    },
    {
      "cell_id": "absolute:integration:string-heavy@4096:ours",
      "worker": 9,
      "order_index": 36,
      "sample": 1,
      "loops": 2,
      "elapsed_ns": 7552208
    },
    {
      "cell_id": "absolute:integration:string-heavy@4096:ours",
      "worker": 9,
      "order_index": 36,
      "sample": 2,
      "loops": 2,
      "elapsed_ns": 7387166
    },
    {
      "cell_id": "absolute:integration:string-heavy@4096:python-toon",
      "worker": 9,
      "order_index": 36,
      "sample": 0,
      "loops": 1,
      "elapsed_ns": 69576083
    },
    {
      "cell_id": "absolute:integration:string-heavy@4096:python-toon",
      "worker": 9,
      "order_index": 36,
      "sample": 1,
      "loops": 1,
      "elapsed_ns": 68787833
    },
    {
      "cell_id": "absolute:integration:string-heavy@4096:python-toon",
      "worker": 9,
      "order_index": 36,
      "sample": 2,
      "loops": 1,
      "elapsed_ns": 69828833
    },
    {
      "cell_id": "absolute:integration:string-heavy@4096:python-toon-cli",
      "worker": 9,
      "order_index": 36,
      "sample": 0,
      "loops": 1,
      "elapsed_ns": 150141375
    },
    {
      "cell_id": "absolute:integration:string-heavy@4096:python-toon-cli",
      "worker": 9,
      "order_index": 36,
      "sample": 1,
      "loops": 1,
      "elapsed_ns": 150190417
    },
    {
      "cell_id": "absolute:integration:string-heavy@4096:python-toon-cli",
      "worker": 9,
      "order_index": 36,
      "sample": 2,
      "loops": 1,
      "elapsed_ns": 150042125
    }
  ]
}
- Evidence: R-owned inference evidence (local file: /var/folders/pt/19wwr80d5bs38zmww5rm2vsc0000gn/T/no-mistakes-evidence/01KZZ1DNBPEZYZM9D6GQF5CPF0/abi3-report-default-r.json)
Evidence: Historical runtime comparison

Representative historical rows produced a 19.86-minute weighted estimate versus 3.10 minutes observed for the verified ABI3 harness—a 16.76-minute estimated reduction.

{"row": "typed@4096", "elapsed_seconds": 62.674}
{"row": "codecs:irregular@4096", "elapsed_seconds": 27.941}
{"row": "integration:irregular@4096", "elapsed_seconds": 28.656}
{"row": "key-cardinality@4096", "elapsed_seconds": 35.299}
{
  "contract": "base harness representative maximum-size rows",
  "base_commit": "f6cfb27b0527f88b6618a6c9c6dcc0ec9cf57c9f",
  "worker_processes_per_row": 11,
  "samples": [
    {
      "row": "typed@4096",
      "elapsed_seconds": 62.674
    },
    {
      "row": "codecs:irregular@4096",
      "elapsed_seconds": 27.941
    },
    {
      "row": "integration:irregular@4096",
      "elapsed_seconds": 28.656
    },
    {
      "row": "key-cardinality@4096",
      "elapsed_seconds": 35.299
    }
  ],
  "size-weighted_37_row_seconds": 1191.547,
  "size-weighted_37_row_minutes": 19.86
}
real 154.65
user 150.06
sys 4.07

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 3 errors
  • 🚨 tests/test_release_workflows.py:252 - This test reads workflow YAML through a regex-based indentation parser and then asserts literal/substrings of run scripts. It can pass without proving that GitHub executes the verified-wheel install, benchmark pipeline, or evidence publication, violating the source-content-only test rule. Replace these assertions—and the same-pattern workflow checks directly in scope—with a real workflow/YAML semantic consumer plus executable seams that assert installed-wheel identity and produced evidence artifacts.

🔧 Fix: Verify release wheel identity and evidence contracts
3 errors still open:

  • 🚨 benches/_panel.py:70 - Intent requires benchmarking “the exact verified ABI3 wheel,” but workers inherit PYTHONPATH and run without -I. A shadow msgspec_toon package can therefore supply different Python wrappers while reusing the verified native extension; the later check compares only the extension digest and would accept it. Launch this shared worker boundary in isolated mode and bind its package/native paths to the verification record.
  • 🚨 benches/analyze_ab.R:224 - Intent requires “family-wise multiplicity control,” but one declared A/B family is split into separate Holm procedures for non-inferiority and improvement claims. For distinct-key-hotfix, the two procedures can jointly exceed alpha while the output advertises one holm adjustment. Adjust one role-appropriate p-value vector across every confirmatory member, or explicitly declare separate families.
  • 🚨 benches/analyze_report.R:264 - The report analyzer similarly runs independent Holm procedures for meets_floor and misses_floor within each declared family, then publishes both as inferential statuses. At the boundary, either directional false classification can total roughly 10% even for a one-comparison family. Use one simultaneous two-sided classification procedure, or mark one direction non-confirmatory.
✅ **Test** - passed

✅ No issues found.

  • uv run --no-sync pytest -q tests/test_performance_evidence.py tests/test_release_artifacts.py tests/test_release_report.py tests/test_release_workflows.py
  • CARGO_TARGET_DIR=<evidence>/cargo-target uv run --no-sync maturin build --release --locked --compatibility pypi --out <evidence>/abi3-wheels -i /opt/homebrew/opt/python@3.13/bin/python3.13
  • scripts/release_artifacts.py manifest, followed by isolated ABI3-environment verify and verify-release-wheel --python-abi cp313-abi3
  • /usr/bin/time -p <verified-abi3-python> -I benches/collect_report.py --python <verified-abi3-python> --seed 20260813 --raw-output <evidence>/abi3-report-default-raw.json --output <evidence>/abi3-report-default-r.json
  • Executed representative base-commit public benchmark rows for typed, codec, integration, and key-cardinality categories and weighted them across the historical 37-row/407-process geometry.
  • Parsed and asserted the generated ABI3 identity, raw timing, R inference, runtime, and normalized GitHub Actions workflow contracts into validation-summary.json.
  • Removed test-created caches and confirmed git status --short is empty.
🔧 **Document** - 2 issues found → auto-fixed ✅
  • ⚠️ BENCHMARKS.md:57 - The generated benchmark report still documents the retired Student t and fixed-order methodology. Regeneration requires fresh benchmark evidence, which this phase forbids.
  • ⚠️ benches/_timing.py:63 - The runtime methodology text still says Python publishes worker means and standard deviations. This phase permits edits only to documentation files and doc comments.

🔧 Fix: Refresh benchmark inference documentation
✅ Re-checked - no issues remain.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

@goblinmode2700
goblinmode2700 merged commit 1eda3a8 into main Aug 14, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant