Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
41 changes: 41 additions & 0 deletions .github/workflows/coverage.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,41 @@
name: Manual branch coverage

on:
workflow_dispatch:

permissions:
contents: read

jobs:
coverage:
runs-on: ubuntu-24.04
timeout-minutes: 20
steps:
- name: Check out repository
uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6

- name: Install Nix
uses: DeterminateSystems/nix-installer-action@ef8a148080ab6020fd15196c2084a2eea5ff2d25 # v22

- name: Synchronize frozen Python environment
run: nix develop --command uv sync --frozen

# coverage.py branch measurement: https://coverage.readthedocs.io/en/7.10.7/branch.html
# uv --with contract: https://docs.astral.sh/uv/reference/cli/#uv-run--with
- name: Run branch coverage
run: nix develop --command uv run --frozen --with coverage==7.10.7 coverage run --branch --source=src -m pytest

- name: Write coverage reports
run: |
nix develop --command uv run --frozen --with coverage==7.10.7 coverage report
nix develop --command uv run --frozen --with coverage==7.10.7 coverage xml -o coverage.xml
nix develop --command uv run --frozen --with coverage==7.10.7 coverage html

- name: Upload coverage evidence
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
with:
name: branch-coverage-${{ github.sha }}
path: |
coverage.xml
htmlcov/
if-no-files-found: error
4 changes: 3 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -76,7 +76,9 @@ uv run --frozen format-bench report --run-dir runs/fair-local

The native robustness suite records pinned Arrow, Vortex, and FastLanes targets. Arrow requires a checkout at the recorded source commit plus binaries in `native/arrow/build`; Vortex and FastLanes require a checkout whose `HEAD` matches the recorded source commit. FastLanes is recorded as project-seeded rather than coverage-guided. Lance, object JSONL, and TsFile have no confirmed official native target and are retained as `UNSUPPORTED` evidence. Missing binaries or mismatched source checkouts never become a silent pass. Select targets with repeated `--target` options and set the run budget with `--duration-seconds` and `--artifact-budget-mib`.

Robustness reports also aggregate each target's case denominator, pass/fail outcomes, crash and timeout counts, incomplete reasons, duration p50, and artifact/source identities. These are reliability evidence, not a cross-lane score.
Robustness reports also aggregate each target's case denominator, pass/fail outcomes, crash and timeout counts, incomplete reasons, duration p50, and artifact/source identities. For generated artifact mutations only, reports add deterministic descriptive counts for the denominator, completed cases, failures, crashes, timeouts, unsupported cases, incomplete cases, and completed percentage. Named boundary cases are excluded; this is not a conventional mutation score and has no ranking or gate effect. The percentage is `N/A` when no generated artifact mutations are present. The manual [`Coverage branch report`](.github/workflows/coverage.yml) layers a pinned coverage.py release over the frozen uv project environment; see the primary [coverage.py branch measurement documentation](https://coverage.readthedocs.io/en/7.10.7/branch.html) and [uv one-off dependency contract](https://docs.astral.sh/uv/reference/cli/#uv-run--with).

Coverage thresholds and source-code mutation scores are intentionally not enforced yet, so regressions can remain ungated. Repository maintainers own this accepted risk under [#506](https://github.com/Anionix/data-format-lab/issues/506); it expires after three successful reports on distinct merged `main` commits or on 2026-08-07, whichever comes first. That issue fixes the evidence fields, numeric threshold decision, serial source-mutation pilot, 20-minute budget, and closeout criteria.

```bash
uv run --frozen format-bench run --profile robustness --suite native --dataset github-stars-2026-07-03 \
Expand Down
46 changes: 45 additions & 1 deletion src/format_bench/report.py
Original file line number Diff line number Diff line change
Expand Up @@ -502,9 +502,34 @@ def case_engine(item: dict) -> object:
for item in evidence["cases"]
]
target_summary = evidence.get("target_summary")
if not isinstance(target_summary, dict) or not target_summary:
if (
not isinstance(target_summary, dict)
or not target_summary
or any(
not isinstance(item, dict) or "artifact_mutation" not in item
for item in target_summary.values()
)
):
target_summary = summarize_cases(evidence["cases"])
evidence["target_summary"] = target_summary
# LLM contract: DISCOVERED -> ENCODED -> ROUNDTRIP_VERIFIED -> BENCHMARKED -> REPORTED.
# Mutation counts are descriptive report evidence only; they never affect ranking or gates.
mutation_rows = []
for target, item in sorted(target_summary.items()):
mutation = item.get("artifact_mutation", {})
Comment thread
Anionix marked this conversation as resolved.
mutation_rows.append(
[
target,
mutation.get("denominator", 0),
mutation.get("completed", 0),
mutation.get("failures", 0),
mutation.get("crashes", 0),
mutation.get("timeouts", 0),
mutation.get("unsupported", 0),
mutation.get("incomplete", 0),
mutation.get("completed_pct"),
]
)
target_rows = [
[
target,
Expand Down Expand Up @@ -563,6 +588,25 @@ def case_engine(item: dict) -> object:
target_rows,
),
"",
"### Artifact Mutation Coverage",
"",
"Artifact mutation counts cover only generated cases with a persisted mutation recipe identity; named boundary cases are excluded. This is descriptive reliability evidence, not a mutation score, and has no ranking or gate effect.",
"",
*_table(
[
"Target",
"Denominator",
"Completed",
"Failures",
"Crashes",
"Timeouts",
"Unsupported",
"Incomplete",
"Completed %",
],
mutation_rows,
),
"",
"### Evidence Identities",
"",
*_table(["Target", "Artifact SHA-256", "Source identity"], identity_rows),
Expand Down
66 changes: 66 additions & 0 deletions src/format_bench/robustness/summary.py
Original file line number Diff line number Diff line change
Expand Up @@ -23,9 +23,67 @@
"duration_ms_p50": float | None,
"artifact_sha256": list[str],
"source_identities": list[str],
"artifact_mutation": "ArtifactMutationSummary",
},
)

ArtifactMutationSummary = TypedDict(
"ArtifactMutationSummary",
{
"denominator": int,
"completed": int,
"failures": int,
"crashes": int,
"timeouts": int,
"unsupported": int,
"incomplete": int,
"completed_pct": float | None,
},
)


@dataclass
class _ArtifactMutationAccumulator:
denominator: int = 0
completed: int = 0
failures: int = 0
crashes: int = 0
timeouts: int = 0
unsupported: int = 0
incomplete: int = 0

def observe(self, verdict: str, observed: str) -> None:
self.denominator += 1
if verdict in {"PASS", "FAIL"}:
self.completed += 1
if verdict == "FAIL":
self.failures += 1
if observed == "CRASHED":
self.crashes += 1
elif observed == "TIMED_OUT":
self.timeouts += 1
elif observed == "UNSUPPORTED":
self.unsupported += 1
if verdict == "INCOMPLETE":
self.incomplete += 1

def emit(self) -> ArtifactMutationSummary:
percentage = (
round(100 * self.completed / self.denominator, 3)
if self.denominator
else None
)
return {
"denominator": self.denominator,
"completed": self.completed,
"failures": self.failures,
"crashes": self.crashes,
"timeouts": self.timeouts,
"unsupported": self.unsupported,
"incomplete": self.incomplete,
"completed_pct": percentage,
}


@dataclass
class _TargetAccumulator:
Expand All @@ -43,6 +101,9 @@ class _TargetAccumulator:
durations: list[float] = field(default_factory=list)
artifact_sha256: set[str] = field(default_factory=set)
source_identities: set[str] = field(default_factory=set)
artifact_mutation: _ArtifactMutationAccumulator = field(
default_factory=_ArtifactMutationAccumulator
)

def observe_outcome(self, observed: str) -> None:
if observed == "CRASHED":
Expand Down Expand Up @@ -74,6 +135,7 @@ def emit(self) -> TargetSummary:
),
"artifact_sha256": sorted(self.artifact_sha256),
"source_identities": sorted(self.source_identities),
"artifact_mutation": self.artifact_mutation.emit(),
}


Expand Down Expand Up @@ -110,6 +172,10 @@ def summarize_cases(cases: Sequence[Mapping[str, object]]) -> dict[str, TargetSu
group.cases += 1
verdict = _value(case.get("verdict"))
observed = _value(case.get("observed"))
mutation = case.get("mutation")
recipe_id = mutation.get("recipe_id") if isinstance(mutation, Mapping) else None
if isinstance(recipe_id, str) and recipe_id:
group.artifact_mutation.observe(verdict, observed)
if verdict != "NOT_APPLICABLE":
group.applicable += 1
if verdict == "PASS":
Expand Down
12 changes: 11 additions & 1 deletion tests/test_report.py
Original file line number Diff line number Diff line change
Expand Up @@ -453,16 +453,24 @@ def test_robustness_report_separates_case_contract_and_is_deterministic(
},
"summary": {
"PASS": 1,
"FAIL": 0,
"FAIL": 1,
"NOT_APPLICABLE": 0,
"INCOMPLETE": 0,
},
# A pre-mutation-metrics summary must be refreshed from persisted cases.
"target_summary": {"csv": {"tier": "CORE"}},
"cases": [
{
"target": "csv", "tier": "CORE", "details": {"engine": "coverage-guided"}, "case_id": "rows-1",
"expectation": "MUST_ROUNDTRIP",
"observed": "ROUNDTRIP_EQUAL", "verdict": "PASS",
},
{
"target": "csv", "tier": "CORE", "case_id": "mutation-000",
"mutation": {"recipe_id": "recipe-000", "operation": "truncate"},
"expectation": "MUST_NOT_CRASH",
"observed": "CRASHED", "verdict": "FAIL",
},
],
}
},
Expand All @@ -482,4 +490,6 @@ def test_robustness_report_separates_case_contract_and_is_deterministic(
reported = json.loads((tmp_path / "results.json").read_text())
assert reported["results"]["robustness_v1"]["state"] == "REPORTED"
assert reported["results"]["robustness_v1"]["target_summary"]["csv"]["pass"] == 1
assert "### Artifact Mutation Coverage" in first
assert "| csv | 1 | 1 | 1 | 1 | 0 | 0 | 0 | 100.0 |" in first
assert render_report(tmp_path).read_text() == first
58 changes: 58 additions & 0 deletions tests/test_robustness_summary.py
Original file line number Diff line number Diff line change
Expand Up @@ -35,4 +35,62 @@ def test_summary_ignores_malformed_duration_and_hash_evidence() -> None:
"duration_ms_p50": None,
"artifact_sha256": ["artifact-hash"],
"source_identities": [],
"artifact_mutation": {
"denominator": 0,
"completed": 0,
"failures": 0,
"crashes": 0,
"timeouts": 0,
"unsupported": 0,
"incomplete": 0,
"completed_pct": None,
},
}


def test_summary_counts_only_artifact_mutations_and_reports_nullable_completion() -> None:
summary = summarize_cases(
[
{"target": "csv", "tier": "CORE", "case_id": "rows-1", "verdict": "PASS", "observed": "ROUNDTRIP_EQUAL"},
{
"target": "csv",
"tier": "CORE",
"case_id": "malformed-truncated",
"mutation": {"operation": "truncate"},
"verdict": "PASS",
"observed": "REJECTED",
},
{"target": "csv", "tier": "CORE", "case_id": "mutation-000", "mutation": {"recipe_id": "recipe-000", "operation": "flip"}, "verdict": "PASS", "observed": "ROUNDTRIP_EQUAL"},
{"target": "csv", "tier": "CORE", "case_id": "mutation-001", "mutation": {"recipe_id": "recipe-001", "operation": "truncate"}, "verdict": "FAIL", "observed": "CRASHED"},
{"target": "csv", "tier": "CORE", "case_id": "mutation-002", "mutation": {"recipe_id": "recipe-002", "operation": "flip"}, "verdict": "INCOMPLETE", "observed": "TIMED_OUT"},
{"target": "csv", "tier": "CORE", "case_id": "mutation-003", "mutation": {"recipe_id": "recipe-003", "operation": "flip"}, "verdict": "INCOMPLETE", "observed": "UNSUPPORTED"},
]
)

assert summary["csv"]["artifact_mutation"] == {
"denominator": 4,
"completed": 2,
"failures": 1,
"crashes": 1,
"timeouts": 1,
"unsupported": 1,
"incomplete": 2,
"completed_pct": 50.0,
}


def test_summary_uses_absent_percentage_for_no_artifact_mutations() -> None:
summary = summarize_cases(
[{"target": "csv", "tier": "CORE", "verdict": "PASS", "observed": "ROUNDTRIP_EQUAL"}]
)

assert summary["csv"]["artifact_mutation"] == {
"denominator": 0,
"completed": 0,
"failures": 0,
"crashes": 0,
"timeouts": 0,
"unsupported": 0,
"incomplete": 0,
"completed_pct": None,
}
Loading