Skip to content

Proposal (with PR offer): normative evidentiary tiers + a scheme-property model on top of #10's content_fingerprint / provenance (follow-up to #5) #11

Description

@erik-sv

Summary

#5 surfaced the third-party-sourced (resale) coverage gap, and #10 already implements the two fields agreed there: the provenance enum and the scheme-agnostic content_fingerprint object. #10 intentionally leaves two things to a follow-up: the claims-vs-corroborated-facts section its field descriptions are written to slot into, and any structure on the free-string scheme. This issue proposes both, and offers a PR that builds on #10 rather than replacing it.

The intent is not to favour any technique. It is to give the spec a principled, testable way to say which reported signals can be relied on independently of the reporter, so the numbers the standard produces can carry the weight downstream compensation will put on them.

Motivation

Telemetry is a cooperative measurement layer. It reports what willing participants emit, so by construction it cannot represent a non-emitting actor, and it should not try to. That part of the resale gap is an access-layer concern (RSL / CAP / bot-auth), and the right move in the spec is to name the boundary (8.9), as @landomo suggested.

The numbers it does produce are self-reported and become a basis for compensation. #10's descriptions already say detected and preserved_in_output are inference with no origin-side counterpart (the §7.3 gap), and invite "the broader claims-vs-corroborated-facts section when that lands for v0.1." This is that section, plus the scheme semantics that make the distinction operational.

Proposal

1. Normative evidentiary tiers

Define, in a dedicated evidentiary section, three tiers that any reported signal falls into:

  • Tier 1, Claim: agent-reported, uncorroborated.
  • Tier 2, Origin-corroborated: a matching origin-side event exists.
  • Tier 3, Independently verifiable: checkable by any party against a public trust anchor and a trusted timestamp, with no origin-side counterpart required.

Tier 3 is the only tier that closes the §7.3 gap, because it depends on neither the agent's self-report nor an origin-side event. The tiers are scheme-neutral: any technique that qualifies can reach a given tier, and more than one should. This composes directly with #10: its field-level inference labelling becomes "Tier 1 unless the scheme qualifies higher."

2. Scheme-property model for content_fingerprint

Keep scheme a free string, as in #10. Add declared scheme properties that map to the maximum tier the scheme's signals can claim:

  • content_borne: the signal lives in the content and survives reprocessing, rather than in detachable metadata.
  • identity_bearing: it carries rights-holder identity, not only a match.
  • verifiable: it is cryptographically checkable against a public anchor.

A scheme that is content-borne, identity-bearing, and verifiable is eligible for Tier 3. A matching-only fingerprint (for example a similarity hash without a registry) is Tier 1 or 2. An optional evidence pointer on the object can say where the signal is independently checkable.

3. provenance enum

No change proposed; #10's enum and placement are right. Listed here only for completeness.

Worked example: C2PA / C2PA-text as one Tier-3 scheme

Offered as one worked example, not a dependency, and explicitly inviting other Tier-3 schemes. #10 already lists C2PA text provenance as a candidate scheme; this formalises what that entry means.

C2PA text provenance (the A.7 unstructured-text method) binds the manifest into the text itself, so it is content-borne: it survives copy, plain-text save, encoding and serialization round-trips (UTF-8/UTF-16, JSON), and all Unicode normalization forms (NFC/NFD/NFKC/NFKD). A single contiguous in-band block survives whole-text and aggregated reuse; distributed markers extend survival to the fragment case, so a copied span, even mid-sentence, still resolves to the source work. It is identity-bearing (signed publisher certificate) and verifiable (against the public C2PA trust list), with first-publication anchored by an RFC 3161 timestamp, which also answers ISCC's two honest limits: identity rather than matching, and a first-publication anchor.

The "C2PA is attached metadata, stripped on reprocessing" characterisation in #5 is accurate for a manifest embedded in a media file, but not for in-band text provenance, which is a different mechanism.

Survivability: evidence, not assertion

The PR will include a transform-by-transform survivability matrix (Unicode normalization forms, encoding and serialization, copy-paste through common editors, whitespace and markup sanitizers, fragment and aggregation), backed by test output, so robustness is checkable rather than claimed.

Honest boundaries, stated up front:

Consequence worth encoding: absence is informative

Once a robust, expected mark is in play, its absence on otherwise-matching content is itself a signal, and a grounding system asserting a reliable source it cannot attribute is carrying that risk. That is the demand-side reason a cooperating agent populates these fields honestly: they are its own due-diligence record, not only publisher telemetry.

Out of scope (explicit)

  • Adversary coverage and the resale path itself: that is the access layer, not telemetry. This proposal does not claim to close it; it makes the cooperating channel's signals trustworthy and names the boundary.
  • Semantic or paraphrase-robust fingerprinting: pre-standard today and not relied upon here.

Contribution offer

I am happy to author a follow-on PR (on top of #10) that:

  1. adds the normative evidentiary-tiers section;
  2. adds the scheme-property fields and the property-to-tier mapping to content_fingerprint;
  3. adds a scheme-neutral C2PA-text profile as one worked Tier-3 example;
  4. attaches the survivability matrix.

Co-authorship with @landomo and @jchomat welcome, and additional Tier-3 scheme registrations more so.

Open questions

  • Should the tier be an explicit field, or derived from the scheme's declared properties?
  • Is a small registry warranted for scheme strings and their declared properties and tier?
  • How should the Tier-2 origin-corroboration path relate to the access layer (for example RSL <reporting>)?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    profile-evidenceRouted to the optional evidence profile workstream

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions