Summary
#5 surfaced the third-party-sourced (resale) coverage gap, and #10 already implements the two fields agreed there: the provenance enum and the scheme-agnostic content_fingerprint object. #10 intentionally leaves two things to a follow-up: the claims-vs-corroborated-facts section its field descriptions are written to slot into, and any structure on the free-string scheme. This issue proposes both, and offers a PR that builds on #10 rather than replacing it.
The intent is not to favour any technique. It is to give the spec a principled, testable way to say which reported signals can be relied on independently of the reporter, so the numbers the standard produces can carry the weight downstream compensation will put on them.
Motivation
Telemetry is a cooperative measurement layer. It reports what willing participants emit, so by construction it cannot represent a non-emitting actor, and it should not try to. That part of the resale gap is an access-layer concern (RSL / CAP / bot-auth), and the right move in the spec is to name the boundary (8.9), as @landomo suggested.
The numbers it does produce are self-reported and become a basis for compensation. #10's descriptions already say detected and preserved_in_output are inference with no origin-side counterpart (the §7.3 gap), and invite "the broader claims-vs-corroborated-facts section when that lands for v0.1." This is that section, plus the scheme semantics that make the distinction operational.
Proposal
1. Normative evidentiary tiers
Define, in a dedicated evidentiary section, three tiers that any reported signal falls into:
- Tier 1, Claim: agent-reported, uncorroborated.
- Tier 2, Origin-corroborated: a matching origin-side event exists.
- Tier 3, Independently verifiable: checkable by any party against a public trust anchor and a trusted timestamp, with no origin-side counterpart required.
Tier 3 is the only tier that closes the §7.3 gap, because it depends on neither the agent's self-report nor an origin-side event. The tiers are scheme-neutral: any technique that qualifies can reach a given tier, and more than one should. This composes directly with #10: its field-level inference labelling becomes "Tier 1 unless the scheme qualifies higher."
2. Scheme-property model for content_fingerprint
Keep scheme a free string, as in #10. Add declared scheme properties that map to the maximum tier the scheme's signals can claim:
content_borne: the signal lives in the content and survives reprocessing, rather than in detachable metadata.
identity_bearing: it carries rights-holder identity, not only a match.
verifiable: it is cryptographically checkable against a public anchor.
A scheme that is content-borne, identity-bearing, and verifiable is eligible for Tier 3. A matching-only fingerprint (for example a similarity hash without a registry) is Tier 1 or 2. An optional evidence pointer on the object can say where the signal is independently checkable.
3. provenance enum
No change proposed; #10's enum and placement are right. Listed here only for completeness.
Worked example: C2PA / C2PA-text as one Tier-3 scheme
Offered as one worked example, not a dependency, and explicitly inviting other Tier-3 schemes. #10 already lists C2PA text provenance as a candidate scheme; this formalises what that entry means.
C2PA text provenance (the A.7 unstructured-text method) binds the manifest into the text itself, so it is content-borne: it survives copy, plain-text save, encoding and serialization round-trips (UTF-8/UTF-16, JSON), and all Unicode normalization forms (NFC/NFD/NFKC/NFKD). A single contiguous in-band block survives whole-text and aggregated reuse; distributed markers extend survival to the fragment case, so a copied span, even mid-sentence, still resolves to the source work. It is identity-bearing (signed publisher certificate) and verifiable (against the public C2PA trust list), with first-publication anchored by an RFC 3161 timestamp, which also answers ISCC's two honest limits: identity rather than matching, and a first-publication anchor.
The "C2PA is attached metadata, stripped on reprocessing" characterisation in #5 is accurate for a manifest embedded in a media file, but not for in-band text provenance, which is a different mechanism.
Survivability: evidence, not assertion
The PR will include a transform-by-transform survivability matrix (Unicode normalization forms, encoding and serialization, copy-paste through common editors, whitespace and markup sanitizers, fragment and aggregation), backed by test output, so robustness is checkable rather than claimed.
Honest boundaries, stated up front:
Consequence worth encoding: absence is informative
Once a robust, expected mark is in play, its absence on otherwise-matching content is itself a signal, and a grounding system asserting a reliable source it cannot attribute is carrying that risk. That is the demand-side reason a cooperating agent populates these fields honestly: they are its own due-diligence record, not only publisher telemetry.
Out of scope (explicit)
- Adversary coverage and the resale path itself: that is the access layer, not telemetry. This proposal does not claim to close it; it makes the cooperating channel's signals trustworthy and names the boundary.
- Semantic or paraphrase-robust fingerprinting: pre-standard today and not relied upon here.
Contribution offer
I am happy to author a follow-on PR (on top of #10) that:
- adds the normative evidentiary-tiers section;
- adds the scheme-property fields and the property-to-tier mapping to
content_fingerprint;
- adds a scheme-neutral C2PA-text profile as one worked Tier-3 example;
- attaches the survivability matrix.
Co-authorship with @landomo and @jchomat welcome, and additional Tier-3 scheme registrations more so.
Open questions
- Should the tier be an explicit field, or derived from the scheme's declared properties?
- Is a small registry warranted for
scheme strings and their declared properties and tier?
- How should the Tier-2 origin-corroboration path relate to the access layer (for example RSL
<reporting>)?
Summary
#5 surfaced the third-party-sourced (resale) coverage gap, and #10 already implements the two fields agreed there: the
provenanceenum and the scheme-agnosticcontent_fingerprintobject. #10 intentionally leaves two things to a follow-up: the claims-vs-corroborated-facts section its field descriptions are written to slot into, and any structure on the free-stringscheme. This issue proposes both, and offers a PR that builds on #10 rather than replacing it.The intent is not to favour any technique. It is to give the spec a principled, testable way to say which reported signals can be relied on independently of the reporter, so the numbers the standard produces can carry the weight downstream compensation will put on them.
Motivation
Telemetry is a cooperative measurement layer. It reports what willing participants emit, so by construction it cannot represent a non-emitting actor, and it should not try to. That part of the resale gap is an access-layer concern (RSL / CAP / bot-auth), and the right move in the spec is to name the boundary (8.9), as @landomo suggested.
The numbers it does produce are self-reported and become a basis for compensation. #10's descriptions already say
detectedandpreserved_in_outputare inference with no origin-side counterpart (the §7.3 gap), and invite "the broader claims-vs-corroborated-facts section when that lands for v0.1." This is that section, plus the scheme semantics that make the distinction operational.Proposal
1. Normative evidentiary tiers
Define, in a dedicated evidentiary section, three tiers that any reported signal falls into:
Tier 3 is the only tier that closes the §7.3 gap, because it depends on neither the agent's self-report nor an origin-side event. The tiers are scheme-neutral: any technique that qualifies can reach a given tier, and more than one should. This composes directly with #10: its field-level inference labelling becomes "Tier 1 unless the scheme qualifies higher."
2. Scheme-property model for
content_fingerprintKeep
schemea free string, as in #10. Add declared scheme properties that map to the maximum tier the scheme's signals can claim:content_borne: the signal lives in the content and survives reprocessing, rather than in detachable metadata.identity_bearing: it carries rights-holder identity, not only a match.verifiable: it is cryptographically checkable against a public anchor.A scheme that is content-borne, identity-bearing, and verifiable is eligible for Tier 3. A matching-only fingerprint (for example a similarity hash without a registry) is Tier 1 or 2. An optional
evidencepointer on the object can say where the signal is independently checkable.3.
provenanceenumNo change proposed; #10's enum and placement are right. Listed here only for completeness.
Worked example: C2PA / C2PA-text as one Tier-3 scheme
Offered as one worked example, not a dependency, and explicitly inviting other Tier-3 schemes. #10 already lists C2PA text provenance as a candidate
scheme; this formalises what that entry means.C2PA text provenance (the A.7 unstructured-text method) binds the manifest into the text itself, so it is content-borne: it survives copy, plain-text save, encoding and serialization round-trips (UTF-8/UTF-16, JSON), and all Unicode normalization forms (NFC/NFD/NFKC/NFKD). A single contiguous in-band block survives whole-text and aggregated reuse; distributed markers extend survival to the fragment case, so a copied span, even mid-sentence, still resolves to the source work. It is identity-bearing (signed publisher certificate) and verifiable (against the public C2PA trust list), with first-publication anchored by an RFC 3161 timestamp, which also answers ISCC's two honest limits: identity rather than matching, and a first-publication anchor.
The "C2PA is attached metadata, stripped on reprocessing" characterisation in #5 is accurate for a manifest embedded in a media file, but not for in-band text provenance, which is a different mechanism.
Survivability: evidence, not assertion
The PR will include a transform-by-transform survivability matrix (Unicode normalization forms, encoding and serialization, copy-paste through common editors, whitespace and markup sanitizers, fragment and aggregation), backed by test output, so robustness is checkable rather than claimed.
Honest boundaries, stated up front:
Consequence worth encoding: absence is informative
Once a robust, expected mark is in play, its absence on otherwise-matching content is itself a signal, and a grounding system asserting a reliable source it cannot attribute is carrying that risk. That is the demand-side reason a cooperating agent populates these fields honestly: they are its own due-diligence record, not only publisher telemetry.
Out of scope (explicit)
Contribution offer
I am happy to author a follow-on PR (on top of #10) that:
content_fingerprint;Co-authorship with @landomo and @jchomat welcome, and additional Tier-3 scheme registrations more so.
Open questions
schemestrings and their declared properties and tier?<reporting>)?