Skip to content

Publish Content Telemetry 1.0 - #49

Merged
jalexspringer merged 54 commits into
mainfrom
v1-draft
Sep 2, 2026
Merged

jalexspringer merged 54 commits into
mainfrom
v1-draft

Conversation

@jalexspringer

Copy link
Copy Markdown
Contributor

Version 1.0 replaces the v0.1 preview. The changes accepted through the 12 June - 24 July 2026 public consultation, the release-candidate hardening, and the migration notes are all in SPECIFICATION.md; section 12.1 is the complete migration record for v0.1 implementers.

Headline changes from 0.1:

Conformance suite: 171 fixtures, 15 validated worked examples, 59 mutation checks.

🤖 Generated with Claude Code

jalexspringer and others added 30 commits July 18, 2026 13:45
Consultation issue #2. A SHA-256 of a client IP is a pseudonym, not an
anonymous value: the IPv4 space is small enough to enumerate against a
candidate digest. The field is removed from the edge and origin data
profiles, the schema and the fixtures.

Section 9.1 previously told emitters to "hash or anonymise identifiers
where possible", which is the claim the consultation objected to. It now
points emitters at asn, asn_org and country, which describe the network
path rather than the client.

Event data accepts additional properties, so the schemas cannot reject a
withdrawn field. The conformance suite gains an application-layer rule and
a negative fixture instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Consultation issue #6. Token counts are measured in the emitter's own
tokeniser, so they are not comparable between agents and shift when a
vendor revises a tokeniser. A content owner receiving grounding events
from several agents could not aggregate the only volume measure the
specification offered.

chars_ingested counts Unicode code points in the content placed in the
generation context. Two emitters counting the same text agree.
tokens_ingested stays, described as supplementary, and section 5.5 now
carries the same caveat for query_tokens and response_tokens.

Grounding conformance asks for chars_ingested rather than tokens_ingested.
Fixtures and worked examples carry both.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Consultation issue #22. Section 5.2.3 said a consumer receiving
license_ref "can verify that content usage was licensed". Core neither
resolves nor interprets the value, so it establishes none of that. The
field is part of the emitter's claim: it records which grant the emitter
says applied, not that the grant existed, covered this content, was valid
at the time, or permitted the use.

Also records what the field does not identify. COUNTER Metrics read
license_ref as a possible carrier for the institution deriving access
rights; it is not one. Where a publisher issues one grant per subscriber
the value can work as a proxy inside that publisher's namespace, but
nothing requires it to be typed, stable, or comparable between emitters.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Consultation issues #13 and #5, both documentation dispositions.

Section 1.3 named model training as a non-goal but said nothing about
fine-tuning, embeddings or index construction, so readers could not tell
which side of the line those fell. New section 1.3.1 draws the line at
construction against use: assembling a corpus, fine-tuning, computing
embeddings and building an index are outside scope, while querying such a
store during a response is an ordinary grounding event. It also states
plainly that nothing here reports whether a model was trained on a work.

Section 4.4 gains the limit raised by the scraper-resale path in #5. Where
the intermediary in the middle does not emit, the agent's events are the
only record and carry only what the agent was told. Consumers should treat
the path back to the content owner as unestablished rather than infer it.

The provenance fields for the same issue are PR #10 and are not touched
here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The output constructor's claim that the output artifact contains
identified source content - a quotation, an excerpt, or a full copy,
verbatim or near-verbatim. Reproduction and citation are sibling
output-construction claims: reproduction records the material,
citation records the credit; neither implies the other.

Covers uncredited reproduction in unpresented output (API delivery),
which no existing event can host. Char-first counts mirror the
excerpt_chars/excerpt_tokens pairing on citation data. Builds on the
v1 presentation semantics change and is prepared pending D1.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A source association with no resolvable reference is not a citation.
The JSON Schema now rejects a content_cited event whose content_url
and content_id are both absent or null - stricter than the
application-layer identifier rule that covers content events
generally, because the reference is what makes the credit routable
to an owner.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A reproduction claim is only meaningful for an identified source. The
schema rejects a content_reproduced event whose content_url and
content_id are both absent or null, matching the enforcement added
for content_cited on v1-citation-source-reference; merge that branch
first so the cross-reference in section 6.6 holds.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Renames content_displayed to content_presented, distinguishes content
from source_reference, and adds output/citation/presentation
correlation with migration notes and multimodal fixtures.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A source association with no resolvable reference is not a citation:
content_cited events must carry a non-null content_url or content_id,
enforced by the JSON Schema.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sixth core event: the output constructor's claim that the output
artifact reproduces identified source content, verbatim or
near-verbatim, independent of credit (cited) and delivery
(presented). Includes the source-reference schema enforcement
matching content_cited.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Isolate the cited source-reference anyOf in both negative fixtures
  (they previously also omitted id/output_id and failed on those)
- Update five-stage/three-model/enumeration text that predated the
  sixth event (1.3.1, 5.2 field table, 5.7.5, 7.3, 10.1, 10.2)
- Schema-enforcement prose now names content_reproduced alongside
  content_cited; section 6 intro no longer claims no data field is
  required; terms table gains presentation
- Reorder session-multi-turn events chronologically; correct stale
  fixture descriptions and README event names; batch grounding
  fixture now meets the conformance level it claims
- Add invalid fixtures isolating cited output_id, reproduced id and
  presented presentation_type; add envelope-level parent_session_id
  coverage; pair chars_ingested with tokens_ingested in new fixtures
- tests/README documents the ip_hash application-layer check and no
  longer claims batch-envelope engagement coverage

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Complete v1 grounding provenance and fingerprint semantics
Output-side reuse is reported with content_reproduced on the landed
v1-draft; grounding-time detection stays as the emitter claim.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Add `provenance` and `content_fingerprint` to `content_grounded` (v0.1)
… resolution (#23, #28, #30)

New section 7.4 replaces the v0.1 click manifest with the click context:

- Token pattern ct_[A-Za-z0-9_-]{8,240}, opaque, bound to exactly one
  presentation at mint time; per-click minting for routed surfaces.
- Reserved query parameters ctx_token and ctx_iss with SHOULD-level
  redirect propagation mirroring section 7.2.
- Resolver discovery through the issuer manifest: telemetry.ctx_resolution.
- Resolution returns the engagement, the clicked content's lineage cut by
  content identity across turns, and an optional count-based session
  summary; never the raw session_id, never other owners' events.
- presentation_id becomes conditional: required for agent-reported
  engagements, restored from issuer state for destination reports that
  carry ctx_token; the URL-carried presentation UUID is withdrawn.
- Consumer-custody trade-off recorded (#29).

Schemas, six new/updated fixtures, migration note. 74/74 conformance
checks, 9/9 worked examples.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…livery (#28, PR #38 review)

Two changes from Pedro Santos's review of #38, both from #28.

Contributing sources (7.4.4). Lineage by content identity alone cannot
support the commerce case: a click through to a destination contributed
by another owner's content resolves nothing about the contributor. The
click context regains the v0.1 click manifest's role as a consent-gated
contributing-source set, narrowed to the engaged presentation's turn and
gated per contributing owner's recorded opt-in. Owners without an opt-in
remain visible only as counts. The migration note in 12.1 now records a
scoping plus the retained gate, not a removal.

Owner-scoped delivery (new 7.4.5). The agent-authored click
content_engaged is delivered to the engaged content's owner through 7.3
filtered views on the same terms as grounding and citation events,
session_id included. The 7.4.4 session_id prohibition binds token
resolution (possession of a URL-carried value), not 7.3 delivery to a
verified domain owner. This survives token loss in redirect chains,
removes the resolver dependency for the destination's own join, and
lets a party holding both a contributor's and a destination's
owner-scoped streams match on session_id for attribution. Not token
distribution: only the clicked content's owner receives the event.

Recorded limit renumbered to 7.4.6. Conformance suite 74/74, worked
examples 9/9.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
jalexspringer and others added 24 commits August 17, 2026 08:57
Click context: ctx_token issuance, carriage, discovery and resolution
…s_context, content_depth (#3, #42, #43, #44)

Terms reference (new 5.2.4, #3). Optional event-level terms_ref, distinct
from license_ref: the terms are the basis, the licence is the proof. A
public URL or an opaque identifier both parties can resolve; private
resolution conforms. Core carries the reference and does not resolve,
validate or interpret it, and the terms it points at do not redefine core
event semantics. Event-level only: a session-level default was considered
and deferred because sessions legitimately span content under different
terms.

Occurrence and coverage (#42). Section 4.3 now states one occurrence
boundary per core event: retrieval per completed fetch, grounding per
distinct content item per declared scope, reproduction per source item
per output element, presentation per rendering, engagement per observed
action. "Qualifying occurrence" is defined, a conformance level is stated
to be a capability claim rather than a coverage claim, and coverage
becomes an explicit declaration (new 5.7.6): complete, sampled,
aggregated or selected, over a stated relationship scope, with the rule
for the last three objectively decidable at emission time and disclosed.
Manifests MAY declare the modes machine-readably (8.5). Governing terms
select events and coverage; they do not redefine semantics (SCOPE.md,
the boundary recorded in #4).

Envelope identity. manifest_ref joins the standalone event and event
batch envelopes (7.1), the only envelope field naming the manifest - and
so the domain - under which the emitter claims to report. SHOULD-level
for standalone delivery under settlement or audit obligations.

Session access context (new 5.1.3, #43). Sessions gain a data container
mirroring the event-level field, closing the accidental extension point
at the session root (11.1, 12.1). One core container: access_context,
typed institutional identifiers (ror, saml_entity_id, isni, open
vocabulary) - the institution whose access rights the session used,
never an individual. Populated only where governing terms require it,
placed inside the privacy model in 5.5 with intent/minimal pairing
guidance.

Content depth (6.1, #44). content_depth records how much of a content
record a retrieval reached: metadata, abstract or full, open values,
applicable from any source_role. Depth is what was reachable at
retrieval, independent of what later entered a generation context.

Four new fixtures. Conformance suite 78/78, worked examples 10/10.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…eans nothing (#3)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Terms references, coverage declarations, envelope identity, session access context and content depth
The spec body already makes v1 normative claims (ip_hash withdrawal in
9.1, the content_displayed rename in 12.1) while the header still said
0.1/Preview and every schema pinned schema_version to "0.1" under a
/schema/v0.1/ $id, so a conforming v1 emitter had to declare "0.1" and
the two versions were indistinguishable on the wire.

- Header: Version 1.0 (release candidate draft), status release
  candidate in preparation, feature freeze 21 August 2026.
- All four schemas: $id path v0.1 -> v1, schema_version const -> "1.0";
  manifest const description now says v1 emitters MUST use "1.0".
- 5.7.4 and 12.1 now state explicitly that v1 documents declare "1.0",
  that a v0.1 consumer rejects them under the preview rule, and that a
  v1 consumer rejects "0.1".
- Every fixture and every inline example in SPECIFICATION.md and
  README.md now declares "1.0". invalid/invalid-schema-version.json
  keeps its deliberately wrong "2.0" (still != the new const).
- validate.py / tests/README.md docstrings updated; no hardcoded
  versions or v0.1 paths remain in the test scripts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The table said id was optional and generated by the server if not
provided. The schema requires id on content_cited, content_reproduced
and content_presented, and server-side generation is incompatible with
citation_id and presentation_id, which reference an id the emitter must
already hold when it constructs the referencing event. Reworded to the
table's conditional style (as used by output_id and presentation_id):
required for reproduced/cited/presented, optional elsewhere,
emitter-assigned.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Section 5.7.3 lists data.citation_type among the Citation-emitter
requirements marked schema-enforced, but the content_cited conditional
required only id and output_id - unlike content_reproduced and
content_presented, whose conditionals require their data objects. The
content_cited conditional now requires data with citation_type,
mirroring the reproduced (data.reproduction_type) and presented
(data.presentation_kind/presentation_type) pattern.

Two cited events carried no data and gained citation_type: "reference"
(each is paired with a link presentation of the same source): the
7.1 event-batch example and tests/valid/event-batch-agent.json, plus
the same event in tests/invalid/batch-missing-session-and-ctx-token.json
so that fixture still passes the schema and fails at the application
layer as its description intends. All other cited fixtures and examples
already carried citation_type.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- 8.3: revert 'Presentation name' to 'Display name' for operator.name.
  The display->presentation rename applies to the presentation event,
  not ordinary UI terminology; manifest.json already says Display name.
- telemetry-session.json: ad_rendered description now says 'rendered',
  matching the field name and 5.4.
- 6.5: restate the excerpt pair on the v1 chars-primary hierarchy -
  excerpt_chars is the portable primary measurement under 6.4's
  counting rule, excerpt_tokens the agent-native supplementary one -
  mirroring 6.4's chars_ingested/tokens_ingested wording so 6.6's
  'same pairing' cross-reference holds.
- 6 intro: the profiles are no longer in lifecycle order (6.5 Citation
  precedes 6.6 Reproduction while the lifecycle runs Reproduced then
  Cited), so the intro no longer claims they are; sections keep their
  numbers.
- 5.7: the optional-signals sentence now names reproduction alongside
  presentation and engagement - it likewise sits outside the
  Retrieval/Grounding/Citation ladder as a SHOULD.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…shapes

Mutation-verified review findings, all previously undetected:

- Build every validator with format_checker so format: uuid / date-time /
  uri assertions enforce instead of annotate; both runners hard-error at
  startup if the checker lacks uuid or date-time. Install line becomes
  pip install "jsonschema[format-nongpl]".
- Require _expected_error on every invalid fixture: a substring that must
  appear in the actual error (first schema error message and JSON pointer,
  or the application-layer violation text). A fixture that fails for the
  wrong reason, or carries no pin, now fails the run.
- Reconcile APPLICATION_LAYER_VIOLATIONS keys against invalid/: an entry
  with no matching file fails the run instead of silently dropping the
  expectation.
- Apply privacy field gating (5.5) to turns wherever they appear: session
  documents, event batches, and standalone event envelopes, not only
  session events lists.
- Add referential integrity checks within a session document (6.6-6.8):
  content_engaged.presentation_id must match a content_presented event id,
  and citation_id on content_presented/content_reproduced must match a
  content_cited event id. Standalone envelopes and batch members are
  exempt; the corroborating click-out flow is out of scope here.
- check_examples.py: validate complete bare event objects (type +
  timestamp) against the TelemetryEvent definition instead of skipping
  them; 3 of the 7 skipped fragments are now validated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
New invalid fixtures, each pinned by _expected_error:

- presented-missing-id, presented-missing-output-id, cited-missing-id:
  the required event and output identifiers of sections 6.5 and 6.7.
- cited-missing-citation-type: the schema now requires
  data.citation_type on content_cited.
- reproduction-type-invalid: reproduction_type 'paraphrase' - section
  6.6 says paraphrase is not reproduction.
- presentation-kind-invalid: presentation_kind is a closed two-value
  enum.
- grounded-negative-chars-ingested, reproduced-negative-chars: count
  fields carry minimum 0.
- malformed-parent-session-id: format: uuid, caught only now that the
  runners enforce format assertions.
- privacy-violation-{response-text,query-intent,topics,response-type,
  response-mode,model-id}-at-minimal: one fixture per remaining field
  forbidden at minimal privacy (section 5.5).
- presented/retrieved/engaged-missing-identifier: the section 5.7.5
  identifier rule, previously only tested via content_grounded.
- engaged-presentation-id-unmatched, presented-citation-id-unmatched:
  the new intra-session referential integrity checks (sections 6.7,
  6.8).

New valid fixture:

- session-reproduction-no-grounding: the fifth funnel departure of
  section 4.3, reproduction of memorised content with no grounding
  event.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
mutation_smoke.py copies the schemas and tests/ into a temp directory,
applies each known suite-weakening mutation - dropping the format
checker, gutting the withdrawn-ip-hash fixture, shrinking
CONTENT_EVENT_TYPES and PRIVACY_FORBIDDEN_FIELDS, pointing an engagement
at an all-zeros presentation_id - and confirms validate.py fails under
every one. Each of these previously went undetected. The working tree
is never modified; a surviving mutation fails the script.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…s branch

The rebase onto v1-draft brought in fixtures and an inline example added by
PRs #38, #41 and #45, none of which were swept by the v1 identity commit or
carry the pinned violation this branch now requires.

- CI installed plain `jsonschema`, so the hardened runners' format-checker
  guard aborted the job. The workflow now matches tests/README.md and
  installs `jsonschema[format-nongpl]`.
- 17 fixtures and the §5.1.3 access_context example still declared
  schema_version 0.1. Bumped to 1.0; invalid-schema-version.json keeps its
  deliberately wrong value.
- 8 invalid fixtures had no `_expected_error`. Each now pins its own
  violation, including the two application-layer provenance rules and the
  ctx_token pattern.

Suite: 107/107 fixtures, 13/13 examples, 5/5 mutations caught.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ick-context gate, migration, suite hardening

Prose (SPECIFICATION.md, README, tests/README)
- schema_version is major.minor; 5.7.4, 8.7 and 12 state the one-schema-per-minor
  rule; the 1.0.0-style examples are gone and the const "1.0" stays
- source_role is a MUST on content_retrieved everywhere (5.2.2 aligned with 5.7.1)
- data.scope and citation_type stated as required in 6, 6.4, 6.5; 5.7.1 says
  "every content event"
- 5.7.5 gains the field-placement, distinct-id, same-content join, one-token-one-
  presentation, envelope-token and manifest placement rules
- 7.4.1 token minimum 16 characters and an unguessability MUST; 7.4.4 resolution
  gate defined on the click turn's privacy_level (minimal: engagement and lineage
  only; no turn fields at any level); 7.4.5 destination-owner caveat; 7.4.6
  records token lifetime and requester authentication as v1 limits
- 9.1 states that privacy_level gates the named turn fields only and that opaque
  and extension fields must not carry identity or withheld text
- 12.1: the "URL-carried presentation_id" and published-v0.1 preserved_in_output
  paragraphs corrected; migration notes for citation_type, scope, the non-null
  reference rule, ip_hash, source_role and license_ref semantics added
- Section 8 written in v1 tense; manifest id and endpoint are https-only and id
  sits at the well-known path; tokens_ingested at minimal; share and
  user-supplied agent_navigate engagements; Annex B.5 multi-owner catalogue (#32)

Schemas
- content_grounded requires data.scope; content_depth typed on retrievals;
  ct_ pattern {16,240}; manifest_ref format uri; envelope agent_id nullable like
  the session field; manifest id/endpoint https patterns

Conformance suite
- validate.py: source_role, field placement, duplicate ids, same-content joins,
  shared tokens, envelope ctx_token, root-manifest domains, agent/platform
  ctx_resolution; check_examples.py runs the application-layer rules too
- 66 new invalid and 7 new valid fixtures: every closed enum, every format
  assertion, every cross-shape rule in all three document shapes, origin and
  index retrievals, platform manifest, all engagement/citation/presentation
  types, a destination engagement batch, the v0.1 wire version rejected
- every schema pin is pointer-prefixed; session-minimal carries source_role;
  the one fixture without a description has one
- mutation_smoke.py replays 59 mutations (54 new) and runs in CI

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The v1 lifecycle keeps the content_displayed -> content_presented rename
and drops the separate reproduction event. Output-side reuse reporting
(reproduction_type, reproduced_chars/tokens/hash, the reproduced funnel
departures and per-reproduction counting) leaves core; a future version
or profile can reintroduce it on implementation evidence. The migration
note records that the event existed only on the pre-release draft line.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Discovery/outcome-layer clarification in 1.4 and the consumer vs
  index-emitter role separation in 7.3 (#12)
- Optional manifest identifier_schemes block and co-primary URL /
  content_id owner resolution with defined conflict and failure
  behaviour (#17, #31), with fixtures and placement checks
- The manifest identifies a domain, not a legal person; party-identity
  declarations recorded as deferred (#46)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…aveat

- Version 1.0 / current-specification framing in the spec header,
  README and GOVERNANCE; the consultation section becomes a past-tense
  record and the checkbox list goes away
- GOVERNANCE gains the decision process owed before 1.0 and loses the
  broken request-for-comment anchor
- Cross-language citation guidance in 6.5 (#24 fallback) and the
  bounded-portion caveat restated in 6.6 (#25)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- bot_category becomes purpose: an open enum classifying the access
  (training, inference, search, advertising), with the vendor signal
  mappings moved to informative Annex C and a migration note (#7)
- Generic evidence reference slot on content events (6.8): scheme,
  detached ref plus digest, no status escalation; the status vocabulary
  and trust policy stay with the evidence profile (#8)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
B.6 shows a delegated child session with parent_session_id, closing the
example promised on #1; 5.3 states the namespaced-name rule that 5.7.6
already relied on (#9).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
v1 identity, spec corrections, and conformance-suite hardening
… link forward

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@jalexspringer
jalexspringer merged commit 7106115 into main Sep 2, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants