Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
15 commits
Select commit Hold shift + click to select a range
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 4 additions & 2 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,8 @@ jobs:
- name: Install uv
uses: astral-sh/setup-uv@v5
- name: Run conformance test suite
run: uv run --with jsonschema python tests/validate.py
run: uv run --with "jsonschema[format-nongpl]" python tests/validate.py
- name: Validate worked examples in the spec
run: uv run --with jsonschema python tests/check_examples.py
run: uv run --with "jsonschema[format-nongpl]" python tests/check_examples.py
- name: Replay suite-weakening mutations
run: uv run --with "jsonschema[format-nongpl]" python tests/mutation_smoke.py
18 changes: 10 additions & 8 deletions GOVERNANCE.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

## Status

Content Telemetry is an open specification stewarded by the SPUR Coalition. It is in preview (v0.x); see [SPECIFICATION.md](./SPECIFICATION.md) section 12 for the versioning policy.
Content Telemetry is an open specification stewarded by the SPUR Coalition. Version 1.0 is the current specification; see [SPECIFICATION.md](./SPECIFICATION.md) section 12 for the versioning policy.

## Stewardship

Expand All @@ -18,20 +18,22 @@ The SPUR Coalition stewards the specification and holds this repository. The sta

The SPUR Coalition is a group of publishers and content owners that maintains the Content Telemetry standard. It holds the intellectual property through the preview period and releases the standard under Apache 2.0 from 12 June 2026.

Contributing to the standard does not require membership. The wire format is developed in the open, and anyone - content owner, agent operator, intermediary, or implementer - can take part through the issue tracker and the comment process described in the [README](./README.md#request-for-comment).
Contributing to the standard does not require membership. The wire format is developed in the open, and anyone - content owner, agent operator, intermediary, or implementer - can take part through the issue tracker and the process described in the [README](./README.md#consultation-record).

The standard is maintained by Alex Springer (alex@spurcoalition.org). The decision-making process will be set out before 1.0.
The standard is maintained by Alex Springer (alex@spurcoalition.org).

## Path to 1.0
## Version 1.0 and decisions

The specification is at v0.1 (preview). It moves to 1.0 once the open questions are resolved, the conformance suite is stable, and there are independent interoperable implementations from more than one party. There is no fixed date.
The specification reached 1.0 on 2 September 2026, following the public consultation of 12 June to 24 July 2026 and the release-candidate work recorded on the issue tracker.

Decisions follow the proposal and alignment process in [CONTRIBUTING.md](./CONTRIBUTING.md): anyone may propose a change, a maintainer records the disposition publicly on the issue tracker, and the SPUR Steering Board approves releases. The tracker and pull-request history are the public decision record. Minor versions add optional fields and event types; breaking changes require a major version (SPECIFICATION.md section 12).

## How to participate

- File questions and bugs on the [issue tracker](https://github.com/SPUR-Coalition/telemetry/issues) (see the templates).
- Follow the [post-consultation status](./README.md#consultation-status). The
formal v1 comment window is closed, but concrete bugs and implementation
evidence remain welcome on the issue tracker.
- See the [consultation record](./README.md#consultation-record). The formal
v1 comment window is closed, but concrete bugs and implementation evidence
remain welcome on the issue tracker.
- Propose new capabilities or changes in behaviour as a short human-written note
in [`proposals/`](./proposals/), following
[CONTRIBUTING.md](./CONTRIBUTING.md).
Expand Down
78 changes: 26 additions & 52 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,16 +2,7 @@

**Signal format for AI content usage reporting.**

This is a preview specification. Field names, event types, and schema structure may change before 1.0.

> **Consultation status — 12 August 2026:** The public comment period closed on
> **24 July 2026**. The SPUR Steering Board has approved the v1 direction, and a
> final disposition is now recorded on every consultation thread: each carries
> an outcome label, and accepted core changes sit on the
> [v1 release candidate milestone](https://github.com/SPUR-Coalition/telemetry/milestone/1)
> with a schema freeze targeted for **21 August 2026**. Accepted changes are
> merging on the `v1-draft` integration line. Version 0.1 remains the current
> published preview. See [Consultation status](#consultation-status) below.
**Version 1.0** is the current specification, published 2 September 2026. It replaces the v0.1 preview; the changes and migration steps are recorded in [SPECIFICATION.md section 12.1](./SPECIFICATION.md#121-migration-from-the-v01-preview).

## Contents

Expand All @@ -21,8 +12,8 @@ This is a preview specification. Field names, event types, and schema structure
- [Repo contents](#repo-contents)
- [Example](#example)
- [Relationship to other protocols](#relationship-to-other-protocols)
- [Consultation status](#consultation-status)
- [Open questions in v0.1](#open-questions-in-v01)
- [Consultation record](#consultation-record)
- [Open questions in v1](#open-questions-in-v1)
- [Versioning](#versioning)

## Problem
Expand All @@ -33,12 +24,11 @@ Platforms self-report usage metrics (if they report at all), and content owners

## Telemetry events

Content Telemetry tracks content through six stages:
Content Telemetry tracks content through five stages:

```
Retrieved → content fetched over HTTP (content owner can see this today)
Grounded → content loaded into the agent's generation context
Reproduced → content appearing verbatim or near-verbatim in the response
Cited → content explicitly referenced in the response
Presented → content or a source reference made perceivable on a recipient-facing surface
Engaged → user clicked, copied, shared, or directed the agent to act
Expand All @@ -50,7 +40,6 @@ The gaps between stages show how content was used:

- **Retrieval without grounding** - your content was fetched but not used
- **Grounding without citation** - your content influenced the answer but you got no credit
- **Reproduction without citation** - your content appeared in the answer without credit
- **Citation without engagement** - your content was cited but the user didn't click through

The grounding event captures the boundary "this content entered the agent's generation context." It is architecture-neutral and decoupled from retrieval: content cached by the agent for days still produces a grounding event in every session it influences.
Expand All @@ -61,7 +50,7 @@ Grounding and presentation record different boundary crossings: grounding means

**Post-hoc, not pre-declared.** Events report what actually happened, not what the agent said it would do at request time. An agent cannot reliably declare how it will use content before reading it.

**Observable boundaries, not agent internals.** The six event types mark boundary crossings. What happens between them - the fan-out, relevance evaluation, re-ranking, reasoning chains - is internal to the agent and changes constantly. The spec does not model it.
**Observable boundaries, not agent internals.** The five content event types mark boundary crossings. What happens between them - the fan-out, relevance evaluation, re-ranking, reasoning chains - is internal to the agent and changes constantly. The spec does not model it.

**Multiple observers, one event.** A content retrieval can be reported by the content owner's CDN, the content owner's origin server, and the AI agent independently. The `Content-Telemetry-ID` header correlates these into a single corroborated event. Uncorroborated retrievals (no matching agent event) may indicate an agent that does not yet support the telemetry protocol.

Expand All @@ -74,7 +63,7 @@ Grounding and presentation record different boundary crossings: grounding means
- [telemetry-event-batch.json](./telemetry-event-batch.json) - JSON Schema for event batch envelopes
- [manifest.json](./manifest.json) - JSON Schema for the `.well-known/content-telemetry.json` manifest ([section 8](./SPECIFICATION.md#8-manifest))
- [tests/](./tests/) - conformance test suite
- [GOVERNANCE.md](./GOVERNANCE.md) - stewardship, preview status, relationship to profiles
- [GOVERNANCE.md](./GOVERNANCE.md) - stewardship, versioning status, relationship to profiles
- [LICENSE](./LICENSE) - Apache License 2.0

This repository is the **standard** - the wire format. Publisher-facing accreditation and the SPUR conformance mark are defined separately in the [SPUR Content Telemetry Profile](https://github.com/SPUR-Coalition/telemetry-profile), which references this specification by version. The standard defines the privacy mechanism (section 5.5); whether a profile makes any privacy level binding is the profile's choice. See [GOVERNANCE.md](./GOVERNANCE.md).
Expand All @@ -85,7 +74,7 @@ A user asks an AI agent about UK interest rates. The agent grounds its response

```json
{
"schema_version": "0.1",
"schema_version": "1.0",
"session_id": "660e8400-e29b-41d4-a716-446655440000",
"agent_id": "copilot-v3",
"started_at": "2026-03-28T09:00:00Z",
Expand Down Expand Up @@ -163,57 +152,42 @@ The content owner can derive: FT article `abc123` was in context for the respons

Content Telemetry is focussed on **reporting**, while content **access** protocols (Really Simple Licensing, peek-then-pay, IAB CoMP, bilateral APIs) aim to govern how agents discover and license content. The `license_ref` field on events connects telemetry to whatever access protocol issued the licence, but the schemas are independent - telemetry works with any access protocol, or none.

## Consultation status
## Consultation record

The public comment period ran from **12 June to 24 July 2026** and is now closed.
Thank you to everyone who opened an issue, submitted a pull request, joined a
working session or supplied implementation evidence.
The public comment period ran from **12 June to 24 July 2026**. Thank you to
everyone who opened an issue, submitted a pull request, joined a working
session or supplied implementation evidence.

The consultation produced 29 specification issue threads, three profile issue
threads and five pull requests. The maintainers are now:

- [x] reviewing the full consultation record;
- [x] preparing a proposed disposition for every thread;
- [x] recording the approved dispositions on the issue tracker;
- [ ] completing focused v1 changes and migration fixtures on `v1-draft`;
- [ ] publishing a v1 release candidate for implementer testing; and
- [ ] publishing final v1 only after publisher, intermediary and agent/platform
acceptance cases pass.

Consultation issues remain open while their dispositions are recorded. An open
issue does not mean its proposal has been accepted, and a preparation branch does
not change the published specification. The issue tracker and pull-request
history will remain the public decision record.

Concrete schema, fixture and documentation bugs may still be filed using the
threads and five pull requests. Every thread carries a recorded outcome, and
the issue tracker and pull-request history remain the public decision record.
The accepted changes were merged on the `v1-draft` integration line and
published as version 1.0 on 2 September 2026.

Concrete schema, fixture and documentation bugs may be filed using the
*Schema or example bug* template, and questions or unclear requirements using
*Spec feedback / open question*. For a new capability or change in behaviour,
submit a short human-written note to [`proposals/`](./proposals/) and wait for
explicit maintainer alignment before beginning implementation (see
[CONTRIBUTING.md](./CONTRIBUTING.md)); new design proposals are not
automatically part of the v1 consultation scope. Pull requests remain welcome
for specific fixes. Feedback on accreditation or the conformance mark belongs
on the [profile
[CONTRIBUTING.md](./CONTRIBUTING.md)). Pull requests remain welcome for
specific fixes. Feedback on accreditation or the conformance mark belongs on
the [profile
repository](https://github.com/SPUR-Coalition/telemetry-profile/issues).

Required fields, event types and schema structure may still change before 1.0
(section 12). Version 0.1 remains the current published preview until a later
version is released.

## Open questions in v0.1
## Open questions in v1

This is a preview specification. The following areas are under active discussion and will be refined with implementer input:
The following areas are expected to develop in 1.x minor versions and profiles, with implementer input:

**Grounding boundary.** The spec defines grounding as content entering the generation model's context (sections 4.3 and 6.4). For straightforward RAG pipelines this is clear. For pipelines with multiple processing stages - embedding, re-ranking, summarisation before context insertion - the boundary requires judgement. The spec draws the line at the generation context (not earlier retrieval stages), but edge cases remain. When a re-ranking or summarisation stage is itself a generative model, the multi-step rule in section 6.4 (content entering a sub-agent's generation context is grounded) can pull selection stages back inside the boundary. Input from platform engineering teams building real implementations will sharpen this definition.

**Event volume at scale.** A single deep-research query can produce 100+ retrieval events and dozens of grounding/citation events. The session document format already handles transport - one POST with all events after the session ends, not one request per event. Volume management beyond that (storage, processing, consumer-side aggregation) is an implementation concern, not a protocol gap. Sampling and aggregation are options for future versions but are not in v0.1; the standard sets no default for reporting granularity, leaving it to profiles and deployments.
**Event volume at scale.** A single deep-research query can produce 100+ retrieval events and dozens of grounding/citation events. The session document format already handles transport - one POST with all events after the session ends, not one request per event. Volume management beyond that (storage, processing, consumer-side aggregation) is an implementation concern, not a protocol gap. Version 1 adds an explicit coverage declaration - `complete`, `sampled`, `aggregated` or `selected` (section 5.7.6) - and a manifest field for it (section 8.5); the standard still sets no default for reporting granularity, leaving it to profiles and deployments.

**Verification of grounding and citation.** Grounding and citation events are reported by the agent, which is also the party that may owe compensation under a licence. In v0.1, manifest signing is informational: consumers may verify signatures but are not required to, and the specification defines no required proof binding an event to its emitter (sections 8.4 and 8.9). The events attribution depends on are therefore self-reported by the reporting party. Verifiable credentials and signed events are deferred (section 8.9). One corroboration mechanism works without signing: the `Content-Telemetry-ID` field correlates an agent-reported retrieval with an origin- or edge-reported one (section 7.2), but it covers retrieval only - grounding, reproduction, citation, presentation, and engagement have no independent observer. Signing, even once required, would prove who reported an event, not that the event is true or that all qualifying events were reported. Input is wanted on what a verification layer should cover and where it belongs. Mechanisms that test truthfulness and completeness rather than origin, such as sampled audits or publisher-seeded canary content, are of particular interest.
**Verification of grounding and citation.** Grounding and citation events are reported by the agent, which is also the party that may owe compensation under a licence. In v1, manifest signing is informational: consumers may verify signatures but are not required to, and the specification defines no required proof binding an event to its emitter (sections 8.4 and 8.9). The events attribution depends on are therefore self-reported by the reporting party. Verifiable credentials and signed events are deferred (section 8.9). One corroboration mechanism works without signing: the `Content-Telemetry-ID` field correlates an agent-reported retrieval with an origin- or edge-reported one (section 7.2), but it covers retrieval only - grounding, citation, presentation, and engagement have no independent observer. Signing, even once required, would prove who reported an event, not that the event is true or that all qualifying events were reported. Input is wanted on what a verification layer should cover and where it belongs. Mechanisms that test truthfulness and completeness rather than origin, such as sampled audits or publisher-seeded canary content, are of particular interest.

**Reporting granularity.** The standard sets no default for reporting granularity, leaving it to profiles and deployments (see *Event volume* above). The SPUR profile requires event-level delivery and does not permit aggregation. The open question is whether the standard should say more about sampling and aggregation so that profiles do not each define it separately, and how event-level delivery scales for the highest-volume case. No mechanism is selected in v0.1.
**Reporting granularity.** The standard sets no default for reporting granularity, leaving it to profiles and deployments (see *Event volume* above). The SPUR profile requires event-level delivery and does not permit aggregation. Version 1 answers the first half of the question: coverage modes are defined once, in section 5.7.6, so that profiles reference them rather than each define their own. How event-level delivery scales for the highest-volume case remains open.

## Versioning

This repo tracks the specification version. SDK repos have their own release cadences and declare which spec version they support.

Current spec version: **0.1** (preview)
Current spec version: **1.0**
2 changes: 1 addition & 1 deletion SCOPE.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ Conformance and verification answer five separate questions:
4. **Factual truth and completeness:** did the event happen as claimed, and were all qualifying events reported?
5. **Entitlement:** was the reported use permitted under an applicable grant or agreement?

Events are claims by identified emitters. Evidence applies to a particular assertion. Origin or access evidence can corroborate only what that observer could see; it cannot prove grounding, reproduction, citation, presentation, engagement, truth, completeness or entitlement.
Events are claims by identified emitters. Evidence applies to a particular assertion. Origin or access evidence can corroborate only what that observer could see; it cannot prove grounding, citation, presentation, engagement, truth, completeness or entitlement.

Relationship configuration should avoid profile proliferation. A publisher may require `content_grounded` and `content_cited` events, intent-level topics, event delivery to a named endpoint and a set of aggregate reports. Another may require `content_cited`, `content_presented` and `content_engaged` events with a different privacy level. These are deployment choices backed by governing terms, not publisher-specific protocol profiles. A new profile is justified only when a class of relationships introduces semantics or processing rules that multiple implementations must interpret in the same way.

Expand Down
Loading
Loading