From 35d7f42432aee47d045fe5540d9904846cda5429 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Sat, 18 Jul 2026 13:45:54 +0100 Subject: [PATCH 01/38] Define v1 presentation semantics --- README.md | 22 ++- SPECIFICATION.md | 131 ++++++++++++------ telemetry-session.json | 42 +++++- tests/README.md | 5 +- .../batch-missing-session-and-ctx-token.json | 2 + .../engaged-missing-presentation-id.json | 14 ++ tests/invalid/legacy-content-displayed.json | 14 ++ tests/invalid/presented-missing-kind.json | 16 +++ tests/valid/event-batch-agent.json | 15 ++ .../event-standalone-engaged-ctx-token.json | 1 + tests/valid/event-standalone-presented.json | 20 +++ tests/valid/session-citation-tier.json | 13 +- ...on-funnel-exception-cited-no-grounded.json | 12 +- ...-funnel-exception-presented-no-cited.json} | 10 +- ...nnel-exception-presented-no-grounded.json} | 17 ++- tests/valid/session-multi-turn.json | 12 +- .../session-presentation-multimodal.json | 96 +++++++++++++ tests/validate.py | 2 +- 18 files changed, 374 insertions(+), 70 deletions(-) create mode 100644 tests/invalid/engaged-missing-presentation-id.json create mode 100644 tests/invalid/legacy-content-displayed.json create mode 100644 tests/invalid/presented-missing-kind.json create mode 100644 tests/valid/event-standalone-presented.json rename tests/valid/{session-funnel-exception-displayed-no-cited.json => session-funnel-exception-presented-no-cited.json} (77%) rename tests/valid/{session-funnel-exception-displayed-no-grounded.json => session-funnel-exception-presented-no-grounded.json} (65%) create mode 100644 tests/valid/session-presentation-multimodal.json diff --git a/README.md b/README.md index 451dcd2..d84e06b 100644 --- a/README.md +++ b/README.md @@ -18,7 +18,7 @@ This is a preview specification. Field names, event types, and schema structure ## Problem -AI agents retrieve a content owner's content, use it to generate responses, and sometimes cite it. Content owners currently see an initial retrieval event - HTTP requests hitting their servers or access logs from content repositories. Whether the content actually influenced the response, whether it was cited, whether a user saw the citation, whether they clicked through - is not reported back to content owners. +AI agents retrieve a content owner's content, use it to generate responses, and sometimes cite it. Content owners currently see an initial retrieval event - HTTP requests hitting their servers or access logs from content repositories. Whether the content actually entered a generation context, whether it was cited, whether content or a source reference was made perceivable, and whether anyone interacted with that presentation is not reported back to content owners. Platforms self-report usage metrics (if they report at all), and content owners have no way to verify the numbers or compare across platforms. @@ -30,7 +30,7 @@ Content Telemetry tracks content through five stages: Retrieved → content fetched over HTTP (content owner can see this today) Grounded → content loaded into the agent's generation context Cited → content explicitly referenced in the response - Displayed → user saw it - a reference, or the content embedded in the answer + Presented → content or a source reference made perceivable on a recipient-facing surface Engaged → user clicked, copied, shared, or directed the agent to act ``` @@ -44,7 +44,7 @@ The gaps between stages show how content was used: The grounding event captures the boundary "this content entered the agent's generation context." It is architecture-neutral and decoupled from retrieval: content cached by the agent for days still produces a grounding event in every session it influences. -Grounding and display record two different kinds of influence: grounding means the content influenced the agent, display means it reached the user. The two diverge as agent experiences move beyond the chat window - an agentic browser can render a page to the user that never entered a generation context, reported as a `content_displayed` event with `display_type: embed` and no grounding event. +Grounding and presentation record different boundary crossings: grounding means the content entered a generation context; presentation means content or a source reference was made perceivable on a recipient-facing surface. Presentation does not prove attention. The two diverge as agent experiences move beyond the chat window - an agentic browser can render a page that never entered a generation context, reported as a `content_presented` event with `presentation_kind: content` and `presentation_type: embed` and no grounding event. ## Design principles @@ -101,9 +101,12 @@ A user asks an AI agent about UK interest rates. The agent grounds its response } }, { + "id": "550e8400-e29b-41d4-a716-446655440001", "type": "content_cited", "timestamp": "2026-03-28T09:00:05Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:paragraph:2", "content_url": "https://www.ft.com/content/abc123", "content_id": "ft:abc123", "data": { @@ -112,12 +115,19 @@ A user asks an AI agent about UK interest rates. The agent grounds its response } }, { - "type": "content_displayed", + "id": "550e8400-e29b-41d4-a716-446655440002", + "type": "content_presented", "timestamp": "2026-03-28T09:00:05Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:paragraph:2", + "citation_id": "550e8400-e29b-41d4-a716-446655440001", "content_url": "https://www.ft.com/content/abc123", "content_id": "ft:abc123", - "data": { "display_type": "link" } + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } }, { "type": "turn_completed", @@ -134,7 +144,7 @@ A user asks an AI agent about UK interest rates. The agent grounds its response } ``` -The content owner can derive: FT article `abc123` was in context for the response, cited as a paraphrase, link was displayed, user never clicked, ads were shown alongside. +The content owner can derive: FT article `abc123` was in context for the response, cited as a paraphrase, its link was made perceivable, and no engagement was reported; ads were shown alongside. ## Relationship to other protocols diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 746e4a4..265be15 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -11,7 +11,7 @@ 3. [Terms and definitions](#3-terms-and-definitions) 4. [Concepts](#4-concepts) - roles, sessions, event lifecycle, source roles, content identification 5. [Schema](#5-schema) - session, event, event types, conversation turn, privacy, intent, conformance levels -6. [Data profiles](#6-data-profiles) - retrieval, edge enrichment, origin enrichment, grounding, citation, display, engagement +6. [Data profiles](#6-data-profiles) - retrieval, edge enrichment, origin enrichment, grounding, citation, presentation, engagement 7. [Transport](#7-transport) - delivery formats, Content-Telemetry-ID header 8. [Manifest](#8-manifest) - discovery, schema, operator, keys, telemetry, domains 9. [Privacy](#9-privacy) - data minimisation, recommended levels, retention @@ -169,7 +169,7 @@ Session │ ├── content_retrieved (HTTP layer) │ ├── content_grounded (influence layer) │ ├── content_cited (response layer) -│ ├── content_displayed (UI layer) +│ ├── content_presented (recipient-facing surface) │ ├── turn_completed │ ├── content_engaged (user action layer) │ └── ... @@ -190,19 +190,19 @@ Content moves through five stages during an agent interaction: Grounding is architecture-neutral: same event whether the agent uses RAG, chain-of-thought reasoning, embeddings, or multi-step delegation (see section 6.4 for architecture-specific guidance). Grounding is decoupled from retrieval: content may be grounded from a live fetch, from agent-side cache, or from a pre-loaded index. Only the agent can report grounding events. -3. **Cited** - Content explicitly referenced in the agent's response: quoted, paraphrased, or linked. A subset of grounded content. Content can influence every response in a session without being cited once. +3. **Cited** - An output artifact explicitly associates identified source content with a response, claim, passage, quotation, or other output element. Citation is an output-construction relationship, not evidence that the output was delivered. A subset of grounded content is commonly cited, but a citation can also be emitted without a matching grounding event when an agent produces an uncorroborated or hallucinated source association. -4. **Displayed** - Content presented to the end user: a reference (a link, snippet, inline quote, or preview card) or the content itself embedded in the response surface (an iframe, a page rendered by an agentic browser, an embedded media player). Not all citations result in display (e.g., when the agent uses content internally without surfacing the source). +4. **Presented** - Content or a source reference was rendered, played, spoken, embedded, or otherwise made perceivable on a recipient-facing surface. Presentation does not assert that a person noticed or attended to it. `presentation_kind` distinguishes source content (including a reproduced excerpt or media) from a source reference (such as a link, credit, or card). Not all citations are presented: an output can be stored, suppressed, or passed to another system before delivery. - Grounding and display record two different kinds of influence: grounding records that content influenced the agent, display records that it reached the user. As agent experiences evolve beyond the chat window the two diverge - content can shape an answer whose source the user never sees, and an agent can render content to the user that never entered a generation context (see *Departures from the funnel model* below). + Grounding and presentation record different boundary crossings: grounding records entry into a generation context, while presentation records a recipient-facing delivery occurrence. As agent experiences evolve beyond the chat window the two diverge - content can shape an answer whose source is never presented, and an agent can present content that never entered a generation context (see *Departures from the funnel model* below). -5. **Engaged** - The user acted on displayed or cited content: clicked a link, expanded a preview, copied text, shared the response, or directed the agent to act on the content on their behalf (opening the linked page, retrieving more from the source). Engagement connects the preceding events to down-funnel activity: a click-out carries a `ctx_token` that a destination can resolve to the session's click manifest - the content that influenced the response (section 7.1). +5. **Engaged** - The recipient or agent performed an observable action on a presentation: clicked a link, expanded a preview, copied text, shared the response, or directed the agent to act on the content. It does not imply attention beyond the reported action. `presentation_id` links the action to the exact presentation occurrence; a click-out can also carry a `ctx_token` that a destination resolves to the session's click manifest (section 7.1). ``` Retrieved (HTTP layer, cacheable) → Grounded (influence layer, per-session or per-turn) → Cited (response layer, per-turn) - → Displayed (UI layer, per-turn) + → Presented (recipient-facing surface, per-turn) → Engaged (user action layer) ``` @@ -210,16 +210,16 @@ Each stage is typically a progressively narrower subset. The ratios between stag - **Retrieval-to-grounding** measures content fetched but not used (irrelevant, stale, or a competing source was preferred) - **Grounding-to-citation** measures content that influenced the response without explicit attribution -- **Citation-to-display** measures content attributed internally but not shown to the user -- **Display-to-engagement** measures interactions where the user did or did not visit the source +- **Citation-to-presentation** measures source associations constructed in output but not made perceivable +- **Presentation-to-engagement** measures observable actions on exact presentation occurrences #### Departures from the funnel model Three cases break the strict subset model: -- **Displayed without cited.** An agent may display content references (e.g., a "Sources" sidebar) without citing the content in the response text. In this case, a `content_displayed` event exists with no corresponding `content_cited` event. +- **Presented without cited.** An agent may present content references (e.g., a "Sources" sidebar) without semantically associating them with a response element. In this case, a `content_presented` event exists with no corresponding `content_cited` event. - **Cited without grounded.** A hallucinated citation references content the agent never retrieved or loaded into context. The `content_cited` event has no preceding `content_grounded` event. Telemetry consumers SHOULD treat uncorroborated citations (no matching grounding event) as lower-confidence signals. -- **Displayed without grounded.** An agent can render content to the user without that content entering a generation context: an agentic browser showing a page, an embedded video played in the response surface. A `content_displayed` event (typically `display_type: embed`) exists with no corresponding `content_grounded` event. The content influenced the user directly rather than through the model, and the engagement that follows it is consumption the content owner cannot otherwise observe. +- **Presented without grounded.** An agent can present content without that content entering a generation context: an agentic browser showing a page or an embedded video played on a response surface. A `content_presented` event (typically `presentation_kind: content` and `presentation_type: embed`) exists with no corresponding `content_grounded` event. These cases are valid. Emitters SHOULD produce the events that reflect what actually happened, even when the result does not follow the typical funnel ordering. @@ -230,7 +230,7 @@ Conversation turns overlay this lifecycle: 1. **Turn started** - user submits a query 2. **Turn completed** - agent finishes response -A single grounding event with session scope influences all subsequent turns. Citation, display, and engagement events occur within specific turns. +A single grounding event with session scope influences all subsequent turns. Citation, presentation, and engagement events occur within specific turns. ### 4.4 Source roles @@ -247,7 +247,7 @@ The `origin` and `edge` source roles enable content owners to report AI agent tr A marketplace operating as both emitter and telemetry consumer receives telemetry from platforms (as a consumer), resolves content owner identity from `content_id` or `content_url`, and generates per-content-owner usage reports. The marketplace's own `source_role: index` events provide a corroboration layer - it can cross-reference what it served against what platforms reported using. -`content_grounded`, `content_cited`, and `content_displayed` events are reported by the agent (or agent operator) only. These events describe what happened inside the agent or in the agent's user interface, which is not observable from the content owner's infrastructure. +`content_grounded`, `content_cited`, and `content_presented` events are reported by the agent (or agent operator) only. These events describe what happened inside the agent, during output construction, or on a recipient-facing surface, which is not observable from the content owner's infrastructure. `content_engaged` events are usually reported by the agent for in-product interactions. For a click-out to a landing page, a downstream marketplace, affiliate network, or destination site MAY report a corroborating `content_engaged` event using `ctx_token` in place of `session_id` (section 7.1). @@ -324,6 +324,10 @@ Format: the URL of a manifest served at `/.well-known/content-telemetry.json` un | `type` | EventType | Yes | Event type (see 5.3) | | `timestamp` | datetime | Yes | Event timestamp (UTC) | | `turn_id` | string | No | Associates this event with a conversation turn (see 5.2.1) | +| `output_id` | string | For cited/presented | Opaque output-artifact identifier joining construction to later delivery | +| `output_element_id` | string | No | Opaque element within `output_id`, such as a passage, media track, caption, link, or card | +| `citation_id` | UUID | No | On `content_presented`, the `id` of the associated citation event; absent for uncited presentations | +| `presentation_id` | UUID | For engaged | On `content_engaged`, the `id` of the exact presentation occurrence acted upon | | `source_role` | SourceRole | No | Who is reporting: `origin`, `edge`, `index`, `agent` (see 4.4) | | `content_telemetry_id` | UUID | No | Correlation ID for cross-observer deduplication (see 7.2) | | `content_url` | string | No | Content URL as fetched or canonical URL | @@ -334,7 +338,7 @@ Format: the URL of a manifest served at `/.well-known/content-telemetry.json` un #### 5.2.1 Turn association -The `turn_id` field associates content events with a specific conversation turn. Emitters SHOULD set `turn_id` on `content_cited`, `content_displayed`, and `content_engaged` events. Emitters SHOULD also set `turn_id` on `content_grounded` events when `scope` is `turn`. The corresponding `turn_started` and `turn_completed` events SHOULD carry the same `turn_id`. +The `turn_id` field associates content events with a specific conversation turn. Emitters SHOULD set `turn_id` on `content_cited`, `content_presented`, and `content_engaged` events. Emitters SHOULD also set `turn_id` on `content_grounded` events when `scope` is `turn`. The corresponding `turn_started` and `turn_completed` events SHOULD carry the same `turn_id`. `turn_id` is scoped to the session. Format is emitter-defined (sequential integers, UUIDs, or any opaque string). @@ -356,9 +360,9 @@ The `license_ref` field connects a telemetry event to the content access licence |------|-------------|-----------------| | `content_retrieved` | Content fetched from source | `content_url`, `source_role`, `data.media_type` | | `content_grounded` | Content loaded into agent context | `content_url` or `content_id`, `data.scope`, `data.cached` | -| `content_cited` | Content referenced in response | `content_url`, `data.citation_type`, `data.position` | -| `content_displayed` | Content or a reference to it shown to user | `content_url`, `data.display_type` | -| `content_engaged` | User acted on content | `content_url`, `data.engagement_type` (see 6.7) | +| `content_cited` | Output explicitly associates source content with an output element | `id`, `output_id`, `content_url` or `content_id`, `data.citation_type` | +| `content_presented` | Content or a source reference was made perceivable | `id`, `output_id`, `content_url` or `content_id`, `data.presentation_kind`, `data.presentation_type` | +| `content_engaged` | Observable action on an exact presentation | `presentation_id`, `content_url` or `content_id`, `data.engagement_type` (see 6.7) | #### Conversation events @@ -389,7 +393,7 @@ A conversation turn represents one query-response exchange. Turn data is carried | `query_tokens` | integer | No | Query token count | | `response_tokens` | integer | No | Response token count | | `model_id` | string | No | Model identifier | -| `ad_rendered` | boolean | No | Whether advertising was displayed alongside the response | +| `ad_rendered` | boolean | No | Whether advertising was rendered alongside the response | #### 5.4.1 Response modes @@ -449,7 +453,7 @@ Each level is named for the event it adds: a level proves the emitter produces t | **Grounding** | Above + `content_grounded`, turn events | Content entered the agent's context | Agent with basic instrumentation | | **Citation** | Above + `content_cited` | Content was explicitly referenced in the agent's response | Agent with citation instrumentation | -Display and engagement events are optional lifecycle signals. A Citation emitter SHOULD emit them when applicable (section 5.7.3), but the Citation level proves citation support, not full retrieval-to-engagement coverage. +Presentation and engagement events are optional lifecycle signals. A Citation emitter SHOULD emit them when applicable (section 5.7.3), but the Citation level proves citation support, not full retrieval-to-engagement coverage. #### 5.7.1 Retrieval conformance @@ -481,15 +485,16 @@ Emitters using standalone event delivery (section 7.1) MUST include `agent_id`, A conforming **Citation** emitter MUST satisfy Grounding requirements and also: -- Emit `content_cited` events with `data.citation_type` +- Emit `content_cited` events with `id`, `output_id`, and `data.citation_type` The privacy-level field restriction (section 5.5) applies to Citation emitters as it does to any emitter producing conversation turns; it is inherited through the Grounding requirements above. A Citation emitter SHOULD: -- Emit `content_displayed` and `content_engaged` events when applicable +- Emit `content_presented` and `content_engaged` events when applicable - Include `data.position` on citation events -- Include `data.display_type` on display events +- Include `output_element_id` when the cited or presented element has a stable identity +- Include `citation_id` on a presentation of a cited source association #### 5.7.4 Telemetry consumers @@ -653,14 +658,17 @@ The `unclassified` value for `citation_type` indicates the agent did not classif When `content_hash` is absent or does not match any grounding event's hash (for example, because the agent re-chunked content between grounding and citation), consumers SHOULD fall back to matching on `content_url` or `content_id`, accepting that the correlation may be imprecise when the same content appears in multiple grounding events. -### 6.6 Display data (`content_displayed`) +### 6.6 Presentation data (`content_presented`) | Field | Type | Description | |-------|------|-------------| -| `display_type` | string | How the content or a reference to it was presented (see below) | -| `media_type` | string | Content medium: `text`, `image`, `video`, `audio` (open vocabulary, see 6.1). Defaults to `text` when absent. Most useful on `embed` displays (an embedded video reports `media_type: video`). | +| `presentation_kind` | string | What was made perceivable: `content` or `source_reference` | +| `presentation_type` | string | How it was made perceivable (see below) | +| `media_type` | string | Medium made perceivable: `text`, `image`, `video`, `audio` (open vocabulary, see 6.1). Defaults to `text` when absent. | -#### Display types +`presentation_kind: content` means source content itself, a bounded excerpt, or a derived representation was made perceivable. It does not claim that the whole source was reproduced. `presentation_kind: source_reference` means a credit, identifier, link, card, or other reference to the source was made perceivable. This distinction is independent of modality: a spoken credit is a source reference; played source audio is content. + +#### Presentation types | Value | Description | |-------|-------------| @@ -670,12 +678,15 @@ When `content_hash` is absent or does not match any grounding event's hash (for | `card` | Rich preview card (title, description, image) | | `detail_view` | Expanded or full-content presentation within the agent's own interface | | `embed` | Source content rendered in the response surface: an iframe, a page rendered by an agentic browser, an embedded media player | +| `spoken_credit` | Source reference spoken in an audio output or assistive surface | + +`presentation_kind`, rather than `presentation_type`, determines whether the occurrence carries source content or a source reference. For example, a snippet may be an attributed source reference or an uncredited content excerpt. An embed can occur without a grounding event when the content never entered a generation context (section 4.3, *Departures from the funnel model*). -The first five values present a reference or excerpt within the agent's interface; `embed` presents the source content itself. An `embed` display can occur without a grounding event when the content never entered a generation context (section 4.3, *Departures from the funnel model*). +These are the core values. Platforms with additional presentation surfaces MAY use custom string values. Telemetry consumers MUST tolerate unknown `presentation_type` values. -These are the core values. Platforms with additional presentation surfaces MAY use custom string values. Telemetry consumers MUST tolerate unknown `display_type` values. +Each presentation event MUST have an `id` and `output_id`. When it presents a citation, `citation_id` references that `content_cited` event's `id`; an uncited presentation omits `citation_id`. Repeated presentations of the same source or output element receive distinct event IDs. This allows a later `content_engaged.presentation_id` to identify the exact surface occurrence rather than matching only by URL. -When a session includes `content_displayed` events but no subsequent `content_engaged` events, the user saw a content reference but did not interact with it. Whether this pattern is meaningful depends on the commercial agreement - a per-citation deal may not care about clickthrough, while a traffic-based deal will. This pattern is only detectable from platform-reported `content_displayed` and `content_engaged` events. Retrieval is the only event stage observable from the CDN edge. +When a session includes `content_presented` events but no subsequent `content_engaged` events, the telemetry establishes only that content or a reference was made perceivable and no reported interaction followed. It does not establish human attention. Whether this pattern is meaningful depends on the governing terms. Retrieval remains the only lifecycle stage observable from the CDN edge. ### 6.7 Engagement data (`content_engaged`) @@ -683,7 +694,7 @@ When a session includes `content_displayed` events but no subsequent `content_en |-------|------|-------------| | `engagement_type` | string | Type of interaction (see below) | -The content URL is identified by the event-level `content_url` field (section 5.2), not duplicated in `data`. +The content URL is identified by the event-level `content_url` field (section 5.2), not duplicated in `data`. Every engagement MUST carry `presentation_id`, referencing the exact `content_presented.id` on which the action occurred. Matching on URL alone is insufficient because the same source reference can be presented more than once. #### Engagement types @@ -699,7 +710,7 @@ These are the core values. Extensions MAY define additional engagement actions - `agent_navigate` is the agent-mediated counterpart of a click: the user reached the source through the agent rather than through a browser. Consumers measuring traffic SHOULD count it alongside `link_click`, distinguishing the two where the commercial agreement does. -`link_click` is the primary signal for clickthrough rate calculation. Telemetry consumers can derive per-content-owner and aggregate clickthrough rates from the ratio of `link_click` engagements to `content_displayed` events for the same `content_url`. +`link_click` is the primary signal for clickthrough rate calculation. Telemetry consumers can derive per-content-owner and aggregate clickthrough rates from `link_click` engagements and link presentations, joining each engagement through `presentation_id` rather than URL alone. A `link_click` or `agent_navigate` engagement reported from the landing page after a click-out crosses a trust boundary. Such events carry a `ctx_token` in place of `session_id`, which the telemetry consumer resolves to the originating session's click manifest (see section 7.1). @@ -754,8 +765,10 @@ An event batch carries the same envelope fields with `"document_type": "event_ba "content_url": "https://www.ft.com/content/abc123" }, { + "id": "770e8400-e29b-41d4-a716-446655440301", "type": "content_cited", "timestamp": "2026-01-15T10:30:04Z", + "output_id": "response:1", "source_role": "agent", "content_url": "https://www.ft.com/content/abc123" } @@ -769,7 +782,7 @@ For origin-side emitters at Retrieval conformance level, `session_id` MAY be omi For `content_engaged` events emitted from a landing page after a click-out (typically by a content marketplace, affiliate network, or destination site), `session_id` MAY be replaced by a `ctx_token` field that carries an opaque click-token issued by the originating agent. Telemetry consumers resolve the token to the owning session. This lets a downstream observer report a corroborating engagement event without sharing the session UUID across trust boundaries. An event MUST carry either `session_id` or `ctx_token` at Grounding conformance and above. -**ctx_token resolution.** A telemetry consumer that supports `ctx_token` resolution exposes, for a resolved token, the **click manifest**: the set of `content_grounded`, `content_cited`, and `content_displayed` events belonging to the resolved session, identifying every source that informed the response that produced the click. The manifest is gated by the resolved session's `privacy_level` and by consent. A consumer MUST return the manifest only when the issuing agent has opted in to sharing sessions via click tokens; when the agent opt-in is absent, the consumer MUST NOT disclose the manifest. Within a returned manifest, a source MUST appear only when its content owner has opted in to being visible in click-token lookups; the consumer MUST withhold the events of any content owner whose opt-in is absent while returning the remainder of the manifest. A resolution response MUST NOT include the resolved session's raw `session_id`: the token exists so that the session UUID never crosses the trust boundary, and a resolution response that returned it would undo that. The mechanism by which an agent and a content owner record these opt-ins is operator-defined; the consent gate is normative. +**ctx_token resolution.** A telemetry consumer that supports `ctx_token` resolution exposes, for a resolved token, the **click manifest**: the set of `content_grounded`, `content_cited`, and `content_presented` events belonging to the resolved session, identifying every source that informed the response that produced the click. The manifest is gated by the resolved session's `privacy_level` and by consent. A consumer MUST return the manifest only when the issuing agent has opted in to sharing sessions via click tokens; when the agent opt-in is absent, the consumer MUST NOT disclose the manifest. Within a returned manifest, a source MUST appear only when its content owner has opted in to being visible in click-token lookups; the consumer MUST withhold the events of any content owner whose opt-in is absent while returning the remainder of the manifest. A resolution response MUST NOT include the resolved session's raw `session_id`: the token exists so that the session UUID never crosses the trust boundary, and a resolution response that returned it would undo that. The mechanism by which an agent and a content owner record these opt-ins is operator-defined; the consent gate is normative. The primary schema (`telemetry-session.json`) validates session documents. A standalone event envelope schema (`telemetry-event.json`) validates the event delivery format, and a batch envelope schema (`telemetry-event-batch.json`) validates the event batch format. All three schemas share the `TelemetryEvent` definition. @@ -873,7 +886,7 @@ A manifest MAY declare multiple roles (e.g. `["content_owner", "agent"]`). A mor | Field | Type | Required | Description | |-------|------|----------|-------------| -| `name` | string | Yes | Display name of the operating organisation. | +| `name` | string | Yes | Presentation name of the operating organisation. | | `domain` | string | No | Primary domain. Defaults to the manifest URL's host. | ### 8.4 Keys @@ -1063,7 +1076,7 @@ Content can influence every response in a session without being explicitly cited - At the **grounding** level: counts all content that was in the agent's context, regardless of citation. This captures the full extent of content influence, including silent grounding. - At the **citation** level: counts only explicitly attributed content. Simpler to verify but undercounts content influence. -- At the **display** level: counts only content references shown to users. Narrowest scope, highest confidence. +- At the **presentation** level: counts content or source references made perceivable. It does not prove attention. Content owners and platforms should agree on which level to count at. The telemetry data supports all three; the choice is commercial, not technical. @@ -1126,6 +1139,23 @@ Telemetry consumers MUST tolerate unknown `response_mode` values. ## 12. Versioning +### 12.1 Migration from the v0.1 preview + +V1 replaces `content_displayed` with `content_presented`; emitters MUST NOT send +the old event name on the v1 integration line. Rename `data.display_type` to +`data.presentation_type` and add `data.presentation_kind` with either `content` +or `source_reference`. This is an intentional pre-1.0 breaking change: merely +renaming the event would preserve the visual-only ambiguity and would not say +what crossed the presentation boundary. + +For every `content_cited` event, assign an event `id` and `output_id`. For every +`content_presented` event, assign a distinct event `id` and the `output_id` of the +artifact made perceivable; add `output_element_id` when a stable element identity +exists and `citation_id` when the presentation carries a citation. For every +`content_engaged` event, add `presentation_id` referencing the exact presentation +event. Do not migrate clicks by matching URL alone: repeated presentations of the +same URL are distinct occurrences. + Preview versions (0.x) use two-component version numbers. From 1.0.0 onward, versions follow [semantic versioning](https://semver.org/): - **Major** (1.0.0 → 2.0.0) - breaking changes to required fields @@ -1186,9 +1216,12 @@ A user asks a shopping assistant to compare noise-cancelling headphones. The age } }, { + "id": "770e8400-e29b-41d4-a716-446655440302", "type": "content_cited", "timestamp": "2026-01-15T10:30:05Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:recommendation:1", "content_url": "https://www.wirecutter.com/reviews/best-wireless-headphones", "content_id": "wirecutter:best-wireless-headphones-2026", "data": { @@ -1198,13 +1231,18 @@ A user asks a shopping assistant to compare noise-cancelling headphones. The age } }, { - "type": "content_displayed", + "id": "770e8400-e29b-41d4-a716-446655440303", + "type": "content_presented", "timestamp": "2026-01-15T10:30:05Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:recommendation:1", + "citation_id": "770e8400-e29b-41d4-a716-446655440302", "content_url": "https://www.wirecutter.com/reviews/best-wireless-headphones", "content_id": "wirecutter:best-wireless-headphones-2026", "data": { - "display_type": "link" + "presentation_kind": "source_reference", + "presentation_type": "link" } }, { @@ -1230,6 +1268,7 @@ A user asks a shopping assistant to compare noise-cancelling headphones. The age "type": "content_engaged", "timestamp": "2026-01-15T10:32:00Z", "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440303", "content_url": "https://www.wirecutter.com/reviews/best-wireless-headphones", "content_id": "wirecutter:best-wireless-headphones-2026", "data": { @@ -1321,9 +1360,12 @@ An AI agent previously fetched an FT article and cached it. In a new session, th } }, { + "id": "770e8400-e29b-41d4-a716-446655440304", "type": "content_cited", "timestamp": "2026-03-28T09:00:05Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:paragraph:2", "content_url": "https://www.ft.com/content/abc123", "content_id": "ft:abc123", "data": { @@ -1335,13 +1377,18 @@ An AI agent previously fetched an FT article and cached it. In a new session, th } }, { - "type": "content_displayed", + "id": "770e8400-e29b-41d4-a716-446655440305", + "type": "content_presented", "timestamp": "2026-03-28T09:00:05Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:paragraph:2", + "citation_id": "770e8400-e29b-41d4-a716-446655440304", "content_url": "https://www.ft.com/content/abc123", "content_id": "ft:abc123", "data": { - "display_type": "link" + "presentation_kind": "source_reference", + "presentation_type": "link" } }, { @@ -1368,9 +1415,11 @@ An AI agent previously fetched an FT article and cached it. In a new session, th } }, { + "id": "770e8400-e29b-41d4-a716-446655440306", "type": "content_cited", "timestamp": "2026-03-28T09:01:08Z", "turn_id": "2", + "output_id": "response:2", "content_url": "https://www.ft.com/content/abc123", "content_id": "ft:abc123", "data": { @@ -1420,11 +1469,11 @@ In this session: - 1 article grounded from cache (no `content_retrieved` event - the CDN saw nothing) - 3 turns of conversation - 2 explicit citations (turns 1 and 2) -- 1 display event (link shown in turn 1) -- 0 engagement events (user did not click through to ft.com) +- 1 presentation event (link made perceivable in turn 1) +- 0 engagement events (no click-through was reported) - Advertising was rendered alongside the first response -The content owner can derive: article `ft:abc123` was in context for all turns, cited twice, displayed once, never clicked. The content was 14.5 hours old (cached from previous day). The response was monetised with advertising. +The content owner can derive: article `ft:abc123` was in context for all turns, cited twice, presented once, and had no reported engagement. The content was 14.5 hours old (cached from previous day). The response was monetised with advertising. ### B.4 Minimal privacy level diff --git a/telemetry-session.json b/telemetry-session.json index 9804802..bfe8e4d 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -78,7 +78,27 @@ }, "turn_id": { "type": ["string", "null"], - "description": "Associates this event with a conversation turn. Scoped to the session. SHOULD be set on turn_started, turn_completed, content_cited, content_displayed, content_engaged events, and content_grounded events when scope is turn." + "description": "Associates this event with a conversation turn. Scoped to the session. SHOULD be set on turn_started, turn_completed, content_cited, content_presented, content_engaged events, and content_grounded events when scope is turn." + }, + "output_id": { + "type": "string", + "minLength": 1, + "description": "Opaque identifier for the output artifact. REQUIRED on content_cited and content_presented events so output construction and later presentation can be correlated across services or times." + }, + "output_element_id": { + "type": "string", + "minLength": 1, + "description": "Opaque identifier for the element within output_id that carries the citation or presentation, such as a passage, media track, caption, link, or card." + }, + "citation_id": { + "type": "string", + "format": "uuid", + "description": "The id of the content_cited event associated with this presentation. Valid only on content_presented events and absent when the presentation is not a citation." + }, + "presentation_id": { + "type": "string", + "format": "uuid", + "description": "The id of the exact content_presented event on which the engagement occurred. REQUIRED on content_engaged events." }, "source_role": { "$ref": "#/$defs/SourceRole", @@ -140,6 +160,7 @@ "required": ["type"] }, "then": { + "required": ["id", "output_id"], "properties": { "data": { "properties": { @@ -158,14 +179,17 @@ }, { "if": { - "properties": { "type": { "const": "content_displayed" } }, + "properties": { "type": { "const": "content_presented" } }, "required": ["type"] }, "then": { + "required": ["id", "output_id", "data"], "properties": { "data": { + "required": ["presentation_kind", "presentation_type"], "properties": { - "display_type": { "$ref": "#/$defs/DisplayType" }, + "presentation_kind": { "$ref": "#/$defs/PresentationKind" }, + "presentation_type": { "$ref": "#/$defs/PresentationType" }, "media_type": { "$ref": "#/$defs/MediaType" } } } @@ -178,6 +202,7 @@ "required": ["type"] }, "then": { + "required": ["presentation_id"], "properties": { "data": { "properties": { @@ -223,7 +248,7 @@ "content_retrieved", "content_grounded", "content_cited", - "content_displayed", + "content_presented", "content_engaged", "turn_started", "turn_completed" @@ -334,9 +359,14 @@ "description": "Whether content informed all subsequent responses in the session or a specific turn only", "enum": ["session", "turn"] }, - "DisplayType": { + "PresentationKind": { + "type": "string", + "description": "What was made perceivable: source content itself, including a bounded excerpt or derived representation, or a reference to the source.", + "enum": ["content", "source_reference"] + }, + "PresentationType": { "type": "string", - "description": "How content or a content reference was presented to the user. Core values: link, snippet, inline_quote, card, detail_view, embed. Platforms MAY use custom values for additional presentation surfaces. Consumers MUST tolerate unknown values." + "description": "How content or a source reference was made perceivable. Core values: link, snippet, inline_quote, card, detail_view, embed, spoken_credit. Platforms MAY use custom values for additional presentation surfaces. Consumers MUST tolerate unknown values." } } } diff --git a/tests/README.md b/tests/README.md index b22d382..3ff93d2 100644 --- a/tests/README.md +++ b/tests/README.md @@ -31,8 +31,9 @@ Run from the repository root. Without uv: `pip install jsonschema`, then `python - All three conformance levels (Retrieval, Grounding, Citation) - Standalone event envelopes (CDN edge, agent with session FK) - Privacy level field gating (application-layer conformance) -- Funnel exceptions (displayed-no-cited, cited-no-grounded, displayed-no-grounded) -- Embedded display (`display_type: embed`) and agent-mediated engagement (`agent_navigate`) +- Funnel exceptions (presented-no-cited, cited-no-grounded, presented-no-grounded) +- Text, image, audio, video, suppressed-citation, and repeated-presentation cases +- Exact presentation-to-engagement correlation across session, standalone, and batch envelopes - Multi-turn sessions, cached grounding - Custom response_mode values diff --git a/tests/invalid/batch-missing-session-and-ctx-token.json b/tests/invalid/batch-missing-session-and-ctx-token.json index 6c70711..4ceb7e6 100644 --- a/tests/invalid/batch-missing-session-and-ctx-token.json +++ b/tests/invalid/batch-missing-session-and-ctx-token.json @@ -4,8 +4,10 @@ "schema_version": "0.1", "events": [ { + "id": "990e8400-e29b-41d4-a716-446655440063", "type": "content_cited", "timestamp": "2026-03-28T10:05:00Z", + "output_id": "response:1", "source_role": "agent", "content_url": "https://example.com/article/test" } diff --git a/tests/invalid/engaged-missing-presentation-id.json b/tests/invalid/engaged-missing-presentation-id.json new file mode 100644 index 0000000..8632196 --- /dev/null +++ b/tests/invalid/engaged-missing-presentation-id.json @@ -0,0 +1,14 @@ +{ + "_test_description": "content_engaged must identify the exact presentation occurrence rather than matching only by URL.", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440092", + "started_at": "2026-07-18T13:00:00Z", + "events": [ + { + "type": "content_engaged", + "timestamp": "2026-07-18T13:00:02Z", + "content_id": "publisher:article:1", + "data": { "engagement_type": "link_click" } + } + ] +} diff --git a/tests/invalid/legacy-content-displayed.json b/tests/invalid/legacy-content-displayed.json new file mode 100644 index 0000000..ab93b1e --- /dev/null +++ b/tests/invalid/legacy-content-displayed.json @@ -0,0 +1,14 @@ +{ + "_test_description": "The v1 presentation event replaces the preview content_displayed name.", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440093", + "started_at": "2026-07-18T13:00:00Z", + "events": [ + { + "type": "content_displayed", + "timestamp": "2026-07-18T13:00:03Z", + "content_id": "publisher:article:1", + "data": { "display_type": "link" } + } + ] +} diff --git a/tests/invalid/presented-missing-kind.json b/tests/invalid/presented-missing-kind.json new file mode 100644 index 0000000..27a7594 --- /dev/null +++ b/tests/invalid/presented-missing-kind.json @@ -0,0 +1,16 @@ +{ + "_test_description": "content_presented must distinguish source content from a source reference.", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440090", + "started_at": "2026-07-18T13:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440091", + "type": "content_presented", + "timestamp": "2026-07-18T13:00:01Z", + "output_id": "response:1", + "content_id": "publisher:article:1", + "data": { "presentation_type": "link" } + } + ] +} diff --git a/tests/valid/event-batch-agent.json b/tests/valid/event-batch-agent.json index de25dde..7d7b01b 100644 --- a/tests/valid/event-batch-agent.json +++ b/tests/valid/event-batch-agent.json @@ -20,10 +20,25 @@ "content_url": "https://arstechnica.com/science/2026/03/new-battery-chemistry-breakthrough/" }, { + "id": "990e8400-e29b-41d4-a716-446655440062", "type": "content_cited", "timestamp": "2026-03-28T08:20:05Z", + "output_id": "response:1", "source_role": "agent", "content_url": "https://arstechnica.com/science/2026/03/new-battery-chemistry-breakthrough/" + }, + { + "id": "990e8400-e29b-41d4-a716-446655440063", + "type": "content_presented", + "timestamp": "2026-03-28T08:20:06Z", + "output_id": "response:1", + "citation_id": "990e8400-e29b-41d4-a716-446655440062", + "source_role": "agent", + "content_url": "https://arstechnica.com/science/2026/03/new-battery-chemistry-breakthrough/", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } } ] } diff --git a/tests/valid/event-standalone-engaged-ctx-token.json b/tests/valid/event-standalone-engaged-ctx-token.json index 659609c..f73d6af 100644 --- a/tests/valid/event-standalone-engaged-ctx-token.json +++ b/tests/valid/event-standalone-engaged-ctx-token.json @@ -6,6 +6,7 @@ "event": { "type": "content_engaged", "timestamp": "2026-03-28T14:06:00Z", + "presentation_id": "990e8400-e29b-41d4-a716-446655440064", "content_url": "https://www.example-review.com/headphones/best-noise-cancelling", "data": { "engagement_type": "link_click" diff --git a/tests/valid/event-standalone-presented.json b/tests/valid/event-standalone-presented.json new file mode 100644 index 0000000..3ef689f --- /dev/null +++ b/tests/valid/event-standalone-presented.json @@ -0,0 +1,20 @@ +{ + "_test_description": "Standalone presentation envelope using the shared modality-neutral presentation schema.", + "document_type": "event", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440080", + "agent_id": "voice-assistant-v1", + "started_at": "2026-07-18T12:30:00Z", + "event": { + "id": "770e8400-e29b-41d4-a716-446655440081", + "type": "content_presented", + "timestamp": "2026-07-18T12:30:05Z", + "output_id": "audio-response:1", + "content_id": "publisher:article:1", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "spoken_credit", + "media_type": "audio" + } + } +} diff --git a/tests/valid/session-citation-tier.json b/tests/valid/session-citation-tier.json index 24ff208..5734223 100644 --- a/tests/valid/session-citation-tier.json +++ b/tests/valid/session-citation-tier.json @@ -41,9 +41,12 @@ } }, { + "id": "880e8400-e29b-41d4-a716-446655440011", "type": "content_cited", "timestamp": "2026-03-28T14:00:06Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:recommendation:1", "content_url": "https://www.runnersworld.com/gear/best-trail-running-shoes", "content_id": "rw:best-trail-shoes-2026", "data": { @@ -56,13 +59,18 @@ } }, { - "type": "content_displayed", + "id": "880e8400-e29b-41d4-a716-446655440012", + "type": "content_presented", "timestamp": "2026-03-28T14:00:06Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:recommendation:1", + "citation_id": "880e8400-e29b-41d4-a716-446655440011", "content_url": "https://www.runnersworld.com/gear/best-trail-running-shoes", "content_id": "rw:best-trail-shoes-2026", "data": { - "display_type": "card" + "presentation_kind": "source_reference", + "presentation_type": "card" } }, { @@ -88,6 +96,7 @@ "type": "content_engaged", "timestamp": "2026-03-28T14:02:30Z", "turn_id": "1", + "presentation_id": "880e8400-e29b-41d4-a716-446655440012", "content_url": "https://www.runnersworld.com/gear/best-trail-running-shoes", "content_id": "rw:best-trail-shoes-2026", "data": { diff --git a/tests/valid/session-funnel-exception-cited-no-grounded.json b/tests/valid/session-funnel-exception-cited-no-grounded.json index fad0e48..6145718 100644 --- a/tests/valid/session-funnel-exception-cited-no-grounded.json +++ b/tests/valid/session-funnel-exception-cited-no-grounded.json @@ -16,9 +16,12 @@ } }, { + "id": "880e8400-e29b-41d4-a716-446655440021", "type": "content_cited", "timestamp": "2026-03-28T20:00:05Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:claim:1", "content_url": "https://www.ons.gov.uk/peoplepopulationandcommunity/populationandmigration", "data": { "citation_type": "reference", @@ -27,12 +30,17 @@ } }, { - "type": "content_displayed", + "id": "880e8400-e29b-41d4-a716-446655440022", + "type": "content_presented", "timestamp": "2026-03-28T20:00:05Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:claim:1", + "citation_id": "880e8400-e29b-41d4-a716-446655440021", "content_url": "https://www.ons.gov.uk/peoplepopulationandcommunity/populationandmigration", "data": { - "display_type": "link" + "presentation_kind": "source_reference", + "presentation_type": "link" } }, { diff --git a/tests/valid/session-funnel-exception-displayed-no-cited.json b/tests/valid/session-funnel-exception-presented-no-cited.json similarity index 77% rename from tests/valid/session-funnel-exception-displayed-no-cited.json rename to tests/valid/session-funnel-exception-presented-no-cited.json index d5ed91e..eb29a80 100644 --- a/tests/valid/session-funnel-exception-displayed-no-cited.json +++ b/tests/valid/session-funnel-exception-presented-no-cited.json @@ -1,5 +1,5 @@ { - "_test_description": "Funnel exception: content_displayed without content_cited. This is the Sources sidebar pattern - the agent shows source links without explicitly citing the content in the response text. Valid per section 4.3.", + "_test_description": "Funnel exception: content_presented without content_cited. This is the Sources sidebar pattern - the agent makes source links perceivable without semantically associating them with an output element. Valid per section 4.3.", "schema_version": "0.1", "session_id": "660e8400-e29b-41d4-a716-446655440010", "agent_id": "search-assistant-v1", @@ -32,12 +32,16 @@ } }, { - "type": "content_displayed", + "id": "880e8400-e29b-41d4-a716-446655440031", + "type": "content_presented", "timestamp": "2026-03-28T19:00:06Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "sources:1", "content_url": "https://www.kingarthurbaking.com/recipes/sourdough-starter-maintenance", "data": { - "display_type": "link" + "presentation_kind": "source_reference", + "presentation_type": "link" } }, { diff --git a/tests/valid/session-funnel-exception-displayed-no-grounded.json b/tests/valid/session-funnel-exception-presented-no-grounded.json similarity index 65% rename from tests/valid/session-funnel-exception-displayed-no-grounded.json rename to tests/valid/session-funnel-exception-presented-no-grounded.json index 308350f..856afc4 100644 --- a/tests/valid/session-funnel-exception-displayed-no-grounded.json +++ b/tests/valid/session-funnel-exception-presented-no-grounded.json @@ -1,5 +1,5 @@ { - "_test_description": "Funnel exception: content_displayed without content_grounded. An agentic browser renders a publisher's page to the user (display_type: embed) without the content entering a generation context, and the user directs the agent to open a second page (engagement_type: agent_navigate). Valid per section 4.3. Also exercises media_type on display data and the open display_type/engagement_type vocabularies.", + "_test_description": "Funnel exception: content_presented without content_grounded. An agentic browser renders a publisher's page on a recipient-facing surface (presentation_type: embed) without the content entering a generation context, and the recipient directs the agent to open a second page (engagement_type: agent_navigate). Valid per section 4.3. Also exercises media_type on presentation data and the open presentation_type/engagement_type vocabularies.", "schema_version": "0.1", "session_id": "660e8400-e29b-41d4-a716-446655440020", "agent_id": "browser-agent-v1", @@ -22,22 +22,28 @@ "content_url": "https://www.fitnessmedia.example/guides/home-workouts" }, { - "type": "content_displayed", + "id": "880e8400-e29b-41d4-a716-446655440041", + "type": "content_presented", "timestamp": "2026-03-29T11:00:02Z", "turn_id": "1", + "output_id": "browser-view:1", "content_url": "https://www.fitnessmedia.example/guides/home-workouts", "data": { - "display_type": "embed", + "presentation_kind": "content", + "presentation_type": "embed", "media_type": "text" } }, { - "type": "content_displayed", + "id": "880e8400-e29b-41d4-a716-446655440042", + "type": "content_presented", "timestamp": "2026-03-29T11:00:04Z", "turn_id": "1", + "output_id": "browser-view:1", "content_url": "https://www.fitnessmedia.example/videos/beginner-routine", "data": { - "display_type": "embed", + "presentation_kind": "content", + "presentation_type": "embed", "media_type": "video" } }, @@ -56,6 +62,7 @@ "type": "content_engaged", "timestamp": "2026-03-29T11:01:00Z", "turn_id": "1", + "presentation_id": "880e8400-e29b-41d4-a716-446655440041", "content_url": "https://www.fitnessmedia.example/guides/home-workouts", "data": { "engagement_type": "agent_navigate" diff --git a/tests/valid/session-multi-turn.json b/tests/valid/session-multi-turn.json index 27439ee..bcc0b9b 100644 --- a/tests/valid/session-multi-turn.json +++ b/tests/valid/session-multi-turn.json @@ -35,9 +35,11 @@ } }, { + "id": "880e8400-e29b-41d4-a716-446655440051", "type": "content_cited", "timestamp": "2026-03-28T16:00:08Z", "turn_id": "1", + "output_id": "response:1", "content_url": "https://www.bbc.co.uk/news/science-environment-68234567", "content_id": "bbc:68234567", "data": { @@ -47,12 +49,16 @@ } }, { - "type": "content_displayed", + "id": "880e8400-e29b-41d4-a716-446655440052", + "type": "content_presented", "timestamp": "2026-03-28T16:00:08Z", "turn_id": "1", + "output_id": "response:1", + "citation_id": "880e8400-e29b-41d4-a716-446655440051", "content_url": "https://www.bbc.co.uk/news/science-environment-68234567", "data": { - "display_type": "link" + "presentation_kind": "source_reference", + "presentation_type": "link" } }, { @@ -100,9 +106,11 @@ } }, { + "id": "880e8400-e29b-41d4-a716-446655440053", "type": "content_cited", "timestamp": "2026-03-28T16:05:12Z", "turn_id": "3", + "output_id": "response:3", "content_url": "https://www.bbc.co.uk/news/science-environment-68234567", "content_id": "bbc:68234567", "data": { diff --git a/tests/valid/session-presentation-multimodal.json b/tests/valid/session-presentation-multimodal.json new file mode 100644 index 0000000..c3a8a3d --- /dev/null +++ b/tests/valid/session-presentation-multimodal.json @@ -0,0 +1,96 @@ +{ + "_test_description": "Modality-neutral citation and presentation: an image excerpt is presented as content, an audio output speaks a source credit, a video is embedded without citation, a constructed citation is suppressed, and a repeated link presentation receives the engagement.", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440070", + "agent_id": "multimodal-assistant-v1", + "started_at": "2026-07-18T12:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-07-18T12:00:01Z", + "content_id": "museum:image:42", + "data": { "scope": "turn", "cached": true, "media_type": "image" } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440071", + "type": "content_cited", + "timestamp": "2026-07-18T12:00:04Z", + "turn_id": "1", + "output_id": "output:mixed:1", + "output_element_id": "image:region:credit", + "content_id": "museum:image:42", + "data": { "citation_type": "reference", "media_type": "image" } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440072", + "type": "content_presented", + "timestamp": "2026-07-18T12:00:05Z", + "turn_id": "1", + "output_id": "output:mixed:1", + "output_element_id": "image:region:1", + "content_id": "museum:image:42", + "data": { "presentation_kind": "content", "presentation_type": "embed", "media_type": "image" } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440073", + "type": "content_presented", + "timestamp": "2026-07-18T12:00:06Z", + "turn_id": "1", + "output_id": "output:mixed:1", + "output_element_id": "audio:credit:1", + "citation_id": "770e8400-e29b-41d4-a716-446655440071", + "content_id": "museum:image:42", + "data": { "presentation_kind": "source_reference", "presentation_type": "spoken_credit", "media_type": "audio" } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440074", + "type": "content_presented", + "timestamp": "2026-07-18T12:00:07Z", + "turn_id": "1", + "output_id": "output:mixed:1", + "output_element_id": "video:1", + "content_id": "archive:video:7", + "data": { "presentation_kind": "content", "presentation_type": "embed", "media_type": "video" } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440075", + "type": "content_cited", + "timestamp": "2026-07-18T12:00:08Z", + "turn_id": "1", + "output_id": "output:suppressed:1", + "output_element_id": "caption:1", + "content_id": "audio:interview:9", + "data": { "citation_type": "reference", "media_type": "audio" } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440076", + "type": "content_presented", + "timestamp": "2026-07-18T12:00:09Z", + "turn_id": "1", + "output_id": "output:mixed:1", + "output_element_id": "sources:link:1", + "citation_id": "770e8400-e29b-41d4-a716-446655440071", + "content_id": "museum:image:42", + "data": { "presentation_kind": "source_reference", "presentation_type": "link", "media_type": "text" } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440077", + "type": "content_presented", + "timestamp": "2026-07-18T12:00:10Z", + "turn_id": "1", + "output_id": "output:mixed:1", + "output_element_id": "sidebar:link:1", + "citation_id": "770e8400-e29b-41d4-a716-446655440071", + "content_id": "museum:image:42", + "data": { "presentation_kind": "source_reference", "presentation_type": "link", "media_type": "text" } + }, + { + "type": "content_engaged", + "timestamp": "2026-07-18T12:00:12Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440077", + "content_id": "museum:image:42", + "data": { "engagement_type": "link_click" } + } + ] +} diff --git a/tests/validate.py b/tests/validate.py index ad51e01..67ac9e8 100644 --- a/tests/validate.py +++ b/tests/validate.py @@ -103,7 +103,7 @@ # turn_completed are turn events, not content events, and are exempt. CONTENT_EVENT_TYPES = { "content_retrieved", "content_grounded", "content_cited", - "content_displayed", "content_engaged", + "content_presented", "content_engaged", } # Fields that MUST NOT appear at each privacy level (section 5.5). From d3144c19314082514a5bcec6db371fd1fe1f7794 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Tue, 28 Jul 2026 18:13:59 +0100 Subject: [PATCH 02/38] Withdraw ip_hash and correct the data-minimisation text Consultation issue #2. A SHA-256 of a client IP is a pseudonym, not an anonymous value: the IPv4 space is small enough to enumerate against a candidate digest. The field is removed from the edge and origin data profiles, the schema and the fixtures. Section 9.1 previously told emitters to "hash or anonymise identifiers where possible", which is the claim the consultation objected to. It now points emitters at asn, asn_org and country, which describe the network path rather than the client. Event data accepts additional properties, so the schemas cannot reject a withdrawn field. The conformance suite gains an application-layer rule and a negative fixture instead. Co-Authored-By: Claude Opus 5 (1M context) --- SPECIFICATION.md | 11 +++++---- telemetry-session.json | 3 +-- tests/invalid/withdrawn-ip-hash.json | 21 ++++++++++++++++ tests/valid/event-standalone-edge.json | 3 +-- tests/valid/session-retrieval-tier.json | 3 +-- tests/validate.py | 32 +++++++++++++++++++++++++ 6 files changed, 62 insertions(+), 11 deletions(-) create mode 100644 tests/invalid/withdrawn-ip-hash.json diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 746e4a4..9197b4d 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -544,7 +544,6 @@ CDN and edge network integrations SHOULD include these fields: | `asn` | integer | Client AS number | | `asn_org` | string | Client AS organisation name | | `country` | string | ISO 3166-1 alpha-2 country code | -| `ip_hash` | string | SHA-256 of client IP (`sha256:{hex}`) | #### Bot categories @@ -565,7 +564,6 @@ Emitting a `training`-category `content_retrieved` event is permitted but non-at | Field | Type | Description | |-------|------|-------------| | `user_agent` | string | Request User-Agent header | -| `ip_hash` | string | SHA-256 of client IP | | `response_status` | integer | HTTP response status code | ### 6.4 Grounding data (`content_grounded`) @@ -1009,8 +1007,12 @@ The following are deferred to later versions: Emitters SHOULD: - Use the minimum `privacy_level` necessary -- Hash or anonymise identifiers where possible - Use coarse `topics` values that do not identify sensitive categories (health, political or religious affiliation, sexuality) +- Carry network-level context about the request rather than about the client: `asn`, `asn_org` and `country` describe the network path and do not identify the individual behind it + +Hashing does not anonymise a value drawn from a space small enough to enumerate. The entire IPv4 address space can be hashed and compared against a candidate digest on commodity hardware, so a hashed IP address is a pseudonym rather than an anonymous value and should be treated as personal data. Version 0.1 defined an `ip_hash` field in the edge and origin data profiles (sections 6.2 and 6.3). Version 1 withdraws it, and emitters MUST NOT populate it. + +The schemas cannot enforce this: event `data` accepts additional properties by design, so a withdrawn field validates as an ordinary extension. The conformance suite checks it at the application layer instead. ### 9.2 Recommended levels @@ -1276,8 +1278,7 @@ A content owner's CDN detects an AI agent fetching content. The agent also repor "ja4": "t13d1517h2_8daaf6152771_02e4c6ae3e16", "asn": 14618, "asn_org": "Anthropic", - "country": "US", - "ip_hash": "sha256:d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5" + "country": "US" } } ``` diff --git a/telemetry-session.json b/telemetry-session.json index 9804802..b2ddf18 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -207,8 +207,7 @@ "ja4": { "type": "string", "description": "JA4 TLS client fingerprint" }, "asn": { "type": "integer", "description": "Client AS number" }, "asn_org": { "type": "string", "description": "Client AS organisation name" }, - "country": { "type": "string", "pattern": "^[A-Z]{2}$", "description": "ISO 3166-1 alpha-2 country code" }, - "ip_hash": { "type": "string", "pattern": "^sha256:[a-f0-9]{64}$", "description": "SHA-256 of client IP (sha256:{hex})" } + "country": { "type": "string", "pattern": "^[A-Z]{2}$", "description": "ISO 3166-1 alpha-2 country code" } } } } diff --git a/tests/invalid/withdrawn-ip-hash.json b/tests/invalid/withdrawn-ip-hash.json new file mode 100644 index 0000000..100b3ae --- /dev/null +++ b/tests/invalid/withdrawn-ip-hash.json @@ -0,0 +1,21 @@ +{ + "_test_description": "Standalone event envelope from a CDN carrying ip_hash in data. The field was withdrawn in v1 (section 9.1): hashing does not anonymise a value drawn from a space small enough to enumerate. Passes JSON Schema because event data accepts additional properties, so the rule is enforced at the application layer.", + "document_type": "event", + "schema_version": "0.1", + "event": { + "type": "content_retrieved", + "timestamp": "2026-03-28T08:15:00Z", + "source_role": "edge", + "content_telemetry_id": "990e8400-e29b-41d4-a716-446655440051", + "content_url": "https://www.telegraph.co.uk/business/2026/03/28/ftse-100-markets-live", + "data": { + "user_agent": "PerplexityBot/1.0", + "bot_category": "inference", + "response_status": 200, + "asn": 396982, + "asn_org": "Perplexity AI", + "country": "US", + "ip_hash": "sha256:e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6" + } + } +} diff --git a/tests/valid/event-standalone-edge.json b/tests/valid/event-standalone-edge.json index 4acec96..018f44b 100644 --- a/tests/valid/event-standalone-edge.json +++ b/tests/valid/event-standalone-edge.json @@ -18,8 +18,7 @@ "ja4": "t13d1516h2_5b57614c22b0_7cb938dcc8ab", "asn": 396982, "asn_org": "Perplexity AI", - "country": "US", - "ip_hash": "sha256:e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6" + "country": "US" } } } diff --git a/tests/valid/session-retrieval-tier.json b/tests/valid/session-retrieval-tier.json index 494f0bb..b17669d 100644 --- a/tests/valid/session-retrieval-tier.json +++ b/tests/valid/session-retrieval-tier.json @@ -20,8 +20,7 @@ "ja4": "t13d1517h2_8daaf6152771_02e4c6ae3e16", "asn": 14618, "asn_org": "Anthropic", - "country": "US", - "ip_hash": "sha256:a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2" + "country": "US" } } ] diff --git a/tests/validate.py b/tests/validate.py index ad51e01..3909491 100644 --- a/tests/validate.py +++ b/tests/validate.py @@ -96,6 +96,19 @@ "Violates section 8.6: every entry MUST be the manifest's own host or a " "subdomain of it. Consumers reject the manifest as malformed (section 8.7)." ), + "withdrawn-ip-hash.json": ( + "Retrieval event carries ip_hash in data. " + "Violates section 9.1: the field was withdrawn in v1 and emitters " + "MUST NOT populate it." + ), +} + +# Fields withdrawn in v1 that emitters MUST NOT populate (section 9.1). The +# schemas cannot catch these: event `data` accepts additional properties by +# design, so a withdrawn field validates as an ordinary extension unless it is +# checked here. +WITHDRAWN_EVENT_DATA_FIELDS = { + "ip_hash": "withdrawn in v1; hashing does not anonymise an IP address", } # Event types that carry content and therefore require an identifier @@ -312,12 +325,31 @@ def check_session_or_ctx_token(data): return [] +def check_withdrawn_fields(data): + """ + Check that no event carries a field withdrawn in v1 (section 9.1). + Returns a list of violation descriptions. + """ + violations = [] + for event in _iter_events(data): + event_data = event.get("data") + if not isinstance(event_data, dict): + continue + for field, reason in WITHDRAWN_EVENT_DATA_FIELDS.items(): + if field in event_data: + violations.append( + f"Event '{event.get('type')}' carries '{field}' in data ({reason})" + ) + return violations + + def check_application_layer(data): """Run every application-layer conformance rule and return all violations.""" return ( check_privacy_conformance(data) + check_content_identifier(data) + check_session_or_ctx_token(data) + + check_withdrawn_fields(data) ) From f65888961b30cf90432e7fe31655b58d9e95bedd Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Tue, 28 Jul 2026 18:15:19 +0100 Subject: [PATCH 03/38] Add chars_ingested and make token counts supplementary Consultation issue #6. Token counts are measured in the emitter's own tokeniser, so they are not comparable between agents and shift when a vendor revises a tokeniser. A content owner receiving grounding events from several agents could not aggregate the only volume measure the specification offered. chars_ingested counts Unicode code points in the content placed in the generation context. Two emitters counting the same text agree. tokens_ingested stays, described as supplementary, and section 5.5 now carries the same caveat for query_tokens and response_tokens. Grounding conformance asks for chars_ingested rather than tokens_ingested. Fixtures and worked examples carry both. Co-Authored-By: Claude Opus 5 (1M context) --- README.md | 1 + SPECIFICATION.md | 15 +++++++++++---- telemetry-session.json | 3 ++- tests/valid/session-cached-grounding.json | 1 + tests/valid/session-citation-tier.json | 1 + tests/valid/session-custom-response-mode.json | 1 + ...ssion-funnel-exception-displayed-no-cited.json | 1 + tests/valid/session-grounding-tier.json | 1 + tests/valid/session-multi-turn.json | 1 + 9 files changed, 20 insertions(+), 5 deletions(-) diff --git a/README.md b/README.md index 451dcd2..99d5697 100644 --- a/README.md +++ b/README.md @@ -86,6 +86,7 @@ A user asks an AI agent about UK interest rates. The agent grounds its response "data": { "scope": "session", "cached": true, + "chars_ingested": 12800, "tokens_ingested": 3200, "content_last_modified": "2026-03-27T18:30:00Z" } diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 746e4a4..a8c9668 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -415,7 +415,7 @@ These are the recommended values. Platforms with additional product surfaces (co An emitter that populates a conversation turn MUST NOT include a field above that turn's declared `privacy_level` - for example, `query_text` MUST NOT be present when `privacy_level` is `intent` or `minimal`. This restriction is a property of `privacy_level` itself: it applies wherever conversation turns are emitted, independent of the emitter's conformance level. -**Token counts** includes `query_tokens` and `response_tokens`. These are available at all levels because they are needed for token-based counting models and do not reveal user intent or platform strategy. +**Token counts** includes `query_tokens` and `response_tokens`. These are available at all levels because they are needed for token-based counting models and do not reveal user intent or platform strategy. They carry the same portability limit as `tokens_ingested` (section 6.4): both are measured in the emitter's own tokeniser and are not comparable between agents. **Response classification** includes `response_type` (e.g., `"recommendation"`, `"explanation"`). Available at `intent` level and above, as it can reveal the nature of the user's query. @@ -473,7 +473,7 @@ A conforming **Grounding** emitter MUST satisfy Retrieval requirements and also: - Emit `turn_started` and `turn_completed` events with `privacy_level` - Restrict conversation turn fields to the declared `privacy_level` (section 5.5) -A Grounding emitter SHOULD include `data.tokens_ingested` and `data.cached` on grounding events. +A Grounding emitter SHOULD include `data.chars_ingested` and `data.cached` on grounding events, and MAY add `data.tokens_ingested` alongside them (section 6.4). Emitters using standalone event delivery (section 7.1) MUST include `agent_id`, `started_at`, and either `session_id` or, for click-out engagement events, `ctx_token` on the standalone event envelope to satisfy Grounding conformance. @@ -574,13 +574,18 @@ Emitting a `training`-category `content_retrieved` event is permitted but non-at |-------|------|-------------| | `scope` | string | Influence scope: `session` or `turn` (see below) | | `cached` | boolean | Content served from agent-side cache rather than a live fetch | -| `tokens_ingested` | integer | Token count of content placed in the generation context (see below) | +| `chars_ingested` | integer | Character count of content placed in the generation context (see below) | +| `tokens_ingested` | integer | Token count of the same content, supplementary (see below) | | `content_version` | string | Content version identifier (ETag, revision ID, CMS version) | | `content_last_modified` | datetime | When the content was last modified at source | | `content_hash` | string | SHA-256 of the content as ingested (`sha256:{hex}`) | | `media_type` | string | Content medium: `text`, `image`, `video`, `audio` (open vocabulary, see 6.1) | -`tokens_ingested` counts tokens actually placed in the generation model's context. For chunked retrieval, count only the tokens used, not the full source document. The token count uses the generation model's tokeniser (the model identified in `model_id` on the corresponding `turn_completed` event), not the retrieval or embedding model's tokeniser. +Both fields measure the content actually placed in the generation model's context. For chunked retrieval, count only the portion used, not the full source document. + +`chars_ingested` counts Unicode code points in that content. It is the portable measure: two emitters counting the same text agree, so a content owner can compare volumes across agents and over time without knowing which model produced the number. + +`tokens_ingested` counts the same content in the generation model's tokeniser (the model identified in `model_id` on the corresponding `turn_completed` event), not the retrieval or embedding model's tokeniser. It is supplementary. Token counts are model-specific, change when a vendor revises a tokeniser, and are not comparable between agents, so a consumer cannot aggregate them across emitters or treat a difference as a difference in volume. Emitters SHOULD send `chars_ingested` where they send `tokens_ingested`, and consumers that receive only token counts SHOULD record which model produced them. #### Grounding scope @@ -1180,6 +1185,7 @@ A user asks a shopping assistant to compare noise-cancelling headphones. The age "data": { "scope": "session", "cached": false, + "chars_ingested": 16800, "tokens_ingested": 4200, "content_last_modified": "2026-01-10T14:00:00Z", "media_type": "text" @@ -1304,6 +1310,7 @@ An AI agent previously fetched an FT article and cached it. In a new session, th "data": { "scope": "session", "cached": true, + "chars_ingested": 12800, "tokens_ingested": 3200, "content_last_modified": "2026-03-27T18:30:00Z", "content_hash": "sha256:a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2", diff --git a/telemetry-session.json b/telemetry-session.json index 9804802..3cf2227 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -124,7 +124,8 @@ "properties": { "scope": { "$ref": "#/$defs/GroundingScope" }, "cached": { "type": "boolean" }, - "tokens_ingested": { "type": "integer", "minimum": 0 }, + "chars_ingested": { "type": "integer", "minimum": 0, "description": "Unicode code points of content placed in the generation context. Portable across emitters; preferred over tokens_ingested (section 6.4)." }, + "tokens_ingested": { "type": "integer", "minimum": 0, "description": "Token count of the same content in the emitter's own tokeniser. Supplementary: not comparable between emitters (section 6.4)." }, "content_version": { "type": "string" }, "content_last_modified": { "type": "string", "format": "date-time" }, "content_hash": { "type": "string", "pattern": "^sha256:[a-f0-9]{64}$" }, diff --git a/tests/valid/session-cached-grounding.json b/tests/valid/session-cached-grounding.json index 59e8b09..a9821f1 100644 --- a/tests/valid/session-cached-grounding.json +++ b/tests/valid/session-cached-grounding.json @@ -15,6 +15,7 @@ "data": { "scope": "session", "cached": true, + "chars_ingested": 11200, "tokens_ingested": 2800, "content_last_modified": "2026-03-27T16:00:00Z", "content_hash": "sha256:c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4", diff --git a/tests/valid/session-citation-tier.json b/tests/valid/session-citation-tier.json index 24ff208..791e852 100644 --- a/tests/valid/session-citation-tier.json +++ b/tests/valid/session-citation-tier.json @@ -34,6 +34,7 @@ "data": { "scope": "session", "cached": false, + "chars_ingested": 24800, "tokens_ingested": 6200, "content_last_modified": "2026-03-20T10:00:00Z", "content_hash": "sha256:b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3", diff --git a/tests/valid/session-custom-response-mode.json b/tests/valid/session-custom-response-mode.json index 9d7b3a6..fb5e627 100644 --- a/tests/valid/session-custom-response-mode.json +++ b/tests/valid/session-custom-response-mode.json @@ -28,6 +28,7 @@ "data": { "scope": "turn", "cached": false, + "chars_ingested": 18000, "tokens_ingested": 4500 } }, diff --git a/tests/valid/session-funnel-exception-displayed-no-cited.json b/tests/valid/session-funnel-exception-displayed-no-cited.json index d5ed91e..dede6f4 100644 --- a/tests/valid/session-funnel-exception-displayed-no-cited.json +++ b/tests/valid/session-funnel-exception-displayed-no-cited.json @@ -28,6 +28,7 @@ "data": { "scope": "turn", "cached": false, + "chars_ingested": 7200, "tokens_ingested": 1800 } }, diff --git a/tests/valid/session-grounding-tier.json b/tests/valid/session-grounding-tier.json index d3002ca..576abea 100644 --- a/tests/valid/session-grounding-tier.json +++ b/tests/valid/session-grounding-tier.json @@ -30,6 +30,7 @@ "data": { "scope": "session", "cached": false, + "chars_ingested": 20400, "tokens_ingested": 5100, "content_last_modified": "2026-03-15T08:00:00Z", "media_type": "text" diff --git a/tests/valid/session-multi-turn.json b/tests/valid/session-multi-turn.json index 27439ee..8ccca4d 100644 --- a/tests/valid/session-multi-turn.json +++ b/tests/valid/session-multi-turn.json @@ -20,6 +20,7 @@ "data": { "scope": "session", "cached": false, + "chars_ingested": 9600, "tokens_ingested": 2400, "media_type": "text" } From dcdeb1857c0781f447d6657e2610a1d4cba3216f Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Tue, 28 Jul 2026 18:16:05 +0100 Subject: [PATCH 04/38] Correct what license_ref establishes Consultation issue #22. Section 5.2.3 said a consumer receiving license_ref "can verify that content usage was licensed". Core neither resolves nor interprets the value, so it establishes none of that. The field is part of the emitter's claim: it records which grant the emitter says applied, not that the grant existed, covered this content, was valid at the time, or permitted the use. Also records what the field does not identify. COUNTER Metrics read license_ref as a possible carrier for the institution deriving access rights; it is not one. Where a publisher issues one grant per subscriber the value can work as a proxy inside that publisher's namespace, but nothing requires it to be typed, stable, or comparable between emitters. Co-Authored-By: Claude Opus 5 (1M context) --- SPECIFICATION.md | 10 +++++++--- telemetry-session.json | 2 +- 2 files changed, 8 insertions(+), 4 deletions(-) diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 746e4a4..e05300f 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -328,7 +328,7 @@ Format: the URL of a manifest served at `/.well-known/content-telemetry.json` un | `content_telemetry_id` | UUID | No | Correlation ID for cross-observer deduplication (see 7.2) | | `content_url` | string | No | Content URL as fetched or canonical URL | | `content_id` | string | No | Content owner's stable content identifier (see 4.5) | -| `license_ref` | string | No | Reference to the licence under which content was accessed | +| `license_ref` | string | No | Reference to a licence or grant the emitter associates with this event (see 5.2.3) | | `turn` | ConversationTurn | No | Conversation data (for turn events) | | `data` | object | No | Type-specific metadata (see section 6) | @@ -346,7 +346,11 @@ The `source_role` field SHOULD be set on `content_retrieved` events. When multip #### 5.2.3 Licence reference -The `license_ref` field connects a telemetry event to the content access licence that authorised it. The format depends on the access protocol: a JWT `jti` claim, a CoMP package ID, or any opaque identifier that both parties can resolve. When present, telemetry consumers can verify that content usage was licensed. +The `license_ref` field associates a telemetry event with a licence or grant the emitter references. The format depends on the access protocol: a JWT `jti` claim, a CoMP package ID, or any opaque identifier that both parties can resolve. + +Core does not resolve, validate or interpret the reference. `license_ref` is part of the emitter's claim about the event: it records which grant the emitter says applied. It does not establish that the grant existed, that it covered this content, that it was valid at the time of use, or that the use was permitted. A consumer that needs any of those has to check the issuer's own records, or use evidence defined outside this specification. + +`license_ref` also does not identify the party whose entitlement was used. Where a publisher issues one grant per subscriber the value may work as a proxy for that subscriber, but only within the issuing publisher's namespace: nothing here requires the value to be typed, stable across sessions, or comparable between emitters. ### 5.3 Event types @@ -618,7 +622,7 @@ The `cached` field distinguishes live fetches from cached reuse. A live fetch pr Telemetry consumers may weight cached and live groundings differently. An agent may cache an article for days or weeks, grounding it in multiple sessions from a single retrieval. A single retrieval produces one `content_retrieved` event but potentially many `content_grounded` events across subsequent sessions. -Agents SHOULD preserve the `license_ref` from the original retrieval when emitting cached grounding events. Without this, telemetry consumers cannot link cached usage to the licence that authorised the original access. +Agents SHOULD preserve the `license_ref` from the original retrieval when emitting cached grounding events. Without this, telemetry consumers cannot link cached usage to the grant referenced at the original access. #### Freshness and verification diff --git a/telemetry-session.json b/telemetry-session.json index 9804802..79cc60c 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -100,7 +100,7 @@ }, "license_ref": { "type": ["string", "null"], - "description": "Reference to the content access licence (JWT jti, CoMP package ID, or opaque identifier)" + "description": "Reference to a licence or grant the emitter associates with this event (JWT jti, CoMP package ID, or opaque identifier). Core does not resolve or interpret it. See section 5.2.3." }, "turn": { "oneOf": [{ "$ref": "#/$defs/ConversationTurn" }, { "type": "null" }], From 84a9b41e4877ab513b2fb2df5da2d284870dd683 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Tue, 28 Jul 2026 18:29:23 +0100 Subject: [PATCH 05/38] State the inference-time scope and the non-emitting-intermediary limit Consultation issues #13 and #5, both documentation dispositions. Section 1.3 named model training as a non-goal but said nothing about fine-tuning, embeddings or index construction, so readers could not tell which side of the line those fell. New section 1.3.1 draws the line at construction against use: assembling a corpus, fine-tuning, computing embeddings and building an index are outside scope, while querying such a store during a response is an ordinary grounding event. It also states plainly that nothing here reports whether a model was trained on a work. Section 4.4 gains the limit raised by the scraper-resale path in #5. Where the intermediary in the middle does not emit, the agent's events are the only record and carry only what the agent was told. Consumers should treat the path back to the content owner as unestablished rather than infer it. The provenance fields for the same issue are PR #10 and are not touched here. Co-Authored-By: Claude Opus 5 (1M context) --- SPECIFICATION.md | 18 ++++++++++++++++-- 1 file changed, 16 insertions(+), 2 deletions(-) diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 746e4a4..9905a9e 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -6,7 +6,7 @@ ## Contents -1. [Introduction](#1-introduction) - problem, goals, non-goals, relationship to access protocols, conventions +1. [Introduction](#1-introduction) - problem, goals, non-goals, inference-time scope, relationship to access protocols, conventions 2. [Normative references](#2-normative-references) 3. [Terms and definitions](#3-terms-and-definitions) 4. [Concepts](#4-concepts) - roles, sessions, event lifecycle, source roles, content identification @@ -53,9 +53,17 @@ Content Telemetry does not: - Mandate specific privacy policies (left to agreements between parties) - Require specific transport protocols (HTTP, gRPC, etc. all valid) - Define content access or licensing protocols (see 1.4) -- Model content usage for model training. The five-stage lifecycle covers inference-time usage only. The `bot_category` field on retrieval events (section 6.2) can distinguish training crawls from inference fetches, but training-specific telemetry is out of scope. +- Model what a system does with content other than at inference time. Assembling a training corpus, training or fine-tuning a model, computing embeddings and constructing a retrieval index are all outside scope (see 1.3.1). - Define accreditation tiers, conformance marks, or community-specific conformance requirements. These belong in profiles layered on this specification (see [GOVERNANCE.md](./GOVERNANCE.md)). +#### 1.3.1 Inference-time scope + +The five-stage lifecycle reports content use observable at inference time: identified content entered a generation context for a particular response, and what the resulting output did with it. + +Retrieval is the boundary case. A crawl whose purpose is training or index building can be reported as a `content_retrieved` event, and `bot_category` (section 6.2) distinguishes it, but the event is non-attributable: no grounding, citation, presentation or engagement follows it. What the system then does with the content, whether it enters a training corpus, a fine-tuning set, an embedding store or a search index, is outside this specification. Nothing here reports that a model was trained on a work, and a conforming implementation says nothing either way about it. + +Using such a store at inference time is inside scope. When an index built over a content owner's material is queried during a response and returns content that grounds the answer, that is a `content_grounded` event like any other, with `source_role: index` on the retrieval that served it (section 4.4). The line is between constructing a derived artefact and using one to answer a query, not whether an index was involved. + ### 1.4 Relationship to content access protocols Content access protocols govern how AI agents discover and license content. Examples include peek-then-pay (HTTP 203 previews with JWT licensing), IAB CoMP (content package negotiation), and bilateral API agreements. @@ -253,6 +261,12 @@ A marketplace operating as both emitter and telemetry consumer receives telemetr When multiple observers report the same retrieval, events are correlated using the `Content-Telemetry-ID` header (see section 7.2). A retrieval corroborated by multiple sources is a stronger signal than either alone. An uncorroborated origin- or edge-reported retrieval (no matching agent event) may indicate a scraper that does not support the telemetry protocol, or missing header propagation. +#### Supply paths with a non-emitting intermediary + +Telemetry cannot describe a supply path whose middle does not emit. Where an agent obtains content from an intermediary that is not a telemetry participant, the agent's events are the only record, and they carry what the agent was told: usually a URL or identifier supplied by that intermediary. Core provides no way to establish from telemetry alone that the intermediary held the content lawfully, or that the content owner ever served it. + +An origin or edge event correlated by `Content-Telemetry-ID` is what closes the gap, and it exists only where the content owner observed the original request. Where it is absent, a consumer SHOULD treat the path back to the content owner as unestablished rather than infer it from the agent's report. This is a limit of the observation model, not a defect in the emitter: an agent reporting honestly cannot supply evidence about a party it did not observe. + ### 4.5 Content identification Events identify content using at least one of two fields: From 70d6a2a7acc2a79ae01fc02ccd3c645a2948176c Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Wed, 29 Jul 2026 12:54:12 +0100 Subject: [PATCH 06/38] Narrow ip_hash prohibition to the v1 migration rule --- SPECIFICATION.md | 2 +- tests/validate.py | 20 ++++++++++---------- 2 files changed, 11 insertions(+), 11 deletions(-) diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 9197b4d..36e6ffa 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -1012,7 +1012,7 @@ Emitters SHOULD: Hashing does not anonymise a value drawn from a space small enough to enumerate. The entire IPv4 address space can be hashed and compared against a candidate digest on commodity hardware, so a hashed IP address is a pseudonym rather than an anonymous value and should be treated as personal data. Version 0.1 defined an `ip_hash` field in the edge and origin data profiles (sections 6.2 and 6.3). Version 1 withdraws it, and emitters MUST NOT populate it. -The schemas cannot enforce this: event `data` accepts additional properties by design, so a withdrawn field validates as an ordinary extension. The conformance suite checks it at the application layer instead. +The schemas cannot enforce this v1 migration rule: event `data` accepts additional properties by design, so `ip_hash` would otherwise validate as an ordinary extension. The conformance suite therefore checks this specific prohibition at the application layer. This does not establish a general registry of withdrawn extension names. ### 9.2 Recommended levels diff --git a/tests/validate.py b/tests/validate.py index 3909491..a1755ec 100644 --- a/tests/validate.py +++ b/tests/validate.py @@ -103,12 +103,12 @@ ), } -# Fields withdrawn in v1 that emitters MUST NOT populate (section 9.1). The -# schemas cannot catch these: event `data` accepts additional properties by -# design, so a withdrawn field validates as an ordinary extension unless it is -# checked here. -WITHDRAWN_EVENT_DATA_FIELDS = { - "ip_hash": "withdrawn in v1; hashing does not anonymise an IP address", +# V0.1 fields prohibited by the v1 migration rule (section 9.1). This is a +# compatibility check for the v1 transition, not a general registry of every +# field the specification may ever withdraw. The schemas cannot catch it: +# event `data` accepts additional properties by design. +V1_PROHIBITED_V0_1_EVENT_DATA_FIELDS = { + "ip_hash": "prohibited by the v1 migration rule; hashing does not anonymise an IP address", } # Event types that carry content and therefore require an identifier @@ -325,9 +325,9 @@ def check_session_or_ctx_token(data): return [] -def check_withdrawn_fields(data): +def check_v1_migration_prohibitions(data): """ - Check that no event carries a field withdrawn in v1 (section 9.1). + Check v1's explicit prohibitions on fields carried forward from v0.1. Returns a list of violation descriptions. """ violations = [] @@ -335,7 +335,7 @@ def check_withdrawn_fields(data): event_data = event.get("data") if not isinstance(event_data, dict): continue - for field, reason in WITHDRAWN_EVENT_DATA_FIELDS.items(): + for field, reason in V1_PROHIBITED_V0_1_EVENT_DATA_FIELDS.items(): if field in event_data: violations.append( f"Event '{event.get('type')}' carries '{field}' in data ({reason})" @@ -349,7 +349,7 @@ def check_application_layer(data): check_privacy_conformance(data) + check_content_identifier(data) + check_session_or_ctx_token(data) - + check_withdrawn_fields(data) + + check_v1_migration_prohibitions(data) ) From 5d98aab52adad296bcc4bbfe47059c664f60f8f1 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Wed, 29 Jul 2026 12:54:32 +0100 Subject: [PATCH 07/38] Define the portable character count precisely --- SPECIFICATION.md | 4 ++-- telemetry-session.json | 2 +- 2 files changed, 3 insertions(+), 3 deletions(-) diff --git a/SPECIFICATION.md b/SPECIFICATION.md index a8c9668..083d103 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -415,7 +415,7 @@ These are the recommended values. Platforms with additional product surfaces (co An emitter that populates a conversation turn MUST NOT include a field above that turn's declared `privacy_level` - for example, `query_text` MUST NOT be present when `privacy_level` is `intent` or `minimal`. This restriction is a property of `privacy_level` itself: it applies wherever conversation turns are emitted, independent of the emitter's conformance level. -**Token counts** includes `query_tokens` and `response_tokens`. These are available at all levels because they are needed for token-based counting models and do not reveal user intent or platform strategy. They carry the same portability limit as `tokens_ingested` (section 6.4): both are measured in the emitter's own tokeniser and are not comparable between agents. +**Token counts** includes `query_tokens` and `response_tokens`. These are available at all levels because they are needed for token-based counting models and do not reveal user intent or platform strategy. They carry the same portability limit as `tokens_ingested` (section 6.4): both are measured in the emitter's own tokeniser and are not comparable between agents. Version 1 does not define corresponding turn-level character counts; `chars_ingested` measures source content placed in a generation context, not query or response length. **Response classification** includes `response_type` (e.g., `"recommendation"`, `"explanation"`). Available at `intent` level and above, as it can reveal the nature of the user's query. @@ -583,7 +583,7 @@ Emitting a `training`-category `content_retrieved` event is permitted but non-at Both fields measure the content actually placed in the generation model's context. For chunked retrieval, count only the portion used, not the full source document. -`chars_ingested` counts Unicode code points in that content. It is the portable measure: two emitters counting the same text agree, so a content owner can compare volumes across agents and over time without knowing which model produced the number. +`chars_ingested` counts Unicode code points in the exact text placed in context. Count the string as ingested: an emitter MUST NOT apply Unicode normalisation solely to calculate this field. It is the portable measure: two emitters that ingest the same code-point sequence agree, so a content owner can compare volumes across agents and over time without knowing which model produced the number. Different normalised representations remain different ingested sequences and may therefore produce different counts. `tokens_ingested` counts the same content in the generation model's tokeniser (the model identified in `model_id` on the corresponding `turn_completed` event), not the retrieval or embedding model's tokeniser. It is supplementary. Token counts are model-specific, change when a vendor revises a tokeniser, and are not comparable between agents, so a consumer cannot aggregate them across emitters or treat a difference as a difference in volume. Emitters SHOULD send `chars_ingested` where they send `tokens_ingested`, and consumers that receive only token counts SHOULD record which model produced them. diff --git a/telemetry-session.json b/telemetry-session.json index 3cf2227..22f84ba 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -124,7 +124,7 @@ "properties": { "scope": { "$ref": "#/$defs/GroundingScope" }, "cached": { "type": "boolean" }, - "chars_ingested": { "type": "integer", "minimum": 0, "description": "Unicode code points of content placed in the generation context. Portable across emitters; preferred over tokens_ingested (section 6.4)." }, + "chars_ingested": { "type": "integer", "minimum": 0, "description": "Unicode code points in the exact text placed in the generation context, without normalising solely for counting. Portable across emitters; preferred over tokens_ingested (section 6.4)." }, "tokens_ingested": { "type": "integer", "minimum": 0, "description": "Token count of the same content in the emitter's own tokeniser. Supplementary: not comparable between emitters (section 6.4)." }, "content_version": { "type": "string" }, "content_last_modified": { "type": "string", "format": "date-time" }, From bdca66680ee94df9518da27a1d3c2819bf52cfc4 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Wed, 29 Jul 2026 12:55:16 +0100 Subject: [PATCH 08/38] Add optional parent session correlation --- SPECIFICATION.md | 18 +++++++++++-- telemetry-event-batch.json | 5 ++++ telemetry-event.json | 5 ++++ telemetry-session.json | 5 ++++ tests/valid/session-multi-agent-child.json | 31 ++++++++++++++++++++++ 5 files changed, 62 insertions(+), 2 deletions(-) create mode 100644 tests/valid/session-multi-agent-child.json diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 746e4a4..2bc5992 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -288,6 +288,7 @@ Additional content metadata - version, last-modified timestamp, content hash, me |-------|------|----------|-------------| | `schema_version` | string | Yes | Schema version (e.g., "0.1") | | `session_id` | UUID | Yes | Unique session identifier | +| `parent_session_id` | UUID | No | Immediate parent session that delegated work to this session | | `agent_id` | string | No | Responding agent identifier | | `content_scope` | string | No | Opaque content collection identifier (see 5.1.1) | | `manifest_ref` | string | No | Manifest reference (see 5.1.2 and section 8) | @@ -297,6 +298,19 @@ Additional content metadata - version, last-modified timestamp, content hash, me | `document_type` | string | No | `"session"` for session documents (see section 7.1 for the standalone event and event batch formats) | | `events` | Event[] | No | Ordered list of events | +`parent_session_id` links a delegated session to its immediate parent without +requiring an emitter to disclose the agent system's full internal topology. A +child session retains its own `session_id`, events and conformance obligations. +Emitters MAY omit the link when the relationship is unavailable or its disclosure +is not appropriate. Consumers MUST NOT infer that an unlinked session had no +parent. + +The event boundary does not change in a multi-agent system. Content entering a +sub-agent's generation context is grounded in the child session. A source +reference that appears only in the sub-agent's response to its orchestrator is +not thereby a citation or presentation to the end user; those events require the +corresponding relationship or presentation in the recipient-facing output. + #### 5.1.1 Content scope The `content_scope` field is an opaque identifier that groups sessions by their content access context. Implementers define its meaning: @@ -717,7 +731,7 @@ The schema supports three delivery formats: **Event batch.** Multiple events sharing one session context, delivered together. The envelope carries the same fields as a standalone event, with an `events` array in place of the single `event`. Suitable for emitters that buffer events and flush periodically: edge platforms aggregating detections across requests, or SDKs batching events within a session. -A standalone event carries `document_type`, `schema_version`, and optionally `session_id` alongside the event fields. The `document_type` field distinguishes standalone events from session documents: +A standalone event carries `document_type`, `schema_version`, and optionally `session_id` and `parent_session_id` alongside the event fields. The `document_type` field distinguishes standalone events from session documents: ```json { @@ -739,7 +753,7 @@ A standalone event carries `document_type`, `schema_version`, and optionally `se } ``` -An event batch carries the same envelope fields with `"document_type": "event_batch"` and an `events` array. Envelope-level fields (`session_id`, `ctx_token`, `agent_id`, `started_at`) apply to every event in the batch; events belonging to different sessions MUST be delivered in separate batches or as session documents. +An event batch carries the same envelope fields with `"document_type": "event_batch"` and an `events` array. Envelope-level fields (`session_id`, `parent_session_id`, `ctx_token`, `agent_id`, `started_at`) apply to every event in the batch; events belonging to different sessions MUST be delivered in separate batches or as session documents. ```json { diff --git a/telemetry-event-batch.json b/telemetry-event-batch.json index 3738c14..588c95f 100644 --- a/telemetry-event-batch.json +++ b/telemetry-event-batch.json @@ -21,6 +21,11 @@ "format": "uuid", "description": "Session identifier, applying to every event in the batch. MAY be omitted by origin-side emitters with no session context. REQUIRED for emitters at Grounding conformance or above (unless ctx_token is carried instead - see section 7.1). Events belonging to different sessions MUST be delivered in separate batches." }, + "parent_session_id": { + "type": "string", + "format": "uuid", + "description": "Optional identifier of the immediate parent session that delegated work to this session. Applies to every event in the batch. See section 5.1." + }, "ctx_token": { "type": "string", "description": "Opaque click-token issued by the originating agent, carried in place of session_id on batches of content_engaged events emitted from a landing page after a click-out. Applies to every event in the batch and is resolved by the telemetry consumer to the owning session. See section 7.1. An event MUST carry either session_id or ctx_token at Grounding conformance and above." diff --git a/telemetry-event.json b/telemetry-event.json index a7c80a5..fb21ff7 100644 --- a/telemetry-event.json +++ b/telemetry-event.json @@ -21,6 +21,11 @@ "format": "uuid", "description": "Session identifier. MAY be omitted by origin-side emitters with no session context. REQUIRED for emitters at Grounding conformance or above (unless ctx_token is carried instead - see section 7.1)." }, + "parent_session_id": { + "type": "string", + "format": "uuid", + "description": "Optional identifier of the immediate parent session that delegated work to this session. See section 5.1." + }, "ctx_token": { "type": "string", "description": "Opaque click-token issued by the originating agent, carried on content_engaged events emitted from a landing page after a click-out in place of session_id. Resolved by the telemetry consumer to the owning session (the click manifest). See section 7.1. An event MUST carry either session_id or ctx_token at Grounding conformance and above." diff --git a/telemetry-session.json b/telemetry-session.json index 9804802..0be7200 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -27,6 +27,11 @@ "format": "uuid", "description": "Unique session identifier" }, + "parent_session_id": { + "type": "string", + "format": "uuid", + "description": "Optional identifier of the immediate parent session that delegated work to this session. See section 5.1." + }, "agent_id": { "type": ["string", "null"], "description": "Responding agent identifier" diff --git a/tests/valid/session-multi-agent-child.json b/tests/valid/session-multi-agent-child.json new file mode 100644 index 0000000..1f17456 --- /dev/null +++ b/tests/valid/session-multi-agent-child.json @@ -0,0 +1,31 @@ +{ + "_test_description": "A delegated child session links to its immediate parent. The source enters the child agent's generation context and is grounded there; no citation or presentation is emitted merely because the child returns an internal response to its orchestrator.", + "document_type": "session", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440021", + "parent_session_id": "660e8400-e29b-41d4-a716-446655440020", + "agent_id": "research-subagent-v1", + "started_at": "2026-07-29T09:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-07-29T09:00:02Z", + "turn_id": "child-turn-1", + "content_url": "https://example.org/research/source", + "data": { + "scope": "turn", + "cached": false, + "tokens_ingested": 900 + } + }, + { + "type": "turn_completed", + "timestamp": "2026-07-29T09:00:05Z", + "turn_id": "child-turn-1", + "turn": { + "privacy_level": "minimal", + "response_tokens": 180 + } + } + ] +} From ff1555d3a4cafbfe9c76a0b39e2d613d4315b6d1 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Tue, 4 Aug 2026 08:13:32 +0100 Subject: [PATCH 09/38] Add content_reproduced to the core lifecycle The output constructor's claim that the output artifact contains identified source content - a quotation, an excerpt, or a full copy, verbatim or near-verbatim. Reproduction and citation are sibling output-construction claims: reproduction records the material, citation records the credit; neither implies the other. Covers uncredited reproduction in unpresented output (API delivery), which no existing event can host. Char-first counts mirror the excerpt_chars/excerpt_tokens pairing on citation data. Builds on the v1 presentation semantics change and is prepared pending D1. Co-Authored-By: Claude Fable 5 --- README.md | 8 +- SPECIFICATION.md | 78 +++++++++--- telemetry-session.json | 36 +++++- tests/README.md | 1 + .../invalid/reproduced-missing-output-id.json | 15 +++ tests/invalid/reproduced-missing-type.json | 16 +++ .../session-reproduction-credited-quote.json | 114 ++++++++++++++++++ .../session-reproduction-uncredited.json | 70 +++++++++++ tests/validate.py | 4 +- 9 files changed, 317 insertions(+), 25 deletions(-) create mode 100644 tests/invalid/reproduced-missing-output-id.json create mode 100644 tests/invalid/reproduced-missing-type.json create mode 100644 tests/valid/session-reproduction-credited-quote.json create mode 100644 tests/valid/session-reproduction-uncredited.json diff --git a/README.md b/README.md index d84e06b..b01c18c 100644 --- a/README.md +++ b/README.md @@ -24,11 +24,12 @@ Platforms self-report usage metrics (if they report at all), and content owners ## Telemetry events -Content Telemetry tracks content through five stages: +Content Telemetry tracks content through six stages: ``` Retrieved → content fetched over HTTP (content owner can see this today) Grounded → content loaded into the agent's generation context + Reproduced → content appearing verbatim or near-verbatim in the response Cited → content explicitly referenced in the response Presented → content or a source reference made perceivable on a recipient-facing surface Engaged → user clicked, copied, shared, or directed the agent to act @@ -40,6 +41,7 @@ The gaps between stages show how content was used: - **Retrieval without grounding** - your content was fetched but not used - **Grounding without citation** - your content influenced the answer but you got no credit +- **Reproduction without citation** - your content appeared in the answer without credit - **Citation without engagement** - your content was cited but the user didn't click through The grounding event captures the boundary "this content entered the agent's generation context." It is architecture-neutral and decoupled from retrieval: content cached by the agent for days still produces a grounding event in every session it influences. @@ -50,7 +52,7 @@ Grounding and presentation record different boundary crossings: grounding means **Post-hoc, not pre-declared.** Events report what actually happened, not what the agent said it would do at request time. An agent cannot reliably declare how it will use content before reading it. -**Observable boundaries, not agent internals.** The five event types mark boundary crossings. What happens between them - the fan-out, relevance evaluation, re-ranking, reasoning chains - is internal to the agent and changes constantly. The spec does not model it. +**Observable boundaries, not agent internals.** The six event types mark boundary crossings. What happens between them - the fan-out, relevance evaluation, re-ranking, reasoning chains - is internal to the agent and changes constantly. The spec does not model it. **Multiple observers, one event.** A content retrieval can be reported by the content owner's CDN, the content owner's origin server, and the AI agent independently. The `Content-Telemetry-ID` header correlates these into a single corroborated event. Uncorroborated retrievals (no matching agent event) may indicate an agent that does not yet support the telemetry protocol. @@ -160,7 +162,7 @@ Comment is most useful on: - The [open questions below](#open-questions-in-v01). - Whether the conformance and privacy levels (sections 5.5 and 5.7) are implementable as written by a team building an emitter or consumer. -- How the five-stage event model fits real agent architectures (section 6.4). +- How the six-stage event model fits real agent architectures (section 6.4). - Anything that would require an implementer to depend on a particular operator or service to participate. The standard should be implementable from the public schemas alone. - Any worked example that does not validate against its schema, or any mismatch between the prose and the schemas. diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 265be15..5dc0616 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -11,7 +11,7 @@ 3. [Terms and definitions](#3-terms-and-definitions) 4. [Concepts](#4-concepts) - roles, sessions, event lifecycle, source roles, content identification 5. [Schema](#5-schema) - session, event, event types, conversation turn, privacy, intent, conformance levels -6. [Data profiles](#6-data-profiles) - retrieval, edge enrichment, origin enrichment, grounding, citation, presentation, engagement +6. [Data profiles](#6-data-profiles) - retrieval, edge enrichment, origin enrichment, grounding, citation, reproduction, presentation, engagement 7. [Transport](#7-transport) - delivery formats, Content-Telemetry-ID header 8. [Manifest](#8-manifest) - discovery, schema, operator, keys, telemetry, domains 9. [Privacy](#9-privacy) - data minimisation, recommended levels, retention @@ -53,7 +53,7 @@ Content Telemetry does not: - Mandate specific privacy policies (left to agreements between parties) - Require specific transport protocols (HTTP, gRPC, etc. all valid) - Define content access or licensing protocols (see 1.4) -- Model content usage for model training. The five-stage lifecycle covers inference-time usage only. The `bot_category` field on retrieval events (section 6.2) can distinguish training crawls from inference fetches, but training-specific telemetry is out of scope. +- Model content usage for model training. The six-stage lifecycle covers inference-time usage only. The `bot_category` field on retrieval events (section 6.2) can distinguish training crawls from inference fetches, but training-specific telemetry is out of scope. - Define accreditation tiers, conformance marks, or community-specific conformance requirements. These belong in profiles layered on this specification (see [GOVERNANCE.md](./GOVERNANCE.md)). ### 1.4 Relationship to content access protocols @@ -97,6 +97,7 @@ For the purposes of this specification, the following terms apply. | **content owner** | entity that owns or licences content accessed by an AI agent | | **agent operator** | entity running the AI agent that uses content | | **grounding** | content entering the generation model's context, the boundary where content can directly influence output (section 4.3) | +| **reproduction** | verbatim or near-verbatim appearance of identified source content in an output artifact, independent of credit and delivery (section 4.3) | | **source role** | classification of the observer reporting a retrieval event: `origin`, `edge`, `index`, or `agent` (section 4.4) | | **privacy level** | data sharing tier controlling which conversation fields are populated: `full`, `summary`, `intent`, or `minimal` (section 5.5) | | **conformance level** | emitter capability tier: Retrieval, Grounding, or Citation (section 5.7) | @@ -168,6 +169,7 @@ Session │ ├── turn_started │ ├── content_retrieved (HTTP layer) │ ├── content_grounded (influence layer) +│ ├── content_reproduced (response layer) │ ├── content_cited (response layer) │ ├── content_presented (recipient-facing surface) │ ├── turn_completed @@ -180,7 +182,7 @@ These are the event types a session can contain, not a strict ordering: events a ### 4.3 Event lifecycle -Content moves through five stages during an agent interaction: +Content moves through six stages during an agent interaction: 1. **Retrieved** - Content fetched over HTTP from an origin server, CDN, marketplace, or index. This is an infrastructure event observable by the content owner's infrastructure (origin server, edge network) and the agent. A retrieval may be cached by the agent for use across multiple sessions. @@ -190,36 +192,44 @@ Content moves through five stages during an agent interaction: Grounding is architecture-neutral: same event whether the agent uses RAG, chain-of-thought reasoning, embeddings, or multi-step delegation (see section 6.4 for architecture-specific guidance). Grounding is decoupled from retrieval: content may be grounded from a live fetch, from agent-side cache, or from a pre-loaded index. Only the agent can report grounding events. -3. **Cited** - An output artifact explicitly associates identified source content with a response, claim, passage, quotation, or other output element. Citation is an output-construction relationship, not evidence that the output was delivered. A subset of grounded content is commonly cited, but a citation can also be emitted without a matching grounding event when an agent produces an uncorroborated or hallucinated source association. +3. **Reproduced** - The output artifact contains identified source content: a quotation, an excerpt, or a full copy, verbatim or near-verbatim. Reproduction is an output-construction claim by the system that built the output. It is independent of credit and of delivery: reproduced content may or may not also be cited, and the artifact may or may not later be presented. An uncredited excerpt in a response delivered through an API produces a `content_reproduced` event and nothing else - without this event, that use would be unreportable. -4. **Presented** - Content or a source reference was rendered, played, spoken, embedded, or otherwise made perceivable on a recipient-facing surface. Presentation does not assert that a person noticed or attended to it. `presentation_kind` distinguishes source content (including a reproduced excerpt or media) from a source reference (such as a link, credit, or card). Not all citations are presented: an output can be stored, suppressed, or passed to another system before delivery. +4. **Cited** - An output artifact explicitly associates identified source content with a response, claim, passage, quotation, or other output element. Citation is an output-construction relationship, not evidence that the output was delivered. A subset of grounded content is commonly cited, but a citation can also be emitted without a matching grounding event when an agent produces an uncorroborated or hallucinated source association. + + Reproduction and citation are sibling claims about the same artifact: reproduction records the material, citation records the credit. Neither implies the other. A credited quotation produces both events; an uncredited excerpt produces only a reproduction; a reference citation with no quoted material produces only a citation. + +5. **Presented** - Content or a source reference was rendered, played, spoken, embedded, or otherwise made perceivable on a recipient-facing surface. Presentation does not assert that a person noticed or attended to it. `presentation_kind` distinguishes source content (including a reproduced excerpt or media) from a source reference (such as a link, credit, or card). Not all citations are presented: an output can be stored, suppressed, or passed to another system before delivery. Grounding and presentation record different boundary crossings: grounding records entry into a generation context, while presentation records a recipient-facing delivery occurrence. As agent experiences evolve beyond the chat window the two diverge - content can shape an answer whose source is never presented, and an agent can present content that never entered a generation context (see *Departures from the funnel model* below). -5. **Engaged** - The recipient or agent performed an observable action on a presentation: clicked a link, expanded a preview, copied text, shared the response, or directed the agent to act on the content. It does not imply attention beyond the reported action. `presentation_id` links the action to the exact presentation occurrence; a click-out can also carry a `ctx_token` that a destination resolves to the session's click manifest (section 7.1). +6. **Engaged** - The recipient or agent performed an observable action on a presentation: clicked a link, expanded a preview, copied text, shared the response, or directed the agent to act on the content. It does not imply attention beyond the reported action. `presentation_id` links the action to the exact presentation occurrence; a click-out can also carry a `ctx_token` that a destination resolves to the session's click manifest (section 7.1). ``` Retrieved (HTTP layer, cacheable) → Grounded (influence layer, per-session or per-turn) - → Cited (response layer, per-turn) + → Reproduced (response layer, per-turn) ⎫ sibling output-construction + → Cited (response layer, per-turn) ⎭ claims; neither implies the other → Presented (recipient-facing surface, per-turn) → Engaged (user action layer) ``` -Each stage is typically a progressively narrower subset. The ratios between stages are meaningful for potential attribution: +Each stage after retrieval is typically a progressively narrower subset, except that reproduction and citation are siblings at the response layer rather than steps in the chain. The ratios between stages are meaningful for potential attribution: - **Retrieval-to-grounding** measures content fetched but not used (irrelevant, stale, or a competing source was preferred) - **Grounding-to-citation** measures content that influenced the response without explicit attribution +- **Reproduction-to-citation** measures reproduced material without an explicit source association - the uncredited-reproduction rate - **Citation-to-presentation** measures source associations constructed in output but not made perceivable - **Presentation-to-engagement** measures observable actions on exact presentation occurrences #### Departures from the funnel model -Three cases break the strict subset model: +Five cases break the strict subset model: - **Presented without cited.** An agent may present content references (e.g., a "Sources" sidebar) without semantically associating them with a response element. In this case, a `content_presented` event exists with no corresponding `content_cited` event. - **Cited without grounded.** A hallucinated citation references content the agent never retrieved or loaded into context. The `content_cited` event has no preceding `content_grounded` event. Telemetry consumers SHOULD treat uncorroborated citations (no matching grounding event) as lower-confidence signals. - **Presented without grounded.** An agent can present content without that content entering a generation context: an agentic browser showing a page or an embedded video played on a response surface. A `content_presented` event (typically `presentation_kind: content` and `presentation_type: embed`) exists with no corresponding `content_grounded` event. +- **Reproduced without cited.** An output contains source material with no explicit source association: an uncredited excerpt. The `content_reproduced` event exists with no corresponding `content_cited` event. This is the primary case the reproduction event exists to record. +- **Reproduced without grounded.** An output can reproduce content that never entered this session's generation context, most commonly content the model memorised during training. When the emitter can identify the source, the `content_reproduced` event stands without a grounding event; telemetry consumers SHOULD treat it, like an uncorroborated citation, as a lower-confidence signal. These cases are valid. Emitters SHOULD produce the events that reflect what actually happened, even when the result does not follow the typical funnel ordering. @@ -230,7 +240,7 @@ Conversation turns overlay this lifecycle: 1. **Turn started** - user submits a query 2. **Turn completed** - agent finishes response -A single grounding event with session scope influences all subsequent turns. Citation, presentation, and engagement events occur within specific turns. +A single grounding event with session scope influences all subsequent turns. Reproduction, citation, presentation, and engagement events occur within specific turns. ### 4.4 Source roles @@ -247,7 +257,7 @@ The `origin` and `edge` source roles enable content owners to report AI agent tr A marketplace operating as both emitter and telemetry consumer receives telemetry from platforms (as a consumer), resolves content owner identity from `content_id` or `content_url`, and generates per-content-owner usage reports. The marketplace's own `source_role: index` events provide a corroboration layer - it can cross-reference what it served against what platforms reported using. -`content_grounded`, `content_cited`, and `content_presented` events are reported by the agent (or agent operator) only. These events describe what happened inside the agent, during output construction, or on a recipient-facing surface, which is not observable from the content owner's infrastructure. +`content_grounded`, `content_reproduced`, `content_cited`, and `content_presented` events are reported by the agent (or agent operator) only. These events describe what happened inside the agent, during output construction, or on a recipient-facing surface, which is not observable from the content owner's infrastructure. A third party that detects reproduced content in a delivered output is corroborating or contradicting the emitter's claims, not observing construction; detection results belong to verification tooling, not to these event types. `content_engaged` events are usually reported by the agent for in-product interactions. For a click-out to a landing page, a downstream marketplace, affiliate network, or destination site MAY report a corroborating `content_engaged` event using `ctx_token` in place of `session_id` (section 7.1). @@ -338,7 +348,7 @@ Format: the URL of a manifest served at `/.well-known/content-telemetry.json` un #### 5.2.1 Turn association -The `turn_id` field associates content events with a specific conversation turn. Emitters SHOULD set `turn_id` on `content_cited`, `content_presented`, and `content_engaged` events. Emitters SHOULD also set `turn_id` on `content_grounded` events when `scope` is `turn`. The corresponding `turn_started` and `turn_completed` events SHOULD carry the same `turn_id`. +The `turn_id` field associates content events with a specific conversation turn. Emitters SHOULD set `turn_id` on `content_reproduced`, `content_cited`, `content_presented`, and `content_engaged` events. Emitters SHOULD also set `turn_id` on `content_grounded` events when `scope` is `turn`. The corresponding `turn_started` and `turn_completed` events SHOULD carry the same `turn_id`. `turn_id` is scoped to the session. Format is emitter-defined (sequential integers, UUIDs, or any opaque string). @@ -360,9 +370,10 @@ The `license_ref` field connects a telemetry event to the content access licence |------|-------------|-----------------| | `content_retrieved` | Content fetched from source | `content_url`, `source_role`, `data.media_type` | | `content_grounded` | Content loaded into agent context | `content_url` or `content_id`, `data.scope`, `data.cached` | +| `content_reproduced` | Source content appears verbatim or near-verbatim in an output artifact | `id`, `output_id`, `content_url` or `content_id`, `data.reproduction_type` | | `content_cited` | Output explicitly associates source content with an output element | `id`, `output_id`, `content_url` or `content_id`, `data.citation_type` | | `content_presented` | Content or a source reference was made perceivable | `id`, `output_id`, `content_url` or `content_id`, `data.presentation_kind`, `data.presentation_type` | -| `content_engaged` | Observable action on an exact presentation | `presentation_id`, `content_url` or `content_id`, `data.engagement_type` (see 6.7) | +| `content_engaged` | Observable action on an exact presentation | `presentation_id`, `content_url` or `content_id`, `data.engagement_type` (see 6.8) | #### Conversation events @@ -492,6 +503,7 @@ The privacy-level field restriction (section 5.5) applies to Citation emitters a A Citation emitter SHOULD: - Emit `content_presented` and `content_engaged` events when applicable +- Emit `content_reproduced` events whenever the response reproduces identified source content, including alongside every `direct_quote` citation (section 6.6) - Include `data.position` on citation events - Include `output_element_id` when the cited or presented element has a stable identity - Include `citation_id` on a presentation of a cited source association @@ -530,7 +542,7 @@ When the reporter is the agent (`source_role: agent`), the following fields are `media_type` on retrieval events allows content owners to see what types of content are being fetched, independent of whether those retrievals result in grounding or citation. Defaults to `text` when absent. -`text`, `image`, `video`, and `audio` are the core values. Emitters MAY use custom string values for media outside the core set (for example `3d` or `dataset`). Telemetry consumers MUST tolerate unknown `media_type` values. This rule applies to `media_type` on every event type that carries it (sections 6.4, 6.5, 6.6). +`text`, `image`, `video`, and `audio` are the core values. Emitters MAY use custom string values for media outside the core set (for example `3d` or `dataset`). Telemetry consumers MUST tolerate unknown `media_type` values. This rule applies to `media_type` on every event type that carries it (sections 6.4, 6.5, 6.6, 6.7). ### 6.2 Edge enrichment (`content_retrieved` + `source_role: edge`) @@ -658,7 +670,32 @@ The `unclassified` value for `citation_type` indicates the agent did not classif When `content_hash` is absent or does not match any grounding event's hash (for example, because the agent re-chunked content between grounding and citation), consumers SHOULD fall back to matching on `content_url` or `content_id`, accepting that the correlation may be imprecise when the same content appears in multiple grounding events. -### 6.6 Presentation data (`content_presented`) +A `direct_quote` citation records the credit; the reproduced material itself is recorded by a companion `content_reproduced` event (section 6.6). Emitters SHOULD emit both, sharing `output_element_id`, so that reproduction totals can be computed over reproduction events alone without unioning citation types. + +### 6.6 Reproduction data (`content_reproduced`) + +| Field | Type | Description | +|-------|------|-------------| +| `reproduction_type` | string | Fidelity of the reproduction: `verbatim`, `near_verbatim`, `unclassified` | +| `media_type` | string | Content medium: `text`, `image`, `video`, `audio` (open vocabulary, see 6.1). Defaults to `text` when absent. | +| `reproduced_chars` | integer | Unicode code points in the reproduced span as it appears in the output | +| `reproduced_tokens` | integer | Token count of the reproduced span (supplementary) | +| `reproduced_hash` | string | SHA-256 of the reproduced span as it appears in the output (`sha256:{hex}`) | +| `content_hash` | string | SHA-256 matching the corresponding `content_grounded` event (`sha256:{hex}`) | + +A `content_reproduced` event is the output constructor's claim that the output artifact contains identified source content: a quotation, an excerpt, or a full copy. It records construction only. Whether the reproduction was credited is recorded by a `content_cited` event; whether it reached a person is recorded by a `content_presented` event. The event exists so that reproduction remains reportable when neither occurred - an uncredited excerpt in a response delivered through an API is the canonical case, and produces a `content_reproduced` event with no citation and no presentation. + +Only the system that constructed the output emits this event (section 4.4). + +`verbatim` means the reproduced span matches the source exactly. `near_verbatim` means it differs only by bounded surface edits: truncation, elision marks, whitespace, punctuation, or case. Paraphrase is not reproduction - a credited paraphrase is a `content_cited` event with `citation_type: paraphrase`, and an uncredited one is silent grounding (section 10.2). A translation is likewise not a reproduction of the source text. `unclassified` indicates the emitter identified reproduced content without classifying its fidelity. Reproduction is not limited to excerpts: a full copy is the same event whose span is the entire work. + +`reproduced_chars` counts Unicode code points in the exact reproduced span as it appears in the output. `reproduced_tokens` is the agent-native supplementary measurement, following the same pairing as `excerpt_chars` and `excerpt_tokens` (section 6.5). `reproduced_hash` is the SHA-256 of the reproduced span as produced - for a `verbatim` reproduction it matches a hash of the corresponding source span, and for a credited quotation of the same span it equals the citation's `excerpt_hash`. `content_hash` correlates the reproduction to the grounding event it drew from, with the same fallback rules as citation data (section 6.5). + +Each `content_reproduced` event MUST have an `id` and `output_id`, and SHOULD carry `output_element_id` identifying the passage or element containing the reproduction. When the reproduction is credited, the event MAY carry `citation_id` referencing the crediting `content_cited` event, and the two SHOULD share `output_element_id`. When reproduced content is later made perceivable, that occurrence is a `content_presented` event with `presentation_kind: content` sharing the same `output_id`; presentation does not re-assert reproduction, and reproduction does not assert presentation. + +For non-text media, `media_type` and `reproduced_hash` identify the reproduced material. Finer-grained portion references for time-based and spatial media (time ranges, regions, segments) are not defined in this version. + +### 6.7 Presentation data (`content_presented`) | Field | Type | Description | |-------|------|-------------| @@ -688,7 +725,7 @@ Each presentation event MUST have an `id` and `output_id`. When it presents a ci When a session includes `content_presented` events but no subsequent `content_engaged` events, the telemetry establishes only that content or a reference was made perceivable and no reported interaction followed. It does not establish human attention. Whether this pattern is meaningful depends on the governing terms. Retrieval remains the only lifecycle stage observable from the CDN edge. -### 6.7 Engagement data (`content_engaged`) +### 6.8 Engagement data (`content_engaged`) | Field | Type | Description | |-------|------|-------------| @@ -1067,6 +1104,7 @@ Whether this constitutes one royalty event, three, or ten depends on the commerc | Per-grounding | One event per article entering context per session | Access-based or flat-fee licensing ("you used our content") | | Per-citation | One event per explicit reference in a response | Performance-based licensing ("you cited our content") | | Per-turn-influenced | One event per turn where content was in context | Usage-based licensing ("our content informed N answers") | +| Per-reproduction | One event per reproduced portion appearing in output | Excerpt-based licensing ("N characters of our content appeared in answers") | The `content_grounded` event with `scope: session` plus the count of subsequent `turn_completed` events provides the inputs for all three models without requiring the schema to embed a commercial opinion. @@ -1075,6 +1113,7 @@ The `content_grounded` event with `scope: session` plus the count of subsequent Content can influence every response in a session without being explicitly cited. A common royalty formula (individual content owner usage / total content owner usage x royalty rate) can be applied at any level of the funnel: - At the **grounding** level: counts all content that was in the agent's context, regardless of citation. This captures the full extent of content influence, including silent grounding. +- At the **reproduction** level: counts source material appearing in the output, credited or not. This captures verbatim reuse that citation-level counting misses, with `reproduced_chars` providing a magnitude. - At the **citation** level: counts only explicitly attributed content. Simpler to verify but undercounts content influence. - At the **presentation** level: counts content or source references made perceivable. It does not prove attention. @@ -1156,6 +1195,13 @@ exists and `citation_id` when the presentation carries a citation. For every event. Do not migrate clicks by matching URL alone: repeated presentations of the same URL are distinct occurrences. +V1 adds `content_reproduced`. The v0.1 preview has no equivalent: verbatim reuse +was inferable only from `direct_quote` citations, which conflate the credit with +the material and cannot record an uncredited copy. When migrating historical +v0.1 data, consumers MAY treat a `direct_quote` citation as an implied +reproduction of its excerpt. V1 emitters record reproduction explicitly and +SHOULD NOT rely on that inference. + Preview versions (0.x) use two-component version numbers. From 1.0.0 onward, versions follow [semantic versioning](https://semver.org/): - **Major** (1.0.0 → 2.0.0) - breaking changes to required fields diff --git a/telemetry-session.json b/telemetry-session.json index bfe8e4d..d10b249 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -78,22 +78,22 @@ }, "turn_id": { "type": ["string", "null"], - "description": "Associates this event with a conversation turn. Scoped to the session. SHOULD be set on turn_started, turn_completed, content_cited, content_presented, content_engaged events, and content_grounded events when scope is turn." + "description": "Associates this event with a conversation turn. Scoped to the session. SHOULD be set on turn_started, turn_completed, content_reproduced, content_cited, content_presented, content_engaged events, and content_grounded events when scope is turn." }, "output_id": { "type": "string", "minLength": 1, - "description": "Opaque identifier for the output artifact. REQUIRED on content_cited and content_presented events so output construction and later presentation can be correlated across services or times." + "description": "Opaque identifier for the output artifact. REQUIRED on content_reproduced, content_cited, and content_presented events so output construction and later presentation can be correlated across services or times." }, "output_element_id": { "type": "string", "minLength": 1, - "description": "Opaque identifier for the element within output_id that carries the citation or presentation, such as a passage, media track, caption, link, or card." + "description": "Opaque identifier for the element within output_id that carries the reproduction, citation, or presentation, such as a passage, media track, caption, link, or card." }, "citation_id": { "type": "string", "format": "uuid", - "description": "The id of the content_cited event associated with this presentation. Valid only on content_presented events and absent when the presentation is not a citation." + "description": "The id of the content_cited event associated with this presentation or reproduction. Valid only on content_presented and content_reproduced events; absent when the presentation is not a citation or the reproduction is uncredited." }, "presentation_id": { "type": "string", @@ -177,6 +177,28 @@ } } }, + { + "if": { + "properties": { "type": { "const": "content_reproduced" } }, + "required": ["type"] + }, + "then": { + "required": ["id", "output_id", "data"], + "properties": { + "data": { + "required": ["reproduction_type"], + "properties": { + "reproduction_type": { "$ref": "#/$defs/ReproductionType" }, + "media_type": { "$ref": "#/$defs/MediaType" }, + "reproduced_chars": { "type": "integer", "minimum": 0, "description": "Unicode code points in the reproduced span as it appears in the output" }, + "reproduced_tokens": { "type": "integer", "minimum": 0 }, + "reproduced_hash": { "type": "string", "pattern": "^sha256:[a-f0-9]{64}$", "description": "SHA-256 of the reproduced span as it appears in the output (sha256:{hex})" }, + "content_hash": { "type": "string", "pattern": "^sha256:[a-f0-9]{64}$", "description": "SHA-256 matching the corresponding content_grounded event. When the agent chunked the source, this is the chunk hash." } + } + } + } + } + }, { "if": { "properties": { "type": { "const": "content_presented" } }, @@ -247,6 +269,7 @@ "enum": [ "content_retrieved", "content_grounded", + "content_reproduced", "content_cited", "content_presented", "content_engaged", @@ -340,6 +363,11 @@ "description": "How content was used in the response", "enum": ["direct_quote", "paraphrase", "reference", "contradiction", "unclassified"] }, + "ReproductionType": { + "type": "string", + "description": "Fidelity of a reproduction. near_verbatim differs from the source only by bounded surface edits: truncation, elision marks, whitespace, punctuation, or case.", + "enum": ["verbatim", "near_verbatim", "unclassified"] + }, "CitationPosition": { "type": "string", "description": "Prominence of citation in response", diff --git a/tests/README.md b/tests/README.md index 3ff93d2..49c1e3f 100644 --- a/tests/README.md +++ b/tests/README.md @@ -32,6 +32,7 @@ Run from the repository root. Without uv: `pip install jsonschema`, then `python - Standalone event envelopes (CDN edge, agent with session FK) - Privacy level field gating (application-layer conformance) - Funnel exceptions (presented-no-cited, cited-no-grounded, presented-no-grounded) +- Reproduction cases: credited quotation (reproduction + direct_quote citation sharing an output element) and uncredited reproduction in unpresented API output - Text, image, audio, video, suppressed-citation, and repeated-presentation cases - Exact presentation-to-engagement correlation across session, standalone, and batch envelopes - Multi-turn sessions, cached grounding diff --git a/tests/invalid/reproduced-missing-output-id.json b/tests/invalid/reproduced-missing-output-id.json new file mode 100644 index 0000000..4f5405b --- /dev/null +++ b/tests/invalid/reproduced-missing-output-id.json @@ -0,0 +1,15 @@ +{ + "_test_description": "content_reproduced must identify the output artifact containing the reproduction: id and output_id are required so reproduction can be correlated with citation and presentation.", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440130", + "started_at": "2026-08-01T11:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440131", + "type": "content_reproduced", + "timestamp": "2026-08-01T11:00:01Z", + "content_id": "publisher:article:2", + "data": { "reproduction_type": "verbatim", "reproduced_chars": 500 } + } + ] +} diff --git a/tests/invalid/reproduced-missing-type.json b/tests/invalid/reproduced-missing-type.json new file mode 100644 index 0000000..a5c00fb --- /dev/null +++ b/tests/invalid/reproduced-missing-type.json @@ -0,0 +1,16 @@ +{ + "_test_description": "content_reproduced must classify the fidelity of the reproduction: data.reproduction_type is required.", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440120", + "started_at": "2026-08-01T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440121", + "type": "content_reproduced", + "timestamp": "2026-08-01T10:00:01Z", + "output_id": "response:1", + "content_id": "publisher:article:1", + "data": { "reproduced_chars": 300 } + } + ] +} diff --git a/tests/valid/session-reproduction-credited-quote.json b/tests/valid/session-reproduction-credited-quote.json new file mode 100644 index 0000000..ab32217 --- /dev/null +++ b/tests/valid/session-reproduction-credited-quote.json @@ -0,0 +1,114 @@ +{ + "_test_description": "Credited quotation: the same span produces a content_reproduced event (the material) and a content_cited direct_quote event (the credit), sharing output_element_id, with reproduced_hash equal to the citation's excerpt_hash. The quote is then presented inline. Demonstrates that reproduction and citation are sibling output-construction claims joined to one presentation.", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440111", + "agent_id": "research-assistant-v5", + "started_at": "2026-08-01T15:00:00Z", + "ended_at": "2026-08-01T15:00:06Z", + "events": [ + { + "type": "turn_started", + "timestamp": "2026-08-01T15:00:00Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "query_intent": "research", + "topics": ["press freedom", "media law"] + } + }, + { + "type": "content_retrieved", + "timestamp": "2026-08-01T15:00:01Z", + "source_role": "agent", + "content_telemetry_id": "880e8400-e29b-41d4-a716-446655440112", + "content_url": "https://www.example-journal.org/analysis/press-freedom-ruling" + }, + { + "type": "content_grounded", + "timestamp": "2026-08-01T15:00:02Z", + "turn_id": "1", + "content_url": "https://www.example-journal.org/analysis/press-freedom-ruling", + "content_id": "examplejournal:2026:press-freedom-ruling", + "data": { + "scope": "turn", + "cached": false, + "tokens_ingested": 2400, + "content_hash": "sha256:e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6", + "media_type": "text" + } + }, + { + "id": "880e8400-e29b-41d4-a716-446655440113", + "type": "content_cited", + "timestamp": "2026-08-01T15:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:quote:1", + "content_url": "https://www.example-journal.org/analysis/press-freedom-ruling", + "content_id": "examplejournal:2026:press-freedom-ruling", + "data": { + "citation_type": "direct_quote", + "excerpt_tokens": 54, + "excerpt_chars": 236, + "excerpt_hash": "sha256:f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1", + "position": "primary", + "media_type": "text", + "url_verified": true + } + }, + { + "id": "880e8400-e29b-41d4-a716-446655440114", + "type": "content_reproduced", + "timestamp": "2026-08-01T15:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:quote:1", + "citation_id": "880e8400-e29b-41d4-a716-446655440113", + "content_url": "https://www.example-journal.org/analysis/press-freedom-ruling", + "content_id": "examplejournal:2026:press-freedom-ruling", + "data": { + "reproduction_type": "verbatim", + "media_type": "text", + "reproduced_chars": 236, + "reproduced_tokens": 54, + "reproduced_hash": "sha256:f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1", + "content_hash": "sha256:e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6" + } + }, + { + "id": "880e8400-e29b-41d4-a716-446655440115", + "type": "content_presented", + "timestamp": "2026-08-01T15:00:05Z", + "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:quote:1", + "citation_id": "880e8400-e29b-41d4-a716-446655440113", + "content_url": "https://www.example-journal.org/analysis/press-freedom-ruling", + "content_id": "examplejournal:2026:press-freedom-ruling", + "data": { + "presentation_kind": "content", + "presentation_type": "inline_quote", + "media_type": "text" + } + }, + { + "type": "turn_completed", + "timestamp": "2026-08-01T15:00:06Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "query_intent": "research", + "response_type": "analysis", + "response_mode": "standard", + "topics": ["press freedom", "media law"], + "content_urls_retrieved": [ + "https://www.example-journal.org/analysis/press-freedom-ruling" + ], + "content_urls_cited": [ + "https://www.example-journal.org/analysis/press-freedom-ruling" + ], + "response_tokens": 380 + } + } + ] +} diff --git a/tests/valid/session-reproduction-uncredited.json b/tests/valid/session-reproduction-uncredited.json new file mode 100644 index 0000000..10a6fc8 --- /dev/null +++ b/tests/valid/session-reproduction-uncredited.json @@ -0,0 +1,70 @@ +{ + "_test_description": "Uncredited reproduction in unpresented output: an API-delivered response contains a verbatim excerpt of grounded content with no citation and no presentation. The content_reproduced event is the only record that source material appears in the output. Valid per section 4.3 (reproduced without cited).", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440101", + "agent_id": "answers-api-v3", + "started_at": "2026-08-01T09:00:00Z", + "ended_at": "2026-08-01T09:00:04Z", + "events": [ + { + "type": "turn_started", + "timestamp": "2026-08-01T09:00:00Z", + "turn_id": "1", + "turn": { + "privacy_level": "minimal", + "query_tokens": 42 + } + }, + { + "type": "content_retrieved", + "timestamp": "2026-08-01T09:00:01Z", + "source_role": "agent", + "content_telemetry_id": "880e8400-e29b-41d4-a716-446655440102", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + }, + { + "type": "content_grounded", + "timestamp": "2026-08-01T09:00:02Z", + "turn_id": "1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "content_id": "examplenews:2026:rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "tokens_ingested": 1800, + "content_hash": "sha256:c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4", + "media_type": "text" + } + }, + { + "id": "880e8400-e29b-41d4-a716-446655440103", + "type": "content_reproduced", + "timestamp": "2026-08-01T09:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:para:2", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "content_id": "examplenews:2026:rate-decision-analysis", + "data": { + "reproduction_type": "verbatim", + "media_type": "text", + "reproduced_chars": 412, + "reproduced_tokens": 96, + "reproduced_hash": "sha256:d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5", + "content_hash": "sha256:c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4" + } + }, + { + "type": "turn_completed", + "timestamp": "2026-08-01T09:00:04Z", + "turn_id": "1", + "turn": { + "privacy_level": "minimal", + "response_tokens": 240, + "content_urls_retrieved": [ + "https://www.example-news.com/economy/rate-decision-analysis" + ] + } + } + ] +} diff --git a/tests/validate.py b/tests/validate.py index 67ac9e8..421dd1b 100644 --- a/tests/validate.py +++ b/tests/validate.py @@ -102,8 +102,8 @@ # (content_url or content_id) under section 5.7.5. turn_started and # turn_completed are turn events, not content events, and are exempt. CONTENT_EVENT_TYPES = { - "content_retrieved", "content_grounded", "content_cited", - "content_presented", "content_engaged", + "content_retrieved", "content_grounded", "content_reproduced", + "content_cited", "content_presented", "content_engaged", } # Fields that MUST NOT appear at each privacy level (section 5.5). From 5da947958b5ee95e647ee538f04fe3513f8a6edc Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Tue, 4 Aug 2026 08:17:45 +0100 Subject: [PATCH 10/38] Require a resolvable source reference on content_cited A source association with no resolvable reference is not a citation. The JSON Schema now rejects a content_cited event whose content_url and content_id are both absent or null - stricter than the application-layer identifier rule that covers content events generally, because the reference is what makes the credit routable to an owner. Co-Authored-By: Claude Fable 5 --- SPECIFICATION.md | 8 ++++++-- telemetry-session.json | 4 ++++ tests/README.md | 1 + .../cited-missing-source-reference.json | 18 ++++++++++++++++++ .../invalid/cited-null-source-reference.json | 19 +++++++++++++++++++ 5 files changed, 48 insertions(+), 2 deletions(-) create mode 100644 tests/invalid/cited-missing-source-reference.json create mode 100644 tests/invalid/cited-null-source-reference.json diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 746e4a4..6821618 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -192,6 +192,8 @@ Content moves through five stages during an agent interaction: 3. **Cited** - Content explicitly referenced in the agent's response: quoted, paraphrased, or linked. A subset of grounded content. Content can influence every response in a session without being cited once. + A citation MUST carry a resolvable reference to the source it associates: a `content_url` or a `content_id`. A source association with no resolvable reference is not a citation and MUST NOT be emitted as `content_cited`. Unlike other content events, where the identifier requirement is an application-layer rule (section 5.7.5), for `content_cited` it is enforced by the JSON Schema. + 4. **Displayed** - Content presented to the end user: a reference (a link, snippet, inline quote, or preview card) or the content itself embedded in the response surface (an iframe, a page rendered by an agentic browser, an embedded media player). Not all citations result in display (e.g., when the agent uses content internally without surfacing the source). Grounding and display record two different kinds of influence: grounding records that content influenced the agent, display records that it reached the user. As agent experiences evolve beyond the chat window the two diverge - content can shape an answer whose source the user never sees, and an agent can render content to the user that never entered a generation context (see *Departures from the funnel model* below). @@ -481,7 +483,7 @@ Emitters using standalone event delivery (section 7.1) MUST include `agent_id`, A conforming **Citation** emitter MUST satisfy Grounding requirements and also: -- Emit `content_cited` events with `data.citation_type` +- Emit `content_cited` events with `data.citation_type` and a non-null `content_url` or `content_id` (schema-enforced; section 6.5) The privacy-level field restriction (section 5.5) applies to Citation emitters as it does to any emitter producing conversation turns; it is inherited through the Grounding requirements above. @@ -504,7 +506,7 @@ A conforming **telemetry consumer** MUST: The JSON Schema (`telemetry-session.json`) validates structure and types but cannot express every conformance rule. The following are normative requirements verified at the application layer, not by schema validation: -- At least one of `content_url` or `content_id` MUST be present on every content event (section 4.5). +- At least one of `content_url` or `content_id` MUST be present on every content event (section 4.5). For `content_cited` events this requirement is additionally enforced by the JSON Schema, which rejects a citation whose reference is absent or null (section 6.5). - An event MUST carry either `session_id` or `ctx_token` at Grounding conformance and above (section 7.1). - Conversation-turn fields MUST NOT exceed the turn's declared `privacy_level` (section 5.5). - The conformance-level requirements (sections 5.7.1 to 5.7.3) are cumulative. @@ -639,6 +641,8 @@ Agents SHOULD preserve the `license_ref` from the original retrieval when emitti | `content_hash` | string | SHA-256 matching the corresponding `content_grounded` event (`sha256:{hex}`). When the agent chunked the source, this is the chunk hash, not the full document hash. | | `url_verified` | boolean | Whether the cited URL was verified to resolve to matching content | +A citation MUST carry a resolvable source reference: a non-null `content_url` or `content_id` at the event level. This is what distinguishes a citation from vague attribution - the credit names a source that owner routing (section 7.3) can resolve. The JSON Schema enforces this for `content_cited` events; an association the emitter cannot resolve to a URL or identifier is not reportable as a citation. This is stricter than the application-layer identifier rule that applies to content events generally (section 5.7.5). + `media_type` identifies the content medium. Defaults to `text` when absent. `excerpt_tokens` is the agent-native measurement. `excerpt_chars` provides the same information in a unit familiar to content owners and licensors. Emitters SHOULD include both when available. diff --git a/telemetry-session.json b/telemetry-session.json index 9804802..f023839 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -140,6 +140,10 @@ "required": ["type"] }, "then": { + "anyOf": [ + { "required": ["content_url"], "properties": { "content_url": { "type": "string" } } }, + { "required": ["content_id"], "properties": { "content_id": { "type": "string" } } } + ], "properties": { "data": { "properties": { diff --git a/tests/README.md b/tests/README.md index b22d382..1835231 100644 --- a/tests/README.md +++ b/tests/README.md @@ -28,6 +28,7 @@ Run from the repository root. Without uv: `pip install jsonschema`, then `python - Event required fields (`type`, `timestamp`) - Turn required fields (`privacy_level`) - Enum validation (event types, privacy levels, source roles, schema version) +- Citation source-reference requirement (content_cited rejected when content_url/content_id are missing or null) - All three conformance levels (Retrieval, Grounding, Citation) - Standalone event envelopes (CDN edge, agent with session FK) - Privacy level field gating (application-layer conformance) diff --git a/tests/invalid/cited-missing-source-reference.json b/tests/invalid/cited-missing-source-reference.json new file mode 100644 index 0000000..babc716 --- /dev/null +++ b/tests/invalid/cited-missing-source-reference.json @@ -0,0 +1,18 @@ +{ + "_test_description": "content_cited event carrying neither content_url nor content_id. A source association with no resolvable reference is not a citation; unlike other content events, the JSON Schema enforces the identifier requirement for content_cited (section 6.5).", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440140", + "agent_id": "research-assistant-v1", + "started_at": "2026-08-01T12:00:00Z", + "events": [ + { + "type": "content_cited", + "timestamp": "2026-08-01T12:00:01Z", + "turn_id": "1", + "data": { + "citation_type": "reference", + "position": "primary" + } + } + ] +} diff --git a/tests/invalid/cited-null-source-reference.json b/tests/invalid/cited-null-source-reference.json new file mode 100644 index 0000000..ba4cce0 --- /dev/null +++ b/tests/invalid/cited-null-source-reference.json @@ -0,0 +1,19 @@ +{ + "_test_description": "content_cited event with content_url explicitly null and no content_id. Presence of a null reference does not satisfy the citation reference requirement: the schema demands a non-null content_url or content_id on content_cited (section 6.5).", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440141", + "agent_id": "research-assistant-v1", + "started_at": "2026-08-01T12:30:00Z", + "events": [ + { + "type": "content_cited", + "timestamp": "2026-08-01T12:30:01Z", + "turn_id": "1", + "content_url": null, + "data": { + "citation_type": "direct_quote", + "excerpt_chars": 180 + } + } + ] +} From 501bd53198d3f0e647ff544712928441e6e5a6b7 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Tue, 4 Aug 2026 08:18:22 +0100 Subject: [PATCH 11/38] Require a resolvable source reference on content_reproduced A reproduction claim is only meaningful for an identified source. The schema rejects a content_reproduced event whose content_url and content_id are both absent or null, matching the enforcement added for content_cited on v1-citation-source-reference; merge that branch first so the cross-reference in section 6.6 holds. Co-Authored-By: Claude Fable 5 --- SPECIFICATION.md | 2 ++ telemetry-session.json | 4 ++++ .../reproduced-missing-source-reference.json | 15 +++++++++++++++ 3 files changed, 21 insertions(+) create mode 100644 tests/invalid/reproduced-missing-source-reference.json diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 5dc0616..3d17db5 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -691,6 +691,8 @@ Only the system that constructed the output emits this event (section 4.4). `reproduced_chars` counts Unicode code points in the exact reproduced span as it appears in the output. `reproduced_tokens` is the agent-native supplementary measurement, following the same pairing as `excerpt_chars` and `excerpt_tokens` (section 6.5). `reproduced_hash` is the SHA-256 of the reproduced span as produced - for a `verbatim` reproduction it matches a hash of the corresponding source span, and for a credited quotation of the same span it equals the citation's `excerpt_hash`. `content_hash` correlates the reproduction to the grounding event it drew from, with the same fallback rules as citation data (section 6.5). +The claim is only meaningful for an identified source, so a `content_reproduced` event MUST carry a non-null `content_url` or `content_id`. The JSON Schema enforces this, as it does for `content_cited`; reproduction the emitter cannot attribute to an identified source is not reportable as an event. + Each `content_reproduced` event MUST have an `id` and `output_id`, and SHOULD carry `output_element_id` identifying the passage or element containing the reproduction. When the reproduction is credited, the event MAY carry `citation_id` referencing the crediting `content_cited` event, and the two SHOULD share `output_element_id`. When reproduced content is later made perceivable, that occurrence is a `content_presented` event with `presentation_kind: content` sharing the same `output_id`; presentation does not re-assert reproduction, and reproduction does not assert presentation. For non-text media, `media_type` and `reproduced_hash` identify the reproduced material. Finer-grained portion references for time-based and spatial media (time ranges, regions, segments) are not defined in this version. diff --git a/telemetry-session.json b/telemetry-session.json index d10b249..91a4aad 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -184,6 +184,10 @@ }, "then": { "required": ["id", "output_id", "data"], + "anyOf": [ + { "required": ["content_url"], "properties": { "content_url": { "type": "string" } } }, + { "required": ["content_id"], "properties": { "content_id": { "type": "string" } } } + ], "properties": { "data": { "required": ["reproduction_type"], diff --git a/tests/invalid/reproduced-missing-source-reference.json b/tests/invalid/reproduced-missing-source-reference.json new file mode 100644 index 0000000..4c33aa5 --- /dev/null +++ b/tests/invalid/reproduced-missing-source-reference.json @@ -0,0 +1,15 @@ +{ + "_test_description": "content_reproduced event carrying neither content_url nor content_id. A reproduction claim is only meaningful for an identified source; the JSON Schema requires a non-null content_url or content_id on content_reproduced (section 6.6).", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440150", + "started_at": "2026-08-01T13:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440151", + "type": "content_reproduced", + "timestamp": "2026-08-01T13:00:01Z", + "output_id": "response:1", + "data": { "reproduction_type": "verbatim", "reproduced_chars": 250 } + } + ] +} From 900f3569bedaf6321ebac9115239e29daa876820 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Tue, 4 Aug 2026 08:36:56 +0100 Subject: [PATCH 12/38] Align reproduced_chars with the chars_ingested counting rule Co-Authored-By: Claude Fable 5 --- SPECIFICATION.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 6b59afd..183e01e 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -728,7 +728,7 @@ Only the system that constructed the output emits this event (section 4.4). `verbatim` means the reproduced span matches the source exactly. `near_verbatim` means it differs only by bounded surface edits: truncation, elision marks, whitespace, punctuation, or case. Paraphrase is not reproduction - a credited paraphrase is a `content_cited` event with `citation_type: paraphrase`, and an uncredited one is silent grounding (section 10.2). A translation is likewise not a reproduction of the source text. `unclassified` indicates the emitter identified reproduced content without classifying its fidelity. Reproduction is not limited to excerpts: a full copy is the same event whose span is the entire work. -`reproduced_chars` counts Unicode code points in the exact reproduced span as it appears in the output. `reproduced_tokens` is the agent-native supplementary measurement, following the same pairing as `excerpt_chars` and `excerpt_tokens` (section 6.5). `reproduced_hash` is the SHA-256 of the reproduced span as produced - for a `verbatim` reproduction it matches a hash of the corresponding source span, and for a credited quotation of the same span it equals the citation's `excerpt_hash`. `content_hash` correlates the reproduction to the grounding event it drew from, with the same fallback rules as citation data (section 6.5). +`reproduced_chars` counts Unicode code points in the exact reproduced span as it appears in the output, under the same counting rule as `chars_ingested` (section 6.4): no normalisation applied solely for counting. `reproduced_tokens` is the agent-native supplementary measurement, following the same pairing as `excerpt_chars` and `excerpt_tokens` (section 6.5). `reproduced_hash` is the SHA-256 of the reproduced span as produced - for a `verbatim` reproduction it matches a hash of the corresponding source span, and for a credited quotation of the same span it equals the citation's `excerpt_hash`. `content_hash` correlates the reproduction to the grounding event it drew from, with the same fallback rules as citation data (section 6.5). The claim is only meaningful for an identified source, so a `content_reproduced` event MUST carry a non-null `content_url` or `content_id`. The JSON Schema enforces this, as it does for `content_cited`; reproduction the emitter cannot attribute to an identified source is not reportable as an event. From 0c7bf897a46fe5240b30ef15acbfb59a84288094 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Tue, 4 Aug 2026 08:54:44 +0100 Subject: [PATCH 13/38] Fix QA-scrub findings across spec, fixtures and docs - Isolate the cited source-reference anyOf in both negative fixtures (they previously also omitted id/output_id and failed on those) - Update five-stage/three-model/enumeration text that predated the sixth event (1.3.1, 5.2 field table, 5.7.5, 7.3, 10.1, 10.2) - Schema-enforcement prose now names content_reproduced alongside content_cited; section 6 intro no longer claims no data field is required; terms table gains presentation - Reorder session-multi-turn events chronologically; correct stale fixture descriptions and README event names; batch grounding fixture now meets the conformance level it claims - Add invalid fixtures isolating cited output_id, reproduced id and presented presentation_type; add envelope-level parent_session_id coverage; pair chars_ingested with tokens_ingested in new fixtures - tests/README documents the ip_hash application-layer check and no longer claims batch-envelope engagement coverage Co-Authored-By: Claude Fable 5 --- README.md | 2 +- SPECIFICATION.md | 23 ++++++++++--------- tests/README.md | 3 ++- tests/check_examples.py | 1 + tests/invalid/cited-missing-output-id.json | 15 ++++++++++++ .../cited-missing-source-reference.json | 2 ++ .../invalid/cited-null-source-reference.json | 2 ++ .../presented-missing-presentation-type.json | 16 +++++++++++++ tests/invalid/reproduced-missing-id.json | 15 ++++++++++++ tests/valid/event-batch-agent.json | 6 ++++- .../valid/event-standalone-child-session.json | 20 ++++++++++++++++ tests/valid/session-cached-grounding.json | 2 +- tests/valid/session-citation-tier.json | 2 +- tests/valid/session-multi-agent-child.json | 1 + tests/valid/session-multi-turn.json | 22 +++++++++--------- .../session-reproduction-credited-quote.json | 1 + .../session-reproduction-uncredited.json | 1 + 17 files changed, 107 insertions(+), 27 deletions(-) create mode 100644 tests/invalid/cited-missing-output-id.json create mode 100644 tests/invalid/presented-missing-presentation-type.json create mode 100644 tests/invalid/reproduced-missing-id.json create mode 100644 tests/valid/event-standalone-child-session.json diff --git a/README.md b/README.md index 5ddb0b1..30f2cef 100644 --- a/README.md +++ b/README.md @@ -181,7 +181,7 @@ This is a preview specification. The following areas are under active discussion **Event volume at scale.** A single deep-research query can produce 100+ retrieval events and dozens of grounding/citation events. The session document format already handles transport - one POST with all events after the session ends, not one request per event. Volume management beyond that (storage, processing, consumer-side aggregation) is an implementation concern, not a protocol gap. Sampling and aggregation are options for future versions but are not in v0.1; the standard sets no default for reporting granularity, leaving it to profiles and deployments. -**Verification of grounding and citation.** Grounding and citation events are reported by the agent, which is also the party that may owe compensation under a licence. In v0.1, manifest signing is informational: consumers may verify signatures but are not required to, and the specification defines no required proof binding an event to its emitter (sections 8.4 and 8.9). The events attribution depends on are therefore self-reported by the reporting party. Verifiable credentials and signed events are deferred (section 8.9). One corroboration mechanism works without signing: the `Content-Telemetry-ID` field correlates an agent-reported retrieval with an origin- or edge-reported one (section 7.2), but it covers retrieval only - grounding, citation, display, and engagement have no independent observer. Signing, even once required, would prove who reported an event, not that the event is true or that all qualifying events were reported. Input is wanted on what a verification layer should cover and where it belongs. Mechanisms that test truthfulness and completeness rather than origin, such as sampled audits or publisher-seeded canary content, are of particular interest. +**Verification of grounding and citation.** Grounding and citation events are reported by the agent, which is also the party that may owe compensation under a licence. In v0.1, manifest signing is informational: consumers may verify signatures but are not required to, and the specification defines no required proof binding an event to its emitter (sections 8.4 and 8.9). The events attribution depends on are therefore self-reported by the reporting party. Verifiable credentials and signed events are deferred (section 8.9). One corroboration mechanism works without signing: the `Content-Telemetry-ID` field correlates an agent-reported retrieval with an origin- or edge-reported one (section 7.2), but it covers retrieval only - grounding, reproduction, citation, presentation, and engagement have no independent observer. Signing, even once required, would prove who reported an event, not that the event is true or that all qualifying events were reported. Input is wanted on what a verification layer should cover and where it belongs. Mechanisms that test truthfulness and completeness rather than origin, such as sampled audits or publisher-seeded canary content, are of particular interest. **Reporting granularity.** The standard sets no default for reporting granularity, leaving it to profiles and deployments (see *Event volume* above). The SPUR profile requires event-level delivery and does not permit aggregation. The open question is whether the standard should say more about sampling and aggregation so that profiles do not each define it separately, and how event-level delivery scales for the highest-volume case. No mechanism is selected in v0.1. diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 183e01e..527c862 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -2,7 +2,7 @@ **Version:** 0.1 **Status:** Preview -**Last updated:** 2026-06-11 +**Last updated:** 2026-08-04 ## Contents @@ -58,9 +58,9 @@ Content Telemetry does not: #### 1.3.1 Inference-time scope -The five-stage lifecycle reports content use observable at inference time: identified content entered a generation context for a particular response, and what the resulting output did with it. +The six-stage lifecycle reports content use observable at inference time: identified content entered a generation context for a particular response, and what the resulting output did with it. -Retrieval is the boundary case. A crawl whose purpose is training or index building can be reported as a `content_retrieved` event, and `bot_category` (section 6.2) distinguishes it, but the event is non-attributable: no grounding, citation, presentation or engagement follows it. What the system then does with the content, whether it enters a training corpus, a fine-tuning set, an embedding store or a search index, is outside this specification. Nothing here reports that a model was trained on a work, and a conforming implementation says nothing either way about it. +Retrieval is the boundary case. A crawl whose purpose is training or index building can be reported as a `content_retrieved` event, and `bot_category` (section 6.2) distinguishes it, but the event is non-attributable: no grounding, reproduction, citation, presentation or engagement follows it. What the system then does with the content, whether it enters a training corpus, a fine-tuning set, an embedding store or a search index, is outside this specification. Nothing here reports that a model was trained on a work, and a conforming implementation says nothing either way about it. Using such a store at inference time is inside scope. When an index built over a content owner's material is queried during a response and returns content that grounds the answer, that is a `content_grounded` event like any other, with `source_role: index` on the retrieval that served it (section 4.4). The line is between constructing a derived artefact and using one to answer a query, not whether an index was involved. @@ -106,6 +106,7 @@ For the purposes of this specification, the following terms apply. | **agent operator** | entity running the AI agent that uses content | | **grounding** | content entering the generation model's context, the boundary where content can directly influence output (section 4.3) | | **reproduction** | verbatim or near-verbatim appearance of identified source content in an output artifact, independent of credit and delivery (section 4.3) | +| **presentation** | content or a source reference made perceivable on a recipient-facing surface (section 4.3) | | **source role** | classification of the observer reporting a retrieval event: `origin`, `edge`, `index`, or `agent` (section 4.4) | | **privacy level** | data sharing tier controlling which conversation fields are populated: `full`, `summary`, `intent`, or `minimal` (section 5.5) | | **conformance level** | emitter capability tier: Retrieval, Grounding, or Citation (section 5.7) | @@ -204,7 +205,7 @@ Content moves through six stages during an agent interaction: 4. **Cited** - An output artifact explicitly associates identified source content with a response, claim, passage, quotation, or other output element. Citation is an output-construction relationship, not evidence that the output was delivered. A subset of grounded content is commonly cited, but a citation can also be emitted without a matching grounding event when an agent produces an uncorroborated or hallucinated source association. - A citation MUST carry a resolvable reference to the source it associates: a `content_url` or a `content_id`. A source association with no resolvable reference is not a citation and MUST NOT be emitted as `content_cited`. Unlike other content events, where the identifier requirement is an application-layer rule (section 5.7.5), for `content_cited` it is enforced by the JSON Schema. + A citation MUST carry a resolvable reference to the source it associates: a `content_url` or a `content_id`. A source association with no resolvable reference is not a citation and MUST NOT be emitted as `content_cited`. Unlike other content events, where the identifier requirement is an application-layer rule (section 5.7.5), for `content_cited` and `content_reproduced` it is enforced by the JSON Schema. Reproduction and citation are sibling claims about the same artifact: reproduction records the material, citation records the credit. Neither implies the other. A credited quotation produces both events; an uncredited excerpt produces only a reproduction; a reference citation with no quoted material produces only a citation. @@ -364,9 +365,9 @@ Format: the URL of a manifest served at `/.well-known/content-telemetry.json` un | `type` | EventType | Yes | Event type (see 5.3) | | `timestamp` | datetime | Yes | Event timestamp (UTC) | | `turn_id` | string | No | Associates this event with a conversation turn (see 5.2.1) | -| `output_id` | string | For cited/presented | Opaque output-artifact identifier joining construction to later delivery | +| `output_id` | string | For reproduced/cited/presented | Opaque output-artifact identifier joining construction to later delivery | | `output_element_id` | string | No | Opaque element within `output_id`, such as a passage, media track, caption, link, or card | -| `citation_id` | UUID | No | On `content_presented`, the `id` of the associated citation event; absent for uncited presentations | +| `citation_id` | UUID | No | On `content_presented` or `content_reproduced`, the `id` of the associated citation event; absent for uncited presentations and uncredited reproductions | | `presentation_id` | UUID | For engaged | On `content_engaged`, the `id` of the exact presentation occurrence acted upon | | `source_role` | SourceRole | No | Who is reporting: `origin`, `edge`, `index`, `agent` (see 4.4) | | `content_telemetry_id` | UUID | No | Correlation ID for cross-observer deduplication (see 7.2) | @@ -555,7 +556,7 @@ A conforming **telemetry consumer** MUST: The JSON Schema (`telemetry-session.json`) validates structure and types but cannot express every conformance rule. The following are normative requirements verified at the application layer, not by schema validation: -- At least one of `content_url` or `content_id` MUST be present on every content event (section 4.5). For `content_cited` events this requirement is additionally enforced by the JSON Schema, which rejects a citation whose reference is absent or null (section 6.5). +- At least one of `content_url` or `content_id` MUST be present on every content event (section 4.5). For `content_cited` and `content_reproduced` events this requirement is additionally enforced by the JSON Schema, which rejects an event whose reference is absent or null (sections 6.5, 6.6). - An event MUST carry either `session_id` or `ctx_token` at Grounding conformance and above (section 7.1). - Conversation-turn fields MUST NOT exceed the turn's declared `privacy_level` (section 5.5). - The conformance-level requirements (sections 5.7.1 to 5.7.3) are cumulative. @@ -564,7 +565,7 @@ The `tests/` directory provides an informative reference suite for these rules. ## 6. Data profiles -The `data` field on events carries type-specific metadata. These profiles document the recommended fields by event type and source role, in lifecycle order. None are required, but emitting them enables richer attribution. +The `data` field on events carries type-specific metadata. These profiles document the recommended fields by event type and source role, in lifecycle order. None are required except where a section states otherwise (`reproduction_type` in 6.6; `presentation_kind` and `presentation_type` in 6.7), but emitting them enables richer attribution. ### 6.1 Retrieved content metadata (`content_retrieved`) @@ -917,7 +918,7 @@ Any party may operate a consumer: an agent operator, a licensing intermediary, o **Content owner resolution.** Telemetry consumers resolve content owner identity from `content_url` domains. Content owners register and verify their domains with the telemetry consumer; the consumer maps incoming event URLs to the owning organisation. This is the primary resolution path and requires `content_url` to be present on events. Events identified only by `content_id` (e.g., cached groundings where the URL was not preserved, or marketplace API content with no canonical URL) cannot be resolved by domain alone. Telemetry consumers SHOULD support `content_id` prefix-based resolution as a secondary path when content owners register their identifier schemes, but this is not yet a normative requirement. -**Cross-consumer correlation.** Origin-side emitters and agent-side emitters MAY use different telemetry consumers. A content owner's CDN sends retrieval events to one telemetry consumer; an agent sends sessions to another. The `content_telemetry_id` field (section 7.2) correlates the same retrieval across consumers - both sides share the same UUID from the HTTP request. This correlation operates at the retrieval level only. Grounding, citation, and engagement events have no independent origin-side counterpart to correlate against. +**Cross-consumer correlation.** Origin-side emitters and agent-side emitters MAY use different telemetry consumers. A content owner's CDN sends retrieval events to one telemetry consumer; an agent sends sessions to another. The `content_telemetry_id` field (section 7.2) correlates the same retrieval across consumers - both sides share the same UUID from the HTTP request. This correlation operates at the retrieval level only. Grounding, reproduction, citation, presentation, and engagement events have no independent origin-side counterpart to correlate against. ## 8. Manifest @@ -1151,7 +1152,7 @@ Whether this constitutes one royalty event, three, or ten depends on the commerc | Per-turn-influenced | One event per turn where content was in context | Usage-based licensing ("our content informed N answers") | | Per-reproduction | One event per reproduced portion appearing in output | Excerpt-based licensing ("N characters of our content appeared in answers") | -The `content_grounded` event with `scope: session` plus the count of subsequent `turn_completed` events provides the inputs for all three models without requiring the schema to embed a commercial opinion. +The `content_grounded` event with `scope: session` plus the count of subsequent `turn_completed` events provides the inputs for the first three models; the per-reproduction model additionally draws on `content_reproduced` events and their `reproduced_chars`. None of the four requires the schema to embed a commercial opinion. ### 10.2 Grounding without citation @@ -1162,7 +1163,7 @@ Content can influence every response in a session without being explicitly cited - At the **citation** level: counts only explicitly attributed content. Simpler to verify but undercounts content influence. - At the **presentation** level: counts content or source references made perceivable. It does not prove attention. -Content owners and platforms should agree on which level to count at. The telemetry data supports all three; the choice is commercial, not technical. +Content owners and platforms should agree on which level to count at. The telemetry data supports all four; the choice is commercial, not technical. ## 11. Extensibility diff --git a/tests/README.md b/tests/README.md index 1d50b8d..1948ed2 100644 --- a/tests/README.md +++ b/tests/README.md @@ -35,7 +35,7 @@ Run from the repository root. Without uv: `pip install jsonschema`, then `python - Funnel exceptions (presented-no-cited, cited-no-grounded, presented-no-grounded) - Reproduction cases: credited quotation (reproduction + direct_quote citation sharing an output element) and uncredited reproduction in unpresented API output - Text, image, audio, video, suppressed-citation, and repeated-presentation cases -- Exact presentation-to-engagement correlation across session, standalone, and batch envelopes +- Exact presentation-to-engagement correlation across session and standalone envelopes - Multi-turn sessions, cached grounding - Custom response_mode values @@ -49,5 +49,6 @@ Some rules cannot be expressed in JSON Schema alone. These are tested as applica - `content_url` or `content_id` requirement on every content event (section 5.7.5) - `session_id` or `ctx_token` on a standalone event or event batch envelope at Grounding conformance and above (sections 5.7.5, 7.1) - Manifest rejection rules: duplicate `keys[].id`, and `domains` entries that are not the manifest's own host or a subdomain of it (sections 8.6, 8.7) +- Withdrawn `ip_hash` prohibition on `content_retrieved` data (section 9.1 migration rule) Valid fixtures must pass both JSON Schema and these checks; `invalid/` fixtures that pass JSON Schema but fail a check are documented in `validate.py`. The `agent_id`-at-Grounding requirement is not fixture-tested: it depends on the emitter's declared conformance level, which the fixtures do not carry. diff --git a/tests/check_examples.py b/tests/check_examples.py index 3a0e8a4..7870f4d 100644 --- a/tests/check_examples.py +++ b/tests/check_examples.py @@ -8,6 +8,7 @@ session document -> telemetry-session.json standalone event -> telemetry-event.json + event batch -> telemetry-event-batch.json manifest -> manifest.json Fragments (a bare event object, a single turn, a one-field snippet) are not diff --git a/tests/invalid/cited-missing-output-id.json b/tests/invalid/cited-missing-output-id.json new file mode 100644 index 0000000..014e7bd --- /dev/null +++ b/tests/invalid/cited-missing-output-id.json @@ -0,0 +1,15 @@ +{ + "_test_description": "content_cited event with id and a resolvable source reference but no output_id. output_id is required on citation events so output construction can be correlated with later presentation.", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440160", + "started_at": "2026-08-01T14:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440161", + "type": "content_cited", + "timestamp": "2026-08-01T14:00:01Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { "citation_type": "reference" } + } + ] +} diff --git a/tests/invalid/cited-missing-source-reference.json b/tests/invalid/cited-missing-source-reference.json index babc716..d987fbc 100644 --- a/tests/invalid/cited-missing-source-reference.json +++ b/tests/invalid/cited-missing-source-reference.json @@ -6,9 +6,11 @@ "started_at": "2026-08-01T12:00:00Z", "events": [ { + "id": "770e8400-e29b-41d4-a716-446655440142", "type": "content_cited", "timestamp": "2026-08-01T12:00:01Z", "turn_id": "1", + "output_id": "response:1", "data": { "citation_type": "reference", "position": "primary" diff --git a/tests/invalid/cited-null-source-reference.json b/tests/invalid/cited-null-source-reference.json index ba4cce0..2344da6 100644 --- a/tests/invalid/cited-null-source-reference.json +++ b/tests/invalid/cited-null-source-reference.json @@ -6,9 +6,11 @@ "started_at": "2026-08-01T12:30:00Z", "events": [ { + "id": "770e8400-e29b-41d4-a716-446655440143", "type": "content_cited", "timestamp": "2026-08-01T12:30:01Z", "turn_id": "1", + "output_id": "response:1", "content_url": null, "data": { "citation_type": "direct_quote", diff --git a/tests/invalid/presented-missing-presentation-type.json b/tests/invalid/presented-missing-presentation-type.json new file mode 100644 index 0000000..bacdd0e --- /dev/null +++ b/tests/invalid/presented-missing-presentation-type.json @@ -0,0 +1,16 @@ +{ + "_test_description": "content_presented event carrying presentation_kind but no presentation_type. Both are required: the kind says what crossed the presentation boundary, the type says how it was made perceivable.", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440163", + "started_at": "2026-08-01T15:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440164", + "type": "content_presented", + "timestamp": "2026-08-01T15:00:01Z", + "output_id": "response:1", + "content_id": "publisher:article:4", + "data": { "presentation_kind": "source_reference" } + } + ] +} diff --git a/tests/invalid/reproduced-missing-id.json b/tests/invalid/reproduced-missing-id.json new file mode 100644 index 0000000..b322854 --- /dev/null +++ b/tests/invalid/reproduced-missing-id.json @@ -0,0 +1,15 @@ +{ + "_test_description": "content_reproduced event with output_id and a resolvable source reference but no event id. id is required on reproduction events so a crediting citation or verification result can reference the exact reproduction claim.", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440162", + "started_at": "2026-08-01T14:30:00Z", + "events": [ + { + "type": "content_reproduced", + "timestamp": "2026-08-01T14:30:01Z", + "output_id": "response:1", + "content_id": "publisher:article:3", + "data": { "reproduction_type": "verbatim", "reproduced_chars": 320 } + } + ] +} diff --git a/tests/valid/event-batch-agent.json b/tests/valid/event-batch-agent.json index 7d7b01b..08e605c 100644 --- a/tests/valid/event-batch-agent.json +++ b/tests/valid/event-batch-agent.json @@ -17,7 +17,11 @@ "type": "content_grounded", "timestamp": "2026-03-28T08:20:03Z", "source_role": "agent", - "content_url": "https://arstechnica.com/science/2026/03/new-battery-chemistry-breakthrough/" + "content_url": "https://arstechnica.com/science/2026/03/new-battery-chemistry-breakthrough/", + "data": { + "scope": "session", + "chars_ingested": 5200 + } }, { "id": "990e8400-e29b-41d4-a716-446655440062", diff --git a/tests/valid/event-standalone-child-session.json b/tests/valid/event-standalone-child-session.json new file mode 100644 index 0000000..d136d62 --- /dev/null +++ b/tests/valid/event-standalone-child-session.json @@ -0,0 +1,20 @@ +{ + "_test_description": "Standalone event envelope from a child agent session in a multi-agent topology: parent_session_id on the envelope identifies the orchestrating session. The grounded content entered the sub-agent's generation context.", + "document_type": "event", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440170", + "parent_session_id": "660e8400-e29b-41d4-a716-446655440171", + "agent_id": "research-subagent-v2", + "started_at": "2026-08-01T16:00:00Z", + "event": { + "type": "content_grounded", + "timestamp": "2026-08-01T16:00:02Z", + "content_url": "https://www.example-journal.org/analysis/press-freedom-ruling", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 4100, + "tokens_ingested": 1000 + } + } +} diff --git a/tests/valid/session-cached-grounding.json b/tests/valid/session-cached-grounding.json index a9821f1..5344e57 100644 --- a/tests/valid/session-cached-grounding.json +++ b/tests/valid/session-cached-grounding.json @@ -1,5 +1,5 @@ { - "_test_description": "Grounding from agent-side cache with no retrieval event. license_ref preserved from the original retrieval, as recommended by section 6.6.", + "_test_description": "Grounding from agent-side cache with no retrieval event. license_ref preserved from the original retrieval, as recommended by section 6.4.", "schema_version": "0.1", "session_id": "660e8400-e29b-41d4-a716-446655440004", "agent_id": "copilot-v3", diff --git a/tests/valid/session-citation-tier.json b/tests/valid/session-citation-tier.json index 3036e22..b8f9743 100644 --- a/tests/valid/session-citation-tier.json +++ b/tests/valid/session-citation-tier.json @@ -1,5 +1,5 @@ { - "_test_description": "Citation conformance level with optional display and engagement lifecycle signals.", + "_test_description": "Citation conformance level with optional presentation and engagement lifecycle signals.", "schema_version": "0.1", "session_id": "660e8400-e29b-41d4-a716-446655440003", "agent_id": "shopping-assistant-v2", diff --git a/tests/valid/session-multi-agent-child.json b/tests/valid/session-multi-agent-child.json index 1f17456..d333791 100644 --- a/tests/valid/session-multi-agent-child.json +++ b/tests/valid/session-multi-agent-child.json @@ -15,6 +15,7 @@ "data": { "scope": "turn", "cached": false, + "chars_ingested": 3600, "tokens_ingested": 900 } }, diff --git a/tests/valid/session-multi-turn.json b/tests/valid/session-multi-turn.json index 424034a..a220bf1 100644 --- a/tests/valid/session-multi-turn.json +++ b/tests/valid/session-multi-turn.json @@ -1,11 +1,21 @@ { - "_test_description": "Session-scoped grounding, 3 turns, citations in turns 1 and 3, zero-click outcome (browse with new_query exit). Turn 2 has no citation - the grounded content was in context but not explicitly referenced.", + "_test_description": "Session-scoped grounding, 3 turns, citations in turns 1 and 3, zero-click outcome. Turn 2 has no citation - the grounded content was in context but not explicitly referenced.", "schema_version": "0.1", "session_id": "660e8400-e29b-41d4-a716-446655440006", "agent_id": "copilot-v3", "started_at": "2026-03-28T16:00:00Z", "ended_at": "2026-03-28T16:10:00Z", "events": [ + { + "type": "turn_started", + "timestamp": "2026-03-28T16:00:00Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "query_intent": "question", + "topics": ["fusion energy", "ITER"] + } + }, { "type": "content_retrieved", "timestamp": "2026-03-28T16:00:01Z", @@ -25,16 +35,6 @@ "media_type": "text" } }, - { - "type": "turn_started", - "timestamp": "2026-03-28T16:00:00Z", - "turn_id": "1", - "turn": { - "privacy_level": "intent", - "query_intent": "question", - "topics": ["fusion energy", "ITER"] - } - }, { "id": "880e8400-e29b-41d4-a716-446655440051", "type": "content_cited", diff --git a/tests/valid/session-reproduction-credited-quote.json b/tests/valid/session-reproduction-credited-quote.json index ab32217..29628ab 100644 --- a/tests/valid/session-reproduction-credited-quote.json +++ b/tests/valid/session-reproduction-credited-quote.json @@ -32,6 +32,7 @@ "data": { "scope": "turn", "cached": false, + "chars_ingested": 9800, "tokens_ingested": 2400, "content_hash": "sha256:e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6", "media_type": "text" diff --git a/tests/valid/session-reproduction-uncredited.json b/tests/valid/session-reproduction-uncredited.json index 10a6fc8..3c031e0 100644 --- a/tests/valid/session-reproduction-uncredited.json +++ b/tests/valid/session-reproduction-uncredited.json @@ -31,6 +31,7 @@ "data": { "scope": "turn", "cached": false, + "chars_ingested": 7200, "tokens_ingested": 1800, "content_hash": "sha256:c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4", "media_type": "text" From b89e827d62d943c48023ff0fbcf93e1d87912941 Mon Sep 17 00:00:00 2001 From: Erik Sv Date: Wed, 12 Aug 2026 14:44:42 +0000 Subject: [PATCH 14/38] Complete v1 grounding provenance and fingerprint semantics --- SPECIFICATION.md | 41 +++++++++++++++++++ telemetry-session.json | 31 +++++++++++++- tests/README.md | 4 +- ...rounding-fingerprint-missing-detected.json | 21 ++++++++++ ...nding-fingerprint-preserved-in-output.json | 23 +++++++++++ .../grounding-provenance-cached-conflict.json | 18 ++++++++ .../event-batch-grounding-provenance.json | 36 ++++++++++++++++ ...vent-standalone-grounding-fingerprint.json | 26 ++++++++++++ tests/valid/session-cached-grounding.json | 8 +++- tests/validate.py | 37 +++++++++++++++++ 10 files changed, 242 insertions(+), 3 deletions(-) create mode 100644 tests/invalid/grounding-fingerprint-missing-detected.json create mode 100644 tests/invalid/grounding-fingerprint-preserved-in-output.json create mode 100644 tests/invalid/grounding-provenance-cached-conflict.json create mode 100644 tests/valid/event-batch-grounding-provenance.json create mode 100644 tests/valid/event-standalone-grounding-fingerprint.json diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 527c862..b3e960e 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -560,6 +560,8 @@ The JSON Schema (`telemetry-session.json`) validates structure and types but can - An event MUST carry either `session_id` or `ctx_token` at Grounding conformance and above (section 7.1). - Conversation-turn fields MUST NOT exceed the turn's declared `privacy_level` (section 5.5). - The conformance-level requirements (sections 5.7.1 to 5.7.3) are cumulative. +- When `content_grounded.data.provenance` is `agent_fetched`, `data.cached` MUST be `false`; when it is `agent_cached`, `data.cached` MUST be `true` (section 6.4). +- `content_grounded.data.content_fingerprint` MUST NOT contain `preserved_in_output`; output-side reuse is reported with `content_reproduced` (sections 6.4 and 12.1). The `tests/` directory provides an informative reference suite for these rules. A consumer that receives a privacy-violating turn (e.g., `query_text` present at `minimal` level) SHOULD strip the offending fields rather than reject the document carrying them. @@ -624,12 +626,14 @@ Emitting a `training`-category `content_retrieved` event is permitted but non-at |-------|------|-------------| | `scope` | string | Influence scope: `session` or `turn` (see below) | | `cached` | boolean | Content served from agent-side cache rather than a live fetch | +| `provenance` | string | How content reached the context: `agent_fetched`, `agent_cached`, or `third_party_sourced` (see below) | | `chars_ingested` | integer | Character count of content placed in the generation context (see below) | | `tokens_ingested` | integer | Token count of the same content, supplementary (see below) | | `content_version` | string | Content version identifier (ETag, revision ID, CMS version) | | `content_last_modified` | datetime | When the content was last modified at source | | `content_hash` | string | SHA-256 of the content as ingested (`sha256:{hex}`) | | `media_type` | string | Content medium: `text`, `image`, `video`, `audio` (open vocabulary, see 6.1) | +| `content_fingerprint` | object | Agent-reported detection of a fingerprint or provenance signal in the grounded content (see below) | Both fields measure the content actually placed in the generation model's context. For chunked retrieval, count only the portion used, not the full source document. @@ -637,6 +641,34 @@ Both fields measure the content actually placed in the generation model's contex `tokens_ingested` counts the same content in the generation model's tokeniser (the model identified in `model_id` on the corresponding `turn_completed` event), not the retrieval or embedding model's tokeniser. It is supplementary. Token counts are model-specific, change when a vendor revises a tokeniser, and are not comparable between agents, so a consumer cannot aggregate them across emitters or treat a difference as a difference in volume. Emitters SHOULD send `chars_ingested` where they send `tokens_ingested`, and consumers that receive only token counts SHOULD record which model produced them. +#### Provenance and content fingerprints + +`provenance` describes the delivery path by which the grounded representation reached the agent: + +| Value | Description | +|-------|-------------| +| `agent_fetched` | The agent obtained the representation directly from the publisher or from publisher-authorised origin or edge infrastructure for this session | +| `agent_cached` | The agent reused a representation it had obtained before this session | +| `third_party_sourced` | The representation reached the agent through an intermediary rather than through a direct publisher-authorised retrieval by the agent in this session | + +The field describes delivery path, not evidence quality. An emitter declaring Grounding or Citation conformance SHOULD include it when the path is known. It remains optional because an agent may not be able to distinguish its own earlier fetch from intermediary delivery. Consumers MUST NOT infer a value when it is absent. + +Emitters MUST keep `provenance` and `cached` consistent: `agent_fetched` requires `cached: false`, and `agent_cached` requires `cached: true`. `third_party_sourced` leaves `cached` unconstrained because an intermediary-sourced representation may be used immediately or cached by the agent before grounding. + +`content_fingerprint` contains: + +| Field | Type | Required | Description | +|-------|------|----------|-------------| +| `scheme` | string | Yes | Open identifier for the fingerprint or provenance scheme checked | +| `detected` | boolean | Yes | Emitter claim that the scheme's signal was found in the exact grounded representation | +| `value` | string | No | Scheme-defined fingerprint or identifier value, when the scheme produces one | + +`detected` reports a grounding-time observation by the emitter. It does not establish that the signal is authentic, identify who applied it, prove that the content was used later in the output, or raise the evidentiary status of the grounding event. Those questions require profile-defined evidence and consumer trust policy outside the core schema. + +Output-side reuse is reported with `content_reproduced`, not a fingerprint-preservation field on `content_grounded`. A consumer MAY compare a grounding fingerprint with evidence about a reproduced output, but the two remain separate assertions about separate lifecycle stages. + +`scheme` is an open identifier. Emitters SHOULD use a globally collision-resistant value. Core does not register schemes, interpret `value`, or assign capabilities or evidence status from a scheme identifier. A profile MAY define scheme-specific processing rules. + #### Grounding scope | Value | Description | @@ -1248,6 +1280,15 @@ v0.1 data, consumers MAY treat a `direct_quote` citation as an implied reproduction of its excerpt. V1 emitters record reproduction explicitly and SHOULD NOT rely on that inference. +V1 grounding fingerprints report detection only. A preview implementation that +used `data.content_fingerprint.preserved_in_output` removes that field during +migration. When its value was `true` and the historical record contains enough +information to populate every required `content_reproduced` field, the +implementation MAY also create the corresponding reproduction event. It MUST +NOT synthesize a reproduction event from `false` or incomplete historical data. +A grounding event MAY retain `content_fingerprint.scheme`, `detected`, and a +scheme-defined `value`. + Preview versions (0.x) use two-component version numbers. From 1.0.0 onward, versions follow [semantic versioning](https://semver.org/): - **Major** (1.0.0 → 2.0.0) - breaking changes to required fields diff --git a/telemetry-session.json b/telemetry-session.json index 1693fdc..41c29e3 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -149,12 +149,14 @@ "properties": { "scope": { "$ref": "#/$defs/GroundingScope" }, "cached": { "type": "boolean" }, + "provenance": { "$ref": "#/$defs/SourceProvenance" }, "chars_ingested": { "type": "integer", "minimum": 0, "description": "Unicode code points in the exact text placed in the generation context, without normalising solely for counting. Portable across emitters; preferred over tokens_ingested (section 6.4)." }, "tokens_ingested": { "type": "integer", "minimum": 0, "description": "Token count of the same content in the emitter's own tokeniser. Supplementary: not comparable between emitters (section 6.4)." }, "content_version": { "type": "string" }, "content_last_modified": { "type": "string", "format": "date-time" }, "content_hash": { "type": "string", "pattern": "^sha256:[a-f0-9]{64}$" }, - "media_type": { "$ref": "#/$defs/MediaType" } + "media_type": { "$ref": "#/$defs/MediaType" }, + "content_fingerprint": { "$ref": "#/$defs/ContentFingerprint" } } } } @@ -400,6 +402,33 @@ "description": "Whether content informed all subsequent responses in the session or a specific turn only", "enum": ["session", "turn"] }, + "SourceProvenance": { + "type": "string", + "description": "How the grounded representation reached the agent. Describes delivery path, not evidence quality.", + "enum": ["agent_fetched", "agent_cached", "third_party_sourced"] + }, + "ContentFingerprint": { + "type": "object", + "description": "Emitter-reported detection of a fingerprint or provenance signal in the exact grounded representation. Core does not interpret schemes or assign evidence status from detection.", + "required": ["scheme", "detected"], + "properties": { + "scheme": { + "type": "string", + "minLength": 1, + "description": "Open identifier for the fingerprint or provenance scheme checked. Globally collision-resistant values are recommended." + }, + "detected": { + "type": "boolean", + "description": "Emitter claim that the scheme's signal was found in the grounded representation" + }, + "value": { + "type": "string", + "minLength": 1, + "description": "Optional scheme-defined fingerprint or identifier value" + } + }, + "additionalProperties": true + }, "PresentationKind": { "type": "string", "description": "What was made perceivable: source content itself, including a bounded excerpt or derived representation, or a reference to the source.", diff --git a/tests/README.md b/tests/README.md index 1948ed2..ed2a846 100644 --- a/tests/README.md +++ b/tests/README.md @@ -36,7 +36,8 @@ Run from the repository root. Without uv: `pip install jsonschema`, then `python - Reproduction cases: credited quotation (reproduction + direct_quote citation sharing an output element) and uncredited reproduction in unpresented API output - Text, image, audio, video, suppressed-citation, and repeated-presentation cases - Exact presentation-to-engagement correlation across session and standalone envelopes -- Multi-turn sessions, cached grounding +- Multi-turn sessions and cached grounding +- Grounding provenance paths and generic fingerprint detection across session, standalone-event, and event-batch envelopes - Custom response_mode values Each test file has a `_test_description` field explaining what it demonstrates. @@ -50,5 +51,6 @@ Some rules cannot be expressed in JSON Schema alone. These are tested as applica - `session_id` or `ctx_token` on a standalone event or event batch envelope at Grounding conformance and above (sections 5.7.5, 7.1) - Manifest rejection rules: duplicate `keys[].id`, and `domains` entries that are not the manifest's own host or a subdomain of it (sections 8.6, 8.7) - Withdrawn `ip_hash` prohibition on `content_retrieved` data (section 9.1 migration rule) +- Grounding provenance/cache consistency and the prohibition on `preserved_in_output` in `content_fingerprint` (sections 5.7.5, 6.4, 12.1) Valid fixtures must pass both JSON Schema and these checks; `invalid/` fixtures that pass JSON Schema but fail a check are documented in `validate.py`. The `agent_id`-at-Grounding requirement is not fixture-tested: it depends on the emitter's declared conformance level, which the fixtures do not carry. diff --git a/tests/invalid/grounding-fingerprint-missing-detected.json b/tests/invalid/grounding-fingerprint-missing-detected.json new file mode 100644 index 0000000..d1b713e --- /dev/null +++ b/tests/invalid/grounding-fingerprint-missing-detected.json @@ -0,0 +1,21 @@ +{ + "_test_description": "Grounding event carries content_fingerprint without required detected. Must fail JSON Schema.", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440182", + "started_at": "2026-08-12T10:20:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-12T10:20:01Z", + "content_id": "publisher:article:99", + "data": { + "scope": "turn", + "cached": false, + "provenance": "agent_fetched", + "content_fingerprint": { + "scheme": "https://example.org/fingerprint/v1" + } + } + } + ] +} diff --git a/tests/invalid/grounding-fingerprint-preserved-in-output.json b/tests/invalid/grounding-fingerprint-preserved-in-output.json new file mode 100644 index 0000000..b8fb38e --- /dev/null +++ b/tests/invalid/grounding-fingerprint-preserved-in-output.json @@ -0,0 +1,23 @@ +{ + "_test_description": "Grounding fingerprint uses withdrawn preserved_in_output instead of a content_reproduced event. Passes the extensible schema but violates the v1 migration rule.", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440183", + "started_at": "2026-08-12T10:30:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-12T10:30:01Z", + "content_id": "publisher:article:100", + "data": { + "scope": "turn", + "cached": false, + "provenance": "agent_fetched", + "content_fingerprint": { + "scheme": "https://example.org/fingerprint/v1", + "detected": true, + "preserved_in_output": true + } + } + } + ] +} diff --git a/tests/invalid/grounding-provenance-cached-conflict.json b/tests/invalid/grounding-provenance-cached-conflict.json new file mode 100644 index 0000000..fcffe65 --- /dev/null +++ b/tests/invalid/grounding-provenance-cached-conflict.json @@ -0,0 +1,18 @@ +{ + "_test_description": "Grounding event declares agent_fetched with cached true. Passes JSON Schema but violates section 6.4 provenance consistency.", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440184", + "started_at": "2026-08-12T10:40:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-12T10:40:01Z", + "content_id": "publisher:article:101", + "data": { + "scope": "turn", + "cached": true, + "provenance": "agent_fetched" + } + } + ] +} diff --git a/tests/valid/event-batch-grounding-provenance.json b/tests/valid/event-batch-grounding-provenance.json new file mode 100644 index 0000000..946a431 --- /dev/null +++ b/tests/valid/event-batch-grounding-provenance.json @@ -0,0 +1,36 @@ +{ + "_test_description": "Batch envelope carrying grounding events from cached and third-party-sourced representations, exercising the shared provenance schema across batch delivery.", + "document_type": "event_batch", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440181", + "agent_id": "enterprise-rag-v1", + "started_at": "2026-08-12T10:10:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-12T10:10:01Z", + "content_id": "publisher:cached:7", + "data": { + "scope": "session", + "cached": true, + "provenance": "agent_cached", + "chars_ingested": 900 + } + }, + { + "type": "content_grounded", + "timestamp": "2026-08-12T10:10:02Z", + "content_id": "repository:item:9", + "data": { + "scope": "turn", + "cached": false, + "provenance": "third_party_sourced", + "chars_ingested": 500, + "content_fingerprint": { + "scheme": "https://example.org/fingerprint/v1", + "detected": false + } + } + } + ] +} diff --git a/tests/valid/event-standalone-grounding-fingerprint.json b/tests/valid/event-standalone-grounding-fingerprint.json new file mode 100644 index 0000000..70eefe4 --- /dev/null +++ b/tests/valid/event-standalone-grounding-fingerprint.json @@ -0,0 +1,26 @@ +{ + "_test_description": "Standalone grounding event from a publisher-authorised live fetch with an emitter-reported generic fingerprint detection. The event schema is shared by session, standalone and batch envelopes.", + "document_type": "event", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440180", + "agent_id": "research-agent-v1", + "started_at": "2026-08-12T10:00:00Z", + "event": { + "type": "content_grounded", + "timestamp": "2026-08-12T10:00:01Z", + "content_url": "https://publisher.example/articles/42", + "content_id": "publisher:article:42", + "data": { + "scope": "turn", + "cached": false, + "provenance": "agent_fetched", + "chars_ingested": 2400, + "content_hash": "sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa", + "content_fingerprint": { + "scheme": "https://example.org/fingerprint/v1", + "detected": true, + "value": "fp_42" + } + } + } +} diff --git a/tests/valid/session-cached-grounding.json b/tests/valid/session-cached-grounding.json index 5344e57..8f8347a 100644 --- a/tests/valid/session-cached-grounding.json +++ b/tests/valid/session-cached-grounding.json @@ -1,5 +1,5 @@ { - "_test_description": "Grounding from agent-side cache with no retrieval event. license_ref preserved from the original retrieval, as recommended by section 6.4.", + "_test_description": "Session document with agent-cached grounding and emitter-reported generic fingerprint detection. license_ref is preserved from the original retrieval, as recommended by section 6.4.", "schema_version": "0.1", "session_id": "660e8400-e29b-41d4-a716-446655440004", "agent_id": "copilot-v3", @@ -15,10 +15,16 @@ "data": { "scope": "session", "cached": true, + "provenance": "agent_cached", "chars_ingested": 11200, "tokens_ingested": 2800, "content_last_modified": "2026-03-27T16:00:00Z", "content_hash": "sha256:c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4", + "content_fingerprint": { + "scheme": "org.example.fingerprint.v1", + "detected": true, + "value": "fp_cached_42" + }, "media_type": "text" } }, diff --git a/tests/validate.py b/tests/validate.py index 310d052..5790c6f 100644 --- a/tests/validate.py +++ b/tests/validate.py @@ -55,6 +55,11 @@ # manifest's own host or a subdomain of it (section 8.6). JSON Schema # cannot compare values across array items or against the manifest's id. # +# 5. Grounding provenance and fingerprint migration (sections 5.7.5, 6.4, 12.1): +# agent_fetched requires cached false; agent_cached requires cached true. +# content_fingerprint MUST NOT carry preserved_in_output because output-side +# reuse is represented by content_reproduced. +# # Not checked here: agent_id at Grounding/Citation conformance (section # 5.7) depends on the emitter's declared conformance level, which fixtures do # not carry, so it is out of scope for the fixture suite. @@ -101,6 +106,14 @@ "Violates section 9.1: the field was withdrawn in v1 and emitters " "MUST NOT populate it." ), + "grounding-fingerprint-preserved-in-output.json": ( + "Grounding fingerprint carries preserved_in_output. " + "Violates section 6.4: output-side reuse is reported as content_reproduced." + ), + "grounding-provenance-cached-conflict.json": ( + "Grounding event declares agent_fetched with cached true. " + "Violates section 6.4: agent_fetched requires cached false." + ), } # V0.1 fields prohibited by the v1 migration rule (section 9.1). This is a @@ -343,6 +356,29 @@ def check_v1_migration_prohibitions(data): return violations +def check_grounding_provenance(data): + """Check the provenance/cached pairings required by section 6.4.""" + violations = [] + for event in _iter_events(data): + if event.get("type") != "content_grounded": + continue + event_data = event.get("data") + if not isinstance(event_data, dict): + continue + provenance = event_data.get("provenance") + cached = event_data.get("cached") + if provenance == "agent_fetched" and cached is not False: + violations.append("content_grounded with agent_fetched does not carry cached false") + if provenance == "agent_cached" and cached is not True: + violations.append("content_grounded with agent_cached does not carry cached true") + fingerprint = event_data.get("content_fingerprint") + if isinstance(fingerprint, dict) and "preserved_in_output" in fingerprint: + violations.append( + "content_fingerprint carries preserved_in_output; use content_reproduced" + ) + return violations + + def check_application_layer(data): """Run every application-layer conformance rule and return all violations.""" return ( @@ -350,6 +386,7 @@ def check_application_layer(data): + check_content_identifier(data) + check_session_or_ctx_token(data) + check_v1_migration_prohibitions(data) + + check_grounding_provenance(data) ) From f106913fa4c78dec7c5c7c7908f9278c4a96b2ce Mon Sep 17 00:00:00 2001 From: Leandro Oliva Date: Mon, 22 Jun 2026 16:46:15 +0200 Subject: [PATCH 15/38] Add test fixture for provenance and content_fingerprint --- .../event-standalone-grounded-provenance.json | 25 +++++++++++++++++++ 1 file changed, 25 insertions(+) create mode 100644 tests/valid/event-standalone-grounded-provenance.json diff --git a/tests/valid/event-standalone-grounded-provenance.json b/tests/valid/event-standalone-grounded-provenance.json new file mode 100644 index 0000000..fa9e1b9 --- /dev/null +++ b/tests/valid/event-standalone-grounded-provenance.json @@ -0,0 +1,25 @@ +{ + "document_type": "event", + "schema_version": "0.1", + "session_id": "550e8400-e29b-41d4-a716-446655440000", + "agent_id": "agent-example", + "started_at": "2026-06-20T10:00:00Z", + "event": { + "id": "6ba7b810-9dad-11d1-80b4-00c04fd430c8", + "type": "content_grounded", + "timestamp": "2026-06-20T10:00:01Z", + "content_url": "https://example.com/article", + "content_id": "ISCC:KACYPXW445FTYNJ3", + "data": { + "scope": "turn", + "cached": true, + "provenance": "third_party_sourced", + "tokens_ingested": 512, + "content_fingerprint": { + "scheme": "iscc", + "detected": true, + "preserved_in_output": false + } + } + } +} From b5876a9c5c162682c036374a683aee8bfa9f4722 Mon Sep 17 00:00:00 2001 From: Alex Springer Date: Mon, 17 Aug 2026 08:55:44 +0100 Subject: [PATCH 16/38] Drop preserved_in_output from the standalone provenance fixture Output-side reuse is reported with content_reproduced on the landed v1-draft; grounding-time detection stays as the emitter claim. Co-Authored-By: Claude Fable 5 --- tests/valid/event-standalone-grounded-provenance.json | 3 +-- 1 file changed, 1 insertion(+), 2 deletions(-) diff --git a/tests/valid/event-standalone-grounded-provenance.json b/tests/valid/event-standalone-grounded-provenance.json index fa9e1b9..96f3c08 100644 --- a/tests/valid/event-standalone-grounded-provenance.json +++ b/tests/valid/event-standalone-grounded-provenance.json @@ -17,8 +17,7 @@ "tokens_ingested": 512, "content_fingerprint": { "scheme": "iscc", - "detected": true, - "preserved_in_output": false + "detected": true } } } From 8731299dea60455545d8d1f11e69308cda179844 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Wed, 12 Aug 2026 10:48:53 +0100 Subject: [PATCH 17/38] Define the click context: ctx_token issuance, carriage, discovery and resolution (#23, #28, #30) New section 7.4 replaces the v0.1 click manifest with the click context: - Token pattern ct_[A-Za-z0-9_-]{8,240}, opaque, bound to exactly one presentation at mint time; per-click minting for routed surfaces. - Reserved query parameters ctx_token and ctx_iss with SHOULD-level redirect propagation mirroring section 7.2. - Resolver discovery through the issuer manifest: telemetry.ctx_resolution. - Resolution returns the engagement, the clicked content's lineage cut by content identity across turns, and an optional count-based session summary; never the raw session_id, never other owners' events. - presentation_id becomes conditional: required for agent-reported engagements, restored from issuer state for destination reports that carry ctx_token; the URL-carried presentation UUID is withdrawn. - Consumer-custody trade-off recorded (#29). Schemas, six new/updated fixtures, migration note. 74/74 conformance checks, 9/9 worked examples. Co-Authored-By: Claude Fable 5 --- SPECIFICATION.md | 87 +++++++++++++++++-- manifest.json | 6 ++ telemetry-event-batch.json | 25 +++++- telemetry-event.json | 24 ++++- telemetry-session.json | 21 ++++- tests/invalid/ctx-token-bad-pattern.json | 12 +++ ...dalone-missing-presentation-and-token.json | 14 +++ .../event-standalone-engaged-ctx-token.json | 3 +- .../event-standalone-engaged-redirect.json | 14 +++ .../valid/manifest-agent-ctx-resolution.json | 12 +++ tests/valid/session-earlier-turn-click.json | 53 +++++++++++ .../session-repeated-link-engagement.json | 54 ++++++++++++ 12 files changed, 307 insertions(+), 18 deletions(-) create mode 100644 tests/invalid/ctx-token-bad-pattern.json create mode 100644 tests/invalid/engaged-standalone-missing-presentation-and-token.json create mode 100644 tests/valid/event-standalone-engaged-redirect.json create mode 100644 tests/valid/manifest-agent-ctx-resolution.json create mode 100644 tests/valid/session-earlier-turn-click.json create mode 100644 tests/valid/session-repeated-link-engagement.json diff --git a/SPECIFICATION.md b/SPECIFICATION.md index b3e960e..4121219 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -12,7 +12,7 @@ 4. [Concepts](#4-concepts) - roles, sessions, event lifecycle, source roles, content identification 5. [Schema](#5-schema) - session, event, event types, conversation turn, privacy, intent, conformance levels 6. [Data profiles](#6-data-profiles) - retrieval, edge enrichment, origin enrichment, grounding, citation, reproduction, presentation, engagement -7. [Transport](#7-transport) - delivery formats, Content-Telemetry-ID header +7. [Transport](#7-transport) - delivery formats, Content-Telemetry-ID header, routing, click context 8. [Manifest](#8-manifest) - discovery, schema, operator, keys, telemetry, domains 9. [Privacy](#9-privacy) - data minimisation, recommended levels, retention 10. [Attribution](#10-attribution) - counting semantics, grounding without citation @@ -213,7 +213,7 @@ Content moves through six stages during an agent interaction: Grounding and presentation record different boundary crossings: grounding records entry into a generation context, while presentation records a recipient-facing delivery occurrence. As agent experiences evolve beyond the chat window the two diverge - content can shape an answer whose source is never presented, and an agent can present content that never entered a generation context (see *Departures from the funnel model* below). -6. **Engaged** - The recipient or agent performed an observable action on a presentation: clicked a link, expanded a preview, copied text, shared the response, or directed the agent to act on the content. It does not imply attention beyond the reported action. `presentation_id` links the action to the exact presentation occurrence; a click-out can also carry a `ctx_token` that a destination resolves to the session's click manifest (section 7.1). +6. **Engaged** - The recipient or agent performed an observable action on a presentation: clicked a link, expanded a preview, copied text, shared the response, or directed the agent to act on the content. It does not imply attention beyond the reported action. `presentation_id` links the action to the exact presentation occurrence; a click-out can also carry a `ctx_token` that a destination resolves to the click context (section 7.4). ``` Retrieved (HTTP layer, cacheable) @@ -368,7 +368,8 @@ Format: the URL of a manifest served at `/.well-known/content-telemetry.json` un | `output_id` | string | For reproduced/cited/presented | Opaque output-artifact identifier joining construction to later delivery | | `output_element_id` | string | No | Opaque element within `output_id`, such as a passage, media track, caption, link, or card | | `citation_id` | UUID | No | On `content_presented` or `content_reproduced`, the `id` of the associated citation event; absent for uncited presentations and uncredited reproductions | -| `presentation_id` | UUID | For engaged | On `content_engaged`, the `id` of the exact presentation occurrence acted upon | +| `presentation_id` | UUID | For engaged | On agent-reported `content_engaged`, the `id` of the exact presentation occurrence acted upon. Destination-reported events carrying an envelope `ctx_token` omit it (section 7.4) | +| `ctx_token` | string | No | On agent-reported `content_engaged`, the click token minted for this engagement's presentation, recorded so destination reports join to it (section 7.4) | | `source_role` | SourceRole | No | Who is reporting: `origin`, `edge`, `index`, `agent` (see 4.4) | | `content_telemetry_id` | UUID | No | Correlation ID for cross-observer deduplication (see 7.2) | | `content_url` | string | No | Content URL as fetched or canonical URL | @@ -805,7 +806,7 @@ When a session includes `content_presented` events but no subsequent `content_en |-------|------|-------------| | `engagement_type` | string | Type of interaction (see below) | -The content URL is identified by the event-level `content_url` field (section 5.2), not duplicated in `data`. Every engagement MUST carry `presentation_id`, referencing the exact `content_presented.id` on which the action occurred. Matching on URL alone is insufficient because the same source reference can be presented more than once. +The content URL is identified by the event-level `content_url` field (section 5.2), not duplicated in `data`. Every agent-reported engagement MUST carry `presentation_id`, referencing the exact `content_presented.id` on which the action occurred. Matching on URL alone is insufficient because the same source reference can be presented more than once. A destination-reported engagement carries `ctx_token` on its envelope instead: the destination cannot know the presentation UUID, and the telemetry consumer restores the binding from the token at resolution (section 7.4). #### Engagement types @@ -823,7 +824,7 @@ These are the core values. Extensions MAY define additional engagement actions - `link_click` is the primary signal for clickthrough rate calculation. Telemetry consumers can derive per-content-owner and aggregate clickthrough rates from `link_click` engagements and link presentations, joining each engagement through `presentation_id` rather than URL alone. -A `link_click` or `agent_navigate` engagement reported from the landing page after a click-out crosses a trust boundary. Such events carry a `ctx_token` in place of `session_id`, which the telemetry consumer resolves to the originating session's click manifest (see section 7.1). +A `link_click` or `agent_navigate` engagement reported from the landing page after a click-out crosses a trust boundary. Such events carry a `ctx_token` in place of `session_id`, which the telemetry consumer resolves to the click context (see section 7.4). ## 7. Transport @@ -891,9 +892,7 @@ Session documents use `"document_type": "session"`. When `document_type` is abse For origin-side emitters at Retrieval conformance level, `session_id` MAY be omitted when the content owner has no session context. Telemetry consumers correlate these events with agent-reported sessions using the `content_telemetry_id` field. -For `content_engaged` events emitted from a landing page after a click-out (typically by a content marketplace, affiliate network, or destination site), `session_id` MAY be replaced by a `ctx_token` field that carries an opaque click-token issued by the originating agent. Telemetry consumers resolve the token to the owning session. This lets a downstream observer report a corroborating engagement event without sharing the session UUID across trust boundaries. An event MUST carry either `session_id` or `ctx_token` at Grounding conformance and above. - -**ctx_token resolution.** A telemetry consumer that supports `ctx_token` resolution exposes, for a resolved token, the **click manifest**: the set of `content_grounded`, `content_cited`, and `content_presented` events belonging to the resolved session, identifying every source that informed the response that produced the click. The manifest is gated by the resolved session's `privacy_level` and by consent. A consumer MUST return the manifest only when the issuing agent has opted in to sharing sessions via click tokens; when the agent opt-in is absent, the consumer MUST NOT disclose the manifest. Within a returned manifest, a source MUST appear only when its content owner has opted in to being visible in click-token lookups; the consumer MUST withhold the events of any content owner whose opt-in is absent while returning the remainder of the manifest. A resolution response MUST NOT include the resolved session's raw `session_id`: the token exists so that the session UUID never crosses the trust boundary, and a resolution response that returned it would undo that. The mechanism by which an agent and a content owner record these opt-ins is operator-defined; the consent gate is normative. +For `content_engaged` events emitted from a landing page after a click-out (typically by a content marketplace, affiliate network, or destination site), `session_id` MAY be replaced by a `ctx_token` field that carries an opaque click token issued by the originating agent. This lets a downstream observer report a corroborating engagement event without sharing the session UUID across trust boundaries. An event MUST carry either `session_id` or `ctx_token` at Grounding conformance and above. Token issuance, carriage, binding, resolver discovery, and the resolution response are defined in section 7.4. The primary schema (`telemetry-session.json`) validates session documents. A standalone event envelope schema (`telemetry-event.json`) validates the event delivery format, and a batch envelope schema (`telemetry-event-batch.json`) validates the event batch format. All three schemas share the `TelemetryEvent` definition. @@ -952,6 +951,66 @@ Any party may operate a consumer: an agent operator, a licensing intermediary, o **Cross-consumer correlation.** Origin-side emitters and agent-side emitters MAY use different telemetry consumers. A content owner's CDN sends retrieval events to one telemetry consumer; an agent sends sessions to another. The `content_telemetry_id` field (section 7.2) correlates the same retrieval across consumers - both sides share the same UUID from the HTTP request. This correlation operates at the retrieval level only. Grounding, reproduction, citation, presentation, and engagement events have no independent origin-side counterpart to correlate against. +### 7.4 Click context (`ctx_token`) + +A click-out is the moment content usage becomes traffic the destination can observe. The click token lets the destination corroborate that moment and learn what produced it, without receiving the session UUID or any other publisher's activity. + +#### 7.4.1 Token issuance + +A `ctx_token` is an opaque token minted by the originating agent. Its value MUST match `^ct_[A-Za-z0-9_-]{8,240}$` and MUST NOT encode content, session, or user identifiers recoverable without the issuer's state. + +A token MUST be bound to exactly one `content_presented` occurrence at mint time. The same URL presented twice receives two tokens; a token observed on two presentations is malformed issuance and consumers MUST NOT resolve it. Surfaces that route outbound navigation through the agent SHOULD mint per click, additionally binding the token to the resulting `content_engaged` event. Direct-link surfaces mint per presentation; repeated clicks on one presentation then share a token, and are distinguished at resolution by the destination's event timestamps. + +The token-to-presentation binding is issuer state. It never travels in the URL: destinations do not receive `presentation_id`, and the consumer restores the binding at resolution. The agent SHOULD record the minted token on its own `content_engaged` event (the event-level `ctx_token` field) so the consumer can join destination reports to it. + +#### 7.4.2 Carriage and redirects + +Agents that decorate outbound link URLs MUST use the reserved query parameters `ctx_token` (the token) and `ctx_iss` (the issuer locator, section 7.4.3), and MUST NOT place other telemetry data in the URL. + +Parties operating redirects SHOULD propagate both parameters through same-domain redirect hops, mirroring the `Content-Telemetry-ID` redirect guidance in section 7.2. Destinations relying on redirect-based routing SHOULD capture the parameters at the earliest point in the chain. This is a transport and correlation convention: it does not claim that every intermediary preserved the value, and it does not enforce downstream behaviour. + +#### 7.4.3 Resolver discovery + +`ctx_iss` carries the issuer manifest locator: a host, optionally with a path prefix, identifying a well-known manifest location (section 8.1). `ctx_iss=example.com/agents/search` resolves to `https://example.com/agents/search/.well-known/content-telemetry.json`. That manifest declares the resolution endpoint in `telemetry.ctx_resolution` (section 8.5). + +The token stays opaque and carries no routing; the locator travels alongside it. No central registry is required or defined. + +#### 7.4.4 Resolution response - the click context + +A telemetry consumer that supports resolution exposes, for a presented token, the **click context**: + +1. **The engagement.** The `content_engaged` occurrence(s) bound to the token, including the presentation record the token restores: `presentation_id`, `output_id`, `output_element_id` where present, and timestamps. +2. **The lineage of the clicked content, selected by content identity.** The resolved session's `content_retrieved`, `content_grounded`, `content_reproduced`, `content_cited`, and `content_presented` events whose `content_url` or `content_id` identify the same content as the clicked reference - across all turns. The cut is by content identity, not by turn or click timestamp: a click in turn 5 on content grounded in turn 2 resolves that content's full lineage. +3. **An optional privacy-bounded session summary.** Event counts by type and a distinct-source count. Counts, not events, and no content identifiers of other owners. + +A resolution response MUST NOT include the resolved session's raw `session_id`: the token exists so the session UUID never crosses the trust boundary. A resolution response MUST NOT include events for content other than the clicked content; cross-content session detail is a reporting concern, delivered through publisher-filtered views (section 7.3), not through per-click resolution. A consumer MUST resolve a token only when the issuing agent has opted in to click-token resolution, and the response is gated by the resolved session's `privacy_level`. The mechanism by which the opt-in is recorded is operator-defined; the gate is normative. + +A worked resolution response (informative): + +``` +{ + "engagement": { + "engagement_type": "link_click", + "timestamp": "2026-08-10T09:02:31Z", + "presentation_id": "880e8400-e29b-41d4-a716-446655440213", + "output_id": "response:2" + }, + "lineage": [ + { "type": "content_grounded", "timestamp": "2026-08-10T09:00:01Z", "turn_id": "1", "data": { "chars_ingested": 9400 } }, + { "type": "content_cited", "timestamp": "2026-08-10T09:00:04Z", "turn_id": "1", "data": { "citation_type": "paraphrase", "position": "primary" } }, + { "type": "content_presented", "timestamp": "2026-08-10T09:00:04Z", "turn_id": "1", "data": { "presentation_kind": "source_reference", "presentation_type": "link" } }, + { "type": "content_presented", "timestamp": "2026-08-10T09:02:10Z", "turn_id": "2", "data": { "presentation_kind": "source_reference", "presentation_type": "link" } } + ], + "session_summary": { "turns": 2, "distinct_sources": 3, "events_by_type": { "content_grounded": 4, "content_cited": 3, "content_presented": 5 } } +} +``` + +The response shape above is informative in v1; the constraints in this section are normative. A response schema can follow implementation evidence during the release-candidate window. + +#### 7.4.5 Recorded limit: consumer custody + +Resolution depends on the telemetry consumer the agent chose, because that consumer holds the session. Grounding and citation events precede the click and cannot carry its later token, and distributing tokens to every contributing content owner after the fact would weaken the privacy boundary this section maintains. Publisher-derived tokens would require a federation and key-management design; that belongs in a later attribution or evidence profile. Core v1 mitigates the dependency with resolver discoverability (7.4.3) and exact click binding (7.4.1). + ## 8. Manifest Content owners, agents, and platforms publish a manifest declaring their identity and telemetry endpoints. The `manifest_ref` field on session documents (5.1.2) and the routing logic for origin-side emitters (7.3) resolve to manifests defined in this section. @@ -1017,6 +1076,7 @@ Public keys used to sign telemetry events emitted by this participant. Per-event |-------|------|----------|-------------| | `endpoint` | string | Yes | HTTPS URL. For agents and platforms, the outbound submission endpoint. For content owners, the inbound destination for events about the content owner's content. | | `conformance_level` | string | No | Conformance level advertised by this participant's own emitter(s). One of `retrieval`, `grounding`, `citation` (see 5.7). | +| `ctx_resolution` | string | No | HTTPS URL of the click-token resolution endpoint operated by or for this participant (see 7.4). Valid on `agent` and `platform` manifests. | `conformance_level` is informational. It advertises the level of telemetry the manifest's participant emits. It does **not** constrain what an inbound `endpoint` accepts - an endpoint accepts whatever events it is configured to accept, regardless of any level declared here - and it places **no requirement** on other emitters. On a `content_owner` manifest it describes only the events the owner's own infrastructure emits (typically a CDN edge worker at `retrieval`); it says nothing about what agents or platforms report about the owner's content, which those parties advertise in their own manifests. A `content_owner` manifest SHOULD omit `conformance_level` unless the owner operates its own emitter. There is no field for a content owner to *request* a minimum level from agents; consumers tolerate events from any level (see 5.7), and the protocol does not give a manifest a way to demand more. @@ -1289,6 +1349,17 @@ NOT synthesize a reproduction event from `false` or incomplete historical data. A grounding event MAY retain `content_fingerprint.scheme`, `detected`, and a scheme-defined `value`. +V1 narrows `ctx_token` resolution. The v0.1 click manifest returned every +source that informed the resolved session, gated by per-owner opt-in; the v1 +click context (section 7.4) returns the engagement, the clicked content's +lineage by content identity, and at most a count-based session summary. +Consumers implementing v0.1 resolution narrow their response shape and MUST +NOT return other owners' events through per-click resolution. Token values +gain the `ct_` pattern; presentation binding moves from the URL-carried +`presentation_id` a destination could never legitimately know to issuer state +restored at resolution. + + Preview versions (0.x) use two-component version numbers. From 1.0.0 onward, versions follow [semantic versioning](https://semver.org/): - **Major** (1.0.0 → 2.0.0) - breaking changes to required fields diff --git a/manifest.json b/manifest.json index 435a5e2..ddaa8d2 100644 --- a/manifest.json +++ b/manifest.json @@ -82,6 +82,12 @@ "type": "string", "enum": ["retrieval", "grounding", "citation"], "description": "Conformance level advertised by this participant's own emitter(s). Informational. (Sections 8.5, 5.7)" + }, + "ctx_resolution": { + "type": "string", + "format": "uri", + "pattern": "^https://", + "description": "HTTPS URL of the click-token resolution endpoint operated by or for this participant. Valid on agent and platform manifests. Destinations resolve ctx_iss to this manifest and present the token here. (Section 7.4)" } }, "description": "Telemetry endpoint declaration. (Section 8.5)" diff --git a/telemetry-event-batch.json b/telemetry-event-batch.json index 588c95f..2c2d987 100644 --- a/telemetry-event-batch.json +++ b/telemetry-event-batch.json @@ -28,7 +28,8 @@ }, "ctx_token": { "type": "string", - "description": "Opaque click-token issued by the originating agent, carried in place of session_id on batches of content_engaged events emitted from a landing page after a click-out. Applies to every event in the batch and is resolved by the telemetry consumer to the owning session. See section 7.1. An event MUST carry either session_id or ctx_token at Grounding conformance and above." + "pattern": "^ct_[A-Za-z0-9_-]{8,240}$", + "description": "Opaque click token issued by the originating agent, carried in place of session_id on batches of content_engaged events emitted from a landing page after a click-out. Applies to every event in the batch and is resolved by the telemetry consumer to the click context (section 7.4). An event MUST carry either session_id or ctx_token at Grounding conformance and above." }, "agent_id": { "type": "string", @@ -47,5 +48,25 @@ }, "description": "The telemetry events in the batch" } - } + }, + "allOf": [ + { + "if": { + "not": { "required": ["ctx_token"] } + }, + "then": { + "properties": { + "events": { + "items": { + "if": { + "properties": { "type": { "const": "content_engaged" } }, + "required": ["type"] + }, + "then": { "required": ["presentation_id"] } + } + } + } + } + } + ] } diff --git a/telemetry-event.json b/telemetry-event.json index fb21ff7..6ec0b53 100644 --- a/telemetry-event.json +++ b/telemetry-event.json @@ -28,7 +28,8 @@ }, "ctx_token": { "type": "string", - "description": "Opaque click-token issued by the originating agent, carried on content_engaged events emitted from a landing page after a click-out in place of session_id. Resolved by the telemetry consumer to the owning session (the click manifest). See section 7.1. An event MUST carry either session_id or ctx_token at Grounding conformance and above." + "pattern": "^ct_[A-Za-z0-9_-]{8,240}$", + "description": "Opaque click token issued by the originating agent, carried on content_engaged events emitted from a landing page after a click-out in place of session_id. Resolved by the telemetry consumer to the click context (section 7.4). An event MUST carry either session_id or ctx_token at Grounding conformance and above." }, "agent_id": { "type": "string", @@ -43,5 +44,24 @@ "$ref": "telemetry-session.json#/$defs/TelemetryEvent", "description": "The telemetry event" } - } + }, + "allOf": [ + { + "if": { + "required": ["event"], + "properties": { + "event": { + "properties": { "type": { "const": "content_engaged" } }, + "required": ["type"] + } + }, + "not": { "required": ["ctx_token"] } + }, + "then": { + "properties": { + "event": { "required": ["presentation_id"] } + } + } + } + ] } diff --git a/telemetry-session.json b/telemetry-session.json index 41c29e3..2f833a6 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -57,9 +57,18 @@ "events": { "type": "array", "items": { - "$ref": "#/$defs/TelemetryEvent" + "allOf": [ + { "$ref": "#/$defs/TelemetryEvent" }, + { + "if": { + "properties": { "type": { "const": "content_engaged" } }, + "required": ["type"] + }, + "then": { "required": ["presentation_id"] } + } + ] }, - "description": "Ordered list of events in the session (chronological by timestamp)" + "description": "Ordered list of events in the session (chronological by timestamp). Session documents are agent-reported, so content_engaged events here always carry presentation_id; the ctx_token relaxation applies only to standalone/batch envelopes (section 7.4)." } }, "$defs": { @@ -103,7 +112,12 @@ "presentation_id": { "type": "string", "format": "uuid", - "description": "The id of the exact content_presented event on which the engagement occurred. REQUIRED on content_engaged events." + "description": "The id of the exact content_presented event on which the engagement occurred. REQUIRED on agent-reported content_engaged events. Destination-reported events carrying ctx_token omit it; the consumer restores the binding at resolution (section 7.4)." + }, + "ctx_token": { + "type": "string", + "pattern": "^ct_[A-Za-z0-9_-]{8,240}$", + "description": "The click token the agent minted for this engagement's presentation, recorded so the consumer can join destination-reported events to it. Valid only on content_engaged events. See section 7.4." }, "source_role": { "$ref": "#/$defs/SourceRole", @@ -240,7 +254,6 @@ "required": ["type"] }, "then": { - "required": ["presentation_id"], "properties": { "data": { "properties": { diff --git a/tests/invalid/ctx-token-bad-pattern.json b/tests/invalid/ctx-token-bad-pattern.json new file mode 100644 index 0000000..b286ce5 --- /dev/null +++ b/tests/invalid/ctx-token-bad-pattern.json @@ -0,0 +1,12 @@ +{ + "_test_description": "ctx_token failing the value pattern: tokens are opaque but MUST match ^ct_[A-Za-z0-9_-]{8,240}$ so they survive URL carriage and are recognisable in logs (section 7.4). This value has no ct_ prefix and contains reserved characters.", + "document_type": "event", + "schema_version": "0.1", + "ctx_token": "session=660e8400!", + "event": { + "type": "content_engaged", + "timestamp": "2026-08-10T11:20:00Z", + "content_url": "https://news.example.com/markets/rate-decision", + "data": { "engagement_type": "link_click" } + } +} diff --git a/tests/invalid/engaged-standalone-missing-presentation-and-token.json b/tests/invalid/engaged-standalone-missing-presentation-and-token.json new file mode 100644 index 0000000..fb7e85e --- /dev/null +++ b/tests/invalid/engaged-standalone-missing-presentation-and-token.json @@ -0,0 +1,14 @@ +{ + "_test_description": "Agent-reported standalone content_engaged with session_id but no presentation_id. The presentation_id relaxation applies only to envelopes carrying ctx_token; an emitter that knows the session knows the presentation and MUST bind to it (sections 6.8, 7.4).", + "document_type": "event", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440230", + "agent_id": "assistant-v2", + "started_at": "2026-08-10T12:00:00Z", + "event": { + "type": "content_engaged", + "timestamp": "2026-08-10T12:00:30Z", + "content_url": "https://news.example.com/markets/rate-decision", + "data": { "engagement_type": "link_click" } + } +} diff --git a/tests/valid/event-standalone-engaged-ctx-token.json b/tests/valid/event-standalone-engaged-ctx-token.json index f73d6af..ce90f25 100644 --- a/tests/valid/event-standalone-engaged-ctx-token.json +++ b/tests/valid/event-standalone-engaged-ctx-token.json @@ -1,12 +1,11 @@ { - "_test_description": "Standalone content_engaged (link_click) reported from a landing page after a click-out. Carries ctx_token in place of session_id (section 7.1): the destination corroborates the click without receiving the session UUID. The telemetry consumer resolves ctx_token to the originating session's click manifest.", + "_test_description": "Standalone content_engaged (link_click) reported from a landing page after a click-out. Carries ctx_token in place of session_id and no presentation_id: the destination cannot know the presentation UUID, and the consumer restores the token's presentation binding at resolution (section 7.4).", "document_type": "event", "schema_version": "0.1", "ctx_token": "ct_9f3a1c7e2b8d4a06", "event": { "type": "content_engaged", "timestamp": "2026-03-28T14:06:00Z", - "presentation_id": "990e8400-e29b-41d4-a716-446655440064", "content_url": "https://www.example-review.com/headphones/best-noise-cancelling", "data": { "engagement_type": "link_click" diff --git a/tests/valid/event-standalone-engaged-redirect.json b/tests/valid/event-standalone-engaged-redirect.json new file mode 100644 index 0000000..a7ba533 --- /dev/null +++ b/tests/valid/event-standalone-engaged-redirect.json @@ -0,0 +1,14 @@ +{ + "_test_description": "Destination-reported click after a redirect chain: the presented short link redirected to the canonical URL, and the ctx_token and ctx_iss query parameters were propagated through the same-domain hops (section 7.4). The destination reports the canonical URL it serves; correlation runs through the token, not the URL.", + "document_type": "event", + "schema_version": "0.1", + "ctx_token": "ct_77b41f0ac93e5d28", + "event": { + "type": "content_engaged", + "timestamp": "2026-08-10T11:15:42Z", + "content_url": "https://shop.example.com/products/anc-headphones-x9", + "data": { + "engagement_type": "link_click" + } + } +} diff --git a/tests/valid/manifest-agent-ctx-resolution.json b/tests/valid/manifest-agent-ctx-resolution.json new file mode 100644 index 0000000..096fc7c --- /dev/null +++ b/tests/valid/manifest-agent-ctx-resolution.json @@ -0,0 +1,12 @@ +{ + "_test_description": "Agent manifest declaring a click-token resolution endpoint in telemetry.ctx_resolution. A destination receiving ctx_iss=searchco.com/agents/web-search resolves this manifest and presents the token at the declared endpoint (sections 7.4, 8.5).", + "schema_version": "0.1", + "id": "https://searchco.com/agents/web-search/.well-known/content-telemetry.json", + "roles": ["agent"], + "operator": { "name": "SearchCo" }, + "telemetry": { + "endpoint": "https://telemetry.example.com/v1/events", + "conformance_level": "citation", + "ctx_resolution": "https://telemetry.example.com/v1/ctx/resolve" + } +} diff --git a/tests/valid/session-earlier-turn-click.json b/tests/valid/session-earlier-turn-click.json new file mode 100644 index 0000000..a6b3ee5 --- /dev/null +++ b/tests/valid/session-earlier-turn-click.json @@ -0,0 +1,53 @@ +{ + "_test_description": "A click in a later turn on a presentation from an earlier turn: the content_engaged event carries the turn of the click but binds by presentation_id to the turn-1 presentation. Resolution returns the clicked content's lineage across turns by content identity, not a turn or timestamp cut (section 7.4).", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440220", + "agent_id": "assistant-v2", + "started_at": "2026-08-10T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-10T10:00:01Z", + "turn_id": "1", + "content_id": "publisher:feature:512", + "data": { "scope": "session", "chars_ingested": 15200 } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440221", + "type": "content_cited", + "timestamp": "2026-08-10T10:00:05Z", + "turn_id": "1", + "output_id": "response:1", + "content_id": "publisher:feature:512", + "data": { "citation_type": "reference", "position": "primary" } + }, + { + "id": "880e8400-e29b-41d4-a716-446655440222", + "type": "content_presented", + "timestamp": "2026-08-10T10:00:05Z", + "turn_id": "1", + "output_id": "response:1", + "citation_id": "770e8400-e29b-41d4-a716-446655440221", + "content_id": "publisher:feature:512", + "content_url": "https://publisher.example.com/features/512", + "data": { "presentation_kind": "source_reference", "presentation_type": "link" } + }, + { + "type": "content_grounded", + "timestamp": "2026-08-10T10:03:00Z", + "turn_id": "3", + "content_id": "otherpub:brief:77", + "data": { "scope": "turn", "chars_ingested": 2100 } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-10T10:04:12Z", + "turn_id": "3", + "presentation_id": "880e8400-e29b-41d4-a716-446655440222", + "ctx_token": "ct_e5a0b6d92c17f480", + "content_id": "publisher:feature:512", + "content_url": "https://publisher.example.com/features/512", + "data": { "engagement_type": "link_click" } + } + ] +} diff --git a/tests/valid/session-repeated-link-engagement.json b/tests/valid/session-repeated-link-engagement.json new file mode 100644 index 0000000..19669c7 --- /dev/null +++ b/tests/valid/session-repeated-link-engagement.json @@ -0,0 +1,54 @@ +{ + "_test_description": "The same URL presented twice in one session: two content_presented events with distinct ids and distinct minted click tokens, and a content_engaged bound by presentation_id to the second occurrence. Matching on URL alone could not tell the two presentations apart (sections 6.8, 7.4).", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440210", + "agent_id": "assistant-v2", + "started_at": "2026-08-10T09:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-10T09:00:01Z", + "turn_id": "1", + "content_url": "https://news.example.com/markets/rate-decision", + "content_id": "newsex:article:8841", + "data": { "scope": "session", "chars_ingested": 9400 } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440211", + "type": "content_cited", + "timestamp": "2026-08-10T09:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://news.example.com/markets/rate-decision", + "data": { "citation_type": "paraphrase", "position": "primary" } + }, + { + "id": "880e8400-e29b-41d4-a716-446655440212", + "type": "content_presented", + "timestamp": "2026-08-10T09:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "citation_id": "770e8400-e29b-41d4-a716-446655440211", + "content_url": "https://news.example.com/markets/rate-decision", + "data": { "presentation_kind": "source_reference", "presentation_type": "link" } + }, + { + "id": "880e8400-e29b-41d4-a716-446655440213", + "type": "content_presented", + "timestamp": "2026-08-10T09:02:10Z", + "turn_id": "2", + "output_id": "response:2", + "content_url": "https://news.example.com/markets/rate-decision", + "data": { "presentation_kind": "source_reference", "presentation_type": "link" } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-10T09:02:31Z", + "turn_id": "2", + "presentation_id": "880e8400-e29b-41d4-a716-446655440213", + "ctx_token": "ct_r2d1a9c44be07f31", + "content_url": "https://news.example.com/markets/rate-decision", + "data": { "engagement_type": "link_click" } + } + ] +} From af939c9d42eeaa53fdfbf7b1381482eb0bae3c47 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Thu, 13 Aug 2026 21:01:37 +0100 Subject: [PATCH 18/38] Click context: gated contributing sources; owner-scoped engagement delivery (#28, PR #38 review) Two changes from Pedro Santos's review of #38, both from #28. Contributing sources (7.4.4). Lineage by content identity alone cannot support the commerce case: a click through to a destination contributed by another owner's content resolves nothing about the contributor. The click context regains the v0.1 click manifest's role as a consent-gated contributing-source set, narrowed to the engaged presentation's turn and gated per contributing owner's recorded opt-in. Owners without an opt-in remain visible only as counts. The migration note in 12.1 now records a scoping plus the retained gate, not a removal. Owner-scoped delivery (new 7.4.5). The agent-authored click content_engaged is delivered to the engaged content's owner through 7.3 filtered views on the same terms as grounding and citation events, session_id included. The 7.4.4 session_id prohibition binds token resolution (possession of a URL-carried value), not 7.3 delivery to a verified domain owner. This survives token loss in redirect chains, removes the resolver dependency for the destination's own join, and lets a party holding both a contributor's and a destination's owner-scoped streams match on session_id for attribution. Not token distribution: only the clicked content's owner receives the event. Recorded limit renumbered to 7.4.6. Conformance suite 74/74, worked examples 9/9. Co-Authored-By: Claude Fable 5 --- SPECIFICATION.md | 33 +++++++++++++++++++++++++-------- 1 file changed, 25 insertions(+), 8 deletions(-) diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 4121219..fbf71a2 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -824,7 +824,7 @@ These are the core values. Extensions MAY define additional engagement actions - `link_click` is the primary signal for clickthrough rate calculation. Telemetry consumers can derive per-content-owner and aggregate clickthrough rates from `link_click` engagements and link presentations, joining each engagement through `presentation_id` rather than URL alone. -A `link_click` or `agent_navigate` engagement reported from the landing page after a click-out crosses a trust boundary. Such events carry a `ctx_token` in place of `session_id`, which the telemetry consumer resolves to the click context (see section 7.4). +A `link_click` or `agent_navigate` engagement reported from the landing page after a click-out crosses a trust boundary. Such events carry a `ctx_token` in place of `session_id`, which the telemetry consumer resolves to the click context (see section 7.4). The agent-authored engagement itself reaches the engaged content's owner through owner-scoped routing whether or not the token survived the redirect chain (section 7.4.5). ## 7. Transport @@ -981,9 +981,10 @@ A telemetry consumer that supports resolution exposes, for a presented token, th 1. **The engagement.** The `content_engaged` occurrence(s) bound to the token, including the presentation record the token restores: `presentation_id`, `output_id`, `output_element_id` where present, and timestamps. 2. **The lineage of the clicked content, selected by content identity.** The resolved session's `content_retrieved`, `content_grounded`, `content_reproduced`, `content_cited`, and `content_presented` events whose `content_url` or `content_id` identify the same content as the clicked reference - across all turns. The cut is by content identity, not by turn or click timestamp: a click in turn 5 on content grounded in turn 2 resolves that content's full lineage. -3. **An optional privacy-bounded session summary.** Event counts by type and a distinct-source count. Counts, not events, and no content identifiers of other owners. +3. **The contributing sources, gated per owner.** The sources that informed the response the click came from: `content_grounded` events in scope for the engaged presentation's turn (including session-scoped groundings) and that turn's `content_cited` and `content_presented` events, for content other than the clicked content. A contributing owner's events appear only when that owner has opted in to contributing-source disclosure with the resolving consumer; owners without a recorded opt-in are visible only through the counts in the session summary. This is the component that supports multi-citation attribution when the clicked content is not the contributing content - a click through to a commerce destination whose recommendation a publisher's review produced - and it restores the consent-gated role of the v0.1 click manifest (section 12.1), scoped to the click's provenance rather than the whole session. +4. **An optional privacy-bounded session summary.** Event counts by type and a distinct-source count. Counts, not events, and no content identifiers of owners not disclosed above. -A resolution response MUST NOT include the resolved session's raw `session_id`: the token exists so the session UUID never crosses the trust boundary. A resolution response MUST NOT include events for content other than the clicked content; cross-content session detail is a reporting concern, delivered through publisher-filtered views (section 7.3), not through per-click resolution. A consumer MUST resolve a token only when the issuing agent has opted in to click-token resolution, and the response is gated by the resolved session's `privacy_level`. The mechanism by which the opt-in is recorded is operator-defined; the gate is normative. +A resolution response MUST NOT include the resolved session's raw `session_id`: the token exists so the session UUID never crosses the trust boundary. Outside the contributing-source component, a resolution response MUST NOT include events for content other than the clicked content; within it, a response MUST NOT include events for an owner without a recorded contributing-source opt-in. Whole-session cross-content detail remains a reporting concern, delivered through publisher-filtered views (section 7.3), not through per-click resolution. A consumer MUST resolve a token only when the issuing agent has opted in to click-token resolution, and the response is gated by the resolved session's `privacy_level`. The mechanisms by which the issuer and contributing-owner opt-ins are recorded are operator-defined; both gates are normative. A worked resolution response (informative): @@ -1001,13 +1002,28 @@ A worked resolution response (informative): { "type": "content_presented", "timestamp": "2026-08-10T09:00:04Z", "turn_id": "1", "data": { "presentation_kind": "source_reference", "presentation_type": "link" } }, { "type": "content_presented", "timestamp": "2026-08-10T09:02:10Z", "turn_id": "2", "data": { "presentation_kind": "source_reference", "presentation_type": "link" } } ], + "contributing_sources": [ + { + "content_url": "https://publisher-a.example/heaters/space-heater-review", + "events": [ + { "type": "content_grounded", "timestamp": "2026-08-10T09:02:08Z", "turn_id": "2", "data": { "chars_ingested": 7200 } }, + { "type": "content_cited", "timestamp": "2026-08-10T09:02:10Z", "turn_id": "2", "data": { "citation_type": "paraphrase", "position": "primary" } } + ] + } + ], "session_summary": { "turns": 2, "distinct_sources": 3, "events_by_type": { "content_grounded": 4, "content_cited": 3, "content_presented": 5 } } } ``` The response shape above is informative in v1; the constraints in this section are normative. A response schema can follow implementation evidence during the release-candidate window. -#### 7.4.5 Recorded limit: consumer custody +#### 7.4.5 Owner-scoped delivery of the engagement + +The agent-authored click `content_engaged` is a session event like any other: it reaches content owners through routing and aggregation (section 7.3), independent of whether the token in the URL survived the redirect chain. A telemetry consumer that provides owner-filtered views MUST include the agent-authored `content_engaged` event in the filtered view of the engaged content's owner, on the same terms as `content_grounded` and `content_cited` events - including the session identifier that owner-scoped delivery carries. The `session_id` prohibition in section 7.4.4 binds token resolution, where the requesting party is authenticated by nothing more than possession of a URL-carried value; it does not bind section 7.3 delivery to an owner whose domain registration the consumer has verified. + +This delivery is deliberately redundant with the token path. It notifies the destination owner of the click even when `ctx_token` was stripped in transit; it lets that owner join the click to their own grounded and cited events on `session_id` without calling a resolver; and it lets a party processing owner-scoped streams for both a contributing publisher and a click destination match its clients' events on `session_id` for attribution. The owner's filtered view SHOULD carry the event-level `ctx_token`, so a destination that captured the query parameters at landing can join the URL-channel observation to the server-side event directly. This is not token distribution to contributing owners (the limit recorded in section 7.4.6): only the owner of the clicked content receives the event, and that owner already saw the token in the URL. + +#### 7.4.6 Recorded limit: consumer custody Resolution depends on the telemetry consumer the agent chose, because that consumer holds the session. Grounding and citation events precede the click and cannot carry its later token, and distributing tokens to every contributing content owner after the fact would weaken the privacy boundary this section maintains. Publisher-derived tokens would require a federation and key-management design; that belongs in a later attribution or evidence profile. Core v1 mitigates the dependency with resolver discoverability (7.4.3) and exact click binding (7.4.1). @@ -1352,14 +1368,15 @@ scheme-defined `value`. V1 narrows `ctx_token` resolution. The v0.1 click manifest returned every source that informed the resolved session, gated by per-owner opt-in; the v1 click context (section 7.4) returns the engagement, the clicked content's -lineage by content identity, and at most a count-based session summary. -Consumers implementing v0.1 resolution narrow their response shape and MUST -NOT return other owners' events through per-click resolution. Token values +lineage by content identity, a contributing-source set under the same +per-owner opt-in gate - scoped to the turn the click came from rather than +the whole session - and at most a count-based session summary. Consumers +implementing v0.1 resolution narrow the contributing-source scope accordingly +and MUST NOT return events for owners without a recorded opt-in. Token values gain the `ct_` pattern; presentation binding moves from the URL-carried `presentation_id` a destination could never legitimately know to issuer state restored at resolution. - Preview versions (0.x) use two-component version numbers. From 1.0.0 onward, versions follow [semantic versioning](https://semver.org/): - **Major** (1.0.0 → 2.0.0) - breaking changes to required fields From 01d2a4d74e01aabeb758bb0f44037cec8448d485 Mon Sep 17 00:00:00 2001 From: Alex Springer Date: Thu, 13 Aug 2026 21:39:50 +0100 Subject: [PATCH 19/38] Terms, coverage and identity: terms_ref, occurrence boundaries, access_context, content_depth (#3, #42, #43, #44) Terms reference (new 5.2.4, #3). Optional event-level terms_ref, distinct from license_ref: the terms are the basis, the licence is the proof. A public URL or an opaque identifier both parties can resolve; private resolution conforms. Core carries the reference and does not resolve, validate or interpret it, and the terms it points at do not redefine core event semantics. Event-level only: a session-level default was considered and deferred because sessions legitimately span content under different terms. Occurrence and coverage (#42). Section 4.3 now states one occurrence boundary per core event: retrieval per completed fetch, grounding per distinct content item per declared scope, reproduction per source item per output element, presentation per rendering, engagement per observed action. "Qualifying occurrence" is defined, a conformance level is stated to be a capability claim rather than a coverage claim, and coverage becomes an explicit declaration (new 5.7.6): complete, sampled, aggregated or selected, over a stated relationship scope, with the rule for the last three objectively decidable at emission time and disclosed. Manifests MAY declare the modes machine-readably (8.5). Governing terms select events and coverage; they do not redefine semantics (SCOPE.md, the boundary recorded in #4). Envelope identity. manifest_ref joins the standalone event and event batch envelopes (7.1), the only envelope field naming the manifest - and so the domain - under which the emitter claims to report. SHOULD-level for standalone delivery under settlement or audit obligations. Session access context (new 5.1.3, #43). Sessions gain a data container mirroring the event-level field, closing the accidental extension point at the session root (11.1, 12.1). One core container: access_context, typed institutional identifiers (ror, saml_entity_id, isni, open vocabulary) - the institution whose access rights the session used, never an individual. Populated only where governing terms require it, placed inside the privacy model in 5.5 with intent/minimal pairing guidance. Content depth (6.1, #44). content_depth records how much of a content record a retrieval reached: metadata, abstract or full, open values, applicable from any source_role. Depth is what was reachable at retrieval, independent of what later entered a generation context. Four new fixtures. Conformance suite 78/78, worked examples 10/10. Co-Authored-By: Claude Fable 5 --- SPECIFICATION.md | 96 ++++++++++++++++++- manifest.json | 19 ++++ telemetry-event-batch.json | 4 + telemetry-event.json | 4 + telemetry-session.json | 36 +++++++ ...cess-context-identifier-missing-value.json | 15 +++ .../access-context-identifiers-not-array.json | 13 +++ tests/valid/event-standalone-terms-ref.json | 15 +++ tests/valid/session-access-context.json | 64 +++++++++++++ 9 files changed, 264 insertions(+), 2 deletions(-) create mode 100644 tests/invalid/access-context-identifier-missing-value.json create mode 100644 tests/invalid/access-context-identifiers-not-array.json create mode 100644 tests/valid/event-standalone-terms-ref.json create mode 100644 tests/valid/session-access-context.json diff --git a/SPECIFICATION.md b/SPECIFICATION.md index fbf71a2..c1a0c00 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -111,6 +111,9 @@ For the purposes of this specification, the following terms apply. | **privacy level** | data sharing tier controlling which conversation fields are populated: `full`, `summary`, `intent`, or `minimal` (section 5.5) | | **conformance level** | emitter capability tier: Retrieval, Grounding, or Citation (section 5.7) | | **content scope** | opaque identifier grouping sessions by their content access context (section 5.1.1) | +| **governing terms** | licence, contract or other terms selecting which events a relationship requires and the coverage, cadence, delivery, privacy and reports owed over the shared semantics (sections 5.2.4, 5.7.6) | +| **qualifying occurrence** | occurrence satisfying an event type's core definition and occurrence boundary within the relationship scope reported under (section 5.7.6) | +| **coverage** | declared relationship between qualifying occurrences and emitted events: `complete`, `sampled`, `aggregated` or `selected` (section 5.7.6) | ## 4. Concepts @@ -195,26 +198,38 @@ Content moves through six stages during an agent interaction: 1. **Retrieved** - Content fetched over HTTP from an origin server, CDN, marketplace, or index. This is an infrastructure event observable by the content owner's infrastructure (origin server, edge network) and the agent. A retrieval may be cached by the agent for use across multiple sessions. + One retrieval occurrence is one completed fetch of a content representation as observed by the reporting party: a redirect chain resolving to one representation is one occurrence, and a revalidation returning no new representation (an HTTP 304) is not a new occurrence. Serving content from the agent's own cache is not a new retrieval; the reuse surfaces as grounding (stage 2), not as a repeated `content_retrieved` event. + 2. **Grounded** - Content used in the agent's generation context for this session or turn. The boundary is "this content entered the generation model's context" - the point where content can directly influence the model's output. Content used only for retrieval selection (embedding similarity search, re-ranking scores, routing decisions) without entering the generation context is not grounded. Grounding is architecture-neutral: same event whether the agent uses RAG, chain-of-thought reasoning, embeddings, or multi-step delegation (see section 6.4 for architecture-specific guidance). Grounding is decoupled from retrieval: content may be grounded from a live fetch, from agent-side cache, or from a pre-loaded index. Only the agent can report grounding events. + One grounding occurrence is one distinct content item entering a generation context at the declared `data.scope`: at `session` scope, a content item grounds once per session; at `turn` scope, once per turn it enters. A distinct content item is a distinct `content_id`, or its canonical `content_url` where no stable identifier exists (section 4.5). Continued presence within the declared scope is not a further occurrence; re-entry in a later turn is, when the scope is `turn`, and a change of `content_version` opens a new occurrence at either scope. An emitter that ingests a content item in chunks MAY emit one grounding event per chunk, preserving the chunk-level hashes of section 6.4; events sharing content identity within one scope describe one occurrence, and consumers count occurrences by deduplicating on content identity and scope, not by counting events. + 3. **Reproduced** - The output artifact contains identified source content: a quotation, an excerpt, or a full copy, verbatim or near-verbatim. Reproduction is an output-construction claim by the system that built the output. It is independent of credit and of delivery: reproduced content may or may not also be cited, and the artifact may or may not later be presented. An uncredited excerpt in a response delivered through an API produces a `content_reproduced` event and nothing else - without this event, that use would be unreportable. + One reproduction occurrence is one identified source content item appearing in one output element (or in the output artifact, where no element identity exists), so a credited quotation and its companion citation share an `output_element_id` (section 6.6). Three quoted passages from one source in three elements are three occurrences; repeated appearance within one element is not a further occurrence. + 4. **Cited** - An output artifact explicitly associates identified source content with a response, claim, passage, quotation, or other output element. Citation is an output-construction relationship, not evidence that the output was delivered. A subset of grounded content is commonly cited, but a citation can also be emitted without a matching grounding event when an agent produces an uncorroborated or hallucinated source association. A citation MUST carry a resolvable reference to the source it associates: a `content_url` or a `content_id`. A source association with no resolvable reference is not a citation and MUST NOT be emitted as `content_cited`. Unlike other content events, where the identifier requirement is an application-layer rule (section 5.7.5), for `content_cited` and `content_reproduced` it is enforced by the JSON Schema. Reproduction and citation are sibling claims about the same artifact: reproduction records the material, citation records the credit. Neither implies the other. A credited quotation produces both events; an uncredited excerpt produces only a reproduction; a reference citation with no quoted material produces only a citation. + One citation occurrence is one distinct association between a source and an output element (or the output artifact, where no element identity exists). Associating the same source with three separate output elements produces three citation events; repeating the same association is not a further occurrence. + 5. **Presented** - Content or a source reference was rendered, played, spoken, embedded, or otherwise made perceivable on a recipient-facing surface. Presentation does not assert that a person noticed or attended to it. `presentation_kind` distinguishes source content (including a reproduced excerpt or media) from a source reference (such as a link, credit, or card). Not all citations are presented: an output can be stored, suppressed, or passed to another system before delivery. Grounding and presentation record different boundary crossings: grounding records entry into a generation context, while presentation records a recipient-facing delivery occurrence. As agent experiences evolve beyond the chat window the two diverge - content can shape an answer whose source is never presented, and an agent can present content that never entered a generation context (see *Departures from the funnel model* below). + One presentation occurrence is one rendering of content or a source reference on a recipient-facing surface; the event's `id` names that occurrence. Presenting the same artifact again - on a new surface, or in a new delivery - is a new occurrence. + 6. **Engaged** - The recipient or agent performed an observable action on a presentation: clicked a link, expanded a preview, copied text, shared the response, or directed the agent to act on the content. It does not imply attention beyond the reported action. `presentation_id` links the action to the exact presentation occurrence; a click-out can also carry a `ctx_token` that a destination resolves to the click context (section 7.4). + One engagement occurrence is one observed action on one presentation occurrence. + ``` Retrieved (HTTP layer, cacheable) → Grounded (influence layer, per-session or per-turn) @@ -232,6 +247,8 @@ Each stage after retrieval is typically a progressively narrower subset, except - **Citation-to-presentation** measures source associations constructed in output but not made perceivable - **Presentation-to-engagement** measures observable actions on exact presentation occurrences +These ratios are computed over reported events. They are comparable across emitters, and meaningful as measures of behaviour rather than of reporting, only at known coverage (section 5.7.6). + #### Departures from the funnel model Five cases break the strict subset model: @@ -323,6 +340,7 @@ Additional content metadata - version, last-modified timestamp, content hash, me | `ended_at` | datetime | No | Session end (UTC) | | `conformance_level` | string | No | Informational conformance level advertised by the emitter (see section 5.7). Values: `retrieval`, `grounding`, `citation` | | `document_type` | string | No | `"session"` for session documents (see section 7.1 for the standalone event and event batch formats) | +| `data` | object | No | Session-level extension container, including access context (see 5.1.3) | | `events` | Event[] | No | Ordered list of events | `parent_session_id` links a delegated session to its immediate parent without @@ -357,6 +375,36 @@ The `manifest_ref` field optionally references a manifest (section 8), identifyi Format: the URL of a manifest served at `/.well-known/content-telemetry.json` under a path the participant controls. +#### 5.1.3 Session data and access context + +Sessions carry an optional `data` object mirroring the event-level `data` field (section 11.1): an extension container for session-scoped metadata. Extensions SHOULD namespace custom fields or use containers documented in this specification, and consumers MUST tolerate unknown fields within it. The session root itself is not an extension point: custom top-level siblings of `events` are not defined by this specification, and consumers are not required to preserve or interpret them. + +One container is defined in core. `access_context` records the context from which the session's access rights derive - the institution, not the individual: + +```json +{ + "schema_version": "0.1", + "session_id": "770e8400-e29b-41d4-a716-446655440000", + "content_scope": "consortium-agreement-4471", + "started_at": "2026-08-13T14:02:10Z", + "data": { + "access_context": { + "identifiers": [ + { "scheme": "ror", "value": "https://ror.org/013meh722" }, + { "scheme": "saml_entity_id", "value": "https://idp.example.ac.uk/shibboleth" } + ] + } + }, + "events": [] +} +``` + +`identifiers` is an array of typed identifiers, each a `scheme` and a `value`. `ror`, `saml_entity_id` and `isni` are the core scheme values; emitters MAY use other schemes and telemetry consumers MUST tolerate unknown ones, as with `media_type` (section 6.1). It is an array because access rights arrive through consortia, federated identity and proxies at once: a session may legitimately carry a SAML entity ID and the ROR ID it maps to. + +`access_context` identifies an institution, never an individual. Like everything else in a session it is a claim by the emitter. Where the content owner authenticated the session itself (`source_role` of `origin` or `edge`), the owner already knows the institution and does not need the field; it earns its place where a third-party agent holds the entitlement and asserts the affiliation to the content owner, which is also where the claim is least verifiable. What corroborates an asserted affiliation is verification-layer work, outside core. + +Emitters MUST NOT populate `access_context` unless the governing terms of the relationship require it, and SHOULD pair it with `intent` or `minimal` conversation-turn data (section 5.5): an identified institution combined with query text can come close to identifying an individual at a small subscriber. + ### 5.2 Event | Field | Type | Required | Description | @@ -375,6 +423,7 @@ Format: the URL of a manifest served at `/.well-known/content-telemetry.json` un | `content_url` | string | No | Content URL as fetched or canonical URL | | `content_id` | string | No | Content owner's stable content identifier (see 4.5) | | `license_ref` | string | No | Reference to a licence or grant the emitter associates with this event (see 5.2.3) | +| `terms_ref` | string | No | Reference to the governing terms the emitter associates with this event (see 5.2.4) | | `turn` | ConversationTurn | No | Conversation data (for turn events) | | `data` | object | No | Type-specific metadata (see section 6) | @@ -398,6 +447,14 @@ Core does not resolve, validate or interpret the reference. `license_ref` is par `license_ref` also does not identify the party whose entitlement was used. Where a publisher issues one grant per subscriber the value may work as a proxy for that subscriber, but only within the issuing publisher's namespace: nothing here requires the value to be typed, stable across sessions, or comparable between emitters. +#### 5.2.4 Terms reference + +The `terms_ref` field associates a telemetry event with the governing terms under which it is reported: a licence agreement, a standard-form contract, a tariff, a profile's terms, or any other terms document. Like `license_ref`, the value MAY be a public URL or an opaque identifier that both parties can resolve. Nothing requires the terms to be published: per-relationship terms are often confidential, and an opaque identifier resolved privately is a conforming reference. + +The terms are the basis; the licence is the proof. `license_ref` records which grant the emitter says applied (5.2.3). `terms_ref` records which terms govern the event's commercial consequences and the emitter's reporting obligations. Either may appear without the other: an access outside any grant carries no `license_ref`, and can still carry the `terms_ref` of the terms that attach consequences to that access. + +Core does not resolve, validate or interpret the reference, and `terms_ref` does not redefine core event semantics. Governing terms select which events a relationship requires and at what coverage (section 5.7.6), together with the cadence, delivery, privacy and reports owed (SCOPE.md); the meaning and occurrence boundary of each event remain those defined in sections 4.3 and 6, whatever `terms_ref` points to. + ### 5.3 Event types #### Content events @@ -466,6 +523,8 @@ These are the recommended values. Platforms with additional product surfaces (co An emitter that populates a conversation turn MUST NOT include a field above that turn's declared `privacy_level` - for example, `query_text` MUST NOT be present when `privacy_level` is `intent` or `minimal`. This restriction is a property of `privacy_level` itself: it applies wherever conversation turns are emitted, independent of the emitter's conformance level. +The session-level `access_context` container (section 5.1.3) sits outside this table but inside the privacy model. It is populated only where governing terms require it, and pairing it with `full` or `summary` turn data is discouraged (section 5.1.3): the ladder governs how much of the query and response is visible, `access_context` governs whose access rights the session used, and the two disclosures compound. + **Token counts** includes `query_tokens` and `response_tokens`. These are available at all levels because they are needed for token-based counting models and do not reveal user intent or platform strategy. They carry the same portability limit as `tokens_ingested` (section 6.4): both are measured in the emitter's own tokeniser and are not comparable between agents. Version 1 does not define corresponding turn-level character counts; `chars_ingested` measures source content placed in a generation context, not query or response length. **Response classification** includes `response_type` (e.g., `"recommendation"`, `"explanation"`). Available at `intent` level and above, as it can reveal the nature of the user's query. @@ -492,7 +551,7 @@ These are the core values. Extensions MAY define additional intent category valu Emitters that advertise a standard capability tier use one of three conformance levels. The authoritative declaration lives in the emitter's manifest (section 8). Emitters MAY also include an optional `conformance_level` field on individual session documents; when present it is informational and consumers MUST NOT treat it as a substitute for verifying the manifest's declaration. -Each level is named for the event it adds: a level proves the emitter produces that event and everything below it. These levels describe what an emitter reports, not what a consumer computes from it; attribution - the apportioning of credit across content - is performed by a telemetry consumer at whatever funnel level the parties agree (section 10), and can be computed from grounding alone, without citation. An emitter does not need to reach the Citation level for its telemetry to support attribution. +Each level is named for the event it adds: a level proves the emitter produces that event and everything below it. A level is a capability claim, not a coverage claim: it does not assert that every qualifying occurrence was reported (section 5.7.6). These levels describe what an emitter reports, not what a consumer computes from it; attribution - the apportioning of credit across content - is performed by a telemetry consumer at whatever funnel level the parties agree (section 10), and can be computed from grounding alone, without citation. An emitter does not need to reach the Citation level for its telemetry to support attribution. | Level | Events | What it proves | Typical emitter | |-------|--------|----------------|-----------------| @@ -566,6 +625,25 @@ The JSON Schema (`telemetry-session.json`) validates structure and types but can The `tests/` directory provides an informative reference suite for these rules. A consumer that receives a privacy-violating turn (e.g., `query_text` present at `minimal` level) SHOULD strip the offending fields rather than reject the document carrying them. +#### 5.7.6 Occurrence, qualifying events and coverage + +Each core event type has the meaning and occurrence boundary defined in sections 4.3 and 6. Profiles, deployment configurations and governing terms MUST NOT redefine them. A relationship that needs a different assertion defines a namespaced extension event (sections 5.3 and 11.1); it does not reuse a core type with altered semantics. + +An occurrence is **qualifying** for an emitter when it satisfies the core definition and occurrence boundary of its event type and falls within the relationship scope the emitter reports under - the content, domains or relationships selected by the applicable governing terms or deployment configuration. This is the sense in which the fourth conformance question in SCOPE.md asks whether all qualifying events were reported. + +A conformance level (sections 5.7.1 to 5.7.3) is a capability and record-validity claim, not a coverage claim; it does not assert that every qualifying occurrence was reported. Reporting coverage is a separate, explicit declaration, stated as one of four modes: + +- **complete** - every qualifying occurrence is emitted +- **sampled** - qualifying occurrences are emitted under a stated sampling rule +- **aggregated** - qualifying occurrences are reported only through a stated aggregation rule +- **selected** - only qualifying occurrences satisfying a further stated condition are emitted + +A coverage declaration states its mode together with the relationship scope it applies over. The rule or condition for `sampled`, `aggregated` and `selected` MUST be objectively decidable from information available at emission time and disclosed to the receiving party; a condition the emitter can satisfy or vary at its own discretion is not a stated condition, and a declaration over an undisclosed scope is not a declaration. An emitter reporting under governing terms that state a coverage mode MUST report at that mode, and an emitter MUST NOT declare or describe its reporting as `complete` for an event type unless every qualifying occurrence is emitted. A consumer MUST NOT treat the absence of an event as evidence that no occurrence happened except where complete coverage applies. + +An emitter MAY declare its coverage modes machine-readably in its manifest (`telemetry.coverage`, section 8.5); a manifest declaration is subject to the same rules, and where governing terms and a manifest declaration conflict, the governing terms control the relationship they govern. + +Whether an emitter's reporting in fact met its declared coverage is the completeness question of SCOPE.md's conformance list: it is answered by verification and audit mechanisms outside core, not by the declaration itself. + ## 6. Data profiles The `data` field on events carries type-specific metadata. These profiles document the recommended fields by event type and source role, in lifecycle order. None are required except where a section states otherwise (`reproduction_type` in 6.6; `presentation_kind` and `presentation_type` in 6.7), but emitting them enables richer attribution. @@ -577,11 +655,16 @@ When the reporter is the agent (`source_role: agent`), the following fields are | Field | Type | Description | |-------|------|-------------| | `media_type` | string | Content medium: `text`, `image`, `video`, `audio` (see below) | +| `content_depth` | string | Depth of the content record reached: `metadata`, `abstract`, `full` (see below) | `media_type` on retrieval events allows content owners to see what types of content are being fetched, independent of whether those retrievals result in grounding or citation. Defaults to `text` when absent. `text`, `image`, `video`, and `audio` are the core values. Emitters MAY use custom string values for media outside the core set (for example `3d` or `dataset`). Telemetry consumers MUST tolerate unknown `media_type` values. This rule applies to `media_type` on every event type that carries it (sections 6.4, 6.5, 6.6, 6.7). +`content_depth` records how much of the content record the retrieval reached: `metadata` for a bibliographic or descriptive record only, `abstract` for an abstract or summary record, `full` for the full content record. These are the core values; emitters MAY use custom values and telemetry consumers MUST tolerate unknown ones. The distinction is load-bearing where entitlement gates depth: a retrieval that reached only an abstract and a retrieval of full text from which a single span was later grounded are otherwise indistinguishable at the retrieval layer. Depth records what was reachable at retrieval, independent of what portion of it later entered a generation context. + +Although listed in the agent profile, `content_depth` applies to `content_retrieved` events from any `source_role`. The origin that served the response knows the depth authoritatively, and origin and edge reporters SHOULD include it alongside their fields in sections 6.2 and 6.3 where entitlement gates depth. + ### 6.2 Edge enrichment (`content_retrieved` + `source_role: edge`) CDN and edge network integrations SHOULD include these fields: @@ -862,7 +945,7 @@ A standalone event carries `document_type`, `schema_version`, and optionally `se } ``` -An event batch carries the same envelope fields with `"document_type": "event_batch"` and an `events` array. Envelope-level fields (`session_id`, `parent_session_id`, `ctx_token`, `agent_id`, `started_at`) apply to every event in the batch; events belonging to different sessions MUST be delivered in separate batches or as session documents. +An event batch carries the same envelope fields with `"document_type": "event_batch"` and an `events` array. Envelope-level fields (`session_id`, `parent_session_id`, `ctx_token`, `agent_id`, `started_at`, `manifest_ref`) apply to every event in the batch; events belonging to different sessions MUST be delivered in separate batches or as session documents. ```json { @@ -900,6 +983,8 @@ The primary schema (`telemetry-session.json`) validates session documents. A sta An agent emitter that uses standalone events or event batches for streaming delivery and wants to achieve Grounding or Citation conformance MUST include the optional `agent_id` and `started_at` fields on the envelope. Each envelope MUST also carry `session_id`, except for click-out engagement events where `ctx_token` is used instead. Consumers reconstruct the session from the stream of envelopes sharing the same `session_id`, or resolve the owning session from `ctx_token`. +The optional `manifest_ref` field is available on standalone event and event batch envelopes, mirroring the session-level field (section 5.1.2). It identifies the emitter's manifest where no session document carries one. An event delivered standalone in support of settlement, audit or other obligations under governing terms (section 5.2.4) SHOULD carry `manifest_ref`, since it is the only envelope field that names the manifest - and so the domain - under which the emitter claims to report; verifying that claim uses the manifest mechanisms of section 8. + Origin-side emitters (source role `origin` or `edge`) are not expected to achieve Grounding conformance and do not need these fields. Telemetry consumers MUST accept all three delivery formats, reconstructing sessions from standalone events and event batches where needed. @@ -1093,6 +1178,9 @@ Public keys used to sign telemetry events emitted by this participant. Per-event | `endpoint` | string | Yes | HTTPS URL. For agents and platforms, the outbound submission endpoint. For content owners, the inbound destination for events about the content owner's content. | | `conformance_level` | string | No | Conformance level advertised by this participant's own emitter(s). One of `retrieval`, `grounding`, `citation` (see 5.7). | | `ctx_resolution` | string | No | HTTPS URL of the click-token resolution endpoint operated by or for this participant (see 7.4). Valid on `agent` and `platform` manifests. | +| `coverage` | object | No | Per-event-type coverage declaration: a map from event type to `{ "mode": …, "terms_ref": … }`, where `mode` is one of `complete`, `sampled`, `aggregated`, `selected` (see 5.7.6) and `terms_ref` optionally names the terms stating the rule or condition. | + +`coverage` makes the emitter's declared coverage machine-visible. It is a claim like the rest of the manifest, subject to the rules of section 5.7.6: a `complete` entry asserts that every qualifying occurrence of that event type within the declared relationship scope is emitted, and the other modes are meaningful only with their rule or condition reachable through `terms_ref` or otherwise disclosed to the receiving party. `conformance_level` is informational. It advertises the level of telemetry the manifest's participant emits. It does **not** constrain what an inbound `endpoint` accepts - an endpoint accepts whatever events it is configured to accept, regardless of any level declared here - and it places **no requirement** on other emitters. On a `content_owner` manifest it describes only the events the owner's own infrastructure emits (typically a CDN edge worker at `retrieval`); it says nothing about what agents or platforms report about the owner's content, which those parties advertise in their own manifests. A `content_owner` manifest SHOULD omit `conformance_level` unless the owner operates its own emitter. There is no field for a content owner to *request* a minimum level from agents; consumers tolerate events from any level (see 5.7), and the protocol does not give a manifest a way to demand more. @@ -1292,6 +1380,8 @@ Implementations MAY extend core event types with custom fields in the `data` obj New event types (e.g., a commerce extension's `checkout_completed`) require a schema extension. The core schema validates only the event types listed in section 5.3. +Sessions carry a parallel extension container: the session-level `data` object (section 5.1.3). Session-scoped extension metadata belongs there, not in custom top-level fields on the session document. + ### 11.2 Custom intent categories `query_intent` accepts custom string values beyond the core set. Extensions SHOULD namespace their values to avoid collisions (e.g., `price_check` for a commerce extension). For ad-hoc categories that don't warrant a formal extension, use `other` with details in `topics`. @@ -1377,6 +1467,8 @@ gain the `ct_` pattern; presentation binding moves from the URL-carried `presentation_id` a destination could never legitimately know to issuer state restored at resolution. +V1 tightens occurrence boundaries (section 4.3). Each core event now has a stated occurrence and cardinality: retrieval per completed fetch (a cache serve is not a retrieval), grounding per distinct content item per declared scope (chunk-level events deduplicate to one occurrence by content identity), reproduction per source item per output element, presentation per rendering occurrence, engagement per observed action. Preview emitters that emitted per chunk, per passage, or re-emitted `content_retrieved` on cache serves remain schema-valid but SHOULD re-map to the stated boundaries; consumers comparing preview and v1 volumes should expect counts to shift where emitters previously chose finer or coarser units. Coverage becomes an explicit declaration (section 5.7.6) rather than an implication of conformance level, and the session root is no longer an accidental extension point: session-scoped extension metadata belongs in the session-level `data` container (section 5.1.3), and custom top-level siblings of `events`, which the preview schema tolerated silently, are undefined. + Preview versions (0.x) use two-component version numbers. From 1.0.0 onward, versions follow [semantic versioning](https://semver.org/): - **Major** (1.0.0 → 2.0.0) - breaking changes to required fields diff --git a/manifest.json b/manifest.json index ddaa8d2..fd69a7b 100644 --- a/manifest.json +++ b/manifest.json @@ -88,6 +88,25 @@ "format": "uri", "pattern": "^https://", "description": "HTTPS URL of the click-token resolution endpoint operated by or for this participant. Valid on agent and platform manifests. Destinations resolve ctx_iss to this manifest and present the token here. (Section 7.4)" + }, + "coverage": { + "type": "object", + "description": "Per-event-type coverage declaration, a map from event type to a mode object. A claim subject to the rules of section 5.7.6: complete asserts every qualifying occurrence within the declared relationship scope is emitted. (Sections 8.5, 5.7.6)", + "additionalProperties": { + "type": "object", + "required": ["mode"], + "properties": { + "mode": { + "type": "string", + "enum": ["complete", "sampled", "aggregated", "selected"], + "description": "Coverage mode (section 5.7.6)" + }, + "terms_ref": { + "type": ["string", "null"], + "description": "Reference to the terms stating the sampling, aggregation or selection rule and the relationship scope (section 5.2.4)" + } + } + } } }, "description": "Telemetry endpoint declaration. (Section 8.5)" diff --git a/telemetry-event-batch.json b/telemetry-event-batch.json index 2c2d987..ceef522 100644 --- a/telemetry-event-batch.json +++ b/telemetry-event-batch.json @@ -40,6 +40,10 @@ "format": "date-time", "description": "Session start timestamp (UTC). Mirrors the session-level started_at field. REQUIRED for emitters at Grounding conformance or above when using event batch delivery." }, + "manifest_ref": { + "type": ["string", "null"], + "description": "Manifest reference - the URL of a manifest at /.well-known/content-telemetry.json identifying the emitter. Applies to every event in the batch and mirrors the session-level manifest_ref field. See sections 7.1 and 8." + }, "events": { "type": "array", "minItems": 1, diff --git a/telemetry-event.json b/telemetry-event.json index 6ec0b53..144ff75 100644 --- a/telemetry-event.json +++ b/telemetry-event.json @@ -40,6 +40,10 @@ "format": "date-time", "description": "Session start timestamp (UTC). Mirrors the session-level started_at field. REQUIRED for emitters at Grounding conformance or above when using standalone event delivery." }, + "manifest_ref": { + "type": ["string", "null"], + "description": "Manifest reference - the URL of a manifest at /.well-known/content-telemetry.json identifying the emitter. Mirrors the session-level manifest_ref field. RECOMMENDED on standalone events supporting settlement or audit obligations under governing terms. See sections 7.1 and 8." + }, "event": { "$ref": "telemetry-session.json#/$defs/TelemetryEvent", "description": "The telemetry event" diff --git a/telemetry-session.json b/telemetry-session.json index 2f833a6..9d41324 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -54,6 +54,38 @@ "format": "date-time", "description": "Session end timestamp (UTC)" }, + "data": { + "type": ["object", "null"], + "additionalProperties": true, + "description": "Session-level extension container mirroring the event-level data field (section 5.1.3). Custom fields SHOULD be namespaced; consumers MUST tolerate unknown fields.", + "properties": { + "access_context": { + "type": "object", + "additionalProperties": true, + "description": "Context from which the session's access rights derive - an institution, never an individual. Populated only where the governing terms of the relationship require it. See section 5.1.3.", + "properties": { + "identifiers": { + "type": "array", + "description": "Typed institutional identifiers. An array because access rights can derive through consortia, federated identity and proxies at once.", + "items": { + "type": "object", + "required": ["scheme", "value"], + "properties": { + "scheme": { + "type": "string", + "description": "Identifier scheme. Core values: ror, saml_entity_id, isni. Consumers MUST tolerate unknown schemes." + }, + "value": { + "type": "string", + "description": "Identifier value in the scheme's own format" + } + } + } + } + } + } + } + }, "events": { "type": "array", "items": { @@ -141,6 +173,10 @@ "type": ["string", "null"], "description": "Reference to a licence or grant the emitter associates with this event (JWT jti, CoMP package ID, or opaque identifier). Core does not resolve or interpret it. See section 5.2.3." }, + "terms_ref": { + "type": ["string", "null"], + "description": "Reference to the governing terms the emitter associates with this event (public URL or opaque identifier both parties can resolve). Distinct from license_ref: the terms are the basis, the licence is the proof. Core does not resolve or interpret it, and it does not redefine core event semantics. See section 5.2.4." + }, "turn": { "oneOf": [{ "$ref": "#/$defs/ConversationTurn" }, { "type": "null" }], "description": "Conversation turn data (for turn_started/turn_completed)" diff --git a/tests/invalid/access-context-identifier-missing-value.json b/tests/invalid/access-context-identifier-missing-value.json new file mode 100644 index 0000000..e205a5a --- /dev/null +++ b/tests/invalid/access-context-identifier-missing-value.json @@ -0,0 +1,15 @@ +{ + "_test_description": "Session data.access_context identifier missing its value. The schema requires both scheme and value on every identifier (5.1.3): a scheme with no value identifies nothing and MUST fail validation.", + "document_type": "session", + "schema_version": "0.1", + "session_id": "880e8400-e29b-41d4-a716-446655440090", + "started_at": "2026-08-13T14:02:10Z", + "data": { + "access_context": { + "identifiers": [ + { "scheme": "ror" } + ] + } + }, + "events": [] +} diff --git a/tests/invalid/access-context-identifiers-not-array.json b/tests/invalid/access-context-identifiers-not-array.json new file mode 100644 index 0000000..bdae9b0 --- /dev/null +++ b/tests/invalid/access-context-identifiers-not-array.json @@ -0,0 +1,13 @@ +{ + "_test_description": "Session data.access_context.identifiers as a bare string rather than an array of typed identifiers. The schema requires an array of {scheme, value} objects (5.1.3): free-text institutional identifiers are the failure the typed structure exists to prevent.", + "document_type": "session", + "schema_version": "0.1", + "session_id": "880e8400-e29b-41d4-a716-446655440091", + "started_at": "2026-08-13T14:02:10Z", + "data": { + "access_context": { + "identifiers": "https://ror.org/013meh722" + } + }, + "events": [] +} diff --git a/tests/valid/event-standalone-terms-ref.json b/tests/valid/event-standalone-terms-ref.json new file mode 100644 index 0000000..a6b2c8c --- /dev/null +++ b/tests/valid/event-standalone-terms-ref.json @@ -0,0 +1,15 @@ +{ + "_test_description": "Standalone retrieval event reported under governing terms rather than a grant: an agent reports an access whose consequences are set by a standard-form contract. terms_ref names the governing terms (5.2.4), license_ref is absent because no grant applied, and the envelope carries manifest_ref (7.1) so the record ties the emitter to a domain it controls without a session document. The fee or other consequence is derived by the content owner from this event and the referenced terms; it is not asserted in telemetry.", + "document_type": "event", + "schema_version": "0.1", + "agent_id": "assistant.example.com", + "manifest_ref": "https://assistant.example.com/.well-known/content-telemetry.json", + "event": { + "type": "content_retrieved", + "timestamp": "2026-08-13T15:00:04Z", + "source_role": "agent", + "content_telemetry_id": "990e8400-e29b-41d4-a716-446655440071", + "content_url": "https://example.com/2026/08/13/markets-live", + "terms_ref": "https://terms.example.com/search-only/v2" + } +} diff --git a/tests/valid/session-access-context.json b/tests/valid/session-access-context.json new file mode 100644 index 0000000..9fa7348 --- /dev/null +++ b/tests/valid/session-access-context.json @@ -0,0 +1,64 @@ +{ + "_test_description": "Session carrying the session-level data container (5.1.3): access_context identifies the institution whose entitlement the agent used, populated because the governing terms of the relationship require it, with turn data held at intent level per the pairing guidance in 5.5. The retrieval event records content_depth: full (6.1), preserving the metadata-versus-full-content distinction at the retrieval layer.", + "document_type": "session", + "schema_version": "0.1", + "session_id": "880e8400-e29b-41d4-a716-446655440080", + "agent_id": "scholar-assistant.example.com", + "content_scope": "consortium-agreement-4471", + "started_at": "2026-08-13T14:02:10Z", + "ended_at": "2026-08-13T14:03:44Z", + "data": { + "access_context": { + "identifiers": [ + { "scheme": "ror", "value": "https://ror.org/013meh722" }, + { "scheme": "saml_entity_id", "value": "https://idp.example.ac.uk/shibboleth" } + ] + } + }, + "events": [ + { + "type": "turn_started", + "timestamp": "2026-08-13T14:02:11Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "query_intent": "question", + "topics": ["materials science"] + } + }, + { + "type": "content_retrieved", + "timestamp": "2026-08-13T14:02:14Z", + "source_role": "agent", + "content_telemetry_id": "990e8400-e29b-41d4-a716-446655440081", + "content_url": "https://journals.example.com/article/10.1000/xyz123", + "content_id": "doi:10.1000/xyz123", + "license_ref": "grant-4471-2026", + "data": { + "media_type": "text", + "content_depth": "full" + } + }, + { + "type": "content_grounded", + "timestamp": "2026-08-13T14:02:16Z", + "source_role": "agent", + "turn_id": "1", + "content_id": "doi:10.1000/xyz123", + "data": { + "scope": "turn", + "chars_ingested": 18400 + } + }, + { + "type": "turn_completed", + "timestamp": "2026-08-13T14:03:40Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "query_intent": "question", + "topics": ["materials science"] + } + } + ] +} From 8b7083cb7415a908dfa768a35ca2f8fae62a86bf Mon Sep 17 00:00:00 2001 From: Alex Springer Date: Mon, 17 Aug 2026 09:33:02 +0100 Subject: [PATCH 20/38] State terms_ref carriage: byte-for-byte, stable denotation, absence means nothing (#3) Co-Authored-By: Claude Fable 5 --- SPECIFICATION.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/SPECIFICATION.md b/SPECIFICATION.md index c1a0c00..172e926 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -455,6 +455,8 @@ The terms are the basis; the licence is the proof. `license_ref` records which g Core does not resolve, validate or interpret the reference, and `terms_ref` does not redefine core event semantics. Governing terms select which events a relationship requires and at what coverage (section 5.7.6), together with the cadence, delivery, privacy and reports owed (SCOPE.md); the meaning and occurrence boundary of each event remain those defined in sections 4.3 and 6, whatever `terms_ref` points to. +The reference is carried, not managed. A processor that stores, forwards or transforms a document MUST carry `terms_ref` byte for byte and MUST NOT remove or rewrite it. A `terms_ref` value denotes the same terms for all time: terms that change are referenced by a new value, so the reader of a historical event can still find the terms that governed it. The absence of `terms_ref` carries no meaning; as with coverage (section 5.7.6), silence is not a declaration, and a consumer MUST NOT infer from a missing reference that no terms governed the event. + ### 5.3 Event types #### Content events From 9eee78188375afdd1408076bd3351025f422c6d8 Mon Sep 17 00:00:00 2001 From: Alex Springer Date: Tue, 18 Aug 2026 11:32:38 +0100 Subject: [PATCH 21/38] Tighten prose in 4.3, 5.1.3, 5.2.4, 5.7.6, 6.1 and fixture descriptions Co-Authored-By: Claude Fable 5 --- SPECIFICATION.md | 14 +++++++------- telemetry-session.json | 2 +- .../access-context-identifier-missing-value.json | 2 +- .../access-context-identifiers-not-array.json | 2 +- tests/valid/event-standalone-terms-ref.json | 2 +- tests/valid/session-access-context.json | 2 +- 6 files changed, 12 insertions(+), 12 deletions(-) diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 172e926..282c4f9 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -247,7 +247,7 @@ Each stage after retrieval is typically a progressively narrower subset, except - **Citation-to-presentation** measures source associations constructed in output but not made perceivable - **Presentation-to-engagement** measures observable actions on exact presentation occurrences -These ratios are computed over reported events. They are comparable across emitters, and meaningful as measures of behaviour rather than of reporting, only at known coverage (section 5.7.6). +These ratios are computed over reported events and are comparable across emitters only at known coverage (section 5.7.6). #### Departures from the funnel model @@ -399,9 +399,9 @@ One container is defined in core. `access_context` records the context from whic } ``` -`identifiers` is an array of typed identifiers, each a `scheme` and a `value`. `ror`, `saml_entity_id` and `isni` are the core scheme values; emitters MAY use other schemes and telemetry consumers MUST tolerate unknown ones, as with `media_type` (section 6.1). It is an array because access rights arrive through consortia, federated identity and proxies at once: a session may legitimately carry a SAML entity ID and the ROR ID it maps to. +`identifiers` is an array of typed identifiers, each a `scheme` and a `value`. `ror`, `saml_entity_id` and `isni` are the core scheme values; emitters MAY use other schemes and telemetry consumers MUST tolerate unknown ones, as with `media_type` (section 6.1). Access rights can derive through consortia, federated identity and proxies at once, so a session may carry both a SAML entity ID and the ROR ID it maps to. -`access_context` identifies an institution, never an individual. Like everything else in a session it is a claim by the emitter. Where the content owner authenticated the session itself (`source_role` of `origin` or `edge`), the owner already knows the institution and does not need the field; it earns its place where a third-party agent holds the entitlement and asserts the affiliation to the content owner, which is also where the claim is least verifiable. What corroborates an asserted affiliation is verification-layer work, outside core. +`access_context` identifies an institution, never an individual, and like every session field it is a claim by the emitter. The field serves the third-party agent that holds the entitlement and asserts the affiliation to the content owner; where the owner authenticated the session itself (`source_role` of `origin` or `edge`), it already knows the institution. Corroborating an asserted affiliation is verification-layer work, outside core. Emitters MUST NOT populate `access_context` unless the governing terms of the relationship require it, and SHOULD pair it with `intent` or `minimal` conversation-turn data (section 5.5): an identified institution combined with query text can come close to identifying an individual at a small subscriber. @@ -455,7 +455,7 @@ The terms are the basis; the licence is the proof. `license_ref` records which g Core does not resolve, validate or interpret the reference, and `terms_ref` does not redefine core event semantics. Governing terms select which events a relationship requires and at what coverage (section 5.7.6), together with the cadence, delivery, privacy and reports owed (SCOPE.md); the meaning and occurrence boundary of each event remain those defined in sections 4.3 and 6, whatever `terms_ref` points to. -The reference is carried, not managed. A processor that stores, forwards or transforms a document MUST carry `terms_ref` byte for byte and MUST NOT remove or rewrite it. A `terms_ref` value denotes the same terms for all time: terms that change are referenced by a new value, so the reader of a historical event can still find the terms that governed it. The absence of `terms_ref` carries no meaning; as with coverage (section 5.7.6), silence is not a declaration, and a consumer MUST NOT infer from a missing reference that no terms governed the event. +A processor that stores, forwards or transforms a document MUST carry `terms_ref` byte for byte and MUST NOT remove or rewrite it. A `terms_ref` value denotes the same terms for all time: terms that change are referenced by a new value, so the reader of a historical event can still find the terms that governed it. The absence of `terms_ref` carries no meaning: a consumer MUST NOT infer from a missing reference that no terms governed the event (the same rule as event absence under coverage, section 5.7.6). ### 5.3 Event types @@ -631,7 +631,7 @@ The `tests/` directory provides an informative reference suite for these rules. Each core event type has the meaning and occurrence boundary defined in sections 4.3 and 6. Profiles, deployment configurations and governing terms MUST NOT redefine them. A relationship that needs a different assertion defines a namespaced extension event (sections 5.3 and 11.1); it does not reuse a core type with altered semantics. -An occurrence is **qualifying** for an emitter when it satisfies the core definition and occurrence boundary of its event type and falls within the relationship scope the emitter reports under - the content, domains or relationships selected by the applicable governing terms or deployment configuration. This is the sense in which the fourth conformance question in SCOPE.md asks whether all qualifying events were reported. +An occurrence is **qualifying** for an emitter when it satisfies the core definition and occurrence boundary of its event type and falls within the relationship scope the emitter reports under - the content, domains or relationships selected by the applicable governing terms or deployment configuration. A conformance level (sections 5.7.1 to 5.7.3) is a capability and record-validity claim, not a coverage claim; it does not assert that every qualifying occurrence was reported. Reporting coverage is a separate, explicit declaration, stated as one of four modes: @@ -663,7 +663,7 @@ When the reporter is the agent (`source_role: agent`), the following fields are `text`, `image`, `video`, and `audio` are the core values. Emitters MAY use custom string values for media outside the core set (for example `3d` or `dataset`). Telemetry consumers MUST tolerate unknown `media_type` values. This rule applies to `media_type` on every event type that carries it (sections 6.4, 6.5, 6.6, 6.7). -`content_depth` records how much of the content record the retrieval reached: `metadata` for a bibliographic or descriptive record only, `abstract` for an abstract or summary record, `full` for the full content record. These are the core values; emitters MAY use custom values and telemetry consumers MUST tolerate unknown ones. The distinction is load-bearing where entitlement gates depth: a retrieval that reached only an abstract and a retrieval of full text from which a single span was later grounded are otherwise indistinguishable at the retrieval layer. Depth records what was reachable at retrieval, independent of what portion of it later entered a generation context. +`content_depth` records how much of the content record the retrieval reached: `metadata` for a bibliographic or descriptive record only, `abstract` for an abstract or summary record, `full` for the full content record. These are the core values; emitters MAY use custom values and telemetry consumers MUST tolerate unknown ones. Where entitlement gates depth, a retrieval that reached only an abstract and a retrieval of full text are otherwise indistinguishable at the retrieval layer. Depth records what was reachable at retrieval, independent of what portion later entered a generation context. Although listed in the agent profile, `content_depth` applies to `content_retrieved` events from any `source_role`. The origin that served the response knows the depth authoritatively, and origin and edge reporters SHOULD include it alongside their fields in sections 6.2 and 6.3 where entitlement gates depth. @@ -1469,7 +1469,7 @@ gain the `ct_` pattern; presentation binding moves from the URL-carried `presentation_id` a destination could never legitimately know to issuer state restored at resolution. -V1 tightens occurrence boundaries (section 4.3). Each core event now has a stated occurrence and cardinality: retrieval per completed fetch (a cache serve is not a retrieval), grounding per distinct content item per declared scope (chunk-level events deduplicate to one occurrence by content identity), reproduction per source item per output element, presentation per rendering occurrence, engagement per observed action. Preview emitters that emitted per chunk, per passage, or re-emitted `content_retrieved` on cache serves remain schema-valid but SHOULD re-map to the stated boundaries; consumers comparing preview and v1 volumes should expect counts to shift where emitters previously chose finer or coarser units. Coverage becomes an explicit declaration (section 5.7.6) rather than an implication of conformance level, and the session root is no longer an accidental extension point: session-scoped extension metadata belongs in the session-level `data` container (section 5.1.3), and custom top-level siblings of `events`, which the preview schema tolerated silently, are undefined. +V1 tightens occurrence boundaries (section 4.3). Each core event now has a stated occurrence and cardinality: retrieval per completed fetch (a cache serve is not a retrieval), grounding per distinct content item per declared scope (chunk-level events deduplicate to one occurrence by content identity), reproduction per source item per output element, presentation per rendering occurrence, engagement per observed action. Preview emitters that emitted per chunk, per passage, or re-emitted `content_retrieved` on cache serves remain schema-valid but SHOULD re-map to the stated boundaries; consumers comparing preview and v1 volumes should expect counts to shift where emitters previously chose finer or coarser units. Coverage becomes an explicit declaration (section 5.7.6) rather than an implication of conformance level. Session-scoped extension metadata belongs in the session-level `data` container (section 5.1.3); custom top-level siblings of `events`, accepted by the preview schema, are undefined. Preview versions (0.x) use two-component version numbers. From 1.0.0 onward, versions follow [semantic versioning](https://semver.org/): diff --git a/telemetry-session.json b/telemetry-session.json index 9d41324..ffbfcea 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -66,7 +66,7 @@ "properties": { "identifiers": { "type": "array", - "description": "Typed institutional identifiers. An array because access rights can derive through consortia, federated identity and proxies at once.", + "description": "Typed institutional identifiers. Access rights can derive through consortia, federated identity and proxies at once.", "items": { "type": "object", "required": ["scheme", "value"], diff --git a/tests/invalid/access-context-identifier-missing-value.json b/tests/invalid/access-context-identifier-missing-value.json index e205a5a..6e00daa 100644 --- a/tests/invalid/access-context-identifier-missing-value.json +++ b/tests/invalid/access-context-identifier-missing-value.json @@ -1,5 +1,5 @@ { - "_test_description": "Session data.access_context identifier missing its value. The schema requires both scheme and value on every identifier (5.1.3): a scheme with no value identifies nothing and MUST fail validation.", + "_test_description": "Session access_context identifier missing its value. The schema requires both scheme and value on every identifier (5.1.3).", "document_type": "session", "schema_version": "0.1", "session_id": "880e8400-e29b-41d4-a716-446655440090", diff --git a/tests/invalid/access-context-identifiers-not-array.json b/tests/invalid/access-context-identifiers-not-array.json index bdae9b0..2588b9f 100644 --- a/tests/invalid/access-context-identifiers-not-array.json +++ b/tests/invalid/access-context-identifiers-not-array.json @@ -1,5 +1,5 @@ { - "_test_description": "Session data.access_context.identifiers as a bare string rather than an array of typed identifiers. The schema requires an array of {scheme, value} objects (5.1.3): free-text institutional identifiers are the failure the typed structure exists to prevent.", + "_test_description": "Session access_context.identifiers as a bare string. The schema requires an array of {scheme, value} objects (5.1.3).", "document_type": "session", "schema_version": "0.1", "session_id": "880e8400-e29b-41d4-a716-446655440091", diff --git a/tests/valid/event-standalone-terms-ref.json b/tests/valid/event-standalone-terms-ref.json index a6b2c8c..7293ad0 100644 --- a/tests/valid/event-standalone-terms-ref.json +++ b/tests/valid/event-standalone-terms-ref.json @@ -1,5 +1,5 @@ { - "_test_description": "Standalone retrieval event reported under governing terms rather than a grant: an agent reports an access whose consequences are set by a standard-form contract. terms_ref names the governing terms (5.2.4), license_ref is absent because no grant applied, and the envelope carries manifest_ref (7.1) so the record ties the emitter to a domain it controls without a session document. The fee or other consequence is derived by the content owner from this event and the referenced terms; it is not asserted in telemetry.", + "_test_description": "Standalone retrieval event reported under governing terms with no grant: terms_ref names the terms (5.2.4), license_ref is absent, and the envelope carries manifest_ref (7.1) to tie the emitter to a domain without a session document.", "document_type": "event", "schema_version": "0.1", "agent_id": "assistant.example.com", diff --git a/tests/valid/session-access-context.json b/tests/valid/session-access-context.json index 9fa7348..ee6d3c4 100644 --- a/tests/valid/session-access-context.json +++ b/tests/valid/session-access-context.json @@ -1,5 +1,5 @@ { - "_test_description": "Session carrying the session-level data container (5.1.3): access_context identifies the institution whose entitlement the agent used, populated because the governing terms of the relationship require it, with turn data held at intent level per the pairing guidance in 5.5. The retrieval event records content_depth: full (6.1), preserving the metadata-versus-full-content distinction at the retrieval layer.", + "_test_description": "Session-level data container (5.1.3): access_context identifies the institution whose entitlement the agent used, required by the governing terms, with turn data at intent level per 5.5. The retrieval event carries content_depth: full (6.1).", "document_type": "session", "schema_version": "0.1", "session_id": "880e8400-e29b-41d4-a716-446655440080", From f06c4768f1fb21557e26f765bf29b08d29462dc2 Mon Sep 17 00:00:00 2001 From: Alex Springer Date: Tue, 18 Aug 2026 11:40:47 +0100 Subject: [PATCH 22/38] Remove figurative and borrowed phrasing from 5.2.4, 5.5, 5.7, 5.7.6 Co-Authored-By: Claude Fable 5 --- SPECIFICATION.md | 18 +++++++++--------- telemetry-session.json | 2 +- 2 files changed, 10 insertions(+), 10 deletions(-) diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 282c4f9..8d6ced3 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -111,7 +111,7 @@ For the purposes of this specification, the following terms apply. | **privacy level** | data sharing tier controlling which conversation fields are populated: `full`, `summary`, `intent`, or `minimal` (section 5.5) | | **conformance level** | emitter capability tier: Retrieval, Grounding, or Citation (section 5.7) | | **content scope** | opaque identifier grouping sessions by their content access context (section 5.1.1) | -| **governing terms** | licence, contract or other terms selecting which events a relationship requires and the coverage, cadence, delivery, privacy and reports owed over the shared semantics (sections 5.2.4, 5.7.6) | +| **governing terms** | licence, contract or other terms selecting which events a relationship requires and the coverage, cadence, delivery, privacy and reports owed (sections 5.2.4, 5.7.6) | | **qualifying occurrence** | occurrence satisfying an event type's core definition and occurrence boundary within the relationship scope reported under (section 5.7.6) | | **coverage** | declared relationship between qualifying occurrences and emitted events: `complete`, `sampled`, `aggregated` or `selected` (section 5.7.6) | @@ -206,7 +206,7 @@ Content moves through six stages during an agent interaction: Grounding is architecture-neutral: same event whether the agent uses RAG, chain-of-thought reasoning, embeddings, or multi-step delegation (see section 6.4 for architecture-specific guidance). Grounding is decoupled from retrieval: content may be grounded from a live fetch, from agent-side cache, or from a pre-loaded index. Only the agent can report grounding events. - One grounding occurrence is one distinct content item entering a generation context at the declared `data.scope`: at `session` scope, a content item grounds once per session; at `turn` scope, once per turn it enters. A distinct content item is a distinct `content_id`, or its canonical `content_url` where no stable identifier exists (section 4.5). Continued presence within the declared scope is not a further occurrence; re-entry in a later turn is, when the scope is `turn`, and a change of `content_version` opens a new occurrence at either scope. An emitter that ingests a content item in chunks MAY emit one grounding event per chunk, preserving the chunk-level hashes of section 6.4; events sharing content identity within one scope describe one occurrence, and consumers count occurrences by deduplicating on content identity and scope, not by counting events. + One grounding occurrence is one distinct content item entering a generation context at the declared `data.scope`: at `session` scope, a content item grounds once per session; at `turn` scope, once per turn it enters. A distinct content item is a distinct `content_id`, or its canonical `content_url` where no stable identifier exists (section 4.5). Continued presence within the declared scope is not a further occurrence; re-entry in a later turn is, when the scope is `turn`, and a change of `content_version` is a new occurrence at either scope. An emitter that ingests a content item in chunks MAY emit one grounding event per chunk, preserving the chunk-level hashes of section 6.4; events sharing content identity within one scope describe one occurrence, and consumers count occurrences by deduplicating on content identity and scope, not by counting events. 3. **Reproduced** - The output artifact contains identified source content: a quotation, an excerpt, or a full copy, verbatim or near-verbatim. Reproduction is an output-construction claim by the system that built the output. It is independent of credit and of delivery: reproduced content may or may not also be cited, and the artifact may or may not later be presented. An uncredited excerpt in a response delivered through an API produces a `content_reproduced` event and nothing else - without this event, that use would be unreportable. @@ -451,11 +451,11 @@ Core does not resolve, validate or interpret the reference. `license_ref` is par The `terms_ref` field associates a telemetry event with the governing terms under which it is reported: a licence agreement, a standard-form contract, a tariff, a profile's terms, or any other terms document. Like `license_ref`, the value MAY be a public URL or an opaque identifier that both parties can resolve. Nothing requires the terms to be published: per-relationship terms are often confidential, and an opaque identifier resolved privately is a conforming reference. -The terms are the basis; the licence is the proof. `license_ref` records which grant the emitter says applied (5.2.3). `terms_ref` records which terms govern the event's commercial consequences and the emitter's reporting obligations. Either may appear without the other: an access outside any grant carries no `license_ref`, and can still carry the `terms_ref` of the terms that attach consequences to that access. +`license_ref` records which grant the emitter says applied (5.2.3). `terms_ref` records which terms govern the event's commercial consequences and the emitter's reporting obligations. Either may appear without the other: an access outside any grant carries no `license_ref`, and can still carry the `terms_ref` of the terms that attach consequences to that access. Core does not resolve, validate or interpret the reference, and `terms_ref` does not redefine core event semantics. Governing terms select which events a relationship requires and at what coverage (section 5.7.6), together with the cadence, delivery, privacy and reports owed (SCOPE.md); the meaning and occurrence boundary of each event remain those defined in sections 4.3 and 6, whatever `terms_ref` points to. -A processor that stores, forwards or transforms a document MUST carry `terms_ref` byte for byte and MUST NOT remove or rewrite it. A `terms_ref` value denotes the same terms for all time: terms that change are referenced by a new value, so the reader of a historical event can still find the terms that governed it. The absence of `terms_ref` carries no meaning: a consumer MUST NOT infer from a missing reference that no terms governed the event (the same rule as event absence under coverage, section 5.7.6). +A processor that stores, forwards or transforms a document MUST preserve `terms_ref` unchanged and MUST NOT remove or rewrite it. A `terms_ref` value MUST always refer to the same terms: when terms change, the emitter references them with a new value, so events emitted under earlier terms remain resolvable to them. A consumer MUST NOT read the absence of `terms_ref` as a statement that no terms governed the event (the same rule as event absence under coverage, section 5.7.6). ### 5.3 Event types @@ -525,7 +525,7 @@ These are the recommended values. Platforms with additional product surfaces (co An emitter that populates a conversation turn MUST NOT include a field above that turn's declared `privacy_level` - for example, `query_text` MUST NOT be present when `privacy_level` is `intent` or `minimal`. This restriction is a property of `privacy_level` itself: it applies wherever conversation turns are emitted, independent of the emitter's conformance level. -The session-level `access_context` container (section 5.1.3) sits outside this table but inside the privacy model. It is populated only where governing terms require it, and pairing it with `full` or `summary` turn data is discouraged (section 5.1.3): the ladder governs how much of the query and response is visible, `access_context` governs whose access rights the session used, and the two disclosures compound. +The session-level `access_context` container (section 5.1.3) does not appear in this table but is subject to the privacy model. It is populated only where governing terms require it, and pairing it with `full` or `summary` turn data is discouraged (section 5.1.3): `privacy_level` controls how much of the query and response is visible, `access_context` identifies whose access rights the session used, and populating both makes re-identification easier. **Token counts** includes `query_tokens` and `response_tokens`. These are available at all levels because they are needed for token-based counting models and do not reveal user intent or platform strategy. They carry the same portability limit as `tokens_ingested` (section 6.4): both are measured in the emitter's own tokeniser and are not comparable between agents. Version 1 does not define corresponding turn-level character counts; `chars_ingested` measures source content placed in a generation context, not query or response length. @@ -553,7 +553,7 @@ These are the core values. Extensions MAY define additional intent category valu Emitters that advertise a standard capability tier use one of three conformance levels. The authoritative declaration lives in the emitter's manifest (section 8). Emitters MAY also include an optional `conformance_level` field on individual session documents; when present it is informational and consumers MUST NOT treat it as a substitute for verifying the manifest's declaration. -Each level is named for the event it adds: a level proves the emitter produces that event and everything below it. A level is a capability claim, not a coverage claim: it does not assert that every qualifying occurrence was reported (section 5.7.6). These levels describe what an emitter reports, not what a consumer computes from it; attribution - the apportioning of credit across content - is performed by a telemetry consumer at whatever funnel level the parties agree (section 10), and can be computed from grounding alone, without citation. An emitter does not need to reach the Citation level for its telemetry to support attribution. +Each level is named for the event it adds: a level proves the emitter produces that event and everything below it. A level does not assert that every qualifying occurrence was reported (section 5.7.6). These levels describe what an emitter reports, not what a consumer computes from it; attribution - the apportioning of credit across content - is performed by a telemetry consumer at whatever funnel level the parties agree (section 10), and can be computed from grounding alone, without citation. An emitter does not need to reach the Citation level for its telemetry to support attribution. | Level | Events | What it proves | Typical emitter | |-------|--------|----------------|-----------------| @@ -633,16 +633,16 @@ Each core event type has the meaning and occurrence boundary defined in sections An occurrence is **qualifying** for an emitter when it satisfies the core definition and occurrence boundary of its event type and falls within the relationship scope the emitter reports under - the content, domains or relationships selected by the applicable governing terms or deployment configuration. -A conformance level (sections 5.7.1 to 5.7.3) is a capability and record-validity claim, not a coverage claim; it does not assert that every qualifying occurrence was reported. Reporting coverage is a separate, explicit declaration, stated as one of four modes: +A conformance level (sections 5.7.1 to 5.7.3) does not assert that every qualifying occurrence was reported. Reporting coverage is a separate, explicit declaration, stated as one of four modes: - **complete** - every qualifying occurrence is emitted - **sampled** - qualifying occurrences are emitted under a stated sampling rule - **aggregated** - qualifying occurrences are reported only through a stated aggregation rule - **selected** - only qualifying occurrences satisfying a further stated condition are emitted -A coverage declaration states its mode together with the relationship scope it applies over. The rule or condition for `sampled`, `aggregated` and `selected` MUST be objectively decidable from information available at emission time and disclosed to the receiving party; a condition the emitter can satisfy or vary at its own discretion is not a stated condition, and a declaration over an undisclosed scope is not a declaration. An emitter reporting under governing terms that state a coverage mode MUST report at that mode, and an emitter MUST NOT declare or describe its reporting as `complete` for an event type unless every qualifying occurrence is emitted. A consumer MUST NOT treat the absence of an event as evidence that no occurrence happened except where complete coverage applies. +A coverage declaration states its mode together with the relationship scope it applies over; both MUST be disclosed to the receiving party. The rule or condition for `sampled`, `aggregated` and `selected` MUST be objectively decidable from information available at emission time and MUST NOT depend on the emitter's discretion at the moment of emission. An emitter reporting under governing terms that state a coverage mode MUST report at that mode, and an emitter MUST NOT declare or describe its reporting as `complete` for an event type unless every qualifying occurrence is emitted. A consumer MUST NOT treat the absence of an event as evidence that no occurrence happened except where complete coverage applies. -An emitter MAY declare its coverage modes machine-readably in its manifest (`telemetry.coverage`, section 8.5); a manifest declaration is subject to the same rules, and where governing terms and a manifest declaration conflict, the governing terms control the relationship they govern. +An emitter MAY declare its coverage modes machine-readably in its manifest (`telemetry.coverage`, section 8.5); a manifest declaration is subject to the same rules, and where governing terms and a manifest declaration conflict, the governing terms take precedence for the relationships they cover. Whether an emitter's reporting in fact met its declared coverage is the completeness question of SCOPE.md's conformance list: it is answered by verification and audit mechanisms outside core, not by the declaration itself. diff --git a/telemetry-session.json b/telemetry-session.json index ffbfcea..fb90e62 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -175,7 +175,7 @@ }, "terms_ref": { "type": ["string", "null"], - "description": "Reference to the governing terms the emitter associates with this event (public URL or opaque identifier both parties can resolve). Distinct from license_ref: the terms are the basis, the licence is the proof. Core does not resolve or interpret it, and it does not redefine core event semantics. See section 5.2.4." + "description": "Reference to the governing terms the emitter associates with this event (public URL or opaque identifier both parties can resolve). Distinct from license_ref, which records the grant that applied. Core does not resolve or interpret it, and it does not redefine core event semantics. See section 5.2.4." }, "turn": { "oneOf": [{ "$ref": "#/$defs/ConversationTurn" }, { "type": "null" }], From 74e06fbc2f4f13835677583f1195ac39e83ae2a1 Mon Sep 17 00:00:00 2001 From: Alex Springer Date: Tue, 18 Aug 2026 11:42:36 +0100 Subject: [PATCH 23/38] Add manifest coverage fixtures Co-Authored-By: Claude Fable 5 --- tests/invalid/manifest-coverage-bad-mode.json | 13 +++++++++++++ tests/valid/manifest-coverage-declaration.json | 18 ++++++++++++++++++ 2 files changed, 31 insertions(+) create mode 100644 tests/invalid/manifest-coverage-bad-mode.json create mode 100644 tests/valid/manifest-coverage-declaration.json diff --git a/tests/invalid/manifest-coverage-bad-mode.json b/tests/invalid/manifest-coverage-bad-mode.json new file mode 100644 index 0000000..1e076e6 --- /dev/null +++ b/tests/invalid/manifest-coverage-bad-mode.json @@ -0,0 +1,13 @@ +{ + "_test_description": "Manifest coverage entry with a mode outside the enum. The schema requires one of complete, sampled, aggregated, selected (8.5, 5.7.6).", + "schema_version": "0.1", + "id": "https://assistant.example.com/.well-known/content-telemetry.json", + "roles": ["agent"], + "operator": { "name": "Assistant Example" }, + "telemetry": { + "endpoint": "https://telemetry.assistant.example.com/v1/events", + "coverage": { + "content_grounded": { "mode": "partial" } + } + } +} diff --git a/tests/valid/manifest-coverage-declaration.json b/tests/valid/manifest-coverage-declaration.json new file mode 100644 index 0000000..642d0eb --- /dev/null +++ b/tests/valid/manifest-coverage-declaration.json @@ -0,0 +1,18 @@ +{ + "_test_description": "Agent manifest declaring per-event-type coverage in telemetry.coverage (8.5, 5.7.6): content_grounded reported complete, content_retrieved sampled under a rule stated in the referenced terms.", + "schema_version": "0.1", + "id": "https://assistant.example.com/.well-known/content-telemetry.json", + "roles": ["agent"], + "operator": { "name": "Assistant Example" }, + "telemetry": { + "endpoint": "https://telemetry.assistant.example.com/v1/events", + "conformance_level": "grounding", + "coverage": { + "content_grounded": { "mode": "complete" }, + "content_retrieved": { + "mode": "sampled", + "terms_ref": "https://terms.example.com/reporting/v1" + } + } + } +} From 35b45f4c51fd08b3eca42aa6ce7e71f1eeb31e85 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Wed, 12 Aug 2026 11:10:27 +0100 Subject: [PATCH 24/38] Declare v1 schema identity: schema_version 1.0, /schema/v1/ $ids The spec body already makes v1 normative claims (ip_hash withdrawal in 9.1, the content_displayed rename in 12.1) while the header still said 0.1/Preview and every schema pinned schema_version to "0.1" under a /schema/v0.1/ $id, so a conforming v1 emitter had to declare "0.1" and the two versions were indistinguishable on the wire. - Header: Version 1.0 (release candidate draft), status release candidate in preparation, feature freeze 21 August 2026. - All four schemas: $id path v0.1 -> v1, schema_version const -> "1.0"; manifest const description now says v1 emitters MUST use "1.0". - 5.7.4 and 12.1 now state explicitly that v1 documents declare "1.0", that a v0.1 consumer rejects them under the preview rule, and that a v1 consumer rejects "0.1". - Every fixture and every inline example in SPECIFICATION.md and README.md now declares "1.0". invalid/invalid-schema-version.json keeps its deliberately wrong "2.0" (still != the new const). - validate.py / tests/README.md docstrings updated; no hardcoded versions or v0.1 paths remain in the test scripts. Co-Authored-By: Claude Fable 5 --- README.md | 2 +- SPECIFICATION.md | 36 +++++++++++-------- manifest.json | 6 ++-- telemetry-event-batch.json | 4 +-- telemetry-event.json | 4 +-- telemetry-session.json | 4 +-- tests/README.md | 2 +- tests/invalid/batch-empty-events.json | 2 +- .../batch-missing-session-and-ctx-token.json | 2 +- tests/invalid/cited-missing-output-id.json | 2 +- .../cited-missing-source-reference.json | 2 +- .../invalid/cited-null-source-reference.json | 2 +- .../content-event-missing-identifier.json | 2 +- .../engaged-missing-presentation-id.json | 2 +- tests/invalid/invalid-event-type.json | 2 +- tests/invalid/invalid-privacy-level.json | 2 +- tests/invalid/invalid-schema-version.json | 2 +- tests/invalid/invalid-source-role.json | 2 +- tests/invalid/legacy-content-displayed.json | 2 +- .../manifest-bad-conformance-level.json | 2 +- tests/invalid/manifest-bad-role.json | 2 +- tests/invalid/manifest-duplicate-key-id.json | 2 +- tests/invalid/manifest-foreign-domain.json | 2 +- .../manifest-missing-key-publickey.json | 2 +- tests/invalid/manifest-missing-operator.json | 2 +- tests/invalid/missing-event-timestamp.json | 2 +- tests/invalid/missing-event-type.json | 2 +- tests/invalid/missing-privacy-level.json | 2 +- tests/invalid/missing-session-id.json | 2 +- tests/invalid/missing-started-at.json | 2 +- tests/invalid/presented-missing-kind.json | 2 +- .../presented-missing-presentation-type.json | 2 +- ...vacy-violation-ad-rendered-at-minimal.json | 2 +- .../privacy-violation-query-at-intent.json | 2 +- .../privacy-violation-query-at-minimal.json | 2 +- tests/invalid/reproduced-missing-id.json | 2 +- .../invalid/reproduced-missing-output-id.json | 2 +- .../reproduced-missing-source-reference.json | 2 +- tests/invalid/reproduced-missing-type.json | 2 +- .../standalone-missing-document-type.json | 2 +- tests/invalid/standalone-missing-event.json | 2 +- ...ndalone-missing-session-and-ctx-token.json | 2 +- tests/invalid/withdrawn-ip-hash.json | 2 +- tests/valid/event-batch-agent.json | 2 +- tests/valid/event-batch-edge.json | 2 +- tests/valid/event-standalone-agent.json | 2 +- .../valid/event-standalone-child-session.json | 2 +- tests/valid/event-standalone-edge.json | 2 +- .../event-standalone-engaged-ctx-token.json | 2 +- tests/valid/event-standalone-presented.json | 2 +- tests/valid/manifest-agent-with-keys.json | 2 +- tests/valid/manifest-content-owner-full.json | 2 +- .../valid/manifest-content-owner-minimal.json | 2 +- tests/valid/manifest-multi-role.json | 2 +- tests/valid/session-cached-grounding.json | 2 +- tests/valid/session-citation-tier.json | 2 +- tests/valid/session-custom-media-type.json | 2 +- tests/valid/session-custom-response-mode.json | 2 +- ...on-funnel-exception-cited-no-grounded.json | 2 +- ...n-funnel-exception-presented-no-cited.json | 2 +- ...unnel-exception-presented-no-grounded.json | 2 +- tests/valid/session-grounding-tier.json | 2 +- tests/valid/session-minimal.json | 2 +- tests/valid/session-multi-agent-child.json | 2 +- tests/valid/session-multi-turn.json | 2 +- .../session-presentation-multimodal.json | 2 +- .../session-reproduction-credited-quote.json | 2 +- .../session-reproduction-uncredited.json | 2 +- tests/valid/session-retrieval-tier.json | 2 +- tests/valid/turn-privacy-full.json | 2 +- tests/valid/turn-privacy-intent.json | 2 +- tests/valid/turn-privacy-minimal.json | 2 +- tests/valid/turn-privacy-summary.json | 2 +- tests/validate.py | 2 +- 74 files changed, 100 insertions(+), 92 deletions(-) diff --git a/README.md b/README.md index 2dc598c..607c85b 100644 --- a/README.md +++ b/README.md @@ -85,7 +85,7 @@ A user asks an AI agent about UK interest rates. The agent grounds its response ```json { - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440000", "agent_id": "copilot-v3", "started_at": "2026-03-28T09:00:00Z", diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 8d6ced3..e8a4bae 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -1,8 +1,8 @@ # Content Telemetry Specification -**Version:** 0.1 -**Status:** Preview -**Last updated:** 2026-08-04 +**Version:** 1.0 (release candidate draft) +**Status:** Release candidate in preparation (feature freeze 21 August 2026) +**Last updated:** 2026-08-12 ## Contents @@ -330,7 +330,7 @@ Additional content metadata - version, last-modified timestamp, content hash, me | Field | Type | Required | Description | |-------|------|----------|-------------| -| `schema_version` | string | Yes | Schema version (e.g., "0.1") | +| `schema_version` | string | Yes | Schema version (v1 documents declare "1.0") | | `session_id` | UUID | Yes | Unique session identifier | | `parent_session_id` | UUID | No | Immediate parent session that delegated work to this session | | `agent_id` | string | No | Responding agent identifier | @@ -609,7 +609,7 @@ A Citation emitter SHOULD: A conforming **telemetry consumer** MUST: -- Accept sessions with any `schema_version` that shares the same major version. During the preview period (0.x), consumers MUST accept sessions with the exact same minor version (e.g., a 0.1 consumer accepts 0.1 only). The major-version compatibility rule takes effect from 1.0.0 onward. +- Accept sessions with any `schema_version` that shares the same major version. V1 documents declare `schema_version` `"1.0"`. A v1 consumer MUST reject documents declaring `"0.1"`: v0.1 is a different wire version, not a compatible minor. Conversely, a v0.1 consumer following the preview rule (a 0.x consumer accepts only the exact same minor version, so a 0.1 consumer accepts 0.1 only) rejects documents declaring `"1.0"`. The major-version compatibility rule takes effect from 1.0.0 onward. - Tolerate unknown fields without error - Tolerate events from any conformance level - Accept the session-document, standalone-event, and event-batch delivery formats, reconstructing sessions from standalone events and event batches where needed (see section 7.1) @@ -930,7 +930,7 @@ A standalone event carries `document_type`, `schema_version`, and optionally `se ```json { "document_type": "event", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "550e8400-e29b-41d4-a716-446655440000", "event": { "type": "content_retrieved", @@ -952,7 +952,7 @@ An event batch carries the same envelope fields with `"document_type": "event_ba ```json { "document_type": "event_batch", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "550e8400-e29b-41d4-a716-446655440000", "events": [ { @@ -1143,7 +1143,7 @@ Machine-readable schema: [`./manifest.json`](./manifest.json) (JSON Schema draft | Field | Type | Required | Description | |-------|------|----------|-------------| -| `schema_version` | string | Yes | Manifest schema version. v0.1 emitters MUST use `"0.1"`. | +| `schema_version` | string | Yes | Manifest schema version. v1 emitters MUST use `"1.0"`. | | `id` | string | Yes | The manifest's canonical URL (e.g. `https://example.com/.well-known/content-telemetry.json`). | | `roles` | string[] | Yes | One or more of `content_owner`, `agent`, `platform`. | | `operator` | object | Yes | Operating organisation (see 8.3). | @@ -1213,7 +1213,7 @@ Consumers SHOULD cache resolved manifests respecting the response's `Cache-Contr ```json { - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://example.com/.well-known/content-telemetry.json", "roles": ["content_owner"], "operator": { "name": "Example Media" }, @@ -1228,7 +1228,7 @@ Consumers SHOULD cache resolved manifests respecting the response's `Cache-Contr ```json { - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://searchco.com/agents/web-search/.well-known/content-telemetry.json", "roles": ["agent"], "operator": { "name": "SearchCo" }, @@ -1247,7 +1247,7 @@ Consumers SHOULD cache resolved manifests respecting the response's `Cache-Contr ```json // https://publisher.com/.well-known/content-telemetry.json { - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://publisher.com/.well-known/content-telemetry.json", "roles": ["content_owner"], "operator": { "name": "Publisher Co" }, @@ -1261,7 +1261,7 @@ Consumers SHOULD cache resolved manifests respecting the response's `Cache-Contr ```json // https://publisher.com/agents/assistant/.well-known/content-telemetry.json { - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://publisher.com/agents/assistant/.well-known/content-telemetry.json", "roles": ["agent"], "operator": { "name": "Publisher Co" }, @@ -1426,6 +1426,14 @@ Telemetry consumers MUST tolerate unknown `response_mode` values. ### 12.1 Migration from the v0.1 preview +V1 documents declare `schema_version` `"1.0"`, and the schemas' `$id` URLs move +from `/schema/v0.1/` to `/schema/v1/`. The two versions are distinguishable on +the wire and do not interoperate: a v0.1 consumer, applying the preview rule of +section 5.7.4, rejects a document declaring `"1.0"`, and a v1 consumer rejects a +document declaring `"0.1"`. An emitter moves to v1 by declaring `"1.0"` on +documents that satisfy this section; it MUST NOT declare `"0.1"` on a document +using v1 event types or fields. + V1 replaces `content_displayed` with `content_presented`; emitters MUST NOT send the old event name on the v1 integration line. Rename `data.display_type` to `data.presentation_type` and add `data.presentation_kind` with either `content` @@ -1491,7 +1499,7 @@ A user asks a shopping assistant to compare noise-cancelling headphones. The age ```json { - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "550e8400-e29b-41d4-a716-446655440000", "agent_id": "shopping-assistant-v2", "content_scope": "electronics-reviews", @@ -1644,7 +1652,7 @@ An AI agent previously fetched an FT article and cached it. In a new session, th ```json { - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440000", "agent_id": "copilot-v3", "started_at": "2026-03-28T09:00:00Z", diff --git a/manifest.json b/manifest.json index fd69a7b..c2215ca 100644 --- a/manifest.json +++ b/manifest.json @@ -1,6 +1,6 @@ { "$schema": "https://json-schema.org/draft/2020-12/schema", - "$id": "https://contenttelemetry.org/schema/v0.1/manifest.json", + "$id": "https://contenttelemetry.org/schema/v1/manifest.json", "title": "Content Telemetry Manifest", "description": "Schema for the .well-known/content-telemetry.json manifest defined in section 8 of the Content Telemetry specification. A manifest declares a participant's identity, roles, telemetry endpoint, signing keys, and claimed domains.", "type": "object", @@ -8,8 +8,8 @@ "properties": { "schema_version": { "type": "string", - "const": "0.1", - "description": "Manifest schema version. v0.1 emitters MUST use '0.1'. (Section 8.2)" + "const": "1.0", + "description": "Manifest schema version. v1 emitters MUST use '1.0'. (Section 8.2)" }, "id": { "type": "string", diff --git a/telemetry-event-batch.json b/telemetry-event-batch.json index ceef522..b7a9754 100644 --- a/telemetry-event-batch.json +++ b/telemetry-event-batch.json @@ -1,6 +1,6 @@ { "$schema": "https://json-schema.org/draft/2020-12/schema", - "$id": "https://contenttelemetry.org/schema/v0.1/telemetry-event-batch.json", + "$id": "https://contenttelemetry.org/schema/v1/telemetry-event-batch.json", "title": "Content Telemetry Event Batch", "description": "Schema for batches of Content Telemetry events sharing one session context, delivered outside a session document", "type": "object", @@ -13,7 +13,7 @@ }, "schema_version": { "type": "string", - "const": "0.1", + "const": "1.0", "description": "Content Telemetry schema version" }, "session_id": { diff --git a/telemetry-event.json b/telemetry-event.json index 144ff75..9654c27 100644 --- a/telemetry-event.json +++ b/telemetry-event.json @@ -1,6 +1,6 @@ { "$schema": "https://json-schema.org/draft/2020-12/schema", - "$id": "https://contenttelemetry.org/schema/v0.1/telemetry-event.json", + "$id": "https://contenttelemetry.org/schema/v1/telemetry-event.json", "title": "Content Telemetry Standalone Event", "description": "Schema for standalone Content Telemetry events delivered outside a session document", "type": "object", @@ -13,7 +13,7 @@ }, "schema_version": { "type": "string", - "const": "0.1", + "const": "1.0", "description": "Content Telemetry schema version" }, "session_id": { diff --git a/telemetry-session.json b/telemetry-session.json index fb90e62..e446e41 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -1,6 +1,6 @@ { "$schema": "https://json-schema.org/draft/2020-12/schema", - "$id": "https://contenttelemetry.org/schema/v0.1/telemetry-session.json", + "$id": "https://contenttelemetry.org/schema/v1/telemetry-session.json", "title": "Content Telemetry Session", "description": "Schema for Content Telemetry sessions - tracking content usage in AI agent interactions", "type": "object", @@ -14,7 +14,7 @@ }, "schema_version": { "type": "string", - "const": "0.1", + "const": "1.0", "description": "Content Telemetry schema version" }, "conformance_level": { diff --git a/tests/README.md b/tests/README.md index ed2a846..4ca1a6e 100644 --- a/tests/README.md +++ b/tests/README.md @@ -1,6 +1,6 @@ # Conformance test suite -Tests for the Content Telemetry Specification v0.1. +Tests for the Content Telemetry Specification v1. ## Structure diff --git a/tests/invalid/batch-empty-events.json b/tests/invalid/batch-empty-events.json index f070acb..b11cdd5 100644 --- a/tests/invalid/batch-empty-events.json +++ b/tests/invalid/batch-empty-events.json @@ -1,7 +1,7 @@ { "_test_description": "Event batch envelope with an empty events array. The schema requires minItems 1: an empty batch carries no signal and MUST NOT be delivered.", "document_type": "event_batch", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440008", "events": [] } diff --git a/tests/invalid/batch-missing-session-and-ctx-token.json b/tests/invalid/batch-missing-session-and-ctx-token.json index 4ceb7e6..adc810b 100644 --- a/tests/invalid/batch-missing-session-and-ctx-token.json +++ b/tests/invalid/batch-missing-session-and-ctx-token.json @@ -1,7 +1,7 @@ { "_test_description": "Event batch envelope carrying a content_cited event with neither session_id nor ctx_token on the envelope. Passes JSON Schema (both fields are optional) but violates section 7.1: an event MUST carry one at Grounding conformance and above.", "document_type": "event_batch", - "schema_version": "0.1", + "schema_version": "1.0", "events": [ { "id": "990e8400-e29b-41d4-a716-446655440063", diff --git a/tests/invalid/cited-missing-output-id.json b/tests/invalid/cited-missing-output-id.json index 014e7bd..5f7686b 100644 --- a/tests/invalid/cited-missing-output-id.json +++ b/tests/invalid/cited-missing-output-id.json @@ -1,6 +1,6 @@ { "_test_description": "content_cited event with id and a resolvable source reference but no output_id. output_id is required on citation events so output construction can be correlated with later presentation.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440160", "started_at": "2026-08-01T14:00:00Z", "events": [ diff --git a/tests/invalid/cited-missing-source-reference.json b/tests/invalid/cited-missing-source-reference.json index d987fbc..d619c37 100644 --- a/tests/invalid/cited-missing-source-reference.json +++ b/tests/invalid/cited-missing-source-reference.json @@ -1,6 +1,6 @@ { "_test_description": "content_cited event carrying neither content_url nor content_id. A source association with no resolvable reference is not a citation; unlike other content events, the JSON Schema enforces the identifier requirement for content_cited (section 6.5).", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440140", "agent_id": "research-assistant-v1", "started_at": "2026-08-01T12:00:00Z", diff --git a/tests/invalid/cited-null-source-reference.json b/tests/invalid/cited-null-source-reference.json index 2344da6..754d1c3 100644 --- a/tests/invalid/cited-null-source-reference.json +++ b/tests/invalid/cited-null-source-reference.json @@ -1,6 +1,6 @@ { "_test_description": "content_cited event with content_url explicitly null and no content_id. Presence of a null reference does not satisfy the citation reference requirement: the schema demands a non-null content_url or content_id on content_cited (section 6.5).", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440141", "agent_id": "research-assistant-v1", "started_at": "2026-08-01T12:30:00Z", diff --git a/tests/invalid/content-event-missing-identifier.json b/tests/invalid/content-event-missing-identifier.json index 9f58702..db8507f 100644 --- a/tests/invalid/content-event-missing-identifier.json +++ b/tests/invalid/content-event-missing-identifier.json @@ -1,6 +1,6 @@ { "_test_description": "content_grounded event carrying neither content_url nor content_id. Passes JSON Schema (both are individually optional) but violates section 5.7.5: at least one of content_url or content_id MUST be present on every content event.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440099", "agent_id": "research-assistant-v1", "started_at": "2026-03-28T12:00:00Z", diff --git a/tests/invalid/engaged-missing-presentation-id.json b/tests/invalid/engaged-missing-presentation-id.json index 8632196..a80e22a 100644 --- a/tests/invalid/engaged-missing-presentation-id.json +++ b/tests/invalid/engaged-missing-presentation-id.json @@ -1,6 +1,6 @@ { "_test_description": "content_engaged must identify the exact presentation occurrence rather than matching only by URL.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440092", "started_at": "2026-07-18T13:00:00Z", "events": [ diff --git a/tests/invalid/invalid-event-type.json b/tests/invalid/invalid-event-type.json index af0a4be..589ec33 100644 --- a/tests/invalid/invalid-event-type.json +++ b/tests/invalid/invalid-event-type.json @@ -1,6 +1,6 @@ { "_test_description": "Event with type 'content_summarised' which is not in the EventType enum. Fails JSON Schema validation.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440008", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/invalid-privacy-level.json b/tests/invalid/invalid-privacy-level.json index 6449c79..a4c762d 100644 --- a/tests/invalid/invalid-privacy-level.json +++ b/tests/invalid/invalid-privacy-level.json @@ -1,6 +1,6 @@ { "_test_description": "Turn with privacy_level 'redacted' which is not in the PrivacyLevel enum. Fails JSON Schema validation.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440009", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/invalid-schema-version.json b/tests/invalid/invalid-schema-version.json index 529831f..2099ef8 100644 --- a/tests/invalid/invalid-schema-version.json +++ b/tests/invalid/invalid-schema-version.json @@ -1,5 +1,5 @@ { - "_test_description": "schema_version set to '2.0' which does not match the const '0.1'. Fails JSON Schema validation.", + "_test_description": "schema_version set to '2.0' which does not match the const '1.0'. Fails JSON Schema validation.", "schema_version": "2.0", "session_id": "770e8400-e29b-41d4-a716-446655440010", "started_at": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/invalid-source-role.json b/tests/invalid/invalid-source-role.json index bb8453e..81464df 100644 --- a/tests/invalid/invalid-source-role.json +++ b/tests/invalid/invalid-source-role.json @@ -1,6 +1,6 @@ { "_test_description": "Event with source_role 'cdn' which is not in the SourceRole enum. Fails JSON Schema validation.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440014", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/legacy-content-displayed.json b/tests/invalid/legacy-content-displayed.json index ab93b1e..401644e 100644 --- a/tests/invalid/legacy-content-displayed.json +++ b/tests/invalid/legacy-content-displayed.json @@ -1,6 +1,6 @@ { "_test_description": "The v1 presentation event replaces the preview content_displayed name.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440093", "started_at": "2026-07-18T13:00:00Z", "events": [ diff --git a/tests/invalid/manifest-bad-conformance-level.json b/tests/invalid/manifest-bad-conformance-level.json index 85cb563..ec22c79 100644 --- a/tests/invalid/manifest-bad-conformance-level.json +++ b/tests/invalid/manifest-bad-conformance-level.json @@ -1,6 +1,6 @@ { "_test_description": "Manifest with telemetry.conformance_level 'attribution' which is not one of retrieval, grounding, citation. Fails JSON Schema validation.", - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://searchco.com/agents/web-search/.well-known/content-telemetry.json", "roles": ["agent"], "operator": { "name": "SearchCo" }, diff --git a/tests/invalid/manifest-bad-role.json b/tests/invalid/manifest-bad-role.json index d7208b0..c1c20c4 100644 --- a/tests/invalid/manifest-bad-role.json +++ b/tests/invalid/manifest-bad-role.json @@ -1,6 +1,6 @@ { "_test_description": "Manifest with a role value 'publisher' that is not in the roles enum (content_owner, agent, platform). Fails JSON Schema validation.", - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://example.com/.well-known/content-telemetry.json", "roles": ["publisher"], "operator": { "name": "Example Media" } diff --git a/tests/invalid/manifest-duplicate-key-id.json b/tests/invalid/manifest-duplicate-key-id.json index 16a2d71..ce403f9 100644 --- a/tests/invalid/manifest-duplicate-key-id.json +++ b/tests/invalid/manifest-duplicate-key-id.json @@ -1,6 +1,6 @@ { "_test_description": "Agent manifest with two keys sharing the id 'key-1'. Passes JSON Schema (the entries differ in publicKey) but violates section 8.7: consumers reject a manifest with duplicate keys[].id.", - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://searchco.com/agents/web-search/.well-known/content-telemetry.json", "roles": ["agent"], "operator": { "name": "SearchCo" }, diff --git a/tests/invalid/manifest-foreign-domain.json b/tests/invalid/manifest-foreign-domain.json index 23968f9..ff47db1 100644 --- a/tests/invalid/manifest-foreign-domain.json +++ b/tests/invalid/manifest-foreign-domain.json @@ -1,6 +1,6 @@ { "_test_description": "Content-owner manifest at example.com whose domains array claims othersite.com. Passes JSON Schema but violates section 8.6: every domains entry MUST be the manifest's own host or a subdomain of it. Consumers reject the manifest as malformed (section 8.7).", - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://example.com/.well-known/content-telemetry.json", "roles": ["content_owner"], "operator": { "name": "Example Media" }, diff --git a/tests/invalid/manifest-missing-key-publickey.json b/tests/invalid/manifest-missing-key-publickey.json index dec6aa9..1cb5de6 100644 --- a/tests/invalid/manifest-missing-key-publickey.json +++ b/tests/invalid/manifest-missing-key-publickey.json @@ -1,6 +1,6 @@ { "_test_description": "Manifest with a keys entry missing the required 'publicKey' field. Fails JSON Schema validation.", - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://searchco.com/agents/web-search/.well-known/content-telemetry.json", "roles": ["agent"], "operator": { "name": "SearchCo" }, diff --git a/tests/invalid/manifest-missing-operator.json b/tests/invalid/manifest-missing-operator.json index ceb5df7..b3b9031 100644 --- a/tests/invalid/manifest-missing-operator.json +++ b/tests/invalid/manifest-missing-operator.json @@ -1,6 +1,6 @@ { "_test_description": "Manifest missing the required 'operator' field. Fails JSON Schema validation.", - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://example.com/.well-known/content-telemetry.json", "roles": ["content_owner"] } diff --git a/tests/invalid/missing-event-timestamp.json b/tests/invalid/missing-event-timestamp.json index f626857..45ba33a 100644 --- a/tests/invalid/missing-event-timestamp.json +++ b/tests/invalid/missing-event-timestamp.json @@ -1,6 +1,6 @@ { "_test_description": "Event without timestamp field. Fails JSON Schema validation: timestamp is required on TelemetryEvent.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440005", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/missing-event-type.json b/tests/invalid/missing-event-type.json index d0ea6dc..fb5efad 100644 --- a/tests/invalid/missing-event-type.json +++ b/tests/invalid/missing-event-type.json @@ -1,6 +1,6 @@ { "_test_description": "Event without type field. Fails JSON Schema validation: type is required on TelemetryEvent.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440004", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/missing-privacy-level.json b/tests/invalid/missing-privacy-level.json index 28cab62..1184e23 100644 --- a/tests/invalid/missing-privacy-level.json +++ b/tests/invalid/missing-privacy-level.json @@ -1,6 +1,6 @@ { "_test_description": "Turn object without privacy_level. Fails JSON Schema validation: privacy_level is required on ConversationTurn.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440006", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/missing-session-id.json b/tests/invalid/missing-session-id.json index 4474fa6..d43a75a 100644 --- a/tests/invalid/missing-session-id.json +++ b/tests/invalid/missing-session-id.json @@ -1,6 +1,6 @@ { "_test_description": "Session without session_id. Fails JSON Schema validation: session_id is a required field.", - "schema_version": "0.1", + "schema_version": "1.0", "started_at": "2026-03-28T10:00:00Z", "events": [ { diff --git a/tests/invalid/missing-started-at.json b/tests/invalid/missing-started-at.json index caa221f..ee2fd38 100644 --- a/tests/invalid/missing-started-at.json +++ b/tests/invalid/missing-started-at.json @@ -1,6 +1,6 @@ { "_test_description": "Session without started_at. Fails JSON Schema validation: started_at is a required field.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440003", "events": [ { diff --git a/tests/invalid/presented-missing-kind.json b/tests/invalid/presented-missing-kind.json index 27a7594..1e298d1 100644 --- a/tests/invalid/presented-missing-kind.json +++ b/tests/invalid/presented-missing-kind.json @@ -1,6 +1,6 @@ { "_test_description": "content_presented must distinguish source content from a source reference.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440090", "started_at": "2026-07-18T13:00:00Z", "events": [ diff --git a/tests/invalid/presented-missing-presentation-type.json b/tests/invalid/presented-missing-presentation-type.json index bacdd0e..3ccf305 100644 --- a/tests/invalid/presented-missing-presentation-type.json +++ b/tests/invalid/presented-missing-presentation-type.json @@ -1,6 +1,6 @@ { "_test_description": "content_presented event carrying presentation_kind but no presentation_type. Both are required: the kind says what crossed the presentation boundary, the type says how it was made perceivable.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440163", "started_at": "2026-08-01T15:00:00Z", "events": [ diff --git a/tests/invalid/privacy-violation-ad-rendered-at-minimal.json b/tests/invalid/privacy-violation-ad-rendered-at-minimal.json index 913540d..7cbd248 100644 --- a/tests/invalid/privacy-violation-ad-rendered-at-minimal.json +++ b/tests/invalid/privacy-violation-ad-rendered-at-minimal.json @@ -1,6 +1,6 @@ { "_test_description": "APPLICATION-LAYER VIOLATION: Turn at minimal privacy with ad_rendered present. Passes JSON Schema validation but violates section 5.5: platform metadata (including ad_rendered) is not available at minimal level.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440012", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/privacy-violation-query-at-intent.json b/tests/invalid/privacy-violation-query-at-intent.json index c08d585..8e033c9 100644 --- a/tests/invalid/privacy-violation-query-at-intent.json +++ b/tests/invalid/privacy-violation-query-at-intent.json @@ -1,6 +1,6 @@ { "_test_description": "APPLICATION-LAYER VIOLATION: Turn at intent privacy with query_text present. Passes JSON Schema validation but violates section 5.5: query_text MUST NOT be present at intent level.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440013", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/privacy-violation-query-at-minimal.json b/tests/invalid/privacy-violation-query-at-minimal.json index 950a5bb..31e3bd3 100644 --- a/tests/invalid/privacy-violation-query-at-minimal.json +++ b/tests/invalid/privacy-violation-query-at-minimal.json @@ -1,6 +1,6 @@ { "_test_description": "APPLICATION-LAYER VIOLATION: Turn at minimal privacy with query_text present. Passes JSON Schema validation but violates the privacy field gating rule (section 5.5): query_text MUST NOT be present when privacy_level is minimal.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440011", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/reproduced-missing-id.json b/tests/invalid/reproduced-missing-id.json index b322854..6dad3fb 100644 --- a/tests/invalid/reproduced-missing-id.json +++ b/tests/invalid/reproduced-missing-id.json @@ -1,6 +1,6 @@ { "_test_description": "content_reproduced event with output_id and a resolvable source reference but no event id. id is required on reproduction events so a crediting citation or verification result can reference the exact reproduction claim.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440162", "started_at": "2026-08-01T14:30:00Z", "events": [ diff --git a/tests/invalid/reproduced-missing-output-id.json b/tests/invalid/reproduced-missing-output-id.json index 4f5405b..8140942 100644 --- a/tests/invalid/reproduced-missing-output-id.json +++ b/tests/invalid/reproduced-missing-output-id.json @@ -1,6 +1,6 @@ { "_test_description": "content_reproduced must identify the output artifact containing the reproduction: id and output_id are required so reproduction can be correlated with citation and presentation.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440130", "started_at": "2026-08-01T11:00:00Z", "events": [ diff --git a/tests/invalid/reproduced-missing-source-reference.json b/tests/invalid/reproduced-missing-source-reference.json index 4c33aa5..a9d5a1b 100644 --- a/tests/invalid/reproduced-missing-source-reference.json +++ b/tests/invalid/reproduced-missing-source-reference.json @@ -1,6 +1,6 @@ { "_test_description": "content_reproduced event carrying neither content_url nor content_id. A reproduction claim is only meaningful for an identified source; the JSON Schema requires a non-null content_url or content_id on content_reproduced (section 6.6).", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440150", "started_at": "2026-08-01T13:00:00Z", "events": [ diff --git a/tests/invalid/reproduced-missing-type.json b/tests/invalid/reproduced-missing-type.json index a5c00fb..3fae0ee 100644 --- a/tests/invalid/reproduced-missing-type.json +++ b/tests/invalid/reproduced-missing-type.json @@ -1,6 +1,6 @@ { "_test_description": "content_reproduced must classify the fidelity of the reproduction: data.reproduction_type is required.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440120", "started_at": "2026-08-01T10:00:00Z", "events": [ diff --git a/tests/invalid/standalone-missing-document-type.json b/tests/invalid/standalone-missing-document-type.json index b0a0186..2731dba 100644 --- a/tests/invalid/standalone-missing-document-type.json +++ b/tests/invalid/standalone-missing-document-type.json @@ -1,6 +1,6 @@ { "_test_description": "Envelope-shaped document with an event key but no document_type. Per section 7.1 a document without document_type is treated as a session, and it fails the session schema.", - "schema_version": "0.1", + "schema_version": "1.0", "event": { "type": "content_retrieved", "timestamp": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/standalone-missing-event.json b/tests/invalid/standalone-missing-event.json index c75bdfc..bb91d7b 100644 --- a/tests/invalid/standalone-missing-event.json +++ b/tests/invalid/standalone-missing-event.json @@ -1,6 +1,6 @@ { "_test_description": "Standalone event envelope (document_type 'event') with no event field. Fails the envelope schema, which requires event.", "document_type": "event", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440010" } diff --git a/tests/invalid/standalone-missing-session-and-ctx-token.json b/tests/invalid/standalone-missing-session-and-ctx-token.json index 843c557..c96811b 100644 --- a/tests/invalid/standalone-missing-session-and-ctx-token.json +++ b/tests/invalid/standalone-missing-session-and-ctx-token.json @@ -1,7 +1,7 @@ { "_test_description": "Standalone event envelope carrying neither session_id nor ctx_token. Passes JSON Schema (both are individually optional on the envelope) but violates section 5.7.5: an event MUST carry either session_id or ctx_token at Grounding conformance and above (section 7.1).", "document_type": "event", - "schema_version": "0.1", + "schema_version": "1.0", "event": { "type": "content_grounded", "timestamp": "2026-03-28T08:20:00Z", diff --git a/tests/invalid/withdrawn-ip-hash.json b/tests/invalid/withdrawn-ip-hash.json index 100b3ae..f691b79 100644 --- a/tests/invalid/withdrawn-ip-hash.json +++ b/tests/invalid/withdrawn-ip-hash.json @@ -1,7 +1,7 @@ { "_test_description": "Standalone event envelope from a CDN carrying ip_hash in data. The field was withdrawn in v1 (section 9.1): hashing does not anonymise a value drawn from a space small enough to enumerate. Passes JSON Schema because event data accepts additional properties, so the rule is enforced at the application layer.", "document_type": "event", - "schema_version": "0.1", + "schema_version": "1.0", "event": { "type": "content_retrieved", "timestamp": "2026-03-28T08:15:00Z", diff --git a/tests/valid/event-batch-agent.json b/tests/valid/event-batch-agent.json index 08e605c..7b4deff 100644 --- a/tests/valid/event-batch-agent.json +++ b/tests/valid/event-batch-agent.json @@ -1,7 +1,7 @@ { "_test_description": "Event batch envelope from an agent at Grounding conformance, carrying the session-level fields (session_id, agent_id, started_at) that apply to every event in the batch. The agent buffers events within a session and flushes them together.", "document_type": "event_batch", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440007", "agent_id": "assistant.example.com", "started_at": "2026-03-28T08:19:55Z", diff --git a/tests/valid/event-batch-edge.json b/tests/valid/event-batch-edge.json index 8ae083d..03ee1c5 100644 --- a/tests/valid/event-batch-edge.json +++ b/tests/valid/event-batch-edge.json @@ -1,7 +1,7 @@ { "_test_description": "Event batch envelope from a CDN at Retrieval conformance level. No session_id (content owner has no session context); the edge platform buffers detections across requests and flushes them as one batch. Validated against telemetry-event-batch.json envelope schema.", "document_type": "event_batch", - "schema_version": "0.1", + "schema_version": "1.0", "events": [ { "type": "content_retrieved", diff --git a/tests/valid/event-standalone-agent.json b/tests/valid/event-standalone-agent.json index c5a58a0..095945c 100644 --- a/tests/valid/event-standalone-agent.json +++ b/tests/valid/event-standalone-agent.json @@ -1,7 +1,7 @@ { "_test_description": "Standalone event envelope from an agent, with session_id as a foreign key. The agent reports its own retrieval as a standalone event for streaming delivery.", "document_type": "event", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440006", "event": { "type": "content_retrieved", diff --git a/tests/valid/event-standalone-child-session.json b/tests/valid/event-standalone-child-session.json index d136d62..59c9f37 100644 --- a/tests/valid/event-standalone-child-session.json +++ b/tests/valid/event-standalone-child-session.json @@ -1,7 +1,7 @@ { "_test_description": "Standalone event envelope from a child agent session in a multi-agent topology: parent_session_id on the envelope identifies the orchestrating session. The grounded content entered the sub-agent's generation context.", "document_type": "event", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440170", "parent_session_id": "660e8400-e29b-41d4-a716-446655440171", "agent_id": "research-subagent-v2", diff --git a/tests/valid/event-standalone-edge.json b/tests/valid/event-standalone-edge.json index 018f44b..086bb35 100644 --- a/tests/valid/event-standalone-edge.json +++ b/tests/valid/event-standalone-edge.json @@ -1,7 +1,7 @@ { "_test_description": "Standalone event envelope from a CDN at Retrieval conformance level. No session_id (content owner has no session context). source_role is edge with full edge enrichment data. Validated against telemetry-event.json envelope schema.", "document_type": "event", - "schema_version": "0.1", + "schema_version": "1.0", "event": { "type": "content_retrieved", "timestamp": "2026-03-28T08:15:00Z", diff --git a/tests/valid/event-standalone-engaged-ctx-token.json b/tests/valid/event-standalone-engaged-ctx-token.json index ce90f25..9a0d9cc 100644 --- a/tests/valid/event-standalone-engaged-ctx-token.json +++ b/tests/valid/event-standalone-engaged-ctx-token.json @@ -1,7 +1,7 @@ { "_test_description": "Standalone content_engaged (link_click) reported from a landing page after a click-out. Carries ctx_token in place of session_id and no presentation_id: the destination cannot know the presentation UUID, and the consumer restores the token's presentation binding at resolution (section 7.4).", "document_type": "event", - "schema_version": "0.1", + "schema_version": "1.0", "ctx_token": "ct_9f3a1c7e2b8d4a06", "event": { "type": "content_engaged", diff --git a/tests/valid/event-standalone-presented.json b/tests/valid/event-standalone-presented.json index 3ef689f..58ce6bf 100644 --- a/tests/valid/event-standalone-presented.json +++ b/tests/valid/event-standalone-presented.json @@ -1,7 +1,7 @@ { "_test_description": "Standalone presentation envelope using the shared modality-neutral presentation schema.", "document_type": "event", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440080", "agent_id": "voice-assistant-v1", "started_at": "2026-07-18T12:30:00Z", diff --git a/tests/valid/manifest-agent-with-keys.json b/tests/valid/manifest-agent-with-keys.json index 7b90977..fbb629e 100644 --- a/tests/valid/manifest-agent-with-keys.json +++ b/tests/valid/manifest-agent-with-keys.json @@ -1,6 +1,6 @@ { "_test_description": "Agent manifest served under a path prefix, with an Ed25519 signing key (including expires) and a telemetry endpoint advertising grounding conformance.", - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://searchco.com/agents/web-search/.well-known/content-telemetry.json", "roles": ["agent"], "operator": { "name": "SearchCo" }, diff --git a/tests/valid/manifest-content-owner-full.json b/tests/valid/manifest-content-owner-full.json index b71b116..533fb38 100644 --- a/tests/valid/manifest-content-owner-full.json +++ b/tests/valid/manifest-content-owner-full.json @@ -1,6 +1,6 @@ { "_test_description": "Content-owner manifest with operator.domain, a telemetry endpoint (no conformance_level), and a domains array with literal and wildcard subdomains.", - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://example.com/.well-known/content-telemetry.json", "roles": ["content_owner"], "operator": { "name": "Example Media", "domain": "example.com" }, diff --git a/tests/valid/manifest-content-owner-minimal.json b/tests/valid/manifest-content-owner-minimal.json index bc9a7f1..3fcdae6 100644 --- a/tests/valid/manifest-content-owner-minimal.json +++ b/tests/valid/manifest-content-owner-minimal.json @@ -1,6 +1,6 @@ { "_test_description": "Minimal valid content-owner manifest: schema_version, id, roles, operator only. No keys, telemetry, or domains.", - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://example.com/.well-known/content-telemetry.json", "roles": ["content_owner"], "operator": { "name": "Example Media" } diff --git a/tests/valid/manifest-multi-role.json b/tests/valid/manifest-multi-role.json index 034e65f..8d185c8 100644 --- a/tests/valid/manifest-multi-role.json +++ b/tests/valid/manifest-multi-role.json @@ -1,6 +1,6 @@ { "_test_description": "Single manifest declaring multiple roles (content_owner and agent) with a key, telemetry endpoint, and domains.", - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://publisher.com/.well-known/content-telemetry.json", "roles": ["content_owner", "agent"], "operator": { "name": "Publisher Co" }, diff --git a/tests/valid/session-cached-grounding.json b/tests/valid/session-cached-grounding.json index 8f8347a..11423fe 100644 --- a/tests/valid/session-cached-grounding.json +++ b/tests/valid/session-cached-grounding.json @@ -1,6 +1,6 @@ { "_test_description": "Session document with agent-cached grounding and emitter-reported generic fingerprint detection. license_ref is preserved from the original retrieval, as recommended by section 6.4.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440004", "agent_id": "copilot-v3", "started_at": "2026-03-28T09:00:00Z", diff --git a/tests/valid/session-citation-tier.json b/tests/valid/session-citation-tier.json index b8f9743..c2604d4 100644 --- a/tests/valid/session-citation-tier.json +++ b/tests/valid/session-citation-tier.json @@ -1,6 +1,6 @@ { "_test_description": "Citation conformance level with optional presentation and engagement lifecycle signals.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440003", "agent_id": "shopping-assistant-v2", "content_scope": "electronics-reviews", diff --git a/tests/valid/session-custom-media-type.json b/tests/valid/session-custom-media-type.json index 1d802f9..97e965c 100644 --- a/tests/valid/session-custom-media-type.json +++ b/tests/valid/session-custom-media-type.json @@ -1,6 +1,6 @@ { "_test_description": "Custom media_type value not in the core set (text, image, video, audio). The spec says emitters MAY use custom string values for media outside the core set and telemetry consumers MUST tolerate unknown media_type values. Exercises the custom value on both content_retrieved and content_grounded data.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440010", "agent_id": "cad-assistant-v2", "started_at": "2026-03-28T19:00:00Z", diff --git a/tests/valid/session-custom-response-mode.json b/tests/valid/session-custom-response-mode.json index fb5e627..0d5b72e 100644 --- a/tests/valid/session-custom-response-mode.json +++ b/tests/valid/session-custom-response-mode.json @@ -1,6 +1,6 @@ { "_test_description": "Custom response_mode value not in the recommended set. The spec says platforms MAY use custom string values and telemetry consumers MUST tolerate unknown values.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440009", "agent_id": "podcast-gen-v1", "started_at": "2026-03-28T18:00:00Z", diff --git a/tests/valid/session-funnel-exception-cited-no-grounded.json b/tests/valid/session-funnel-exception-cited-no-grounded.json index 6145718..e84a4f1 100644 --- a/tests/valid/session-funnel-exception-cited-no-grounded.json +++ b/tests/valid/session-funnel-exception-cited-no-grounded.json @@ -1,6 +1,6 @@ { "_test_description": "Funnel exception: content_cited with no preceding content_grounded event. This is a hallucinated citation - the agent references content it never retrieved or loaded into context. Valid per section 4.3. Telemetry consumers SHOULD treat this as a lower-confidence signal.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440011", "agent_id": "assistant-v4", "started_at": "2026-03-28T20:00:00Z", diff --git a/tests/valid/session-funnel-exception-presented-no-cited.json b/tests/valid/session-funnel-exception-presented-no-cited.json index 7c4d219..d4b3f4a 100644 --- a/tests/valid/session-funnel-exception-presented-no-cited.json +++ b/tests/valid/session-funnel-exception-presented-no-cited.json @@ -1,6 +1,6 @@ { "_test_description": "Funnel exception: content_presented without content_cited. This is the Sources sidebar pattern - the agent makes source links perceivable without semantically associating them with an output element. Valid per section 4.3.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440010", "agent_id": "search-assistant-v1", "started_at": "2026-03-28T19:00:00Z", diff --git a/tests/valid/session-funnel-exception-presented-no-grounded.json b/tests/valid/session-funnel-exception-presented-no-grounded.json index 856afc4..9a6de97 100644 --- a/tests/valid/session-funnel-exception-presented-no-grounded.json +++ b/tests/valid/session-funnel-exception-presented-no-grounded.json @@ -1,6 +1,6 @@ { "_test_description": "Funnel exception: content_presented without content_grounded. An agentic browser renders a publisher's page on a recipient-facing surface (presentation_type: embed) without the content entering a generation context, and the recipient directs the agent to open a second page (engagement_type: agent_navigate). Valid per section 4.3. Also exercises media_type on presentation data and the open presentation_type/engagement_type vocabularies.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440020", "agent_id": "browser-agent-v1", "started_at": "2026-03-29T11:00:00Z", diff --git a/tests/valid/session-grounding-tier.json b/tests/valid/session-grounding-tier.json index 576abea..e10f147 100644 --- a/tests/valid/session-grounding-tier.json +++ b/tests/valid/session-grounding-tier.json @@ -1,6 +1,6 @@ { "_test_description": "Grounding conformance level: agent emitting content_retrieved, content_grounded, and turn events with privacy_level. Includes agent_id and data.scope as required by the Grounding tier.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440002", "agent_id": "research-assistant-v1", "started_at": "2026-03-28T12:00:00Z", diff --git a/tests/valid/session-minimal.json b/tests/valid/session-minimal.json index b9c1ee0..39f46f8 100644 --- a/tests/valid/session-minimal.json +++ b/tests/valid/session-minimal.json @@ -1,6 +1,6 @@ { "_test_description": "Bare minimum conforming session: schema_version, session_id, started_at, and one event with type and timestamp.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "550e8400-e29b-41d4-a716-446655440000", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/valid/session-multi-agent-child.json b/tests/valid/session-multi-agent-child.json index d333791..400b4fa 100644 --- a/tests/valid/session-multi-agent-child.json +++ b/tests/valid/session-multi-agent-child.json @@ -1,7 +1,7 @@ { "_test_description": "A delegated child session links to its immediate parent. The source enters the child agent's generation context and is grounded there; no citation or presentation is emitted merely because the child returns an internal response to its orchestrator.", "document_type": "session", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440021", "parent_session_id": "660e8400-e29b-41d4-a716-446655440020", "agent_id": "research-subagent-v1", diff --git a/tests/valid/session-multi-turn.json b/tests/valid/session-multi-turn.json index a220bf1..f3da34b 100644 --- a/tests/valid/session-multi-turn.json +++ b/tests/valid/session-multi-turn.json @@ -1,6 +1,6 @@ { "_test_description": "Session-scoped grounding, 3 turns, citations in turns 1 and 3, zero-click outcome. Turn 2 has no citation - the grounded content was in context but not explicitly referenced.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440006", "agent_id": "copilot-v3", "started_at": "2026-03-28T16:00:00Z", diff --git a/tests/valid/session-presentation-multimodal.json b/tests/valid/session-presentation-multimodal.json index c3a8a3d..61da85e 100644 --- a/tests/valid/session-presentation-multimodal.json +++ b/tests/valid/session-presentation-multimodal.json @@ -1,6 +1,6 @@ { "_test_description": "Modality-neutral citation and presentation: an image excerpt is presented as content, an audio output speaks a source credit, a video is embedded without citation, a constructed citation is suppressed, and a repeated link presentation receives the engagement.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440070", "agent_id": "multimodal-assistant-v1", "started_at": "2026-07-18T12:00:00Z", diff --git a/tests/valid/session-reproduction-credited-quote.json b/tests/valid/session-reproduction-credited-quote.json index 29628ab..d9d21c6 100644 --- a/tests/valid/session-reproduction-credited-quote.json +++ b/tests/valid/session-reproduction-credited-quote.json @@ -1,6 +1,6 @@ { "_test_description": "Credited quotation: the same span produces a content_reproduced event (the material) and a content_cited direct_quote event (the credit), sharing output_element_id, with reproduced_hash equal to the citation's excerpt_hash. The quote is then presented inline. Demonstrates that reproduction and citation are sibling output-construction claims joined to one presentation.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440111", "agent_id": "research-assistant-v5", "started_at": "2026-08-01T15:00:00Z", diff --git a/tests/valid/session-reproduction-uncredited.json b/tests/valid/session-reproduction-uncredited.json index 3c031e0..b6be589 100644 --- a/tests/valid/session-reproduction-uncredited.json +++ b/tests/valid/session-reproduction-uncredited.json @@ -1,6 +1,6 @@ { "_test_description": "Uncredited reproduction in unpresented output: an API-delivered response contains a verbatim excerpt of grounded content with no citation and no presentation. The content_reproduced event is the only record that source material appears in the output. Valid per section 4.3 (reproduced without cited).", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440101", "agent_id": "answers-api-v3", "started_at": "2026-08-01T09:00:00Z", diff --git a/tests/valid/session-retrieval-tier.json b/tests/valid/session-retrieval-tier.json index b17669d..8e0a803 100644 --- a/tests/valid/session-retrieval-tier.json +++ b/tests/valid/session-retrieval-tier.json @@ -1,6 +1,6 @@ { "_test_description": "Retrieval conformance level: content owner CDN emitting content_retrieved with source_role and content_url. No session context beyond the minimum.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440001", "started_at": "2026-03-28T11:00:00Z", "events": [ diff --git a/tests/valid/turn-privacy-full.json b/tests/valid/turn-privacy-full.json index 0cef475..b68a70b 100644 --- a/tests/valid/turn-privacy-full.json +++ b/tests/valid/turn-privacy-full.json @@ -1,6 +1,6 @@ { "_test_description": "turn_completed at full privacy level: query_text, response_text, intent, topics, model_id, ad_rendered all present. Maximum data sharing.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440007", "agent_id": "assistant-v4", "started_at": "2026-03-28T17:00:00Z", diff --git a/tests/valid/turn-privacy-intent.json b/tests/valid/turn-privacy-intent.json index 14a7dac..406a6e1 100644 --- a/tests/valid/turn-privacy-intent.json +++ b/tests/valid/turn-privacy-intent.json @@ -1,6 +1,6 @@ { "_test_description": "turn_completed at intent privacy level: query_intent, topics, response_type, response_mode, model_id, ad_rendered present. No query_text or response_text.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440012", "agent_id": "assistant-v4", "started_at": "2026-03-28T18:30:00Z", diff --git a/tests/valid/turn-privacy-minimal.json b/tests/valid/turn-privacy-minimal.json index 00e6b11..bc5d5db 100644 --- a/tests/valid/turn-privacy-minimal.json +++ b/tests/valid/turn-privacy-minimal.json @@ -1,6 +1,6 @@ { "_test_description": "turn_completed at minimal privacy level: only response_tokens and content_urls present. No intent, topics, query_text, response_text, model_id, ad_rendered, or response_mode.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440008", "agent_id": "assistant-v4", "started_at": "2026-03-28T17:30:00Z", diff --git a/tests/valid/turn-privacy-summary.json b/tests/valid/turn-privacy-summary.json index 38a2e99..44ff95d 100644 --- a/tests/valid/turn-privacy-summary.json +++ b/tests/valid/turn-privacy-summary.json @@ -1,6 +1,6 @@ { "_test_description": "turn_completed at summary privacy level: query_text and response_text are summarised (not verbatim). Includes query_intent, topics, query_tokens, response_tokens, model_id, ad_rendered, response_mode, response_type.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440009", "agent_id": "assistant-v4", "started_at": "2026-03-28T18:00:00Z", diff --git a/tests/validate.py b/tests/validate.py index 5790c6f..a2cea17 100644 --- a/tests/validate.py +++ b/tests/validate.py @@ -1,6 +1,6 @@ #!/usr/bin/env python3 """ -Conformance test runner for Content Telemetry Specification v0.1. +Conformance test runner for Content Telemetry Specification v1. Validates JSON test fixtures against telemetry-session.json, telemetry-event.json, telemetry-event-batch.json, manifest.json, and application-layer conformance rules From 23a4998159fd98b014c3bfebdd3e6125d88053e0 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Wed, 12 Aug 2026 11:10:40 +0100 Subject: [PATCH 25/38] Correct the 5.2 id row: conditionally required, emitter-assigned The table said id was optional and generated by the server if not provided. The schema requires id on content_cited, content_reproduced and content_presented, and server-side generation is incompatible with citation_id and presentation_id, which reference an id the emitter must already hold when it constructs the referencing event. Reworded to the table's conditional style (as used by output_id and presentation_id): required for reproduced/cited/presented, optional elsewhere, emitter-assigned. Co-Authored-By: Claude Fable 5 --- SPECIFICATION.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/SPECIFICATION.md b/SPECIFICATION.md index e8a4bae..81b41f7 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -409,7 +409,7 @@ Emitters MUST NOT populate `access_context` unless the governing terms of the re | Field | Type | Required | Description | |-------|------|----------|-------------| -| `id` | UUID | No | Unique event identifier (generated by server if not provided) | +| `id` | UUID | For reproduced/cited/presented | Emitter-assigned unique event identifier; optional on other event types | | `type` | EventType | Yes | Event type (see 5.3) | | `timestamp` | datetime | Yes | Event timestamp (UTC) | | `turn_id` | string | No | Associates this event with a conversation turn (see 5.2.1) | From eb66e64b69b52ffc7cd6424b043e6a7cfceef472 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Wed, 12 Aug 2026 11:11:19 +0100 Subject: [PATCH 26/38] Require data.citation_type on content_cited in the schema Section 5.7.3 lists data.citation_type among the Citation-emitter requirements marked schema-enforced, but the content_cited conditional required only id and output_id - unlike content_reproduced and content_presented, whose conditionals require their data objects. The content_cited conditional now requires data with citation_type, mirroring the reproduced (data.reproduction_type) and presented (data.presentation_kind/presentation_type) pattern. Two cited events carried no data and gained citation_type: "reference" (each is paired with a link presentation of the same source): the 7.1 event-batch example and tests/valid/event-batch-agent.json, plus the same event in tests/invalid/batch-missing-session-and-ctx-token.json so that fixture still passes the schema and fails at the application layer as its description intends. All other cited fixtures and examples already carried citation_type. Co-Authored-By: Claude Fable 5 --- SPECIFICATION.md | 5 ++++- telemetry-session.json | 3 ++- tests/invalid/batch-missing-session-and-ctx-token.json | 5 ++++- tests/valid/event-batch-agent.json | 5 ++++- 4 files changed, 14 insertions(+), 4 deletions(-) diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 81b41f7..b02dd12 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -967,7 +967,10 @@ An event batch carries the same envelope fields with `"document_type": "event_ba "timestamp": "2026-01-15T10:30:04Z", "output_id": "response:1", "source_role": "agent", - "content_url": "https://www.ft.com/content/abc123" + "content_url": "https://www.ft.com/content/abc123", + "data": { + "citation_type": "reference" + } } ] } diff --git a/telemetry-session.json b/telemetry-session.json index e446e41..5d8c641 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -218,13 +218,14 @@ "required": ["type"] }, "then": { - "required": ["id", "output_id"], + "required": ["id", "output_id", "data"], "anyOf": [ { "required": ["content_url"], "properties": { "content_url": { "type": "string" } } }, { "required": ["content_id"], "properties": { "content_id": { "type": "string" } } } ], "properties": { "data": { + "required": ["citation_type"], "properties": { "citation_type": { "$ref": "#/$defs/CitationType" }, "media_type": { "$ref": "#/$defs/MediaType" }, diff --git a/tests/invalid/batch-missing-session-and-ctx-token.json b/tests/invalid/batch-missing-session-and-ctx-token.json index adc810b..6b3322c 100644 --- a/tests/invalid/batch-missing-session-and-ctx-token.json +++ b/tests/invalid/batch-missing-session-and-ctx-token.json @@ -9,7 +9,10 @@ "timestamp": "2026-03-28T10:05:00Z", "output_id": "response:1", "source_role": "agent", - "content_url": "https://example.com/article/test" + "content_url": "https://example.com/article/test", + "data": { + "citation_type": "reference" + } } ] } diff --git a/tests/valid/event-batch-agent.json b/tests/valid/event-batch-agent.json index 7b4deff..3351c5b 100644 --- a/tests/valid/event-batch-agent.json +++ b/tests/valid/event-batch-agent.json @@ -29,7 +29,10 @@ "timestamp": "2026-03-28T08:20:05Z", "output_id": "response:1", "source_role": "agent", - "content_url": "https://arstechnica.com/science/2026/03/new-battery-chemistry-breakthrough/" + "content_url": "https://arstechnica.com/science/2026/03/new-battery-chemistry-breakthrough/", + "data": { + "citation_type": "reference" + } }, { "id": "990e8400-e29b-41d4-a716-446655440063", From 2757d3ae110702e77a6eece895047d9058335000 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Wed, 12 Aug 2026 11:11:55 +0100 Subject: [PATCH 27/38] Fix prose and description drift in 5.7, 6, 6.5, 8.3 and the schema - 8.3: revert 'Presentation name' to 'Display name' for operator.name. The display->presentation rename applies to the presentation event, not ordinary UI terminology; manifest.json already says Display name. - telemetry-session.json: ad_rendered description now says 'rendered', matching the field name and 5.4. - 6.5: restate the excerpt pair on the v1 chars-primary hierarchy - excerpt_chars is the portable primary measurement under 6.4's counting rule, excerpt_tokens the agent-native supplementary one - mirroring 6.4's chars_ingested/tokens_ingested wording so 6.6's 'same pairing' cross-reference holds. - 6 intro: the profiles are no longer in lifecycle order (6.5 Citation precedes 6.6 Reproduction while the lifecycle runs Reproduced then Cited), so the intro no longer claims they are; sections keep their numbers. - 5.7: the optional-signals sentence now names reproduction alongside presentation and engagement - it likewise sits outside the Retrieval/Grounding/Citation ladder as a SHOULD. Co-Authored-By: Claude Fable 5 --- SPECIFICATION.md | 8 ++++---- telemetry-session.json | 2 +- 2 files changed, 5 insertions(+), 5 deletions(-) diff --git a/SPECIFICATION.md b/SPECIFICATION.md index b02dd12..64c2ec5 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -561,7 +561,7 @@ Each level is named for the event it adds: a level proves the emitter produces t | **Grounding** | Above + `content_grounded`, turn events | Content entered the agent's context | Agent with basic instrumentation | | **Citation** | Above + `content_cited` | Content was explicitly referenced in the agent's response | Agent with citation instrumentation | -Presentation and engagement events are optional lifecycle signals. A Citation emitter SHOULD emit them when applicable (section 5.7.3), but the Citation level proves citation support, not full retrieval-to-engagement coverage. +Reproduction, presentation, and engagement events are optional lifecycle signals outside the Retrieval/Grounding/Citation ladder. A Citation emitter SHOULD emit them when applicable (section 5.7.3), but the Citation level proves citation support, not full retrieval-to-engagement coverage. #### 5.7.1 Retrieval conformance @@ -648,7 +648,7 @@ Whether an emitter's reporting in fact met its declared coverage is the complete ## 6. Data profiles -The `data` field on events carries type-specific metadata. These profiles document the recommended fields by event type and source role, in lifecycle order. None are required except where a section states otherwise (`reproduction_type` in 6.6; `presentation_kind` and `presentation_type` in 6.7), but emitting them enables richer attribution. +The `data` field on events carries type-specific metadata. These profiles document the recommended fields by event type and source role. None are required except where a section states otherwise (`reproduction_type` in 6.6; `presentation_kind` and `presentation_type` in 6.7), but emitting them enables richer attribution. ### 6.1 Retrieved content metadata (`content_retrieved`) @@ -816,7 +816,7 @@ A citation MUST carry a resolvable source reference: a non-null `content_url` or `media_type` identifies the content medium. Defaults to `text` when absent. -`excerpt_tokens` is the agent-native measurement. `excerpt_chars` provides the same information in a unit familiar to content owners and licensors. Emitters SHOULD include both when available. +`excerpt_chars` counts Unicode code points in the cited excerpt under the same counting rule as `chars_ingested` (section 6.4): no normalisation applied solely for counting. It is the portable primary measurement, comparable across emitters and stated in a unit familiar to content owners and licensors. `excerpt_tokens` counts the same excerpt in the generation model's tokeniser; it is the agent-native supplementary measurement, carrying the same portability limits as `tokens_ingested`. Emitters SHOULD send `excerpt_chars` where they send `excerpt_tokens`. `excerpt_hash` is the SHA-256 of the excerpt text as it appears in the agent's response - the exact string the agent produced, not the source text it was derived from. For `direct_quote` citations, a matching hash against the source content confirms verbatim fidelity. For `paraphrase` citations, a non-matching hash is expected; verification tooling can use the hash to confirm which specific excerpt was cited and compare it against known source passages. Emitters SHOULD include `excerpt_hash` when `excerpt_tokens` or `excerpt_chars` is present. The hash uses the same `sha256:{hex}` format as `content_hash`. @@ -1162,7 +1162,7 @@ A manifest MAY declare multiple roles (e.g. `["content_owner", "agent"]`). A mor | Field | Type | Required | Description | |-------|------|----------|-------------| -| `name` | string | Yes | Presentation name of the operating organisation. | +| `name` | string | Yes | Display name of the operating organisation. | | `domain` | string | No | Primary domain. Defaults to the manifest URL's host. | ### 8.4 Keys diff --git a/telemetry-session.json b/telemetry-session.json index 5d8c641..5856b9b 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -401,7 +401,7 @@ }, "ad_rendered": { "type": ["boolean", "null"], - "description": "Whether advertising was displayed alongside the response" + "description": "Whether advertising was rendered alongside the response" } } }, From fed18b0322b7cd3eefa5a1583576acd67201b37f Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Wed, 12 Aug 2026 11:19:29 +0100 Subject: [PATCH 28/38] Harden the conformance runners: formats, pinned errors, all document shapes Mutation-verified review findings, all previously undetected: - Build every validator with format_checker so format: uuid / date-time / uri assertions enforce instead of annotate; both runners hard-error at startup if the checker lacks uuid or date-time. Install line becomes pip install "jsonschema[format-nongpl]". - Require _expected_error on every invalid fixture: a substring that must appear in the actual error (first schema error message and JSON pointer, or the application-layer violation text). A fixture that fails for the wrong reason, or carries no pin, now fails the run. - Reconcile APPLICATION_LAYER_VIOLATIONS keys against invalid/: an entry with no matching file fails the run instead of silently dropping the expectation. - Apply privacy field gating (5.5) to turns wherever they appear: session documents, event batches, and standalone event envelopes, not only session events lists. - Add referential integrity checks within a session document (6.6-6.8): content_engaged.presentation_id must match a content_presented event id, and citation_id on content_presented/content_reproduced must match a content_cited event id. Standalone envelopes and batch members are exempt; the corroborating click-out flow is out of scope here. - check_examples.py: validate complete bare event objects (type + timestamp) against the TelemetryEvent definition instead of skipping them; 3 of the 7 skipped fragments are now validated. Co-Authored-By: Claude Fable 5 --- tests/README.md | 6 +- tests/check_examples.py | 53 ++++- tests/invalid/batch-empty-events.json | 1 + .../invalid/batch-missing-schema-version.json | 1 + .../batch-missing-session-and-ctx-token.json | 1 + tests/invalid/cited-missing-output-id.json | 1 + .../cited-missing-source-reference.json | 1 + .../invalid/cited-null-source-reference.json | 1 + .../content-event-missing-identifier.json | 1 + .../engaged-missing-presentation-id.json | 1 + tests/invalid/invalid-event-type.json | 1 + tests/invalid/invalid-privacy-level.json | 1 + tests/invalid/invalid-schema-version.json | 1 + tests/invalid/invalid-source-role.json | 1 + tests/invalid/legacy-content-displayed.json | 1 + .../manifest-bad-conformance-level.json | 1 + tests/invalid/manifest-bad-role.json | 1 + tests/invalid/manifest-duplicate-key-id.json | 1 + tests/invalid/manifest-foreign-domain.json | 1 + .../manifest-missing-key-publickey.json | 1 + tests/invalid/manifest-missing-operator.json | 1 + tests/invalid/missing-event-timestamp.json | 1 + tests/invalid/missing-event-type.json | 1 + tests/invalid/missing-privacy-level.json | 1 + tests/invalid/missing-schema-version.json | 1 + tests/invalid/missing-session-id.json | 1 + tests/invalid/missing-started-at.json | 1 + tests/invalid/presented-missing-kind.json | 1 + .../presented-missing-presentation-type.json | 1 + ...vacy-violation-ad-rendered-at-minimal.json | 1 + .../privacy-violation-query-at-intent.json | 1 + .../privacy-violation-query-at-minimal.json | 1 + tests/invalid/reproduced-missing-id.json | 1 + .../invalid/reproduced-missing-output-id.json | 1 + .../reproduced-missing-source-reference.json | 1 + tests/invalid/reproduced-missing-type.json | 1 + .../standalone-missing-document-type.json | 1 + tests/invalid/standalone-missing-event.json | 1 + ...ndalone-missing-session-and-ctx-token.json | 1 + tests/invalid/withdrawn-ip-hash.json | 1 + tests/validate.py | 181 +++++++++++++++--- 41 files changed, 242 insertions(+), 36 deletions(-) diff --git a/tests/README.md b/tests/README.md index 4ca1a6e..9c5fef8 100644 --- a/tests/README.md +++ b/tests/README.md @@ -14,11 +14,11 @@ Tests for the Content Telemetry Specification v1. From a clean checkout, with no setup beyond [uv](https://docs.astral.sh/uv/): ```sh -uv run --with jsonschema python tests/validate.py -uv run --with jsonschema python tests/check_examples.py +uv run --with "jsonschema[format-nongpl]" python tests/validate.py +uv run --with "jsonschema[format-nongpl]" python tests/check_examples.py ``` -Run from the repository root. Without uv: `pip install jsonschema`, then `python3 tests/validate.py`. Both commands run in CI on every pull request. +Run from the repository root. Without uv: `pip install "jsonschema[format-nongpl]"`, then `python3 tests/validate.py`. The `format-nongpl` extra pulls in the format validators (`rfc3339-validator` and friends) that make `format: uuid` / `date-time` / `uri` assertions enforce rather than annotate; both scripts hard-error at startup if they are missing. Both commands run in CI on every pull request. `check_examples.py` extracts every fenced `json` block from the spec and README, validates the complete top-level documents (sessions, standalone events, manifests) against the matching schema, and reports the number of fragments it skipped. A worked example that no longer matches its schema fails the build. diff --git a/tests/check_examples.py b/tests/check_examples.py index 7870f4d..9c04fca 100644 --- a/tests/check_examples.py +++ b/tests/check_examples.py @@ -11,13 +11,15 @@ event batch -> telemetry-event-batch.json manifest -> manifest.json -Fragments (a bare event object, a single turn, a one-field snippet) are not -top-level documents and cannot be validated against a top-level schema. They are -counted and listed rather than validated. +A complete bare event object (carrying the required `type` and `timestamp`) is +validated against the TelemetryEvent definition in telemetry-session.json. +Genuine fragments (a single turn, a one-field snippet, an event elided below +its required fields) are not validatable and are counted and listed rather +than validated. Usage: - uv run --with jsonschema python tests/check_examples.py - # or: pip install jsonschema && python tests/check_examples.py + uv run --with "jsonschema[format-nongpl]" python tests/check_examples.py + # or: pip install "jsonschema[format-nongpl]" && python tests/check_examples.py """ import json @@ -29,9 +31,27 @@ from jsonschema import Draft202012Validator from referencing import Registry, Resource except ImportError: - print("ERROR: jsonschema package required. Install with: pip install jsonschema") + print('ERROR: jsonschema package required. Install with: pip install "jsonschema[format-nongpl]"') sys.exit(1) +# Format assertions (uuid, date-time, uri) are annotation-only unless a +# FormatChecker is attached to the validator. Guard at startup so a missing +# optional dependency (rfc3339-validator) hard-errors instead of silently +# downgrading every format assertion to a no-op. +FORMAT_CHECKER = Draft202012Validator.FORMAT_CHECKER +_missing_formats = {"uuid", "date-time"} - set(FORMAT_CHECKER.checkers) +if _missing_formats: + print( + "ERROR: format checker cannot enforce " + + ", ".join(sorted(_missing_formats)) + + '. Install with: pip install "jsonschema[format-nongpl]"' + ) + sys.exit(1) + +# Pseudo-schema name for complete bare event objects in the prose, validated +# against the TelemetryEvent definition rather than a top-level document schema. +EVENT_DEF = "telemetry-session.json#/$defs/TelemetryEvent" + REPO = Path(__file__).resolve().parent.parent SOURCES = [REPO / "SPECIFICATION.md", REPO / "README.md"] FENCE = re.compile(r"```json\n(.*?)\n```", re.DOTALL) @@ -52,10 +72,21 @@ def load_validators(): registry = registry.with_resource( schema.get("$id", name), Resource.from_contents(schema) ) - return { - name: Draft202012Validator(schema, registry=registry) + validators = { + name: Draft202012Validator( + schema, registry=registry, format_checker=FORMAT_CHECKER + ) for name, schema in schemas.items() } + # Bare event objects validate against the TelemetryEvent definition, + # referenced through the session schema's $id so $refs resolve. + session_id = schemas["telemetry-session.json"].get("$id", "") + validators[EVENT_DEF] = Draft202012Validator( + {"$ref": f"{session_id}#/$defs/TelemetryEvent"}, + registry=registry, + format_checker=FORMAT_CHECKER, + ) + return validators def strip_comments(block): @@ -67,7 +98,7 @@ def strip_comments(block): def classify(doc): - """Return the schema a complete document validates against, or None for a fragment.""" + """Return the schema a complete example validates against, or None for a fragment.""" if not isinstance(doc, dict): return None if doc.get("document_type") == "session": @@ -80,6 +111,10 @@ def classify(doc): return "manifest.json" if "session_id" in doc and "started_at" in doc and "events" in doc: return "telemetry-session.json" + if {"type", "timestamp"} <= doc.keys(): + # A bare event object carrying the definition's required fields is a + # complete event, validatable against the TelemetryEvent definition. + return EVENT_DEF return None diff --git a/tests/invalid/batch-empty-events.json b/tests/invalid/batch-empty-events.json index b11cdd5..38a17e4 100644 --- a/tests/invalid/batch-empty-events.json +++ b/tests/invalid/batch-empty-events.json @@ -1,5 +1,6 @@ { "_test_description": "Event batch envelope with an empty events array. The schema requires minItems 1: an empty batch carries no signal and MUST NOT be delivered.", + "_expected_error": "should be non-empty", "document_type": "event_batch", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440008", diff --git a/tests/invalid/batch-missing-schema-version.json b/tests/invalid/batch-missing-schema-version.json index 0b872db..094a850 100644 --- a/tests/invalid/batch-missing-schema-version.json +++ b/tests/invalid/batch-missing-schema-version.json @@ -1,5 +1,6 @@ { "_test_description": "Event batch envelope missing schema_version. The event_batch document_type triggers batch envelope validation, which requires schema_version.", + "_expected_error": "'schema_version' is a required property", "document_type": "event_batch", "session_id": "660e8400-e29b-41d4-a716-446655440009", "events": [ diff --git a/tests/invalid/batch-missing-session-and-ctx-token.json b/tests/invalid/batch-missing-session-and-ctx-token.json index 6b3322c..0c60b89 100644 --- a/tests/invalid/batch-missing-session-and-ctx-token.json +++ b/tests/invalid/batch-missing-session-and-ctx-token.json @@ -1,5 +1,6 @@ { "_test_description": "Event batch envelope carrying a content_cited event with neither session_id nor ctx_token on the envelope. Passes JSON Schema (both fields are optional) but violates section 7.1: an event MUST carry one at Grounding conformance and above.", + "_expected_error": "neither session_id nor ctx_token", "document_type": "event_batch", "schema_version": "1.0", "events": [ diff --git a/tests/invalid/cited-missing-output-id.json b/tests/invalid/cited-missing-output-id.json index 5f7686b..32a3e0c 100644 --- a/tests/invalid/cited-missing-output-id.json +++ b/tests/invalid/cited-missing-output-id.json @@ -1,5 +1,6 @@ { "_test_description": "content_cited event with id and a resolvable source reference but no output_id. output_id is required on citation events so output construction can be correlated with later presentation.", + "_expected_error": "'output_id' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440160", "started_at": "2026-08-01T14:00:00Z", diff --git a/tests/invalid/cited-missing-source-reference.json b/tests/invalid/cited-missing-source-reference.json index d619c37..26e21cf 100644 --- a/tests/invalid/cited-missing-source-reference.json +++ b/tests/invalid/cited-missing-source-reference.json @@ -1,5 +1,6 @@ { "_test_description": "content_cited event carrying neither content_url nor content_id. A source association with no resolvable reference is not a citation; unlike other content events, the JSON Schema enforces the identifier requirement for content_cited (section 6.5).", + "_expected_error": "'content_url' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440140", "agent_id": "research-assistant-v1", diff --git a/tests/invalid/cited-null-source-reference.json b/tests/invalid/cited-null-source-reference.json index 754d1c3..92df5f1 100644 --- a/tests/invalid/cited-null-source-reference.json +++ b/tests/invalid/cited-null-source-reference.json @@ -1,5 +1,6 @@ { "_test_description": "content_cited event with content_url explicitly null and no content_id. Presence of a null reference does not satisfy the citation reference requirement: the schema demands a non-null content_url or content_id on content_cited (section 6.5).", + "_expected_error": "is not of type 'string'", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440141", "agent_id": "research-assistant-v1", diff --git a/tests/invalid/content-event-missing-identifier.json b/tests/invalid/content-event-missing-identifier.json index db8507f..2ff18a2 100644 --- a/tests/invalid/content-event-missing-identifier.json +++ b/tests/invalid/content-event-missing-identifier.json @@ -1,5 +1,6 @@ { "_test_description": "content_grounded event carrying neither content_url nor content_id. Passes JSON Schema (both are individually optional) but violates section 5.7.5: at least one of content_url or content_id MUST be present on every content event.", + "_expected_error": "'content_grounded' carries neither content_url nor content_id", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440099", "agent_id": "research-assistant-v1", diff --git a/tests/invalid/engaged-missing-presentation-id.json b/tests/invalid/engaged-missing-presentation-id.json index a80e22a..0e35747 100644 --- a/tests/invalid/engaged-missing-presentation-id.json +++ b/tests/invalid/engaged-missing-presentation-id.json @@ -1,5 +1,6 @@ { "_test_description": "content_engaged must identify the exact presentation occurrence rather than matching only by URL.", + "_expected_error": "'presentation_id' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440092", "started_at": "2026-07-18T13:00:00Z", diff --git a/tests/invalid/invalid-event-type.json b/tests/invalid/invalid-event-type.json index 589ec33..9948b43 100644 --- a/tests/invalid/invalid-event-type.json +++ b/tests/invalid/invalid-event-type.json @@ -1,5 +1,6 @@ { "_test_description": "Event with type 'content_summarised' which is not in the EventType enum. Fails JSON Schema validation.", + "_expected_error": "'content_summarised' is not one of", "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440008", "started_at": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/invalid-privacy-level.json b/tests/invalid/invalid-privacy-level.json index a4c762d..c0e827c 100644 --- a/tests/invalid/invalid-privacy-level.json +++ b/tests/invalid/invalid-privacy-level.json @@ -1,5 +1,6 @@ { "_test_description": "Turn with privacy_level 'redacted' which is not in the PrivacyLevel enum. Fails JSON Schema validation.", + "_expected_error": "'redacted' is not one of", "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440009", "started_at": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/invalid-schema-version.json b/tests/invalid/invalid-schema-version.json index 2099ef8..e851c2c 100644 --- a/tests/invalid/invalid-schema-version.json +++ b/tests/invalid/invalid-schema-version.json @@ -1,5 +1,6 @@ { "_test_description": "schema_version set to '2.0' which does not match the const '1.0'. Fails JSON Schema validation.", + "_expected_error": "'1.0' was expected", "schema_version": "2.0", "session_id": "770e8400-e29b-41d4-a716-446655440010", "started_at": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/invalid-source-role.json b/tests/invalid/invalid-source-role.json index 81464df..57be3e5 100644 --- a/tests/invalid/invalid-source-role.json +++ b/tests/invalid/invalid-source-role.json @@ -1,5 +1,6 @@ { "_test_description": "Event with source_role 'cdn' which is not in the SourceRole enum. Fails JSON Schema validation.", + "_expected_error": "'cdn' is not one of", "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440014", "started_at": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/legacy-content-displayed.json b/tests/invalid/legacy-content-displayed.json index 401644e..59c6ead 100644 --- a/tests/invalid/legacy-content-displayed.json +++ b/tests/invalid/legacy-content-displayed.json @@ -1,5 +1,6 @@ { "_test_description": "The v1 presentation event replaces the preview content_displayed name.", + "_expected_error": "'content_displayed' is not one of", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440093", "started_at": "2026-07-18T13:00:00Z", diff --git a/tests/invalid/manifest-bad-conformance-level.json b/tests/invalid/manifest-bad-conformance-level.json index ec22c79..5afadc1 100644 --- a/tests/invalid/manifest-bad-conformance-level.json +++ b/tests/invalid/manifest-bad-conformance-level.json @@ -1,5 +1,6 @@ { "_test_description": "Manifest with telemetry.conformance_level 'attribution' which is not one of retrieval, grounding, citation. Fails JSON Schema validation.", + "_expected_error": "'attribution' is not one of", "schema_version": "1.0", "id": "https://searchco.com/agents/web-search/.well-known/content-telemetry.json", "roles": ["agent"], diff --git a/tests/invalid/manifest-bad-role.json b/tests/invalid/manifest-bad-role.json index c1c20c4..646d362 100644 --- a/tests/invalid/manifest-bad-role.json +++ b/tests/invalid/manifest-bad-role.json @@ -1,5 +1,6 @@ { "_test_description": "Manifest with a role value 'publisher' that is not in the roles enum (content_owner, agent, platform). Fails JSON Schema validation.", + "_expected_error": "'publisher' is not one of", "schema_version": "1.0", "id": "https://example.com/.well-known/content-telemetry.json", "roles": ["publisher"], diff --git a/tests/invalid/manifest-duplicate-key-id.json b/tests/invalid/manifest-duplicate-key-id.json index ce403f9..2040a5e 100644 --- a/tests/invalid/manifest-duplicate-key-id.json +++ b/tests/invalid/manifest-duplicate-key-id.json @@ -1,5 +1,6 @@ { "_test_description": "Agent manifest with two keys sharing the id 'key-1'. Passes JSON Schema (the entries differ in publicKey) but violates section 8.7: consumers reject a manifest with duplicate keys[].id.", + "_expected_error": "Duplicate keys[].id", "schema_version": "1.0", "id": "https://searchco.com/agents/web-search/.well-known/content-telemetry.json", "roles": ["agent"], diff --git a/tests/invalid/manifest-foreign-domain.json b/tests/invalid/manifest-foreign-domain.json index ff47db1..db2c2c3 100644 --- a/tests/invalid/manifest-foreign-domain.json +++ b/tests/invalid/manifest-foreign-domain.json @@ -1,5 +1,6 @@ { "_test_description": "Content-owner manifest at example.com whose domains array claims othersite.com. Passes JSON Schema but violates section 8.6: every domains entry MUST be the manifest's own host or a subdomain of it. Consumers reject the manifest as malformed (section 8.7).", + "_expected_error": "'othersite.com' is not the manifest host", "schema_version": "1.0", "id": "https://example.com/.well-known/content-telemetry.json", "roles": ["content_owner"], diff --git a/tests/invalid/manifest-missing-key-publickey.json b/tests/invalid/manifest-missing-key-publickey.json index 1cb5de6..1c49857 100644 --- a/tests/invalid/manifest-missing-key-publickey.json +++ b/tests/invalid/manifest-missing-key-publickey.json @@ -1,5 +1,6 @@ { "_test_description": "Manifest with a keys entry missing the required 'publicKey' field. Fails JSON Schema validation.", + "_expected_error": "'publicKey' is a required property", "schema_version": "1.0", "id": "https://searchco.com/agents/web-search/.well-known/content-telemetry.json", "roles": ["agent"], diff --git a/tests/invalid/manifest-missing-operator.json b/tests/invalid/manifest-missing-operator.json index b3b9031..729f251 100644 --- a/tests/invalid/manifest-missing-operator.json +++ b/tests/invalid/manifest-missing-operator.json @@ -1,5 +1,6 @@ { "_test_description": "Manifest missing the required 'operator' field. Fails JSON Schema validation.", + "_expected_error": "'operator' is a required property", "schema_version": "1.0", "id": "https://example.com/.well-known/content-telemetry.json", "roles": ["content_owner"] diff --git a/tests/invalid/missing-event-timestamp.json b/tests/invalid/missing-event-timestamp.json index 45ba33a..64437ad 100644 --- a/tests/invalid/missing-event-timestamp.json +++ b/tests/invalid/missing-event-timestamp.json @@ -1,5 +1,6 @@ { "_test_description": "Event without timestamp field. Fails JSON Schema validation: timestamp is required on TelemetryEvent.", + "_expected_error": "'timestamp' is a required property", "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440005", "started_at": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/missing-event-type.json b/tests/invalid/missing-event-type.json index fb5efad..07a1c5d 100644 --- a/tests/invalid/missing-event-type.json +++ b/tests/invalid/missing-event-type.json @@ -1,5 +1,6 @@ { "_test_description": "Event without type field. Fails JSON Schema validation: type is required on TelemetryEvent.", + "_expected_error": "'type' is a required property", "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440004", "started_at": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/missing-privacy-level.json b/tests/invalid/missing-privacy-level.json index 1184e23..fc38f4d 100644 --- a/tests/invalid/missing-privacy-level.json +++ b/tests/invalid/missing-privacy-level.json @@ -1,5 +1,6 @@ { "_test_description": "Turn object without privacy_level. Fails JSON Schema validation: privacy_level is required on ConversationTurn.", + "_expected_error": "'privacy_level' is a required property", "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440006", "started_at": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/missing-schema-version.json b/tests/invalid/missing-schema-version.json index 0294251..d275e03 100644 --- a/tests/invalid/missing-schema-version.json +++ b/tests/invalid/missing-schema-version.json @@ -1,5 +1,6 @@ { "_test_description": "Session without schema_version. Fails JSON Schema validation: schema_version is a required field.", + "_expected_error": "'schema_version' is a required property", "session_id": "770e8400-e29b-41d4-a716-446655440001", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/missing-session-id.json b/tests/invalid/missing-session-id.json index d43a75a..20e400a 100644 --- a/tests/invalid/missing-session-id.json +++ b/tests/invalid/missing-session-id.json @@ -1,5 +1,6 @@ { "_test_description": "Session without session_id. Fails JSON Schema validation: session_id is a required field.", + "_expected_error": "'session_id' is a required property", "schema_version": "1.0", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/missing-started-at.json b/tests/invalid/missing-started-at.json index ee2fd38..54f9f63 100644 --- a/tests/invalid/missing-started-at.json +++ b/tests/invalid/missing-started-at.json @@ -1,5 +1,6 @@ { "_test_description": "Session without started_at. Fails JSON Schema validation: started_at is a required field.", + "_expected_error": "'started_at' is a required property", "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440003", "events": [ diff --git a/tests/invalid/presented-missing-kind.json b/tests/invalid/presented-missing-kind.json index 1e298d1..69b4fdf 100644 --- a/tests/invalid/presented-missing-kind.json +++ b/tests/invalid/presented-missing-kind.json @@ -1,5 +1,6 @@ { "_test_description": "content_presented must distinguish source content from a source reference.", + "_expected_error": "'presentation_kind' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440090", "started_at": "2026-07-18T13:00:00Z", diff --git a/tests/invalid/presented-missing-presentation-type.json b/tests/invalid/presented-missing-presentation-type.json index 3ccf305..43858c8 100644 --- a/tests/invalid/presented-missing-presentation-type.json +++ b/tests/invalid/presented-missing-presentation-type.json @@ -1,5 +1,6 @@ { "_test_description": "content_presented event carrying presentation_kind but no presentation_type. Both are required: the kind says what crossed the presentation boundary, the type says how it was made perceivable.", + "_expected_error": "'presentation_type' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440163", "started_at": "2026-08-01T15:00:00Z", diff --git a/tests/invalid/privacy-violation-ad-rendered-at-minimal.json b/tests/invalid/privacy-violation-ad-rendered-at-minimal.json index 7cbd248..c4cf0c1 100644 --- a/tests/invalid/privacy-violation-ad-rendered-at-minimal.json +++ b/tests/invalid/privacy-violation-ad-rendered-at-minimal.json @@ -1,5 +1,6 @@ { "_test_description": "APPLICATION-LAYER VIOLATION: Turn at minimal privacy with ad_rendered present. Passes JSON Schema validation but violates section 5.5: platform metadata (including ad_rendered) is not available at minimal level.", + "_expected_error": "'ad_rendered' present on turn with privacy_level 'minimal'", "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440012", "started_at": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/privacy-violation-query-at-intent.json b/tests/invalid/privacy-violation-query-at-intent.json index 8e033c9..6428d54 100644 --- a/tests/invalid/privacy-violation-query-at-intent.json +++ b/tests/invalid/privacy-violation-query-at-intent.json @@ -1,5 +1,6 @@ { "_test_description": "APPLICATION-LAYER VIOLATION: Turn at intent privacy with query_text present. Passes JSON Schema validation but violates section 5.5: query_text MUST NOT be present at intent level.", + "_expected_error": "'query_text' present on turn with privacy_level 'intent'", "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440013", "started_at": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/privacy-violation-query-at-minimal.json b/tests/invalid/privacy-violation-query-at-minimal.json index 31e3bd3..c0efce5 100644 --- a/tests/invalid/privacy-violation-query-at-minimal.json +++ b/tests/invalid/privacy-violation-query-at-minimal.json @@ -1,5 +1,6 @@ { "_test_description": "APPLICATION-LAYER VIOLATION: Turn at minimal privacy with query_text present. Passes JSON Schema validation but violates the privacy field gating rule (section 5.5): query_text MUST NOT be present when privacy_level is minimal.", + "_expected_error": "'query_text' present on turn with privacy_level 'minimal'", "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440011", "started_at": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/reproduced-missing-id.json b/tests/invalid/reproduced-missing-id.json index 6dad3fb..cff1bc9 100644 --- a/tests/invalid/reproduced-missing-id.json +++ b/tests/invalid/reproduced-missing-id.json @@ -1,5 +1,6 @@ { "_test_description": "content_reproduced event with output_id and a resolvable source reference but no event id. id is required on reproduction events so a crediting citation or verification result can reference the exact reproduction claim.", + "_expected_error": "'id' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440162", "started_at": "2026-08-01T14:30:00Z", diff --git a/tests/invalid/reproduced-missing-output-id.json b/tests/invalid/reproduced-missing-output-id.json index 8140942..7a2a067 100644 --- a/tests/invalid/reproduced-missing-output-id.json +++ b/tests/invalid/reproduced-missing-output-id.json @@ -1,5 +1,6 @@ { "_test_description": "content_reproduced must identify the output artifact containing the reproduction: id and output_id are required so reproduction can be correlated with citation and presentation.", + "_expected_error": "'output_id' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440130", "started_at": "2026-08-01T11:00:00Z", diff --git a/tests/invalid/reproduced-missing-source-reference.json b/tests/invalid/reproduced-missing-source-reference.json index a9d5a1b..c75a3d7 100644 --- a/tests/invalid/reproduced-missing-source-reference.json +++ b/tests/invalid/reproduced-missing-source-reference.json @@ -1,5 +1,6 @@ { "_test_description": "content_reproduced event carrying neither content_url nor content_id. A reproduction claim is only meaningful for an identified source; the JSON Schema requires a non-null content_url or content_id on content_reproduced (section 6.6).", + "_expected_error": "'content_url' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440150", "started_at": "2026-08-01T13:00:00Z", diff --git a/tests/invalid/reproduced-missing-type.json b/tests/invalid/reproduced-missing-type.json index 3fae0ee..25e8e67 100644 --- a/tests/invalid/reproduced-missing-type.json +++ b/tests/invalid/reproduced-missing-type.json @@ -1,5 +1,6 @@ { "_test_description": "content_reproduced must classify the fidelity of the reproduction: data.reproduction_type is required.", + "_expected_error": "'reproduction_type' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440120", "started_at": "2026-08-01T10:00:00Z", diff --git a/tests/invalid/standalone-missing-document-type.json b/tests/invalid/standalone-missing-document-type.json index 2731dba..497f55a 100644 --- a/tests/invalid/standalone-missing-document-type.json +++ b/tests/invalid/standalone-missing-document-type.json @@ -1,5 +1,6 @@ { "_test_description": "Envelope-shaped document with an event key but no document_type. Per section 7.1 a document without document_type is treated as a session, and it fails the session schema.", + "_expected_error": "'session_id' is a required property", "schema_version": "1.0", "event": { "type": "content_retrieved", diff --git a/tests/invalid/standalone-missing-event.json b/tests/invalid/standalone-missing-event.json index bb91d7b..befc66e 100644 --- a/tests/invalid/standalone-missing-event.json +++ b/tests/invalid/standalone-missing-event.json @@ -1,5 +1,6 @@ { "_test_description": "Standalone event envelope (document_type 'event') with no event field. Fails the envelope schema, which requires event.", + "_expected_error": "'event' is a required property", "document_type": "event", "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440010" diff --git a/tests/invalid/standalone-missing-session-and-ctx-token.json b/tests/invalid/standalone-missing-session-and-ctx-token.json index c96811b..a4016e7 100644 --- a/tests/invalid/standalone-missing-session-and-ctx-token.json +++ b/tests/invalid/standalone-missing-session-and-ctx-token.json @@ -1,5 +1,6 @@ { "_test_description": "Standalone event envelope carrying neither session_id nor ctx_token. Passes JSON Schema (both are individually optional on the envelope) but violates section 5.7.5: an event MUST carry either session_id or ctx_token at Grounding conformance and above (section 7.1).", + "_expected_error": "neither session_id nor ctx_token", "document_type": "event", "schema_version": "1.0", "event": { diff --git a/tests/invalid/withdrawn-ip-hash.json b/tests/invalid/withdrawn-ip-hash.json index f691b79..c634676 100644 --- a/tests/invalid/withdrawn-ip-hash.json +++ b/tests/invalid/withdrawn-ip-hash.json @@ -1,5 +1,6 @@ { "_test_description": "Standalone event envelope from a CDN carrying ip_hash in data. The field was withdrawn in v1 (section 9.1): hashing does not anonymise a value drawn from a space small enough to enumerate. Passes JSON Schema because event data accepts additional properties, so the rule is enforced at the application layer.", + "_expected_error": "'ip_hash' in data", "document_type": "event", "schema_version": "1.0", "event": { diff --git a/tests/validate.py b/tests/validate.py index a2cea17..ef0b5b9 100644 --- a/tests/validate.py +++ b/tests/validate.py @@ -11,7 +11,7 @@ document. Usage: - pip install jsonschema + pip install "jsonschema[format-nongpl]" python validate.py """ @@ -25,7 +25,22 @@ from jsonschema import Draft202012Validator, ValidationError from referencing import Registry, Resource except ImportError: - print("ERROR: jsonschema package required. Install with: pip install jsonschema") + print('ERROR: jsonschema package required. Install with: pip install "jsonschema[format-nongpl]"') + sys.exit(1) + +# Format assertions (uuid, date-time, uri) are annotation-only unless a +# FormatChecker is attached to the validator. Every validator constructed in +# this suite MUST pass format_checker=FORMAT_CHECKER; guard at startup so a +# missing optional dependency (rfc3339-validator) hard-errors instead of +# silently downgrading every format assertion to a no-op. +FORMAT_CHECKER = Draft202012Validator.FORMAT_CHECKER +_missing_formats = {"uuid", "date-time"} - set(FORMAT_CHECKER.checkers) +if _missing_formats: + print( + "ERROR: format checker cannot enforce " + + ", ".join(sorted(_missing_formats)) + + '. Install with: pip install "jsonschema[format-nongpl]"' + ) sys.exit(1) @@ -60,6 +75,13 @@ # content_fingerprint MUST NOT carry preserved_in_output because output-side # reuse is represented by content_reproduced. # +# 6. Referential integrity within a session document (sections 6.6-6.8): +# content_engaged.presentation_id references the exact content_presented +# event id, and citation_id on content_presented/content_reproduced +# references a content_cited event id. JSON Schema cannot compare values +# across events. Session documents only: standalone envelopes and batch +# members may reference events delivered elsewhere. +# # Not checked here: agent_id at Grounding/Citation conformance (section # 5.7) depends on the emitter's declared conformance level, which fixtures do # not carry, so it is out of scope for the fixture suite. @@ -188,11 +210,11 @@ def load_schema(schema_path): manifest_schema_id = manifest_schema.get("$id", "") manifest_resource = Resource.from_contents(manifest_schema) registry = registry.with_resource(manifest_schema_id, manifest_resource) - manifest_validator = Draft202012Validator(manifest_schema, registry=registry) + manifest_validator = Draft202012Validator(manifest_schema, registry=registry, format_checker=FORMAT_CHECKER) else: manifest_validator = None - validator = Draft202012Validator(schema, registry=registry) + validator = Draft202012Validator(schema, registry=registry, format_checker=FORMAT_CHECKER) return schema, event_schema, batch_schema, validator, manifest_validator, registry @@ -230,12 +252,12 @@ def validate_standalone_event(data, session_schema, event_schema, registry): just the event body against the TelemetryEvent definition. """ if event_schema is not None: - validator = Draft202012Validator(event_schema, registry=registry) + validator = Draft202012Validator(event_schema, registry=registry, format_checker=FORMAT_CHECKER) errors = list(validator.iter_errors(data)) else: schema_id = session_schema.get("$id", "") wrapper = {"$ref": f"{schema_id}#/$defs/TelemetryEvent"} - validator = Draft202012Validator(wrapper, registry=registry) + validator = Draft202012Validator(wrapper, registry=registry, format_checker=FORMAT_CHECKER) errors = list(validator.iter_errors(data["event"])) return errors @@ -248,11 +270,11 @@ def validate_event_batch(data, session_schema, batch_schema, registry): event body against the TelemetryEvent definition. """ if batch_schema is not None: - validator = Draft202012Validator(batch_schema, registry=registry) + validator = Draft202012Validator(batch_schema, registry=registry, format_checker=FORMAT_CHECKER) return list(validator.iter_errors(data)) schema_id = session_schema.get("$id", "") wrapper = {"$ref": f"{schema_id}#/$defs/TelemetryEvent"} - validator = Draft202012Validator(wrapper, registry=registry) + validator = Draft202012Validator(wrapper, registry=registry, format_checker=FORMAT_CHECKER) errors = [] for event in data.get("events", []): errors.extend(validator.iter_errors(event)) @@ -263,12 +285,15 @@ def check_privacy_conformance(data): """ Check application-layer privacy conformance rules. + The privacy field gating of section 5.5 is a property of privacy_level + itself: it applies wherever conversation turns are emitted - session + documents, event batches, and standalone event envelopes alike. + Returns a list of violation descriptions, empty if conforming. """ violations = [] - events = data.get("events", []) - for event in events: + for event in _iter_events(data): turn = event.get("turn") if turn is None: continue @@ -288,8 +313,9 @@ def check_privacy_conformance(data): def _iter_events(data): - """Yield the content/turn events in a document, whether it is a session - (events list) or a standalone envelope (single event under 'event').""" + """Yield the content/turn events in a document, whatever its shape: a + session or event batch (events list) or a standalone envelope (single + event under 'event').""" if is_standalone_event(data): event = data.get("event") if isinstance(event, dict): @@ -338,6 +364,52 @@ def check_session_or_ctx_token(data): return [] +def check_referential_integrity(data): + """ + Check the intra-document event references of a session document: + + - Every content_engaged.presentation_id MUST reference the exact + content_presented.id on which the action occurred (section 6.8). + - Every citation_id on a content_presented or content_reproduced event + references that content_cited event's id (sections 6.6, 6.7). + + Applies only to session documents, where the referenced events live in + the same document. Standalone envelopes and batch members legitimately + reference events delivered elsewhere (e.g. a click-out engagement carrying + a ctx_token), so they are exempt here; the corroborating click-out flow + is out of scope for this suite. Returns a list of violations. + """ + if is_standalone_event(data) or is_event_batch(data): + return [] + events = data.get("events", []) + presented_ids = { + e.get("id") for e in events + if e.get("type") == "content_presented" and e.get("id") + } + cited_ids = { + e.get("id") for e in events + if e.get("type") == "content_cited" and e.get("id") + } + violations = [] + for event in events: + etype = event.get("type") + if etype == "content_engaged": + pid = event.get("presentation_id") + if pid and pid not in presented_ids: + violations.append( + f"content_engaged presentation_id '{pid}' does not match " + "any content_presented event id in the session" + ) + if etype in ("content_presented", "content_reproduced"): + cid = event.get("citation_id") + if cid and cid not in cited_ids: + violations.append( + f"{etype} citation_id '{cid}' does not match any " + "content_cited event id in the session" + ) + return violations + + def check_v1_migration_prohibitions(data): """ Check v1's explicit prohibitions on fields carried forward from v0.1. @@ -385,6 +457,7 @@ def check_application_layer(data): check_privacy_conformance(data) + check_content_identifier(data) + check_session_or_ctx_token(data) + + check_referential_integrity(data) + check_v1_migration_prohibitions(data) + check_grounding_provenance(data) ) @@ -419,6 +492,30 @@ def check_manifest_application_layer(data): return violations +def schema_error_haystack(errors): + """ + Render the first schema error (deterministically chosen) as searchable + text: its JSON pointer plus its message, recursing into sub-errors of + combinators like anyOf. Invalid fixtures pin their intended violation by + requiring an _expected_error substring to appear in this text. + """ + first = sorted( + errors, + key=lambda e: ([str(p) for p in e.absolute_path], e.message), + )[0] + parts = [] + + def walk(error): + pointer = "/" + "/".join(str(p) for p in error.absolute_path) + parts.append(pointer) + parts.append(error.message) + for sub in error.context or []: + walk(sub) + + walk(first) + return " ".join(parts) + + def run_tests(): """Run all conformance tests and return (passed, failed, results).""" tests_dir = Path(__file__).parent @@ -495,9 +592,21 @@ def run_tests(): data = load_test_file(path) name = path.name desc = data.get("_test_description", "") + expected_error = data.get("_expected_error") is_app_layer = name in APPLICATION_LAYER_VIOLATIONS + # Every invalid fixture must pin its intended violation: a substring + # that must appear in the actual error. Without it, a fixture that + # fails for the wrong reason (e.g. after an unrelated edit) would + # still count as a pass. + if not isinstance(expected_error, str) or not expected_error: + print(f" FAIL {name}") + print(" Fixture missing required _expected_error field") + failed += 1 + results.append((name, False, "missing _expected_error")) + continue + if is_manifest_fixture(path): if manifest_validator is None: print(f" FAIL {name}") @@ -514,12 +623,19 @@ def run_tests(): schema_errors = list(session_validator.iter_errors(data)) if schema_errors: - # Failed JSON Schema - good - msg = schema_errors[0].message - print(f" PASS {name}") - print(f" Schema error: {msg}") - passed += 1 - results.append((name, True, None)) + # Failed JSON Schema - but only for the pinned reason + haystack = schema_error_haystack(schema_errors) + if expected_error in haystack: + print(f" PASS {name}") + print(f" Schema error: {schema_errors[0].message}") + passed += 1 + results.append((name, True, None)) + else: + print(f" FAIL {name}") + print(f" Schema error does not match _expected_error {expected_error!r}") + print(f" Actual: {haystack[:200]}") + failed += 1 + results.append((name, False, "wrong schema error")) elif is_app_layer: # Passes JSON Schema but should fail conformance @@ -528,16 +644,23 @@ def run_tests(): if is_manifest_fixture(path) else check_application_layer(data) ) - if conformance_violations: - print(f" PASS {name} [application-layer]") - print(f" {APPLICATION_LAYER_VIOLATIONS[name]}") - passed += 1 - results.append((name, True, None)) - else: + haystack = "; ".join(conformance_violations) + if not conformance_violations: print(f" FAIL {name}") print(f" Expected application-layer violation but none found") failed += 1 results.append((name, False, "Expected conformance violation")) + elif expected_error not in haystack: + print(f" FAIL {name}") + print(f" Violation does not match _expected_error {expected_error!r}") + print(f" Actual: {haystack[:200]}") + failed += 1 + results.append((name, False, "wrong conformance violation")) + else: + print(f" PASS {name} [application-layer]") + print(f" {APPLICATION_LAYER_VIOLATIONS[name]}") + passed += 1 + results.append((name, True, None)) else: # Should have failed schema but didn't @@ -546,6 +669,16 @@ def run_tests(): failed += 1 results.append((name, False, "Expected schema error")) + # --- Reconcile APPLICATION_LAYER_VIOLATIONS against invalid/ --- + # A dict key with no matching fixture file means an expectation silently + # dropped out of the suite (e.g. a renamed fixture). Fail the run. + invalid_names = {p.name for p in invalid_dir.glob("*.json")} + for name in sorted(set(APPLICATION_LAYER_VIOLATIONS) - invalid_names): + print(f" FAIL {name}") + print(" APPLICATION_LAYER_VIOLATIONS entry has no matching file in invalid/") + failed += 1 + results.append((name, False, "orphaned APPLICATION_LAYER_VIOLATIONS entry")) + # --- Summary --- total = passed + failed print() From 25a13c8f70aaacb4a2c8d62ceea20e53b55714d2 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Wed, 12 Aug 2026 11:22:16 +0100 Subject: [PATCH 29/38] Add fixtures pinning previously unpinned MUSTs New invalid fixtures, each pinned by _expected_error: - presented-missing-id, presented-missing-output-id, cited-missing-id: the required event and output identifiers of sections 6.5 and 6.7. - cited-missing-citation-type: the schema now requires data.citation_type on content_cited. - reproduction-type-invalid: reproduction_type 'paraphrase' - section 6.6 says paraphrase is not reproduction. - presentation-kind-invalid: presentation_kind is a closed two-value enum. - grounded-negative-chars-ingested, reproduced-negative-chars: count fields carry minimum 0. - malformed-parent-session-id: format: uuid, caught only now that the runners enforce format assertions. - privacy-violation-{response-text,query-intent,topics,response-type, response-mode,model-id}-at-minimal: one fixture per remaining field forbidden at minimal privacy (section 5.5). - presented/retrieved/engaged-missing-identifier: the section 5.7.5 identifier rule, previously only tested via content_grounded. - engaged-presentation-id-unmatched, presented-citation-id-unmatched: the new intra-session referential integrity checks (sections 6.7, 6.8). New valid fixture: - session-reproduction-no-grounding: the fifth funnel departure of section 4.3, reproduction of memorised content with no grounding event. Co-Authored-By: Claude Fable 5 --- tests/README.md | 14 ++++-- .../invalid/cited-missing-citation-type.json | 21 ++++++++ tests/invalid/cited-missing-id.json | 20 ++++++++ tests/invalid/engaged-missing-identifier.json | 30 +++++++++++ .../engaged-presentation-id-unmatched.json | 31 ++++++++++++ .../grounded-negative-chars-ingested.json | 19 +++++++ .../invalid/malformed-parent-session-id.json | 17 +++++++ tests/invalid/presentation-kind-invalid.json | 21 ++++++++ .../presented-citation-id-unmatched.json | 34 +++++++++++++ tests/invalid/presented-missing-id.json | 20 ++++++++ .../invalid/presented-missing-identifier.json | 20 ++++++++ .../invalid/presented-missing-output-id.json | 20 ++++++++ ...privacy-violation-model-id-at-minimal.json | 19 +++++++ ...acy-violation-query-intent-at-minimal.json | 19 +++++++ ...cy-violation-response-mode-at-minimal.json | 19 +++++++ ...cy-violation-response-text-at-minimal.json | 19 +++++++ ...cy-violation-response-type-at-minimal.json | 19 +++++++ .../privacy-violation-topics-at-minimal.json | 22 ++++++++ tests/invalid/reproduced-negative-chars.json | 21 ++++++++ tests/invalid/reproduction-type-invalid.json | 21 ++++++++ .../invalid/retrieved-missing-identifier.json | 19 +++++++ .../session-reproduction-no-grounding.json | 50 +++++++++++++++++++ tests/validate.py | 46 +++++++++++++++++ 23 files changed, 536 insertions(+), 5 deletions(-) create mode 100644 tests/invalid/cited-missing-citation-type.json create mode 100644 tests/invalid/cited-missing-id.json create mode 100644 tests/invalid/engaged-missing-identifier.json create mode 100644 tests/invalid/engaged-presentation-id-unmatched.json create mode 100644 tests/invalid/grounded-negative-chars-ingested.json create mode 100644 tests/invalid/malformed-parent-session-id.json create mode 100644 tests/invalid/presentation-kind-invalid.json create mode 100644 tests/invalid/presented-citation-id-unmatched.json create mode 100644 tests/invalid/presented-missing-id.json create mode 100644 tests/invalid/presented-missing-identifier.json create mode 100644 tests/invalid/presented-missing-output-id.json create mode 100644 tests/invalid/privacy-violation-model-id-at-minimal.json create mode 100644 tests/invalid/privacy-violation-query-intent-at-minimal.json create mode 100644 tests/invalid/privacy-violation-response-mode-at-minimal.json create mode 100644 tests/invalid/privacy-violation-response-text-at-minimal.json create mode 100644 tests/invalid/privacy-violation-response-type-at-minimal.json create mode 100644 tests/invalid/privacy-violation-topics-at-minimal.json create mode 100644 tests/invalid/reproduced-negative-chars.json create mode 100644 tests/invalid/reproduction-type-invalid.json create mode 100644 tests/invalid/retrieved-missing-identifier.json create mode 100644 tests/valid/session-reproduction-no-grounding.json diff --git a/tests/README.md b/tests/README.md index 9c5fef8..efdde4a 100644 --- a/tests/README.md +++ b/tests/README.md @@ -28,11 +28,14 @@ Run from the repository root. Without uv: `pip install "jsonschema[format-nongpl - Event required fields (`type`, `timestamp`) - Turn required fields (`privacy_level`) - Enum validation (event types, privacy levels, source roles, schema version) -- Citation source-reference requirement (content_cited rejected when content_url/content_id are missing or null) +- Citation source-reference requirement (content_cited rejected when content_url/content_id are missing or null) and required citation_type +- Required event and output identifiers on cited, reproduced, and presented events +- Closed enums (reproduction_type, presentation_kind) and non-negative counts (chars_ingested, reproduced_chars) +- Format assertions (a malformed parent_session_id fails) - All three conformance levels (Retrieval, Grounding, Citation) - Standalone event envelopes (CDN edge, agent with session FK) -- Privacy level field gating (application-layer conformance) -- Funnel exceptions (presented-no-cited, cited-no-grounded, presented-no-grounded) +- Privacy level field gating (application-layer conformance), one fixture per forbidden field at minimal +- Funnel exceptions (presented-no-cited, cited-no-grounded, presented-no-grounded, reproduced-no-grounded) - Reproduction cases: credited quotation (reproduction + direct_quote citation sharing an output element) and uncredited reproduction in unpresented API output - Text, image, audio, video, suppressed-citation, and repeated-presentation cases - Exact presentation-to-engagement correlation across session and standalone envelopes @@ -40,15 +43,16 @@ Run from the repository root. Without uv: `pip install "jsonschema[format-nongpl - Grounding provenance paths and generic fingerprint detection across session, standalone-event, and event-batch envelopes - Custom response_mode values -Each test file has a `_test_description` field explaining what it demonstrates. +Each test file has a `_test_description` field explaining what it demonstrates. Every `invalid/` fixture also has an `_expected_error` field: a substring that must appear in the actual error (the first schema error's message and JSON pointer, or the application-layer violation text). The runner fails a fixture that fails for a different reason than the one it pins, and fails any invalid fixture missing the field. ## Application-layer conformance Some rules cannot be expressed in JSON Schema alone. These are tested as application-layer conformance checks in `validate.py`: -- Privacy level field gating (e.g. `query_text` MUST NOT be present at `minimal` level) +- Privacy level field gating (e.g. `query_text` MUST NOT be present at `minimal` level), applied to turns wherever they appear: session documents, batches, and standalone envelopes - `content_url` or `content_id` requirement on every content event (section 5.7.5) - `session_id` or `ctx_token` on a standalone event or event batch envelope at Grounding conformance and above (sections 5.7.5, 7.1) +- Referential integrity within a session document: `content_engaged.presentation_id` matches a `content_presented` event id, and `citation_id` on `content_presented`/`content_reproduced` matches a `content_cited` event id (sections 6.6-6.8). Standalone envelopes and batch members are exempt - they may reference events delivered elsewhere. - Manifest rejection rules: duplicate `keys[].id`, and `domains` entries that are not the manifest's own host or a subdomain of it (sections 8.6, 8.7) - Withdrawn `ip_hash` prohibition on `content_retrieved` data (section 9.1 migration rule) - Grounding provenance/cache consistency and the prohibition on `preserved_in_output` in `content_fingerprint` (sections 5.7.5, 6.4, 12.1) diff --git a/tests/invalid/cited-missing-citation-type.json b/tests/invalid/cited-missing-citation-type.json new file mode 100644 index 0000000..9d6427c --- /dev/null +++ b/tests/invalid/cited-missing-citation-type.json @@ -0,0 +1,21 @@ +{ + "_test_description": "content_cited must classify how content was used in the response: data.citation_type is required. Emitters that cannot confidently classify use citation_type 'unclassified' rather than omitting the field (section 6.5).", + "_expected_error": "'citation_type' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440204", + "started_at": "2026-08-05T09:30:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440205", + "type": "content_cited", + "timestamp": "2026-08-05T09:30:01Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "position": "primary", + "url_verified": true + } + } + ] +} diff --git a/tests/invalid/cited-missing-id.json b/tests/invalid/cited-missing-id.json new file mode 100644 index 0000000..633a721 --- /dev/null +++ b/tests/invalid/cited-missing-id.json @@ -0,0 +1,20 @@ +{ + "_test_description": "content_cited event with output_id and a resolvable source reference but no event id. id is required on citation events so a presentation or reproduction can reference the exact citation via citation_id (section 6.5).", + "_expected_error": "'id' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440203", + "started_at": "2026-08-05T09:20:00Z", + "events": [ + { + "type": "content_cited", + "timestamp": "2026-08-05T09:20:01Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "citation_type": "reference", + "position": "primary" + } + } + ] +} diff --git a/tests/invalid/engaged-missing-identifier.json b/tests/invalid/engaged-missing-identifier.json new file mode 100644 index 0000000..aa22303 --- /dev/null +++ b/tests/invalid/engaged-missing-identifier.json @@ -0,0 +1,30 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: content_engaged event carrying neither content_url nor content_id. The engagement references its presentation via presentation_id but still MUST identify the content acted on (section 5.7.5).", + "_expected_error": "'content_engaged' carries neither content_url nor content_id", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440234", + "started_at": "2026-08-05T12:20:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440235", + "type": "content_presented", + "timestamp": "2026-08-05T12:20:01Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-05T12:20:04Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440235", + "data": { + "engagement_type": "link_click" + } + } + ] +} diff --git a/tests/invalid/engaged-presentation-id-unmatched.json b/tests/invalid/engaged-presentation-id-unmatched.json new file mode 100644 index 0000000..7a0aacd --- /dev/null +++ b/tests/invalid/engaged-presentation-id-unmatched.json @@ -0,0 +1,31 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: content_engaged whose presentation_id matches no content_presented event id in the session. Section 6.8 requires presentation_id to reference the exact content_presented.id on which the action occurred; an all-zeros UUID satisfies the schema's format assertion but references nothing.", + "_expected_error": "presentation_id '00000000-0000-0000-0000-000000000000' does not match any content_presented event id", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440236", + "started_at": "2026-08-05T12:30:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440237", + "type": "content_presented", + "timestamp": "2026-08-05T12:30:01Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-05T12:30:04Z", + "turn_id": "1", + "presentation_id": "00000000-0000-0000-0000-000000000000", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + } + } + ] +} diff --git a/tests/invalid/grounded-negative-chars-ingested.json b/tests/invalid/grounded-negative-chars-ingested.json new file mode 100644 index 0000000..dc44c37 --- /dev/null +++ b/tests/invalid/grounded-negative-chars-ingested.json @@ -0,0 +1,19 @@ +{ + "_test_description": "content_grounded with a negative chars_ingested. Ingestion counts are Unicode code point counts and cannot be negative (section 6.4); the schema requires a minimum of 0.", + "_expected_error": "chars_ingested -7200 is less than the minimum", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440210", + "started_at": "2026-08-05T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-05T10:00:01Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "session", + "cached": false, + "chars_ingested": -7200 + } + } + ] +} diff --git a/tests/invalid/malformed-parent-session-id.json b/tests/invalid/malformed-parent-session-id.json new file mode 100644 index 0000000..3d71576 --- /dev/null +++ b/tests/invalid/malformed-parent-session-id.json @@ -0,0 +1,17 @@ +{ + "_test_description": "Session document whose parent_session_id is not a UUID. parent_session_id carries format: uuid (section 5.1); this is only rejected when the validator enforces format assertions, which the conformance runners do.", + "_expected_error": "'orchestrator-main' is not a 'uuid'", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440213", + "parent_session_id": "orchestrator-main", + "agent_id": "research-subagent-v2", + "started_at": "2026-08-05T10:20:00Z", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-05T10:20:01Z", + "source_role": "agent", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + } + ] +} diff --git a/tests/invalid/presentation-kind-invalid.json b/tests/invalid/presentation-kind-invalid.json new file mode 100644 index 0000000..2499eeb --- /dev/null +++ b/tests/invalid/presentation-kind-invalid.json @@ -0,0 +1,21 @@ +{ + "_test_description": "content_presented with presentation_kind 'reference'. presentation_kind is a closed two-value distinction between source content and a reference to the source: only 'content' and 'source_reference' are valid (section 6.7).", + "_expected_error": "'reference' is not one of", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440208", + "started_at": "2026-08-05T09:50:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440209", + "type": "content_presented", + "timestamp": "2026-08-05T09:50:01Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "reference", + "presentation_type": "link" + } + } + ] +} diff --git a/tests/invalid/presented-citation-id-unmatched.json b/tests/invalid/presented-citation-id-unmatched.json new file mode 100644 index 0000000..4b604d2 --- /dev/null +++ b/tests/invalid/presented-citation-id-unmatched.json @@ -0,0 +1,34 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: content_presented whose citation_id matches no content_cited event id in the session. Section 6.7 requires citation_id to reference the presented content_cited event's id; here it points at an all-zeros UUID while the session's actual citation has a different id.", + "_expected_error": "citation_id '00000000-0000-0000-0000-000000000000' does not match any content_cited event id", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440238", + "started_at": "2026-08-05T12:40:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440239", + "type": "content_cited", + "timestamp": "2026-08-05T12:40:01Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "citation_type": "reference", + "position": "primary" + } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440240", + "type": "content_presented", + "timestamp": "2026-08-05T12:40:02Z", + "turn_id": "1", + "output_id": "response:1", + "citation_id": "00000000-0000-0000-0000-000000000000", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + } + ] +} diff --git a/tests/invalid/presented-missing-id.json b/tests/invalid/presented-missing-id.json new file mode 100644 index 0000000..6bdd533 --- /dev/null +++ b/tests/invalid/presented-missing-id.json @@ -0,0 +1,20 @@ +{ + "_test_description": "content_presented event with output_id and presentation data but no event id. id is required on presentation events so a later content_engaged.presentation_id can reference the exact surface occurrence (section 6.7).", + "_expected_error": "'id' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440200", + "started_at": "2026-08-05T09:00:00Z", + "events": [ + { + "type": "content_presented", + "timestamp": "2026-08-05T09:00:01Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + } + ] +} diff --git a/tests/invalid/presented-missing-identifier.json b/tests/invalid/presented-missing-identifier.json new file mode 100644 index 0000000..441d056 --- /dev/null +++ b/tests/invalid/presented-missing-identifier.json @@ -0,0 +1,20 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: content_presented event carrying neither content_url nor content_id. Passes JSON Schema (both are individually optional) but violates section 5.7.5: at least one MUST be present on every content event.", + "_expected_error": "'content_presented' carries neither content_url nor content_id", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440230", + "started_at": "2026-08-05T12:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440231", + "type": "content_presented", + "timestamp": "2026-08-05T12:00:01Z", + "turn_id": "1", + "output_id": "response:1", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + } + ] +} diff --git a/tests/invalid/presented-missing-output-id.json b/tests/invalid/presented-missing-output-id.json new file mode 100644 index 0000000..ff1ac5c --- /dev/null +++ b/tests/invalid/presented-missing-output-id.json @@ -0,0 +1,20 @@ +{ + "_test_description": "content_presented event with an event id and presentation data but no output_id. output_id is required on presentation events so presentation can be correlated with the output artifact it delivers (section 6.7).", + "_expected_error": "'output_id' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440201", + "started_at": "2026-08-05T09:10:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440202", + "type": "content_presented", + "timestamp": "2026-08-05T09:10:01Z", + "turn_id": "1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + } + ] +} diff --git a/tests/invalid/privacy-violation-model-id-at-minimal.json b/tests/invalid/privacy-violation-model-id-at-minimal.json new file mode 100644 index 0000000..f585b04 --- /dev/null +++ b/tests/invalid/privacy-violation-model-id-at-minimal.json @@ -0,0 +1,19 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: Turn at minimal privacy with model_id present. Passes JSON Schema validation but violates the privacy field gating rule (section 5.5): model_id MUST NOT be present when privacy_level is minimal.", + "_expected_error": "'model_id' present on turn with privacy_level 'minimal'", + "schema_version": "1.0", + "session_id": "770e8400-e29b-41d4-a716-446655440225", + "started_at": "2026-08-05T11:00:00Z", + "events": [ + { + "type": "turn_completed", + "timestamp": "2026-08-05T11:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "minimal", + "model_id": "claude-4-sonnet", + "response_tokens": 150 + } + } + ] +} diff --git a/tests/invalid/privacy-violation-query-intent-at-minimal.json b/tests/invalid/privacy-violation-query-intent-at-minimal.json new file mode 100644 index 0000000..e29b62b --- /dev/null +++ b/tests/invalid/privacy-violation-query-intent-at-minimal.json @@ -0,0 +1,19 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: Turn at minimal privacy with query_intent present. Passes JSON Schema validation but violates the privacy field gating rule (section 5.5): query_intent MUST NOT be present when privacy_level is minimal.", + "_expected_error": "'query_intent' present on turn with privacy_level 'minimal'", + "schema_version": "1.0", + "session_id": "770e8400-e29b-41d4-a716-446655440221", + "started_at": "2026-08-05T11:00:00Z", + "events": [ + { + "type": "turn_completed", + "timestamp": "2026-08-05T11:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "minimal", + "query_intent": "purchase_intent", + "response_tokens": 150 + } + } + ] +} diff --git a/tests/invalid/privacy-violation-response-mode-at-minimal.json b/tests/invalid/privacy-violation-response-mode-at-minimal.json new file mode 100644 index 0000000..5cf7c78 --- /dev/null +++ b/tests/invalid/privacy-violation-response-mode-at-minimal.json @@ -0,0 +1,19 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: Turn at minimal privacy with response_mode present. Passes JSON Schema validation but violates the privacy field gating rule (section 5.5): response_mode MUST NOT be present when privacy_level is minimal.", + "_expected_error": "'response_mode' present on turn with privacy_level 'minimal'", + "schema_version": "1.0", + "session_id": "770e8400-e29b-41d4-a716-446655440224", + "started_at": "2026-08-05T11:00:00Z", + "events": [ + { + "type": "turn_completed", + "timestamp": "2026-08-05T11:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "minimal", + "response_mode": "standard", + "response_tokens": 150 + } + } + ] +} diff --git a/tests/invalid/privacy-violation-response-text-at-minimal.json b/tests/invalid/privacy-violation-response-text-at-minimal.json new file mode 100644 index 0000000..7390bbe --- /dev/null +++ b/tests/invalid/privacy-violation-response-text-at-minimal.json @@ -0,0 +1,19 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: Turn at minimal privacy with response_text present. Passes JSON Schema validation but violates the privacy field gating rule (section 5.5): response_text MUST NOT be present when privacy_level is minimal.", + "_expected_error": "'response_text' present on turn with privacy_level 'minimal'", + "schema_version": "1.0", + "session_id": "770e8400-e29b-41d4-a716-446655440220", + "started_at": "2026-08-05T11:00:00Z", + "events": [ + { + "type": "turn_completed", + "timestamp": "2026-08-05T11:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "minimal", + "response_text": "The Baratza Encore ESP is the best grinder under 100 pounds for espresso and filter alike.", + "response_tokens": 150 + } + } + ] +} diff --git a/tests/invalid/privacy-violation-response-type-at-minimal.json b/tests/invalid/privacy-violation-response-type-at-minimal.json new file mode 100644 index 0000000..a7101a2 --- /dev/null +++ b/tests/invalid/privacy-violation-response-type-at-minimal.json @@ -0,0 +1,19 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: Turn at minimal privacy with response_type present. Passes JSON Schema validation but violates the privacy field gating rule (section 5.5): response_type MUST NOT be present when privacy_level is minimal.", + "_expected_error": "'response_type' present on turn with privacy_level 'minimal'", + "schema_version": "1.0", + "session_id": "770e8400-e29b-41d4-a716-446655440223", + "started_at": "2026-08-05T11:00:00Z", + "events": [ + { + "type": "turn_completed", + "timestamp": "2026-08-05T11:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "minimal", + "response_type": "recommendation", + "response_tokens": 150 + } + } + ] +} diff --git a/tests/invalid/privacy-violation-topics-at-minimal.json b/tests/invalid/privacy-violation-topics-at-minimal.json new file mode 100644 index 0000000..a78a26a --- /dev/null +++ b/tests/invalid/privacy-violation-topics-at-minimal.json @@ -0,0 +1,22 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: Turn at minimal privacy with topics present. Passes JSON Schema validation but violates the privacy field gating rule (section 5.5): topics MUST NOT be present when privacy_level is minimal.", + "_expected_error": "'topics' present on turn with privacy_level 'minimal'", + "schema_version": "1.0", + "session_id": "770e8400-e29b-41d4-a716-446655440222", + "started_at": "2026-08-05T11:00:00Z", + "events": [ + { + "type": "turn_completed", + "timestamp": "2026-08-05T11:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "minimal", + "topics": [ + "coffee grinders", + "burr grinders" + ], + "response_tokens": 150 + } + } + ] +} diff --git a/tests/invalid/reproduced-negative-chars.json b/tests/invalid/reproduced-negative-chars.json new file mode 100644 index 0000000..97374c3 --- /dev/null +++ b/tests/invalid/reproduced-negative-chars.json @@ -0,0 +1,21 @@ +{ + "_test_description": "content_reproduced with a negative reproduced_chars. The reproduced span length is a Unicode code point count and cannot be negative (section 6.6); the schema requires a minimum of 0.", + "_expected_error": "reproduced_chars -412 is less than the minimum", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440211", + "started_at": "2026-08-05T10:10:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440212", + "type": "content_reproduced", + "timestamp": "2026-08-05T10:10:01Z", + "turn_id": "1", + "output_id": "response:1", + "content_id": "publisher:article:9", + "data": { + "reproduction_type": "verbatim", + "reproduced_chars": -412 + } + } + ] +} diff --git a/tests/invalid/reproduction-type-invalid.json b/tests/invalid/reproduction-type-invalid.json new file mode 100644 index 0000000..04be830 --- /dev/null +++ b/tests/invalid/reproduction-type-invalid.json @@ -0,0 +1,21 @@ +{ + "_test_description": "content_reproduced with reproduction_type 'paraphrase'. Paraphrase is not reproduction (section 6.6): a credited paraphrase is a content_cited event with citation_type 'paraphrase', and an uncredited one is silent grounding. reproduction_type accepts only verbatim, near_verbatim, unclassified.", + "_expected_error": "'paraphrase' is not one of", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440206", + "started_at": "2026-08-05T09:40:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440207", + "type": "content_reproduced", + "timestamp": "2026-08-05T09:40:01Z", + "turn_id": "1", + "output_id": "response:1", + "content_id": "publisher:article:7", + "data": { + "reproduction_type": "paraphrase", + "reproduced_chars": 280 + } + } + ] +} diff --git a/tests/invalid/retrieved-missing-identifier.json b/tests/invalid/retrieved-missing-identifier.json new file mode 100644 index 0000000..9d9b88a --- /dev/null +++ b/tests/invalid/retrieved-missing-identifier.json @@ -0,0 +1,19 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: content_retrieved event carrying neither content_url nor content_id. Passes JSON Schema (both are individually optional) but violates section 5.7.5: at least one MUST be present on every content event.", + "_expected_error": "'content_retrieved' carries neither content_url nor content_id", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440232", + "started_at": "2026-08-05T12:10:00Z", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-05T12:10:01Z", + "source_role": "agent", + "content_telemetry_id": "880e8400-e29b-41d4-a716-446655440233", + "data": { + "response_status": 200, + "media_type": "text" + } + } + ] +} diff --git a/tests/valid/session-reproduction-no-grounding.json b/tests/valid/session-reproduction-no-grounding.json new file mode 100644 index 0000000..3bc3a78 --- /dev/null +++ b/tests/valid/session-reproduction-no-grounding.json @@ -0,0 +1,50 @@ +{ + "_test_description": "Funnel exception: content_reproduced with no content_grounded event. The response quotes a passage the model memorised during training - the content never entered this session's generation context, so there is no retrieval and no grounding to report. Valid per section 4.3 (reproduced without grounded); telemetry consumers SHOULD treat it, like an uncorroborated citation, as a lower-confidence signal.", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440241", + "agent_id": "assistant-v4", + "started_at": "2026-08-05T15:00:00Z", + "ended_at": "2026-08-05T15:00:03Z", + "events": [ + { + "type": "turn_started", + "timestamp": "2026-08-05T15:00:00Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "query_intent": "question", + "topics": [ + "english literature", + "opening lines" + ] + } + }, + { + "id": "880e8400-e29b-41d4-a716-446655440242", + "type": "content_reproduced", + "timestamp": "2026-08-05T15:00:02Z", + "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:quote:1", + "content_url": "https://www.gutenberg.org/ebooks/1342", + "content_id": "isbn:9780141439518", + "data": { + "reproduction_type": "verbatim", + "media_type": "text", + "reproduced_chars": 122, + "reproduced_hash": "sha256:a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2" + } + }, + { + "type": "turn_completed", + "timestamp": "2026-08-05T15:00:03Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "response_type": "explanation", + "response_mode": "standard", + "response_tokens": 160 + } + } + ] +} diff --git a/tests/validate.py b/tests/validate.py index ef0b5b9..abfc6d7 100644 --- a/tests/validate.py +++ b/tests/validate.py @@ -102,10 +102,56 @@ "Turn at intent privacy includes query_text. " "Violates section 5.5: query_text MUST NOT be present at intent level." ), + "privacy-violation-response-text-at-minimal.json": ( + "Turn at minimal privacy includes response_text. " + "Violates section 5.5: response text MUST NOT be present at minimal level." + ), + "privacy-violation-query-intent-at-minimal.json": ( + "Turn at minimal privacy includes query_intent. " + "Violates section 5.5: intent classification MUST NOT be present at minimal level." + ), + "privacy-violation-topics-at-minimal.json": ( + "Turn at minimal privacy includes topics. " + "Violates section 5.5: topics MUST NOT be present at minimal level." + ), + "privacy-violation-response-type-at-minimal.json": ( + "Turn at minimal privacy includes response_type. " + "Violates section 5.5: response classification MUST NOT be present at minimal level." + ), + "privacy-violation-response-mode-at-minimal.json": ( + "Turn at minimal privacy includes response_mode. " + "Violates section 5.5: platform metadata MUST NOT be present at minimal level." + ), + "privacy-violation-model-id-at-minimal.json": ( + "Turn at minimal privacy includes model_id. " + "Violates section 5.5: platform metadata MUST NOT be present at minimal level." + ), "content-event-missing-identifier.json": ( "content_grounded event has neither content_url nor content_id. " "Violates section 5.7.5: every content event MUST carry at least one." ), + "presented-missing-identifier.json": ( + "content_presented event has neither content_url nor content_id. " + "Violates section 5.7.5: every content event MUST carry at least one." + ), + "retrieved-missing-identifier.json": ( + "content_retrieved event has neither content_url nor content_id. " + "Violates section 5.7.5: every content event MUST carry at least one." + ), + "engaged-missing-identifier.json": ( + "content_engaged event has neither content_url nor content_id. " + "Violates section 5.7.5: every content event MUST carry at least one." + ), + "engaged-presentation-id-unmatched.json": ( + "content_engaged.presentation_id matches no content_presented event id " + "in the session. Violates section 6.8: every engagement references the " + "exact content_presented.id on which the action occurred." + ), + "presented-citation-id-unmatched.json": ( + "content_presented.citation_id matches no content_cited event id in " + "the session. Violates section 6.7: citation_id references the " + "presented content_cited event's id." + ), "standalone-missing-session-and-ctx-token.json": ( "Standalone event envelope has neither session_id nor ctx_token. " "Violates section 5.7.5: an event MUST carry one at Grounding+ (section 7.1)." From f928ea70b60fd2fe3ded66ff559b6bd540dee841 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Wed, 12 Aug 2026 11:23:30 +0100 Subject: [PATCH 30/38] Add a mutation smoke test replaying the review's suite-weakening edits mutation_smoke.py copies the schemas and tests/ into a temp directory, applies each known suite-weakening mutation - dropping the format checker, gutting the withdrawn-ip-hash fixture, shrinking CONTENT_EVENT_TYPES and PRIVACY_FORBIDDEN_FIELDS, pointing an engagement at an all-zeros presentation_id - and confirms validate.py fails under every one. Each of these previously went undetected. The working tree is never modified; a surviving mutation fails the script. Co-Authored-By: Claude Fable 5 --- tests/README.md | 1 + tests/mutation_smoke.py | 128 ++++++++++++++++++++++++++++++++++++++++ 2 files changed, 129 insertions(+) create mode 100644 tests/mutation_smoke.py diff --git a/tests/README.md b/tests/README.md index efdde4a..e764e1a 100644 --- a/tests/README.md +++ b/tests/README.md @@ -8,6 +8,7 @@ Tests for the Content Telemetry Specification v1. - `invalid/` - JSON files that MUST fail validation (either JSON Schema or application-layer conformance) - `validate.py` - Conformance test runner (requires `jsonschema`) - `check_examples.py` - Validates the worked examples in SPECIFICATION.md and README.md against the schemas +- `mutation_smoke.py` - Replays known suite-weakening mutations against a scratch copy and confirms the suite fails under each one ## Running diff --git a/tests/mutation_smoke.py b/tests/mutation_smoke.py new file mode 100644 index 0000000..9783d5c --- /dev/null +++ b/tests/mutation_smoke.py @@ -0,0 +1,128 @@ +#!/usr/bin/env python3 +"""Mutation smoke test for the conformance suite. + +Replays the review's key mutations - each of which the suite previously +missed - against a scratch copy of the repository and confirms validate.py +now fails under every one of them. The working tree is never modified. + +A mutation "survives" when the mutated suite still exits 0; any survivor +is a real detection gap and fails this script. + +Usage: + python tests/mutation_smoke.py +""" + +import json +import shutil +import subprocess +import sys +import tempfile +from pathlib import Path + +REPO = Path(__file__).resolve().parent.parent + + +def mutate_drop_format_checker(root): + """Build every validator without a format checker (formats become no-ops).""" + p = root / "tests" / "validate.py" + src = p.read_text() + mutated = src.replace(", format_checker=FORMAT_CHECKER", "") + assert mutated != src + p.write_text(mutated) + return "malformed-parent-session-id.json" + + +def mutate_gut_withdrawn_ip_hash(root): + """Gut the withdrawn-ip-hash fixture: drop ip_hash and break the event + some other way, so it fails schema for a reason unrelated to its rule.""" + p = root / "tests" / "invalid" / "withdrawn-ip-hash.json" + doc = json.loads(p.read_text()) + del doc["event"]["data"]["ip_hash"] + del doc["event"]["timestamp"] + p.write_text(json.dumps(doc, indent=2) + "\n") + return "withdrawn-ip-hash.json" + + +def mutate_shrink_content_event_types(root): + """Remove content_presented from CONTENT_EVENT_TYPES.""" + p = root / "tests" / "validate.py" + src = p.read_text() + mutated = src.replace('"content_cited", "content_presented", "content_engaged",', + '"content_cited", "content_engaged",') + assert mutated != src + p.write_text(mutated) + return "presented-missing-identifier.json" + + +def mutate_shrink_privacy_forbidden_fields(root): + """Remove topics from the fields forbidden at minimal privacy.""" + p = root / "tests" / "validate.py" + src = p.read_text() + mutated = src.replace('"query_text", "response_text", "query_intent", "topics",', + '"query_text", "response_text", "query_intent",') + assert mutated != src + p.write_text(mutated) + return "privacy-violation-topics-at-minimal.json" + + +def mutate_zero_uuid_engagement(root): + """Point a valid session's engagement at an all-zeros presentation_id.""" + p = root / "tests" / "valid" / "session-citation-tier.json" + doc = json.loads(p.read_text()) + mutated = False + for event in doc["events"]: + if event.get("type") == "content_engaged": + event["presentation_id"] = "00000000-0000-0000-0000-000000000000" + mutated = True + assert mutated + p.write_text(json.dumps(doc, indent=2) + "\n") + return "session-citation-tier.json" + + +MUTATIONS = [ + ("drop format_checker", mutate_drop_format_checker), + ("gut withdrawn-ip-hash fixture", mutate_gut_withdrawn_ip_hash), + ("remove content_presented from CONTENT_EVENT_TYPES", mutate_shrink_content_event_types), + ("shrink PRIVACY_FORBIDDEN_FIELDS", mutate_shrink_privacy_forbidden_fields), + ("engagement at all-zeros presentation_id", mutate_zero_uuid_engagement), +] + + +def run_one(name, mutate): + with tempfile.TemporaryDirectory(prefix="ct-mutation-") as tmp: + root = Path(tmp) / "repo" + root.mkdir() + for schema in ("telemetry-session.json", "telemetry-event.json", + "telemetry-event-batch.json", "manifest.json"): + shutil.copy(REPO / schema, root / schema) + shutil.copytree(REPO / "tests", root / "tests") + expected_fixture = mutate(root) + proc = subprocess.run( + [sys.executable, str(root / "tests" / "validate.py")], + capture_output=True, text=True, + ) + caught = proc.returncode != 0 and f"FAIL {expected_fixture}" in proc.stdout + return caught, expected_fixture, proc + + +def main(): + survivors = 0 + for name, mutate in MUTATIONS: + caught, fixture, proc = run_one(name, mutate) + if caught: + print(f" CAUGHT {name} (failed via {fixture})") + else: + survivors += 1 + print(f" SURVIVED {name} (expected {fixture} to fail)") + print(f" exit={proc.returncode}") + for line in proc.stdout.splitlines()[-5:]: + print(f" {line}") + print() + print("=" * 60) + print(f"SUMMARY: {len(MUTATIONS) - survivors}/{len(MUTATIONS)} mutations caught") + print("=" * 60) + return 1 if survivors else 0 + + +if __name__ == "__main__": + sys.exit(main()) From ff8dc407d85f946b172b9ea4fa9e73eefa56c2e2 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Wed, 12 Aug 2026 11:24:53 +0100 Subject: [PATCH 31/38] Fix stale minimal-level comment: query_tokens is permitted Co-Authored-By: Claude Fable 5 --- tests/validate.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/validate.py b/tests/validate.py index abfc6d7..c9a65cb 100644 --- a/tests/validate.py +++ b/tests/validate.py @@ -53,7 +53,7 @@ # 1. Privacy level field gating (section 5.5): # - At "minimal" level: query_text, response_text, query_intent, topics, # model_id, ad_rendered, response_mode, and response_type MUST NOT be -# present. Only response_tokens and content_urls are allowed. +# present. Only query_tokens, response_tokens and content_urls are allowed. # - At "intent" level: query_text and response_text MUST NOT be present. # # 2. content_url or content_id requirement (section 5.7.5): From 4b4946bae79637992e8713f2218f33f579164b74 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Sat, 22 Aug 2026 12:28:06 +0100 Subject: [PATCH 32/38] Extend v1 identity and error pinning to the fixtures merged after this branch MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The rebase onto v1-draft brought in fixtures and an inline example added by PRs #38, #41 and #45, none of which were swept by the v1 identity commit or carry the pinned violation this branch now requires. - CI installed plain `jsonschema`, so the hardened runners' format-checker guard aborted the job. The workflow now matches tests/README.md and installs `jsonschema[format-nongpl]`. - 17 fixtures and the §5.1.3 access_context example still declared schema_version 0.1. Bumped to 1.0; invalid-schema-version.json keeps its deliberately wrong value. - 8 invalid fixtures had no `_expected_error`. Each now pins its own violation, including the two application-layer provenance rules and the ctx_token pattern. Suite: 107/107 fixtures, 13/13 examples, 5/5 mutations caught. Co-Authored-By: Claude Opus 5 --- .github/workflows/ci.yml | 4 ++-- SPECIFICATION.md | 2 +- .../access-context-identifier-missing-value.json | 7 +++++-- .../access-context-identifiers-not-array.json | 3 ++- tests/invalid/ctx-token-bad-pattern.json | 7 +++++-- ...standalone-missing-presentation-and-token.json | 7 +++++-- .../grounding-fingerprint-missing-detected.json | 3 ++- ...grounding-fingerprint-preserved-in-output.json | 3 ++- .../grounding-provenance-cached-conflict.json | 3 ++- tests/invalid/manifest-coverage-bad-mode.json | 15 +++++++++++---- tests/valid/event-batch-grounding-provenance.json | 2 +- .../valid/event-standalone-engaged-redirect.json | 2 +- .../event-standalone-grounded-provenance.json | 2 +- .../event-standalone-grounding-fingerprint.json | 2 +- tests/valid/event-standalone-terms-ref.json | 2 +- tests/valid/manifest-agent-ctx-resolution.json | 2 +- tests/valid/manifest-coverage-declaration.json | 2 +- tests/valid/session-access-context.json | 2 +- tests/valid/session-earlier-turn-click.json | 2 +- tests/valid/session-repeated-link-engagement.json | 2 +- 20 files changed, 47 insertions(+), 27 deletions(-) diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 9c523ad..d8bd18d 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -17,6 +17,6 @@ jobs: - name: Install uv uses: astral-sh/setup-uv@v5 - name: Run conformance test suite - run: uv run --with jsonschema python tests/validate.py + run: uv run --with "jsonschema[format-nongpl]" python tests/validate.py - name: Validate worked examples in the spec - run: uv run --with jsonschema python tests/check_examples.py + run: uv run --with "jsonschema[format-nongpl]" python tests/check_examples.py diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 64c2ec5..dd39ecf 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -383,7 +383,7 @@ One container is defined in core. `access_context` records the context from whic ```json { - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440000", "content_scope": "consortium-agreement-4471", "started_at": "2026-08-13T14:02:10Z", diff --git a/tests/invalid/access-context-identifier-missing-value.json b/tests/invalid/access-context-identifier-missing-value.json index 6e00daa..f7b38b4 100644 --- a/tests/invalid/access-context-identifier-missing-value.json +++ b/tests/invalid/access-context-identifier-missing-value.json @@ -1,13 +1,16 @@ { "_test_description": "Session access_context identifier missing its value. The schema requires both scheme and value on every identifier (5.1.3).", + "_expected_error": "access_context/identifiers/0 'value' is a required property", "document_type": "session", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "880e8400-e29b-41d4-a716-446655440090", "started_at": "2026-08-13T14:02:10Z", "data": { "access_context": { "identifiers": [ - { "scheme": "ror" } + { + "scheme": "ror" + } ] } }, diff --git a/tests/invalid/access-context-identifiers-not-array.json b/tests/invalid/access-context-identifiers-not-array.json index 2588b9f..226dc7a 100644 --- a/tests/invalid/access-context-identifiers-not-array.json +++ b/tests/invalid/access-context-identifiers-not-array.json @@ -1,7 +1,8 @@ { "_test_description": "Session access_context.identifiers as a bare string. The schema requires an array of {scheme, value} objects (5.1.3).", + "_expected_error": "access_context/identifiers 'https://ror.org/013meh722' is not of type 'array'", "document_type": "session", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "880e8400-e29b-41d4-a716-446655440091", "started_at": "2026-08-13T14:02:10Z", "data": { diff --git a/tests/invalid/ctx-token-bad-pattern.json b/tests/invalid/ctx-token-bad-pattern.json index b286ce5..e1f1560 100644 --- a/tests/invalid/ctx-token-bad-pattern.json +++ b/tests/invalid/ctx-token-bad-pattern.json @@ -1,12 +1,15 @@ { "_test_description": "ctx_token failing the value pattern: tokens are opaque but MUST match ^ct_[A-Za-z0-9_-]{8,240}$ so they survive URL carriage and are recognisable in logs (section 7.4). This value has no ct_ prefix and contains reserved characters.", + "_expected_error": "does not match '^ct_[A-Za-z0-9_-]{8,240}$'", "document_type": "event", - "schema_version": "0.1", + "schema_version": "1.0", "ctx_token": "session=660e8400!", "event": { "type": "content_engaged", "timestamp": "2026-08-10T11:20:00Z", "content_url": "https://news.example.com/markets/rate-decision", - "data": { "engagement_type": "link_click" } + "data": { + "engagement_type": "link_click" + } } } diff --git a/tests/invalid/engaged-standalone-missing-presentation-and-token.json b/tests/invalid/engaged-standalone-missing-presentation-and-token.json index fb7e85e..aab2c10 100644 --- a/tests/invalid/engaged-standalone-missing-presentation-and-token.json +++ b/tests/invalid/engaged-standalone-missing-presentation-and-token.json @@ -1,7 +1,8 @@ { "_test_description": "Agent-reported standalone content_engaged with session_id but no presentation_id. The presentation_id relaxation applies only to envelopes carrying ctx_token; an emitter that knows the session knows the presentation and MUST bind to it (sections 6.8, 7.4).", + "_expected_error": "/event 'presentation_id' is a required property", "document_type": "event", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440230", "agent_id": "assistant-v2", "started_at": "2026-08-10T12:00:00Z", @@ -9,6 +10,8 @@ "type": "content_engaged", "timestamp": "2026-08-10T12:00:30Z", "content_url": "https://news.example.com/markets/rate-decision", - "data": { "engagement_type": "link_click" } + "data": { + "engagement_type": "link_click" + } } } diff --git a/tests/invalid/grounding-fingerprint-missing-detected.json b/tests/invalid/grounding-fingerprint-missing-detected.json index d1b713e..2cff586 100644 --- a/tests/invalid/grounding-fingerprint-missing-detected.json +++ b/tests/invalid/grounding-fingerprint-missing-detected.json @@ -1,6 +1,7 @@ { "_test_description": "Grounding event carries content_fingerprint without required detected. Must fail JSON Schema.", - "schema_version": "0.1", + "_expected_error": "content_fingerprint 'detected' is a required property", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440182", "started_at": "2026-08-12T10:20:00Z", "events": [ diff --git a/tests/invalid/grounding-fingerprint-preserved-in-output.json b/tests/invalid/grounding-fingerprint-preserved-in-output.json index b8fb38e..bd5bfe3 100644 --- a/tests/invalid/grounding-fingerprint-preserved-in-output.json +++ b/tests/invalid/grounding-fingerprint-preserved-in-output.json @@ -1,6 +1,7 @@ { "_test_description": "Grounding fingerprint uses withdrawn preserved_in_output instead of a content_reproduced event. Passes the extensible schema but violates the v1 migration rule.", - "schema_version": "0.1", + "_expected_error": "content_fingerprint carries preserved_in_output; use content_reproduced", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440183", "started_at": "2026-08-12T10:30:00Z", "events": [ diff --git a/tests/invalid/grounding-provenance-cached-conflict.json b/tests/invalid/grounding-provenance-cached-conflict.json index fcffe65..638310c 100644 --- a/tests/invalid/grounding-provenance-cached-conflict.json +++ b/tests/invalid/grounding-provenance-cached-conflict.json @@ -1,6 +1,7 @@ { "_test_description": "Grounding event declares agent_fetched with cached true. Passes JSON Schema but violates section 6.4 provenance consistency.", - "schema_version": "0.1", + "_expected_error": "content_grounded with agent_fetched does not carry cached false", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440184", "started_at": "2026-08-12T10:40:00Z", "events": [ diff --git a/tests/invalid/manifest-coverage-bad-mode.json b/tests/invalid/manifest-coverage-bad-mode.json index 1e076e6..653c5ec 100644 --- a/tests/invalid/manifest-coverage-bad-mode.json +++ b/tests/invalid/manifest-coverage-bad-mode.json @@ -1,13 +1,20 @@ { "_test_description": "Manifest coverage entry with a mode outside the enum. The schema requires one of complete, sampled, aggregated, selected (8.5, 5.7.6).", - "schema_version": "0.1", + "_expected_error": "'partial' is not one of ['complete', 'sampled', 'aggregated', 'selected']", + "schema_version": "1.0", "id": "https://assistant.example.com/.well-known/content-telemetry.json", - "roles": ["agent"], - "operator": { "name": "Assistant Example" }, + "roles": [ + "agent" + ], + "operator": { + "name": "Assistant Example" + }, "telemetry": { "endpoint": "https://telemetry.assistant.example.com/v1/events", "coverage": { - "content_grounded": { "mode": "partial" } + "content_grounded": { + "mode": "partial" + } } } } diff --git a/tests/valid/event-batch-grounding-provenance.json b/tests/valid/event-batch-grounding-provenance.json index 946a431..9184ef2 100644 --- a/tests/valid/event-batch-grounding-provenance.json +++ b/tests/valid/event-batch-grounding-provenance.json @@ -1,7 +1,7 @@ { "_test_description": "Batch envelope carrying grounding events from cached and third-party-sourced representations, exercising the shared provenance schema across batch delivery.", "document_type": "event_batch", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440181", "agent_id": "enterprise-rag-v1", "started_at": "2026-08-12T10:10:00Z", diff --git a/tests/valid/event-standalone-engaged-redirect.json b/tests/valid/event-standalone-engaged-redirect.json index a7ba533..6dd4815 100644 --- a/tests/valid/event-standalone-engaged-redirect.json +++ b/tests/valid/event-standalone-engaged-redirect.json @@ -1,7 +1,7 @@ { "_test_description": "Destination-reported click after a redirect chain: the presented short link redirected to the canonical URL, and the ctx_token and ctx_iss query parameters were propagated through the same-domain hops (section 7.4). The destination reports the canonical URL it serves; correlation runs through the token, not the URL.", "document_type": "event", - "schema_version": "0.1", + "schema_version": "1.0", "ctx_token": "ct_77b41f0ac93e5d28", "event": { "type": "content_engaged", diff --git a/tests/valid/event-standalone-grounded-provenance.json b/tests/valid/event-standalone-grounded-provenance.json index 96f3c08..769c5ba 100644 --- a/tests/valid/event-standalone-grounded-provenance.json +++ b/tests/valid/event-standalone-grounded-provenance.json @@ -1,6 +1,6 @@ { "document_type": "event", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "550e8400-e29b-41d4-a716-446655440000", "agent_id": "agent-example", "started_at": "2026-06-20T10:00:00Z", diff --git a/tests/valid/event-standalone-grounding-fingerprint.json b/tests/valid/event-standalone-grounding-fingerprint.json index 70eefe4..45a5db8 100644 --- a/tests/valid/event-standalone-grounding-fingerprint.json +++ b/tests/valid/event-standalone-grounding-fingerprint.json @@ -1,7 +1,7 @@ { "_test_description": "Standalone grounding event from a publisher-authorised live fetch with an emitter-reported generic fingerprint detection. The event schema is shared by session, standalone and batch envelopes.", "document_type": "event", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440180", "agent_id": "research-agent-v1", "started_at": "2026-08-12T10:00:00Z", diff --git a/tests/valid/event-standalone-terms-ref.json b/tests/valid/event-standalone-terms-ref.json index 7293ad0..4a3ac4e 100644 --- a/tests/valid/event-standalone-terms-ref.json +++ b/tests/valid/event-standalone-terms-ref.json @@ -1,7 +1,7 @@ { "_test_description": "Standalone retrieval event reported under governing terms with no grant: terms_ref names the terms (5.2.4), license_ref is absent, and the envelope carries manifest_ref (7.1) to tie the emitter to a domain without a session document.", "document_type": "event", - "schema_version": "0.1", + "schema_version": "1.0", "agent_id": "assistant.example.com", "manifest_ref": "https://assistant.example.com/.well-known/content-telemetry.json", "event": { diff --git a/tests/valid/manifest-agent-ctx-resolution.json b/tests/valid/manifest-agent-ctx-resolution.json index 096fc7c..0dc8e87 100644 --- a/tests/valid/manifest-agent-ctx-resolution.json +++ b/tests/valid/manifest-agent-ctx-resolution.json @@ -1,6 +1,6 @@ { "_test_description": "Agent manifest declaring a click-token resolution endpoint in telemetry.ctx_resolution. A destination receiving ctx_iss=searchco.com/agents/web-search resolves this manifest and presents the token at the declared endpoint (sections 7.4, 8.5).", - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://searchco.com/agents/web-search/.well-known/content-telemetry.json", "roles": ["agent"], "operator": { "name": "SearchCo" }, diff --git a/tests/valid/manifest-coverage-declaration.json b/tests/valid/manifest-coverage-declaration.json index 642d0eb..e6332c1 100644 --- a/tests/valid/manifest-coverage-declaration.json +++ b/tests/valid/manifest-coverage-declaration.json @@ -1,6 +1,6 @@ { "_test_description": "Agent manifest declaring per-event-type coverage in telemetry.coverage (8.5, 5.7.6): content_grounded reported complete, content_retrieved sampled under a rule stated in the referenced terms.", - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://assistant.example.com/.well-known/content-telemetry.json", "roles": ["agent"], "operator": { "name": "Assistant Example" }, diff --git a/tests/valid/session-access-context.json b/tests/valid/session-access-context.json index ee6d3c4..b2c1422 100644 --- a/tests/valid/session-access-context.json +++ b/tests/valid/session-access-context.json @@ -1,7 +1,7 @@ { "_test_description": "Session-level data container (5.1.3): access_context identifies the institution whose entitlement the agent used, required by the governing terms, with turn data at intent level per 5.5. The retrieval event carries content_depth: full (6.1).", "document_type": "session", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "880e8400-e29b-41d4-a716-446655440080", "agent_id": "scholar-assistant.example.com", "content_scope": "consortium-agreement-4471", diff --git a/tests/valid/session-earlier-turn-click.json b/tests/valid/session-earlier-turn-click.json index a6b3ee5..2227dca 100644 --- a/tests/valid/session-earlier-turn-click.json +++ b/tests/valid/session-earlier-turn-click.json @@ -1,6 +1,6 @@ { "_test_description": "A click in a later turn on a presentation from an earlier turn: the content_engaged event carries the turn of the click but binds by presentation_id to the turn-1 presentation. Resolution returns the clicked content's lineage across turns by content identity, not a turn or timestamp cut (section 7.4).", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440220", "agent_id": "assistant-v2", "started_at": "2026-08-10T10:00:00Z", diff --git a/tests/valid/session-repeated-link-engagement.json b/tests/valid/session-repeated-link-engagement.json index 19669c7..589ee70 100644 --- a/tests/valid/session-repeated-link-engagement.json +++ b/tests/valid/session-repeated-link-engagement.json @@ -1,6 +1,6 @@ { "_test_description": "The same URL presented twice in one session: two content_presented events with distinct ids and distinct minted click tokens, and a content_engaged bound by presentation_id to the second occurrence. Matching on URL alone could not tell the two presentations apart (sections 6.8, 7.4).", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440210", "agent_id": "assistant-v2", "started_at": "2026-08-10T09:00:00Z", From b8a889bf6fb1c42c60bcf465ce5ac3b9b651b091 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Sat, 22 Aug 2026 17:05:45 +0100 Subject: [PATCH 33/38] Close the pre-freeze review findings: versioning, required fields, click-context gate, migration, suite hardening Prose (SPECIFICATION.md, README, tests/README) - schema_version is major.minor; 5.7.4, 8.7 and 12 state the one-schema-per-minor rule; the 1.0.0-style examples are gone and the const "1.0" stays - source_role is a MUST on content_retrieved everywhere (5.2.2 aligned with 5.7.1) - data.scope and citation_type stated as required in 6, 6.4, 6.5; 5.7.1 says "every content event" - 5.7.5 gains the field-placement, distinct-id, same-content join, one-token-one- presentation, envelope-token and manifest placement rules - 7.4.1 token minimum 16 characters and an unguessability MUST; 7.4.4 resolution gate defined on the click turn's privacy_level (minimal: engagement and lineage only; no turn fields at any level); 7.4.5 destination-owner caveat; 7.4.6 records token lifetime and requester authentication as v1 limits - 9.1 states that privacy_level gates the named turn fields only and that opaque and extension fields must not carry identity or withheld text - 12.1: the "URL-carried presentation_id" and published-v0.1 preserved_in_output paragraphs corrected; migration notes for citation_type, scope, the non-null reference rule, ip_hash, source_role and license_ref semantics added - Section 8 written in v1 tense; manifest id and endpoint are https-only and id sits at the well-known path; tokens_ingested at minimal; share and user-supplied agent_navigate engagements; Annex B.5 multi-owner catalogue (#32) Schemas - content_grounded requires data.scope; content_depth typed on retrievals; ct_ pattern {16,240}; manifest_ref format uri; envelope agent_id nullable like the session field; manifest id/endpoint https patterns Conformance suite - validate.py: source_role, field placement, duplicate ids, same-content joins, shared tokens, envelope ctx_token, root-manifest domains, agent/platform ctx_resolution; check_examples.py runs the application-layer rules too - 66 new invalid and 7 new valid fixtures: every closed enum, every format assertion, every cross-shape rule in all three document shapes, origin and index retrievals, platform manifest, all engagement/citation/presentation types, a destination engagement batch, the v0.1 wire version rejected - every schema pin is pointer-prefixed; session-minimal carries source_role; the one fixture without a description has one - mutation_smoke.py replays 59 mutations (54 new) and runs in CI Co-Authored-By: Claude Fable 5 --- .github/workflows/ci.yml | 2 + README.md | 6 +- SPECIFICATION.md | 230 +++++++++++++--- manifest.json | 10 +- telemetry-event-batch.json | 5 +- telemetry-event.json | 5 +- telemetry-session.json | 14 +- tests/README.md | 33 ++- tests/check_examples.py | 18 +- ...cess-context-identifier-missing-value.json | 2 +- .../access-context-identifiers-not-array.json | 2 +- .../invalid/batch-ctx-token-bad-pattern.json | 18 ++ .../batch-ctx-token-non-engagement.json | 28 ++ tests/invalid/batch-empty-events.json | 2 +- ...batch-engaged-missing-presentation-id.json | 20 ++ .../invalid/batch-missing-schema-version.json | 2 +- ...batch-missing-session-mixed-retrieval.json | 24 ++ tests/invalid/citation-id-on-grounded.json | 33 +++ tests/invalid/citation-position-invalid.json | 22 ++ tests/invalid/citation-type-invalid.json | 21 ++ .../invalid/cited-missing-citation-type.json | 2 +- tests/invalid/cited-missing-data.json | 18 ++ tests/invalid/cited-missing-id.json | 2 +- tests/invalid/cited-missing-output-id.json | 6 +- .../cited-missing-source-reference.json | 2 +- .../invalid/cited-negative-excerpt-chars.json | 22 ++ .../invalid/cited-null-source-reference.json | 2 +- tests/invalid/ctx-token-bad-pattern.json | 4 +- tests/invalid/ctx-token-on-grounded.json | 21 ++ tests/invalid/duplicate-event-id.json | 44 ++++ tests/invalid/empty-output-id.json | 22 ++ .../engaged-missing-presentation-id.json | 6 +- ...engaged-presentation-content-mismatch.json | 32 +++ ...aged-presentation-id-no-presentations.json | 30 +++ .../event-level-ctx-token-bad-pattern.json | 33 +++ .../grounded-missing-identifier-batch.json | 20 ++ ...rounded-missing-identifier-standalone.json | 18 ++ .../grounded-negative-chars-ingested.json | 2 +- .../grounded-negative-tokens-ingested.json | 20 ++ ...rounding-fingerprint-missing-detected.json | 2 +- .../grounding-fingerprint-missing-scheme.json | 22 ++ tests/invalid/grounding-missing-data.json | 15 ++ ...ovenance-cached-conflict-agent-cached.json | 20 ++ .../invalid/grounding-provenance-invalid.json | 20 ++ tests/invalid/grounding-scope-invalid.json | 19 ++ tests/invalid/grounding-scope-missing.json | 19 ++ tests/invalid/invalid-event-type.json | 2 +- tests/invalid/invalid-privacy-level.json | 2 +- tests/invalid/invalid-schema-version.json | 2 +- tests/invalid/invalid-source-role.json | 2 +- tests/invalid/legacy-content-displayed.json | 6 +- tests/invalid/legacy-schema-version-0-1.json | 16 ++ tests/invalid/malformed-citation-id.json | 35 +++ tests/invalid/malformed-content-hash.json | 20 ++ .../invalid/malformed-content-urls-cited.json | 22 ++ tests/invalid/malformed-event-id.json | 22 ++ .../invalid/malformed-parent-session-id.json | 2 +- tests/invalid/malformed-presentation-id.json | 32 +++ tests/invalid/malformed-session-id.json | 16 ++ tests/invalid/malformed-started-at.json | 16 ++ .../manifest-bad-conformance-level.json | 10 +- tests/invalid/manifest-bad-key-type.json | 19 ++ tests/invalid/manifest-bad-role.json | 10 +- tests/invalid/manifest-coverage-bad-mode.json | 2 +- .../manifest-coverage-missing-mode.json | 20 ++ .../invalid/manifest-ctx-resolution-http.json | 16 ++ ...ifest-ctx-resolution-on-content-owner.json | 16 ++ .../manifest-domains-on-path-manifest.json | 18 ++ tests/invalid/manifest-duplicate-roles.json | 13 + tests/invalid/manifest-empty-roles.json | 10 + tests/invalid/manifest-endpoint-http.json | 15 ++ tests/invalid/manifest-id-http.json | 12 + tests/invalid/manifest-id-not-a-uri.json | 12 + tests/invalid/manifest-id-not-well-known.json | 12 + tests/invalid/manifest-key-missing-id.json | 18 ++ tests/invalid/manifest-lookalike-domain.json | 16 ++ tests/invalid/manifest-missing-endpoint.json | 15 ++ .../manifest-missing-key-publickey.json | 15 +- tests/invalid/manifest-missing-operator.json | 6 +- tests/invalid/missing-event-timestamp.json | 2 +- tests/invalid/missing-event-type.json | 2 +- tests/invalid/missing-privacy-level.json | 2 +- tests/invalid/missing-schema-version.json | 2 +- tests/invalid/missing-session-id.json | 2 +- tests/invalid/missing-started-at.json | 2 +- tests/invalid/presentation-id-on-cited.json | 35 +++ tests/invalid/presentation-kind-invalid.json | 2 +- .../presented-citation-content-mismatch.json | 35 +++ tests/invalid/presented-missing-data.json | 18 ++ tests/invalid/presented-missing-id.json | 2 +- tests/invalid/presented-missing-kind.json | 6 +- .../invalid/presented-missing-output-id.json | 2 +- .../presented-missing-presentation-type.json | 6 +- ...vacy-violation-query-at-minimal-batch.json | 21 ++ ...violation-query-at-minimal-standalone.json | 19 ++ ...acy-violation-response-text-at-intent.json | 21 ++ .../reproduced-citation-id-unmatched.json | 23 ++ tests/invalid/reproduced-missing-data.json | 18 ++ tests/invalid/reproduced-missing-id.json | 7 +- .../invalid/reproduced-missing-output-id.json | 7 +- .../reproduced-missing-source-reference.json | 7 +- tests/invalid/reproduced-missing-type.json | 6 +- tests/invalid/reproduced-negative-chars.json | 2 +- tests/invalid/reproduction-type-invalid.json | 2 +- tests/invalid/retrieved-bad-country.json | 19 ++ .../retrieved-missing-source-role.json | 15 ++ ...etrieved-response-status-out-of-range.json | 19 ++ .../shared-ctx-token-two-presentations.json | 56 ++++ .../standalone-ctx-token-non-engagement.json | 17 ++ .../standalone-malformed-session-id.json | 13 + .../standalone-missing-document-type.json | 2 +- tests/invalid/standalone-missing-event.json | 2 +- .../standalone-missing-schema-version.json | 12 + tests/invalid/turn-on-content-event.json | 23 ++ tests/invalid/withdrawn-ip-hash-batch.json | 18 ++ tests/invalid/withdrawn-ip-hash-session.json | 20 ++ tests/mutation_smoke.py | 199 +++++++++++++- .../event-batch-destination-engagements.json | 25 ++ .../event-standalone-grounded-provenance.json | 1 + tests/valid/event-standalone-index.json | 18 ++ tests/valid/event-standalone-origin.json | 19 ++ tests/valid/manifest-platform.json | 15 ++ .../valid/session-citation-contradiction.json | 79 ++++++ tests/valid/session-engagement-types.json | 113 ++++++++ tests/valid/session-minimal.json | 3 +- .../valid/session-multi-owner-catalogue.json | 142 ++++++++++ tests/validate.py | 245 +++++++++++++++++- 127 files changed, 2553 insertions(+), 143 deletions(-) create mode 100644 tests/invalid/batch-ctx-token-bad-pattern.json create mode 100644 tests/invalid/batch-ctx-token-non-engagement.json create mode 100644 tests/invalid/batch-engaged-missing-presentation-id.json create mode 100644 tests/invalid/batch-missing-session-mixed-retrieval.json create mode 100644 tests/invalid/citation-id-on-grounded.json create mode 100644 tests/invalid/citation-position-invalid.json create mode 100644 tests/invalid/citation-type-invalid.json create mode 100644 tests/invalid/cited-missing-data.json create mode 100644 tests/invalid/cited-negative-excerpt-chars.json create mode 100644 tests/invalid/ctx-token-on-grounded.json create mode 100644 tests/invalid/duplicate-event-id.json create mode 100644 tests/invalid/empty-output-id.json create mode 100644 tests/invalid/engaged-presentation-content-mismatch.json create mode 100644 tests/invalid/engaged-presentation-id-no-presentations.json create mode 100644 tests/invalid/event-level-ctx-token-bad-pattern.json create mode 100644 tests/invalid/grounded-missing-identifier-batch.json create mode 100644 tests/invalid/grounded-missing-identifier-standalone.json create mode 100644 tests/invalid/grounded-negative-tokens-ingested.json create mode 100644 tests/invalid/grounding-fingerprint-missing-scheme.json create mode 100644 tests/invalid/grounding-missing-data.json create mode 100644 tests/invalid/grounding-provenance-cached-conflict-agent-cached.json create mode 100644 tests/invalid/grounding-provenance-invalid.json create mode 100644 tests/invalid/grounding-scope-invalid.json create mode 100644 tests/invalid/grounding-scope-missing.json create mode 100644 tests/invalid/legacy-schema-version-0-1.json create mode 100644 tests/invalid/malformed-citation-id.json create mode 100644 tests/invalid/malformed-content-hash.json create mode 100644 tests/invalid/malformed-content-urls-cited.json create mode 100644 tests/invalid/malformed-event-id.json create mode 100644 tests/invalid/malformed-presentation-id.json create mode 100644 tests/invalid/malformed-session-id.json create mode 100644 tests/invalid/malformed-started-at.json create mode 100644 tests/invalid/manifest-bad-key-type.json create mode 100644 tests/invalid/manifest-coverage-missing-mode.json create mode 100644 tests/invalid/manifest-ctx-resolution-http.json create mode 100644 tests/invalid/manifest-ctx-resolution-on-content-owner.json create mode 100644 tests/invalid/manifest-domains-on-path-manifest.json create mode 100644 tests/invalid/manifest-duplicate-roles.json create mode 100644 tests/invalid/manifest-empty-roles.json create mode 100644 tests/invalid/manifest-endpoint-http.json create mode 100644 tests/invalid/manifest-id-http.json create mode 100644 tests/invalid/manifest-id-not-a-uri.json create mode 100644 tests/invalid/manifest-id-not-well-known.json create mode 100644 tests/invalid/manifest-key-missing-id.json create mode 100644 tests/invalid/manifest-lookalike-domain.json create mode 100644 tests/invalid/manifest-missing-endpoint.json create mode 100644 tests/invalid/presentation-id-on-cited.json create mode 100644 tests/invalid/presented-citation-content-mismatch.json create mode 100644 tests/invalid/presented-missing-data.json create mode 100644 tests/invalid/privacy-violation-query-at-minimal-batch.json create mode 100644 tests/invalid/privacy-violation-query-at-minimal-standalone.json create mode 100644 tests/invalid/privacy-violation-response-text-at-intent.json create mode 100644 tests/invalid/reproduced-citation-id-unmatched.json create mode 100644 tests/invalid/reproduced-missing-data.json create mode 100644 tests/invalid/retrieved-bad-country.json create mode 100644 tests/invalid/retrieved-missing-source-role.json create mode 100644 tests/invalid/retrieved-response-status-out-of-range.json create mode 100644 tests/invalid/shared-ctx-token-two-presentations.json create mode 100644 tests/invalid/standalone-ctx-token-non-engagement.json create mode 100644 tests/invalid/standalone-malformed-session-id.json create mode 100644 tests/invalid/standalone-missing-schema-version.json create mode 100644 tests/invalid/turn-on-content-event.json create mode 100644 tests/invalid/withdrawn-ip-hash-batch.json create mode 100644 tests/invalid/withdrawn-ip-hash-session.json create mode 100644 tests/valid/event-batch-destination-engagements.json create mode 100644 tests/valid/event-standalone-index.json create mode 100644 tests/valid/event-standalone-origin.json create mode 100644 tests/valid/manifest-platform.json create mode 100644 tests/valid/session-citation-contradiction.json create mode 100644 tests/valid/session-engagement-types.json create mode 100644 tests/valid/session-multi-owner-catalogue.json diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index d8bd18d..2dfce92 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -20,3 +20,5 @@ jobs: run: uv run --with "jsonschema[format-nongpl]" python tests/validate.py - name: Validate worked examples in the spec run: uv run --with "jsonschema[format-nongpl]" python tests/check_examples.py + - name: Replay suite-weakening mutations + run: uv run --with "jsonschema[format-nongpl]" python tests/mutation_smoke.py diff --git a/README.md b/README.md index 607c85b..cccb7d6 100644 --- a/README.md +++ b/README.md @@ -61,7 +61,7 @@ Grounding and presentation record different boundary crossings: grounding means **Post-hoc, not pre-declared.** Events report what actually happened, not what the agent said it would do at request time. An agent cannot reliably declare how it will use content before reading it. -**Observable boundaries, not agent internals.** The six event types mark boundary crossings. What happens between them - the fan-out, relevance evaluation, re-ranking, reasoning chains - is internal to the agent and changes constantly. The spec does not model it. +**Observable boundaries, not agent internals.** The six content event types mark boundary crossings. What happens between them - the fan-out, relevance evaluation, re-ranking, reasoning chains - is internal to the agent and changes constantly. The spec does not model it. **Multiple observers, one event.** A content retrieval can be reported by the content owner's CDN, the content owner's origin server, and the AI agent independently. The `Content-Telemetry-ID` header correlates these into a single corroborated event. Uncorroborated retrievals (no matching agent event) may indicate an agent that does not yet support the telemetry protocol. @@ -206,11 +206,11 @@ This is a preview specification. The following areas are under active discussion **Grounding boundary.** The spec defines grounding as content entering the generation model's context (sections 4.3 and 6.4). For straightforward RAG pipelines this is clear. For pipelines with multiple processing stages - embedding, re-ranking, summarisation before context insertion - the boundary requires judgement. The spec draws the line at the generation context (not earlier retrieval stages), but edge cases remain. When a re-ranking or summarisation stage is itself a generative model, the multi-step rule in section 6.4 (content entering a sub-agent's generation context is grounded) can pull selection stages back inside the boundary. Input from platform engineering teams building real implementations will sharpen this definition. -**Event volume at scale.** A single deep-research query can produce 100+ retrieval events and dozens of grounding/citation events. The session document format already handles transport - one POST with all events after the session ends, not one request per event. Volume management beyond that (storage, processing, consumer-side aggregation) is an implementation concern, not a protocol gap. Sampling and aggregation are options for future versions but are not in v0.1; the standard sets no default for reporting granularity, leaving it to profiles and deployments. +**Event volume at scale.** A single deep-research query can produce 100+ retrieval events and dozens of grounding/citation events. The session document format already handles transport - one POST with all events after the session ends, not one request per event. Volume management beyond that (storage, processing, consumer-side aggregation) is an implementation concern, not a protocol gap. Version 1 adds an explicit coverage declaration - `complete`, `sampled`, `aggregated` or `selected` (section 5.7.6) - and a manifest field for it (section 8.5); the standard still sets no default for reporting granularity, leaving it to profiles and deployments. **Verification of grounding and citation.** Grounding and citation events are reported by the agent, which is also the party that may owe compensation under a licence. In v0.1, manifest signing is informational: consumers may verify signatures but are not required to, and the specification defines no required proof binding an event to its emitter (sections 8.4 and 8.9). The events attribution depends on are therefore self-reported by the reporting party. Verifiable credentials and signed events are deferred (section 8.9). One corroboration mechanism works without signing: the `Content-Telemetry-ID` field correlates an agent-reported retrieval with an origin- or edge-reported one (section 7.2), but it covers retrieval only - grounding, reproduction, citation, presentation, and engagement have no independent observer. Signing, even once required, would prove who reported an event, not that the event is true or that all qualifying events were reported. Input is wanted on what a verification layer should cover and where it belongs. Mechanisms that test truthfulness and completeness rather than origin, such as sampled audits or publisher-seeded canary content, are of particular interest. -**Reporting granularity.** The standard sets no default for reporting granularity, leaving it to profiles and deployments (see *Event volume* above). The SPUR profile requires event-level delivery and does not permit aggregation. The open question is whether the standard should say more about sampling and aggregation so that profiles do not each define it separately, and how event-level delivery scales for the highest-volume case. No mechanism is selected in v0.1. +**Reporting granularity.** The standard sets no default for reporting granularity, leaving it to profiles and deployments (see *Event volume* above). The SPUR profile requires event-level delivery and does not permit aggregation. Version 1 answers the first half of the question: coverage modes are defined once, in section 5.7.6, so that profiles reference them rather than each define their own. How event-level delivery scales for the highest-volume case remains open. ## Versioning diff --git a/SPECIFICATION.md b/SPECIFICATION.md index dd39ecf..7fbdc90 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -330,7 +330,7 @@ Additional content metadata - version, last-modified timestamp, content hash, me | Field | Type | Required | Description | |-------|------|----------|-------------| -| `schema_version` | string | Yes | Schema version (v1 documents declare "1.0") | +| `schema_version` | string | Yes | Schema version as `major.minor` (v1 documents declare "1.0"; see section 12) | | `session_id` | UUID | Yes | Unique session identifier | | `parent_session_id` | UUID | No | Immediate parent session that delegated work to this session | | `agent_id` | string | No | Responding agent identifier | @@ -437,7 +437,7 @@ Content events without a `turn_id` (e.g., `content_grounded` with `scope: sessio #### 5.2.2 Source role -The `source_role` field SHOULD be set on `content_retrieved` events. When multiple systems observe the same retrieval, the `content_telemetry_id` field correlates their events for deduplication. +The `source_role` field MUST be set on `content_retrieved` events (section 5.7.1): without it a consumer cannot tell an agent-reported fetch from an origin- or edge-reported one, or correlate the observers of one retrieval. When multiple systems observe the same retrieval, the `content_telemetry_id` field correlates their events for deduplication. #### 5.2.3 Licence reference @@ -568,7 +568,7 @@ Reproduction, presentation, and engagement events are optional lifecycle signals A conforming **Retrieval** emitter MUST: - Set `source_role` on `content_retrieved` events -- Include at least one of `content_url` or `content_id` on every event +- Include at least one of `content_url` or `content_id` on every content event - Set `type` and `timestamp` on every event This level requires no agent cooperation. Content owners can implement it using CDN edge compute (Cloudflare Workers, Fastly Compute, etc.). @@ -580,7 +580,7 @@ Origin-side emitters operating at the CDN edge SHOULD include `bot_category`, `r A conforming **Grounding** emitter MUST satisfy Retrieval requirements and also: - Produce sessions with `schema_version`, `session_id`, `agent_id`, and `started_at` -- Emit `content_grounded` events with `data.scope` +- Emit `content_grounded` events with `data.scope` (schema-enforced; section 6.4) - Include at least one of `content_url` or `content_id` on every content event - Emit `turn_started` and `turn_completed` events with `privacy_level` - Restrict conversation turn fields to the declared `privacy_level` (section 5.5) @@ -609,7 +609,7 @@ A Citation emitter SHOULD: A conforming **telemetry consumer** MUST: -- Accept sessions with any `schema_version` that shares the same major version. V1 documents declare `schema_version` `"1.0"`. A v1 consumer MUST reject documents declaring `"0.1"`: v0.1 is a different wire version, not a compatible minor. Conversely, a v0.1 consumer following the preview rule (a 0.x consumer accepts only the exact same minor version, so a 0.1 consumer accepts 0.1 only) rejects documents declaring `"1.0"`. The major-version compatibility rule takes effect from 1.0.0 onward. +- Accept documents declaring any `schema_version` with the same major version as the one the consumer implements. `schema_version` is `major.minor` (section 12): v1.0 documents declare `"1.0"`, and the v1.0 schemas accept that value only. Each minor version publishes its own schemas. A consumer implementing 1.y validates a document declaring 1.x, x ≤ y, against the 1.x schemas, and a document declaring a later minor against the latest schemas it implements, tolerating the optional fields that minor added. A v1 consumer MUST reject documents declaring `"0.1"`: v0.1 is a different wire version, not a compatible minor. Conversely, a v0.1 consumer following the preview rule (a 0.x consumer accepts only the exact same minor version, so a 0.1 consumer accepts 0.1 only) rejects documents declaring `"1.0"`. - Tolerate unknown fields without error - Tolerate events from any conformance level - Accept the session-document, standalone-event, and event-batch delivery formats, reconstructing sessions from standalone events and event batches where needed (see section 7.1) @@ -624,6 +624,11 @@ The JSON Schema (`telemetry-session.json`) validates structure and types but can - The conformance-level requirements (sections 5.7.1 to 5.7.3) are cumulative. - When `content_grounded.data.provenance` is `agent_fetched`, `data.cached` MUST be `false`; when it is `agent_cached`, `data.cached` MUST be `true` (section 6.4). - `content_grounded.data.content_fingerprint` MUST NOT contain `preserved_in_output`; output-side reuse is reported with `content_reproduced` (sections 6.4 and 12.1). +- `source_role` MUST be present on every `content_retrieved` event (sections 5.2.2, 5.7.1). +- Fields scoped to an event type MUST NOT appear on other types: `presentation_id` and the event-level `ctx_token` only on `content_engaged`; `citation_id` only on `content_presented` and `content_reproduced`; `turn` only on `turn_started` and `turn_completed` (section 5.2). +- Within a session document, event `id` values MUST be distinct; a `content_engaged.presentation_id` MUST reference a `content_presented` event, and a `citation_id` a `content_cited` event, that identifies the same content - where both events carry `content_id` the values MUST be equal, and likewise for `content_url` (sections 6.7, 6.8); one event-level `ctx_token` MUST NOT appear on engagements bound to two different presentations (section 7.4.1). +- An envelope `ctx_token` (section 7.1) MUST accompany `content_engaged` events only. +- Manifests: `domains` MUST appear only on a manifest served at the domain root, and `telemetry.ctx_resolution` only on a manifest declaring the `agent` or `platform` role (sections 8.5, 8.6). The `tests/` directory provides an informative reference suite for these rules. A consumer that receives a privacy-violating turn (e.g., `query_text` present at `minimal` level) SHOULD strip the offending fields rather than reject the document carrying them. @@ -648,7 +653,7 @@ Whether an emitter's reporting in fact met its declared coverage is the complete ## 6. Data profiles -The `data` field on events carries type-specific metadata. These profiles document the recommended fields by event type and source role. None are required except where a section states otherwise (`reproduction_type` in 6.6; `presentation_kind` and `presentation_type` in 6.7), but emitting them enables richer attribution. +The `data` field on events carries type-specific metadata. These profiles document the recommended fields by event type and source role. None are required except where a section states otherwise - `scope` in 6.4, `citation_type` in 6.5, `reproduction_type` in 6.6, `presentation_kind` and `presentation_type` in 6.7, each enforced by the JSON Schema - but emitting them enables richer attribution. ### 6.1 Retrieved content metadata (`content_retrieved`) @@ -710,7 +715,7 @@ Emitting a `training`-category `content_retrieved` event is permitted but non-at | Field | Type | Description | |-------|------|-------------| -| `scope` | string | Influence scope: `session` or `turn` (see below) | +| `scope` | string | Required. Influence scope: `session` or `turn` (see below) | | `cached` | boolean | Content served from agent-side cache rather than a live fetch | | `provenance` | string | How content reached the context: `agent_fetched`, `agent_cached`, or `third_party_sourced` (see below) | | `chars_ingested` | integer | Character count of content placed in the generation context (see below) | @@ -725,7 +730,7 @@ Both fields measure the content actually placed in the generation model's contex `chars_ingested` counts Unicode code points in the exact text placed in context. Count the string as ingested: an emitter MUST NOT apply Unicode normalisation solely to calculate this field. It is the portable measure: two emitters that ingest the same code-point sequence agree, so a content owner can compare volumes across agents and over time without knowing which model produced the number. Different normalised representations remain different ingested sequences and may therefore produce different counts. -`tokens_ingested` counts the same content in the generation model's tokeniser (the model identified in `model_id` on the corresponding `turn_completed` event), not the retrieval or embedding model's tokeniser. It is supplementary. Token counts are model-specific, change when a vendor revises a tokeniser, and are not comparable between agents, so a consumer cannot aggregate them across emitters or treat a difference as a difference in volume. Emitters SHOULD send `chars_ingested` where they send `tokens_ingested`, and consumers that receive only token counts SHOULD record which model produced them. +`tokens_ingested` counts the same content in the generation model's tokeniser (the model identified in `model_id` on the corresponding `turn_completed` event), not the retrieval or embedding model's tokeniser. It is supplementary. Token counts are model-specific, change when a vendor revises a tokeniser, and are not comparable between agents, so a consumer cannot aggregate them across emitters or treat a difference as a difference in volume. Emitters SHOULD send `chars_ingested` where they send `tokens_ingested`, and consumers that receive only token counts SHOULD record which model produced them. At `minimal` privacy the turn carries no `model_id` (section 5.5); token counts reported at that level name no tokeniser, and consumers SHOULD NOT compare them with counts from any other emitter or model. #### Provenance and content fingerprints @@ -764,6 +769,8 @@ Output-side reuse is reported with `content_reproduced`, not a fingerprint-prese For session-scoped grounding, the number of turns influenced is derivable from the session's `turn_started` events following the grounding event. This avoids redundant per-turn grounding events for content that persists across responses. +`scope` is required on every `content_grounded` event and the JSON Schema enforces it: the occurrence boundary of section 4.3 and every counting model in section 10 depend on knowing whether a grounding informed one response or the rest of the session. + #### Agent architecture and the grounding boundary The grounding event marks the point where content enters the generation model's context - the boundary where content can directly influence the model's output text. Content used only for retrieval selection (embedding similarity search, re-ranking, query routing) without entering the generation context is not grounded. @@ -803,7 +810,7 @@ Agents SHOULD preserve the `license_ref` from the original retrieval when emitti | Field | Type | Description | |-------|------|-------------| -| `citation_type` | string | How content was used: `direct_quote`, `paraphrase`, `reference`, `contradiction`, `unclassified` | +| `citation_type` | string | Required. How content was used: `direct_quote`, `paraphrase`, `reference`, `contradiction`, `unclassified` | | `media_type` | string | Content medium: `text`, `image`, `video`, `audio` (open vocabulary, see 6.1) | | `excerpt_tokens` | integer | Token count of the excerpt used | | `excerpt_chars` | integer | Character count of the excerpt used | @@ -812,7 +819,7 @@ Agents SHOULD preserve the `license_ref` from the original retrieval when emitti | `content_hash` | string | SHA-256 matching the corresponding `content_grounded` event (`sha256:{hex}`). When the agent chunked the source, this is the chunk hash, not the full document hash. | | `url_verified` | boolean | Whether the cited URL was verified to resolve to matching content | -A citation MUST carry a resolvable source reference: a non-null `content_url` or `content_id` at the event level. This is what distinguishes a citation from vague attribution - the credit names a source that owner routing (section 7.3) can resolve. The JSON Schema enforces this for `content_cited` events; an association the emitter cannot resolve to a URL or identifier is not reportable as a citation. This is stricter than the application-layer identifier rule that applies to content events generally (section 5.7.5). +A citation MUST carry a resolvable source reference: a non-null `content_url` or `content_id` at the event level. This is what distinguishes a citation from vague attribution - the credit names a source that owner routing (section 7.3) can resolve. The JSON Schema enforces this for `content_cited` events; an association the emitter cannot resolve to a URL or identifier is not reportable as a citation. This is stricter than the application-layer identifier rule that applies to content events generally (section 5.7.5). `citation_type` is likewise required and schema-enforced; an emitter that cannot classify a citation uses `unclassified` rather than omitting the field. `media_type` identifies the content medium. Defaults to `text` when absent. @@ -881,7 +888,7 @@ For non-text media, `media_type` and `reproduced_hash` identify the reproduced m These are the core values. Platforms with additional presentation surfaces MAY use custom string values. Telemetry consumers MUST tolerate unknown `presentation_type` values. -Each presentation event MUST have an `id` and `output_id`. When it presents a citation, `citation_id` references that `content_cited` event's `id`; an uncited presentation omits `citation_id`. Repeated presentations of the same source or output element receive distinct event IDs. This allows a later `content_engaged.presentation_id` to identify the exact surface occurrence rather than matching only by URL. +Each presentation event MUST have an `id` and `output_id`. When it presents a citation, `citation_id` references that `content_cited` event's `id`, and the two events identify the same content; an uncited presentation omits `citation_id`. Repeated presentations of the same source or output element MUST receive distinct event IDs - event `id` values are unique within a session document. This allows a later `content_engaged.presentation_id` to identify the exact surface occurrence rather than matching only by URL. When a session includes `content_presented` events but no subsequent `content_engaged` events, the telemetry establishes only that content or a reference was made perceivable and no reported interaction followed. It does not establish human attention. Whether this pattern is meaningful depends on the governing terms. Retrieval remains the only lifecycle stage observable from the CDN edge. @@ -891,7 +898,7 @@ When a session includes `content_presented` events but no subsequent `content_en |-------|------|-------------| | `engagement_type` | string | Type of interaction (see below) | -The content URL is identified by the event-level `content_url` field (section 5.2), not duplicated in `data`. Every agent-reported engagement MUST carry `presentation_id`, referencing the exact `content_presented.id` on which the action occurred. Matching on URL alone is insufficient because the same source reference can be presented more than once. A destination-reported engagement carries `ctx_token` on its envelope instead: the destination cannot know the presentation UUID, and the telemetry consumer restores the binding from the token at resolution (section 7.4). +The content URL is identified by the event-level `content_url` field (section 5.2), not duplicated in `data`. Every agent-reported engagement MUST carry `presentation_id`, referencing the exact `content_presented.id` on which the action occurred, and identifies the same content as that presentation (section 5.7.5). Matching on URL alone is insufficient because the same source reference can be presented more than once. A destination-reported engagement carries `ctx_token` on its envelope instead: the destination cannot know the presentation UUID, and the telemetry consumer restores the binding from the token at resolution (section 7.4). #### Engagement types @@ -907,6 +914,8 @@ These are the core values. Extensions MAY define additional engagement actions - `agent_navigate` is the agent-mediated counterpart of a click: the user reached the source through the agent rather than through a browser. Consumers measuring traffic SHOULD count it alongside `link_click`, distinguishing the two where the commercial agreement does. +An action that touches several presentations at once - a `share` of a response containing three source cards - is one engagement occurrence per presentation shared (section 4.3), each bound to its own `presentation_id`; a surface that shares a single card reports one. An `agent_navigate` to a URL the recipient supplied, which no presentation made perceivable, is not an engagement: it begins with a `content_retrieved` event like any other fetch. + `link_click` is the primary signal for clickthrough rate calculation. Telemetry consumers can derive per-content-owner and aggregate clickthrough rates from `link_click` engagements and link presentations, joining each engagement through `presentation_id` rather than URL alone. A `link_click` or `agent_navigate` engagement reported from the landing page after a click-out crosses a trust boundary. Such events carry a `ctx_token` in place of `session_id`, which the telemetry consumer resolves to the click context (see section 7.4). The agent-authored engagement itself reaches the engaged content's owner through owner-scoped routing whether or not the token survived the redirect chain (section 7.4.5). @@ -980,7 +989,7 @@ Session documents use `"document_type": "session"`. When `document_type` is abse For origin-side emitters at Retrieval conformance level, `session_id` MAY be omitted when the content owner has no session context. Telemetry consumers correlate these events with agent-reported sessions using the `content_telemetry_id` field. -For `content_engaged` events emitted from a landing page after a click-out (typically by a content marketplace, affiliate network, or destination site), `session_id` MAY be replaced by a `ctx_token` field that carries an opaque click token issued by the originating agent. This lets a downstream observer report a corroborating engagement event without sharing the session UUID across trust boundaries. An event MUST carry either `session_id` or `ctx_token` at Grounding conformance and above. Token issuance, carriage, binding, resolver discovery, and the resolution response are defined in section 7.4. +For `content_engaged` events emitted from a landing page after a click-out (typically by a content marketplace, affiliate network, or destination site), `session_id` MAY be replaced by a `ctx_token` field that carries an opaque click token issued by the originating agent. This lets a downstream observer report a corroborating engagement event without sharing the session UUID across trust boundaries. An event MUST carry either `session_id` or `ctx_token` at Grounding conformance and above, and an envelope `ctx_token` accompanies `content_engaged` events only: an envelope carrying any other event type carries `session_id`. Token issuance, carriage, binding, resolver discovery, and the resolution response are defined in section 7.4. The primary schema (`telemetry-session.json`) validates session documents. A standalone event envelope schema (`telemetry-event.json`) validates the event delivery format, and a batch envelope schema (`telemetry-event-batch.json`) validates the event batch format. All three schemas share the `TelemetryEvent` definition. @@ -1047,7 +1056,7 @@ A click-out is the moment content usage becomes traffic the destination can obse #### 7.4.1 Token issuance -A `ctx_token` is an opaque token minted by the originating agent. Its value MUST match `^ct_[A-Za-z0-9_-]{8,240}$` and MUST NOT encode content, session, or user identifiers recoverable without the issuer's state. +A `ctx_token` is an opaque token minted by the originating agent. Its value MUST match `^ct_[A-Za-z0-9_-]{16,240}$`, MUST be unguessable - the suffix is drawn from at least 96 bits of cryptographically secure randomness, or is a keyed construction of equivalent strength, so that holding one token gives no way to derive or enumerate another - and MUST NOT encode content, session, or user identifiers recoverable without the issuer's state. A token MUST be bound to exactly one `content_presented` occurrence at mint time. The same URL presented twice receives two tokens; a token observed on two presentations is malformed issuance and consumers MUST NOT resolve it. Surfaces that route outbound navigation through the agent SHOULD mint per click, additionally binding the token to the resulting `content_engaged` event. Direct-link surfaces mint per presentation; repeated clicks on one presentation then share a token, and are distinguished at resolution by the destination's event timestamps. @@ -1074,7 +1083,7 @@ A telemetry consumer that supports resolution exposes, for a presented token, th 3. **The contributing sources, gated per owner.** The sources that informed the response the click came from: `content_grounded` events in scope for the engaged presentation's turn (including session-scoped groundings) and that turn's `content_cited` and `content_presented` events, for content other than the clicked content. A contributing owner's events appear only when that owner has opted in to contributing-source disclosure with the resolving consumer; owners without a recorded opt-in are visible only through the counts in the session summary. This is the component that supports multi-citation attribution when the clicked content is not the contributing content - a click through to a commerce destination whose recommendation a publisher's review produced - and it restores the consent-gated role of the v0.1 click manifest (section 12.1), scoped to the click's provenance rather than the whole session. 4. **An optional privacy-bounded session summary.** Event counts by type and a distinct-source count. Counts, not events, and no content identifiers of owners not disclosed above. -A resolution response MUST NOT include the resolved session's raw `session_id`: the token exists so the session UUID never crosses the trust boundary. Outside the contributing-source component, a resolution response MUST NOT include events for content other than the clicked content; within it, a response MUST NOT include events for an owner without a recorded contributing-source opt-in. Whole-session cross-content detail remains a reporting concern, delivered through publisher-filtered views (section 7.3), not through per-click resolution. A consumer MUST resolve a token only when the issuing agent has opted in to click-token resolution, and the response is gated by the resolved session's `privacy_level`. The mechanisms by which the issuer and contributing-owner opt-ins are recorded are operator-defined; both gates are normative. +A resolution response MUST NOT include the resolved session's raw `session_id`: the token exists so the session UUID never crosses the trust boundary. Outside the contributing-source component, a resolution response MUST NOT include events for content other than the clicked content; within it, a response MUST NOT include events for an owner without a recorded contributing-source opt-in. Whole-session cross-content detail remains a reporting concern, delivered through publisher-filtered views (section 7.3), not through per-click resolution. A consumer MUST resolve a token only when the issuing agent has opted in to click-token resolution. The response is further gated by the `privacy_level` of the turn the engaged presentation belongs to - the turn named by the presentation's `turn_id`, or the turn in progress at its timestamp: at `minimal` the consumer returns the engagement and the lineage (components 1 and 2) only, withholding the contributing-source component and the session summary; at `intent` and above all four components are available. A resolution response never carries conversation-turn fields at any level. The mechanisms by which the issuer and contributing-owner opt-ins are recorded are operator-defined; all three gates are normative. A worked resolution response (informative): @@ -1111,12 +1120,14 @@ The response shape above is informative in v1; the constraints in this section a The agent-authored click `content_engaged` is a session event like any other: it reaches content owners through routing and aggregation (section 7.3), independent of whether the token in the URL survived the redirect chain. A telemetry consumer that provides owner-filtered views MUST include the agent-authored `content_engaged` event in the filtered view of the engaged content's owner, on the same terms as `content_grounded` and `content_cited` events - including the session identifier that owner-scoped delivery carries. The `session_id` prohibition in section 7.4.4 binds token resolution, where the requesting party is authenticated by nothing more than possession of a URL-carried value; it does not bind section 7.3 delivery to an owner whose domain registration the consumer has verified. -This delivery is deliberately redundant with the token path. It notifies the destination owner of the click even when `ctx_token` was stripped in transit; it lets that owner join the click to their own grounded and cited events on `session_id` without calling a resolver; and it lets a party processing owner-scoped streams for both a contributing publisher and a click destination match its clients' events on `session_id` for attribution. The owner's filtered view SHOULD carry the event-level `ctx_token`, so a destination that captured the query parameters at landing can join the URL-channel observation to the server-side event directly. This is not token distribution to contributing owners (the limit recorded in section 7.4.6): only the owner of the clicked content receives the event, and that owner already saw the token in the URL. +This delivery is deliberately redundant with the token path. It notifies the destination owner of the click even when `ctx_token` was stripped in transit; it lets that owner join the click to their own grounded and cited events on `session_id` without calling a resolver; and it lets a party processing owner-scoped streams for both a contributing publisher and a click destination match its clients' events on `session_id` for attribution. The owner's filtered view SHOULD carry the event-level `ctx_token`, so a destination that captured the query parameters at landing can join the URL-channel observation to the server-side event directly. This is not token distribution to contributing owners (the limit recorded in section 7.4.6): only the owner of the clicked content receives the event, and that owner already saw the token in the URL. Where the clicked content's owner is also the destination, that owner therefore holds both the token from the URL and the session identifier from this delivery: the boundary of section 7.4.4 protects sessions from unregistered holders of a URL, not from the verified owner of the content that was clicked. #### 7.4.6 Recorded limit: consumer custody Resolution depends on the telemetry consumer the agent chose, because that consumer holds the session. Grounding and citation events precede the click and cannot carry its later token, and distributing tokens to every contributing content owner after the fact would weaken the privacy boundary this section maintains. Publisher-derived tokens would require a federation and key-management design; that belongs in a later attribution or evidence profile. Core v1 mitigates the dependency with resolver discoverability (7.4.3) and exact click binding (7.4.1). +Token lifetime and requester authentication are likewise not defined in v1: a token resolves for as long as the consumer retains the session, and the resolver authenticates the requester by possession of the token alone (section 7.4.5). A resolution window after issuance, and requester credentials - for example authenticating a destination against the manifest its domain serves (section 8) - belong to the evidence profile, together with the federation design above. + ## 8. Manifest Content owners, agents, and platforms publish a manifest declaring their identity and telemetry endpoints. The `manifest_ref` field on session documents (5.1.2) and the routing logic for origin-side emitters (7.3) resolve to manifests defined in this section. @@ -1138,7 +1149,7 @@ https://example.com/agents/search/.well-known/content-telemetry.json # operated Each manifest is self-contained at its own well-known URL. -Trust derives from TLS and DNS control of the domain. Manifests are unsigned in v0.1. +Trust derives from TLS and DNS control of the domain. Manifests are unsigned in v1 (section 8.9). ### 8.2 Schema @@ -1147,7 +1158,7 @@ Machine-readable schema: [`./manifest.json`](./manifest.json) (JSON Schema draft | Field | Type | Required | Description | |-------|------|----------|-------------| | `schema_version` | string | Yes | Manifest schema version. v1 emitters MUST use `"1.0"`. | -| `id` | string | Yes | The manifest's canonical URL (e.g. `https://example.com/.well-known/content-telemetry.json`). | +| `id` | string | Yes | The manifest's canonical `https://` URL, ending in `/.well-known/content-telemetry.json` (e.g. `https://example.com/.well-known/content-telemetry.json`); the schema rejects other schemes and locations (section 8.1). | | `roles` | string[] | Yes | One or more of `content_owner`, `agent`, `platform`. | | `operator` | object | Yes | Operating organisation (see 8.3). | | `keys` | object[] | No | Public keys for signing telemetry events (see 8.4). | @@ -1167,12 +1178,12 @@ A manifest MAY declare multiple roles (e.g. `["content_owner", "agent"]`). A mor ### 8.4 Keys -Public keys used to sign telemetry events emitted by this participant. Per-event signing is informational in v0.1; consumers MAY verify signatures but are not required to. +Public keys used to sign telemetry events emitted by this participant. Per-event signing remains informational in v1; consumers MAY verify signatures but are not required to. | Field | Type | Required | Description | |-------|------|----------|-------------| | `id` | string | Yes | Key identifier, unique within the manifest. | -| `type` | string | Yes | Key type. v0.1: `Ed25519`. | +| `type` | string | Yes | Key type. v1 defines `Ed25519` only. | | `publicKey` | string | Yes | Multibase-encoded public key (multicodec prefix, base58btc - the same format as `did:key`). | | `expires` | datetime | No | ISO 8601 expiry. | @@ -1180,7 +1191,7 @@ Public keys used to sign telemetry events emitted by this participant. Per-event | Field | Type | Required | Description | |-------|------|----------|-------------| -| `endpoint` | string | Yes | HTTPS URL. For agents and platforms, the outbound submission endpoint. For content owners, the inbound destination for events about the content owner's content. | +| `endpoint` | string | Yes | HTTPS URL (schema-enforced). For agents and platforms, the outbound submission endpoint. For content owners, the inbound destination for events about the content owner's content. | | `conformance_level` | string | No | Conformance level advertised by this participant's own emitter(s). One of `retrieval`, `grounding`, `citation` (see 5.7). | | `ctx_resolution` | string | No | HTTPS URL of the click-token resolution endpoint operated by or for this participant (see 7.4). Valid on `agent` and `platform` manifests. | | `coverage` | object | No | Per-event-type coverage declaration: a map from event type to `{ "mode": …, "terms_ref": … }`, where `mode` is one of `complete`, `sampled`, `aggregated`, `selected` (see 5.7.6) and `terms_ref` optionally names the terms stating the rule or condition. | @@ -1193,9 +1204,9 @@ Public keys used to sign telemetry events emitted by this participant. Per-event The `domains` array MAY appear only on manifests served from the domain root (`https:///.well-known/content-telemetry.json`). Manifests under path prefixes MUST NOT include `domains`. -In v0.1, every entry in `domains` MUST be self-validating: either the manifest's own host, or a subdomain of it (literal `news.example.com` or wildcard `*.example.com`). Control of the apex - proven by serving the manifest at the apex over TLS - implies DNS control of subdomains, so no further validation is needed. A manifest containing entries that are not subdomains of its own host is malformed. +In v1, every entry in `domains` MUST be self-validating: either the manifest's own host, or a subdomain of it (literal `news.example.com` or wildcard `*.example.com`). Control of the apex - proven by serving the manifest at the apex over TLS - implies DNS control of subdomains, so no further validation is needed. A manifest containing entries that are not subdomains of its own host is malformed. -This keeps the v0.1 protocol fully decentralised: every manifest is a self-contained credential, validated by TLS plus the well-known location, with no dependency on consumer-side validation state or any external registry. Cross-apex claims (one operator unifying several unrelated apex domains in a single manifest) are deferred to a later version. +This keeps the v1 protocol fully decentralised: every manifest is a self-contained credential, validated by TLS plus the well-known location, with no dependency on consumer-side validation state or any external registry. Cross-apex claims (one operator unifying several unrelated apex domains in a single manifest) are deferred to a later version. ### 8.7 Consumer behaviour @@ -1203,10 +1214,10 @@ When resolving a manifest from `manifest_ref`, a `content_url` domain, or any ot - **404 or network error.** Treat the participant as unverified. Do not reject telemetry events on this basis alone. - **Invalid JSON or schema validation failure.** Reject the manifest. Treat the participant as unverified. -- **Unknown `schema_version`.** During the v0.x preview period, consumers MUST accept only the exact same minor version. The semver-major compatibility rule applies from 1.0.0 onward (see section 12). +- **Unknown `schema_version`.** Manifests follow the same rule as telemetry documents (section 5.7.4): accept any `1.x` the consumer implements, validating against that minor's schema, and reject `0.x` manifests. During the v0.x preview period consumers accepted only the exact same minor version. - **Duplicate `keys[].id`.** Reject the manifest. - **`domains` entry that is not the manifest's host or a subdomain of it.** Reject the manifest as malformed (see 8.6). -- **Missing `keys` on a manifest referenced by `manifest_ref`.** Not an error in v0.1, since signing is informational. +- **Missing `keys` on a manifest referenced by `manifest_ref`.** Not an error in v1, since signing is informational. Consumers SHOULD cache resolved manifests respecting the response's `Cache-Control` headers. Manifest hosts SHOULD set `Cache-Control: max-age=3600` during onboarding and `max-age=86400` steady-state. @@ -1280,7 +1291,7 @@ Consumers SHOULD cache resolved manifests respecting the response's `Cache-Contr The two manifests live independently at distinct well-known URLs. The content-owner manifest's `domains` and `telemetry` apply to publisher.com's content; the agent manifest's `keys` and `telemetry` apply to events emitted by the assistant. -### 8.9 Out of scope for v0.1 +### 8.9 Out of scope for v1 The following are deferred to later versions: @@ -1290,7 +1301,7 @@ The following are deferred to later versions: - Deployment context, purpose, brand affiliation - Revocation registries - Key rotation procedures beyond the `expires` field -- `did:web` compatibility (the `id` field uses the manifest URL in v0.1) +- `did:web` compatibility (the `id` field uses the manifest URL in v1) - Cross-apex claims (one operator unifying several unrelated apex domains in a single manifest) --- @@ -1309,6 +1320,8 @@ Hashing does not anonymise a value drawn from a space small enough to enumerate. The schemas cannot enforce this v1 migration rule: event `data` accepts additional properties by design, so `ip_hash` would otherwise validate as an ordinary extension. The conformance suite therefore checks this specific prohibition at the application layer. This does not establish a general registry of withdrawn extension names. +`privacy_level` gates the named conversation-turn fields of section 5.5 and nothing else. Extension fields in `turn` or `data`, content identifiers and URLs, `license_ref`, `terms_ref`, `output_id`, `turn_id` and every other opaque string pass through at every level, so their contents are the emitter's responsibility: an emitter MUST NOT use them to carry the end user's identity, or the query and response text that the declared level withholds, and SHOULD apply the minimisation guidance above to them as it does to the named fields. + ### 9.2 Recommended levels | Scenario | Recommended level | @@ -1452,6 +1465,24 @@ exists and `citation_id` when the presentation carries a citation. For every event. Do not migrate clicks by matching URL alone: repeated presentations of the same URL are distinct occurrences. +V1 requires `data.citation_type` on every `content_cited` event and `data.scope` +on every `content_grounded` event; both are schema-enforced. A v0.1 emitter that +omitted them migrates a citation it cannot classify with `citation_type: +unclassified`, and a grounding whose scope it did not record with `scope: turn` +where the event carries a `turn_id` and `scope: session` otherwise. V1 also +rejects a `content_cited` or `content_reproduced` event whose `content_url` and +`content_id` are both absent or null (sections 6.5, 6.6): a v0.1 citation with no +resolvable reference is not migrated as a citation. + +V1 withdraws `ip_hash` from the edge and origin retrieval profiles (section 9.1). +Emitters remove the field and MUST NOT populate it; `asn`, `asn_org` and `country` +remain. V1 also requires `source_role` on every `content_retrieved` event +(section 5.7.1); preview emitters that omitted it add the role they report under. + +`license_ref` keeps its wire form but no longer asserts that the use was licensed +(section 5.2.3): a consumer that read a v0.1 `license_ref` as verification of +entitlement now reads it as the emitter's claim about which grant applied. + V1 adds `content_reproduced`. The v0.1 preview has no equivalent: verbatim reuse was inferable only from `direct_quote` citations, which conflate the credit with the material and cannot record an uncredited copy. When migrating historical @@ -1459,14 +1490,16 @@ v0.1 data, consumers MAY treat a `direct_quote` citation as an implied reproduction of its excerpt. V1 emitters record reproduction explicitly and SHOULD NOT rely on that inference. -V1 grounding fingerprints report detection only. A preview implementation that -used `data.content_fingerprint.preserved_in_output` removes that field during -migration. When its value was `true` and the historical record contains enough -information to populate every required `content_reproduced` field, the -implementation MAY also create the corresponding reproduction event. It MUST -NOT synthesize a reproduction event from `false` or incomplete historical data. -A grounding event MAY retain `content_fingerprint.scheme`, `detected`, and a -scheme-defined `value`. +V1 grounding fingerprints report detection only. The published v0.1 preview +defined no `content_fingerprint` object; the object, and a +`preserved_in_output` field within it, appeared only on the pre-release +`v1-draft` line. An implementation built against that draft removes +`preserved_in_output` during migration. When its value was `true` and the +historical record contains enough information to populate every required +`content_reproduced` field, the implementation MAY also create the corresponding +reproduction event. It MUST NOT synthesize a reproduction event from `false` or +incomplete historical data. A grounding event MAY retain +`content_fingerprint.scheme`, `detected`, and a scheme-defined `value`. V1 narrows `ctx_token` resolution. The v0.1 click manifest returned every source that informed the resolved session, gated by per-owner opt-in; the v1 @@ -1476,19 +1509,20 @@ per-owner opt-in gate - scoped to the turn the click came from rather than the whole session - and at most a count-based session summary. Consumers implementing v0.1 resolution narrow the contributing-source scope accordingly and MUST NOT return events for owners without a recorded opt-in. Token values -gain the `ct_` pattern; presentation binding moves from the URL-carried -`presentation_id` a destination could never legitimately know to issuer state -restored at resolution. +gain the `ct_` pattern and the unguessability rule of section 7.4.1. The binding +from a token to the presentation it was minted for is issuer state: v0.1 defined +no `presentation_id`, and v1 never places one in a URL; a destination reports +the token, and the consumer restores the binding at resolution. V1 tightens occurrence boundaries (section 4.3). Each core event now has a stated occurrence and cardinality: retrieval per completed fetch (a cache serve is not a retrieval), grounding per distinct content item per declared scope (chunk-level events deduplicate to one occurrence by content identity), reproduction per source item per output element, presentation per rendering occurrence, engagement per observed action. Preview emitters that emitted per chunk, per passage, or re-emitted `content_retrieved` on cache serves remain schema-valid but SHOULD re-map to the stated boundaries; consumers comparing preview and v1 volumes should expect counts to shift where emitters previously chose finer or coarser units. Coverage becomes an explicit declaration (section 5.7.6) rather than an implication of conformance level. Session-scoped extension metadata belongs in the session-level `data` container (section 5.1.3); custom top-level siblings of `events`, accepted by the preview schema, are undefined. -Preview versions (0.x) use two-component version numbers. From 1.0.0 onward, versions follow [semantic versioning](https://semver.org/): +Version numbers are `major.minor`, and a document declares the version it was produced under in `schema_version`. From 1.0 onward: -- **Major** (1.0.0 → 2.0.0) - breaking changes to required fields -- **Minor** (1.0.0 → 1.1.0) - new optional fields, new event types -- **Patch** (1.0.0 → 1.0.1) - clarifications +- **Major** (1.0 → 2.0) - breaking changes to required fields or to the meaning of an event +- **Minor** (1.0 → 1.1) - new optional fields and new event types; each minor version publishes its own schemas +- **Patch** - clarifications and errata to prose, examples and fixtures that change no field or constraint; a patch does not change `schema_version` -From 1.0.0 onward, consumers SHOULD accept sessions with any compatible minor version (same major version). During the preview period (0.x) the stricter rule in section 5.7.4 applies: a consumer accepts only the exact same minor version (a 0.1 consumer accepts 0.1 only). +From 1.0 onward, consumers accept documents with any compatible minor version (same major version) as described in section 5.7.4. During the preview period (0.x) the stricter rule applied: a consumer accepted only the exact same minor version (a 0.1 consumer accepts 0.1 only). ## Annex A (normative): JSON Schema @@ -1820,3 +1854,113 @@ The same turn from B.3 at `minimal` privacy. No intent, no topics, no platform m ``` Compare with the `intent` version in B.3: `query_intent`, `topics`, `response_type`, `response_mode`, and `ad_rendered` are all absent. + +### B.5 Multi-owner catalogue under one agreement + +A marketplace intermediary delivers content from many publishers under a single agreement. `content_scope` identifies the agreement, and is the same for every session reported under it; content owner resolution is per event, from each event's `content_url` domain or registered `content_id` prefix (section 7.3). One session, two owners, each event resolving to its own owner. + +```json +{ + "schema_version": "1.0", + "session_id": "990e8400-e29b-41d4-a716-446655440500", + "agent_id": "research-assistant-v5", + "content_scope": "marketplace-agreement-2026-017", + "manifest_ref": "https://assistant.example.com/.well-known/content-telemetry.json", + "started_at": "2026-08-20T09:00:00Z", + "ended_at": "2026-08-20T09:00:09Z", + "events": [ + { + "type": "turn_started", + "timestamp": "2026-08-20T09:00:00Z", + "turn_id": "1", + "turn": { "privacy_level": "intent", "query_intent": "comparison", "topics": ["electric vehicles", "charging"] } + }, + { + "type": "content_retrieved", + "timestamp": "2026-08-20T09:00:01Z", + "source_role": "index", + "content_telemetry_id": "990e8400-e29b-41d4-a716-446655440501", + "content_url": "https://www.autoreview.example/ev/charging-networks-2026", + "content_id": "mkt:autoreview:88213", + "license_ref": "agreement-2026-017:autoreview", + "data": { "media_type": "text", "content_depth": "full" } + }, + { + "type": "content_retrieved", + "timestamp": "2026-08-20T09:00:01Z", + "source_role": "index", + "content_telemetry_id": "990e8400-e29b-41d4-a716-446655440502", + "content_id": "mkt:gridnews:5520", + "license_ref": "agreement-2026-017:gridnews", + "data": { "media_type": "text", "content_depth": "full" } + }, + { + "type": "content_grounded", + "timestamp": "2026-08-20T09:00:02Z", + "turn_id": "1", + "content_url": "https://www.autoreview.example/ev/charging-networks-2026", + "content_id": "mkt:autoreview:88213", + "data": { "scope": "turn", "cached": false, "provenance": "third_party_sourced", "chars_ingested": 11200 } + }, + { + "type": "content_grounded", + "timestamp": "2026-08-20T09:00:02Z", + "turn_id": "1", + "content_id": "mkt:gridnews:5520", + "data": { "scope": "turn", "cached": false, "provenance": "third_party_sourced", "chars_ingested": 6400 } + }, + { + "id": "990e8400-e29b-41d4-a716-446655440503", + "type": "content_cited", + "timestamp": "2026-08-20T09:00:06Z", + "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:networks:1", + "content_url": "https://www.autoreview.example/ev/charging-networks-2026", + "content_id": "mkt:autoreview:88213", + "data": { "citation_type": "paraphrase", "position": "primary" } + }, + { + "id": "990e8400-e29b-41d4-a716-446655440504", + "type": "content_cited", + "timestamp": "2026-08-20T09:00:06Z", + "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:tariffs:1", + "content_id": "mkt:gridnews:5520", + "data": { "citation_type": "reference", "position": "supporting" } + }, + { + "id": "990e8400-e29b-41d4-a716-446655440505", + "type": "content_presented", + "timestamp": "2026-08-20T09:00:07Z", + "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:networks:1", + "citation_id": "990e8400-e29b-41d4-a716-446655440503", + "content_url": "https://www.autoreview.example/ev/charging-networks-2026", + "content_id": "mkt:autoreview:88213", + "data": { "presentation_kind": "source_reference", "presentation_type": "link" } + }, + { + "id": "990e8400-e29b-41d4-a716-446655440506", + "type": "content_presented", + "timestamp": "2026-08-20T09:00:07Z", + "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:tariffs:1", + "citation_id": "990e8400-e29b-41d4-a716-446655440504", + "content_id": "mkt:gridnews:5520", + "data": { "presentation_kind": "source_reference", "presentation_type": "card" } + }, + { + "type": "turn_completed", + "timestamp": "2026-08-20T09:00:09Z", + "turn_id": "1", + "turn": { "privacy_level": "intent", "response_type": "comparison", "response_mode": "standard", "response_tokens": 410 } + } + ] +} +``` + +The telemetry consumer resolves `autoreview.example` by domain registration and the `mkt:gridnews:` prefix by identifier registration, and produces two owner-filtered views: AutoReview sees its retrieval, grounding, citation and link presentation; GridNews sees its own four events and nothing of AutoReview's. Neither view carries the other owner's identifiers, and both carry the shared `content_scope` so the marketplace can reconcile the session against the agreement. The second retrieval has no `content_url` at all - marketplace API content with no canonical URL - and resolves by `content_id` alone. diff --git a/manifest.json b/manifest.json index c2215ca..9ed6b7c 100644 --- a/manifest.json +++ b/manifest.json @@ -14,7 +14,8 @@ "id": { "type": "string", "format": "uri", - "description": "The manifest's canonical URL, e.g. https://example.com/.well-known/content-telemetry.json. (Section 8.2)" + "pattern": "^https://[^?#]+/\\.well-known/content-telemetry\\.json$", + "description": "The manifest's canonical https URL, ending in /.well-known/content-telemetry.json, e.g. https://example.com/.well-known/content-telemetry.json. (Section 8.2)" }, "roles": { "type": "array", @@ -76,7 +77,8 @@ "endpoint": { "type": "string", "format": "uri", - "description": "HTTPS URL. For agents and platforms, the outbound submission endpoint. For content owners, the inbound destination for events about the owner's content. (Section 8.5)" + "pattern": "^https://", + "description": "HTTPS URL (schema-enforced). For agents and platforms, the outbound submission endpoint. For content owners, the inbound destination for events about the owner's content. (Section 8.5)" }, "conformance_level": { "type": "string", @@ -87,7 +89,7 @@ "type": "string", "format": "uri", "pattern": "^https://", - "description": "HTTPS URL of the click-token resolution endpoint operated by or for this participant. Valid on agent and platform manifests. Destinations resolve ctx_iss to this manifest and present the token here. (Section 7.4)" + "description": "HTTPS URL of the click-token resolution endpoint operated by or for this participant. Valid on agent and platform manifests (enforced at the application layer). Destinations resolve ctx_iss to this manifest and present the token here. (Section 7.4)" }, "coverage": { "type": "object", @@ -116,7 +118,7 @@ "items": { "type": "string" }, - "description": "Domains the participant claims authority over. MAY appear only on root manifests; each entry MUST be the manifest's own host or a subdomain of it (literal or wildcard). (Section 8.6)" + "description": "Domains the participant claims authority over. MAY appear only on root manifests (enforced at the application layer); each entry MUST be the manifest's own host or a subdomain of it (literal or wildcard). (Section 8.6)" } } } diff --git a/telemetry-event-batch.json b/telemetry-event-batch.json index b7a9754..4e0b0ea 100644 --- a/telemetry-event-batch.json +++ b/telemetry-event-batch.json @@ -28,11 +28,11 @@ }, "ctx_token": { "type": "string", - "pattern": "^ct_[A-Za-z0-9_-]{8,240}$", + "pattern": "^ct_[A-Za-z0-9_-]{16,240}$", "description": "Opaque click token issued by the originating agent, carried in place of session_id on batches of content_engaged events emitted from a landing page after a click-out. Applies to every event in the batch and is resolved by the telemetry consumer to the click context (section 7.4). An event MUST carry either session_id or ctx_token at Grounding conformance and above." }, "agent_id": { - "type": "string", + "type": ["string", "null"], "description": "Responding agent identifier. REQUIRED for emitters at Grounding conformance or above when using event batch delivery. Mirrors the session-level agent_id field." }, "started_at": { @@ -42,6 +42,7 @@ }, "manifest_ref": { "type": ["string", "null"], + "format": "uri", "description": "Manifest reference - the URL of a manifest at /.well-known/content-telemetry.json identifying the emitter. Applies to every event in the batch and mirrors the session-level manifest_ref field. See sections 7.1 and 8." }, "events": { diff --git a/telemetry-event.json b/telemetry-event.json index 9654c27..5477047 100644 --- a/telemetry-event.json +++ b/telemetry-event.json @@ -28,11 +28,11 @@ }, "ctx_token": { "type": "string", - "pattern": "^ct_[A-Za-z0-9_-]{8,240}$", + "pattern": "^ct_[A-Za-z0-9_-]{16,240}$", "description": "Opaque click token issued by the originating agent, carried on content_engaged events emitted from a landing page after a click-out in place of session_id. Resolved by the telemetry consumer to the click context (section 7.4). An event MUST carry either session_id or ctx_token at Grounding conformance and above." }, "agent_id": { - "type": "string", + "type": ["string", "null"], "description": "Responding agent identifier. REQUIRED for emitters at Grounding conformance or above when using standalone event delivery. Mirrors the session-level agent_id field." }, "started_at": { @@ -42,6 +42,7 @@ }, "manifest_ref": { "type": ["string", "null"], + "format": "uri", "description": "Manifest reference - the URL of a manifest at /.well-known/content-telemetry.json identifying the emitter. Mirrors the session-level manifest_ref field. RECOMMENDED on standalone events supporting settlement or audit obligations under governing terms. See sections 7.1 and 8." }, "event": { diff --git a/telemetry-session.json b/telemetry-session.json index 5856b9b..c950987 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -42,6 +42,7 @@ }, "manifest_ref": { "type": ["string", "null"], + "format": "uri", "description": "Manifest reference - the URL of a manifest at /.well-known/content-telemetry.json. See section 8." }, "started_at": { @@ -112,7 +113,7 @@ "id": { "type": "string", "format": "uuid", - "description": "Unique event identifier" + "description": "Unique event identifier; distinct within a session document (section 6.7)" }, "type": { "$ref": "#/$defs/EventType" @@ -148,12 +149,12 @@ }, "ctx_token": { "type": "string", - "pattern": "^ct_[A-Za-z0-9_-]{8,240}$", - "description": "The click token the agent minted for this engagement's presentation, recorded so the consumer can join destination-reported events to it. Valid only on content_engaged events. See section 7.4." + "pattern": "^ct_[A-Za-z0-9_-]{16,240}$", + "description": "The click token the agent minted for this engagement's presentation, recorded so the consumer can join destination-reported events to it. Valid only on content_engaged events (enforced at the application layer, section 5.7.5). Unguessable, at least 16 characters after the ct_ prefix. See section 7.4." }, "source_role": { "$ref": "#/$defs/SourceRole", - "description": "Who is reporting this event (see section 4.4). SHOULD be set on content_retrieved events." + "description": "Who is reporting this event (see section 4.4). MUST be set on content_retrieved events (sections 5.2.2, 5.7.1; enforced at the application layer)." }, "content_telemetry_id": { "type": ["string", "null"], @@ -194,10 +195,12 @@ "required": ["type"] }, "then": { + "required": ["data"], "properties": { "data": { + "required": ["scope"], "properties": { - "scope": { "$ref": "#/$defs/GroundingScope" }, + "scope": { "$ref": "#/$defs/GroundingScope", "description": "REQUIRED on every content_grounded event: the occurrence boundary and every counting model depend on it (sections 4.3, 6.4, 10)." }, "cached": { "type": "boolean" }, "provenance": { "$ref": "#/$defs/SourceProvenance" }, "chars_ingested": { "type": "integer", "minimum": 0, "description": "Unicode code points in the exact text placed in the generation context, without normalising solely for counting. Portable across emitters; preferred over tokens_ingested (section 6.4)." }, @@ -310,6 +313,7 @@ "data": { "properties": { "media_type": { "$ref": "#/$defs/MediaType" }, + "content_depth": { "type": "string", "description": "Depth of the content record reached: metadata, abstract, full. Open vocabulary; consumers MUST tolerate unknown values (section 6.1)." }, "user_agent": { "type": "string" }, "bot_category": { "type": "string", "description": "Edge platform's bot classification. Recommended values: training, inference, search" }, "bot_name": { "type": "string", "description": "Recognised bot family parsed from the User-Agent (e.g., Claude-User, GPTBot, Perplexity-User). Stable across product variants within a vendor." }, diff --git a/tests/README.md b/tests/README.md index e764e1a..4dde243 100644 --- a/tests/README.md +++ b/tests/README.md @@ -19,9 +19,9 @@ uv run --with "jsonschema[format-nongpl]" python tests/validate.py uv run --with "jsonschema[format-nongpl]" python tests/check_examples.py ``` -Run from the repository root. Without uv: `pip install "jsonschema[format-nongpl]"`, then `python3 tests/validate.py`. The `format-nongpl` extra pulls in the format validators (`rfc3339-validator` and friends) that make `format: uuid` / `date-time` / `uri` assertions enforce rather than annotate; both scripts hard-error at startup if they are missing. Both commands run in CI on every pull request. +Run from the repository root. Without uv: `pip install "jsonschema[format-nongpl]"`, then `python3 tests/validate.py`. `tests/mutation_smoke.py` runs the same way, and in CI. The `format-nongpl` extra pulls in the format validators (`rfc3339-validator` and friends) that make `format: uuid` / `date-time` / `uri` assertions enforce rather than annotate; both scripts hard-error at startup if they are missing. Both commands run in CI on every pull request. -`check_examples.py` extracts every fenced `json` block from the spec and README, validates the complete top-level documents (sessions, standalone events, manifests) against the matching schema, and reports the number of fragments it skipped. A worked example that no longer matches its schema fails the build. +`check_examples.py` extracts every fenced `json` block from the spec and README, validates the complete top-level documents (sessions, standalone events, manifests) against the matching schema and against the application-layer rules of `validate.py`, and reports the number of fragments it skipped. A worked example that no longer matches its schema or violates a conformance rule fails the build. ## What it covers @@ -31,12 +31,18 @@ Run from the repository root. Without uv: `pip install "jsonschema[format-nongpl - Enum validation (event types, privacy levels, source roles, schema version) - Citation source-reference requirement (content_cited rejected when content_url/content_id are missing or null) and required citation_type - Required event and output identifiers on cited, reproduced, and presented events -- Closed enums (reproduction_type, presentation_kind) and non-negative counts (chars_ingested, reproduced_chars) -- Format assertions (a malformed parent_session_id fails) -- All three conformance levels (Retrieval, Grounding, Citation) +- Closed enums (citation_type, position, scope, provenance, reproduction_type, presentation_kind) and non-negative counts +- Required `data.scope` on grounding events; required `source_role` on retrieval events +- Format assertions on every uuid, date-time and uri field (malformed session, event, citation, presentation and parent ids; malformed started_at; malformed turn URL arrays) +- Field placement: presentation_id and event-level ctx_token only on content_engaged, citation_id only on presented/reproduced, turn only on turn events; envelope ctx_token only with engagements +- Session integrity: distinct event ids, engagement/presentation and presentation/citation identify the same content, one token per presentation +- Rejection of documents declaring the v0.1 wire version +- All three conformance levels (Retrieval, Grounding, Citation) and all four source roles (origin, edge, index, agent) +- Every core engagement type, citation type, position and presentation type; a destination-reported engagement batch; a multi-owner catalogue session under one agreement +- Manifests for all three roles, including platform; manifest rejection for http endpoints, path-prefixed domains, ctx_resolution on the wrong role, duplicate or empty roles, missing endpoint and coverage mode - Standalone event envelopes (CDN edge, agent with session FK) -- Privacy level field gating (application-layer conformance), one fixture per forbidden field at minimal -- Funnel exceptions (presented-no-cited, cited-no-grounded, presented-no-grounded, reproduced-no-grounded) +- Privacy level field gating (application-layer conformance), one fixture per forbidden field at minimal and intent, in session, standalone and batch shapes +- Funnel exceptions (presented-no-cited, cited-no-grounded, presented-no-grounded, reproduced-no-cited, reproduced-no-grounded) - Reproduction cases: credited quotation (reproduction + direct_quote citation sharing an output element) and uncredited reproduction in unpresented API output - Text, image, audio, video, suppressed-citation, and repeated-presentation cases - Exact presentation-to-engagement correlation across session and standalone envelopes @@ -44,18 +50,19 @@ Run from the repository root. Without uv: `pip install "jsonschema[format-nongpl - Grounding provenance paths and generic fingerprint detection across session, standalone-event, and event-batch envelopes - Custom response_mode values -Each test file has a `_test_description` field explaining what it demonstrates. Every `invalid/` fixture also has an `_expected_error` field: a substring that must appear in the actual error (the first schema error's message and JSON pointer, or the application-layer violation text). The runner fails a fixture that fails for a different reason than the one it pins, and fails any invalid fixture missing the field. +Each test file has a `_test_description` field explaining what it demonstrates. Every `invalid/` fixture also has an `_expected_error` field: a substring that must appear in the actual error (the first schema error's JSON pointer and message, or the application-layer violation text). Schema pins carry the pointer (`/events/0 'id' is a required property`) so that the same message at a different location does not satisfy them. The runner fails a fixture that fails for a different reason than the one it pins, and fails any invalid fixture missing the field. ## Application-layer conformance Some rules cannot be expressed in JSON Schema alone. These are tested as application-layer conformance checks in `validate.py`: -- Privacy level field gating (e.g. `query_text` MUST NOT be present at `minimal` level), applied to turns wherever they appear: session documents, batches, and standalone envelopes -- `content_url` or `content_id` requirement on every content event (section 5.7.5) +- Privacy level field gating (e.g. `query_text` MUST NOT be present at `minimal` level), applied to turns wherever they appear: session documents, batches, and standalone envelopes - each shape has its own fixture +- `content_url` or `content_id` requirement on every content event, and `source_role` on every `content_retrieved` event (sections 5.2.2, 5.7.5), in every document shape +- Field placement by event type (section 5.7.5): `presentation_id` and event-level `ctx_token` only on `content_engaged`, `citation_id` only on `content_presented`/`content_reproduced`, `turn` only on turn events; an envelope `ctx_token` only with `content_engaged` events - `session_id` or `ctx_token` on a standalone event or event batch envelope at Grounding conformance and above (sections 5.7.5, 7.1) -- Referential integrity within a session document: `content_engaged.presentation_id` matches a `content_presented` event id, and `citation_id` on `content_presented`/`content_reproduced` matches a `content_cited` event id (sections 6.6-6.8). Standalone envelopes and batch members are exempt - they may reference events delivered elsewhere. -- Manifest rejection rules: duplicate `keys[].id`, and `domains` entries that are not the manifest's own host or a subdomain of it (sections 8.6, 8.7) -- Withdrawn `ip_hash` prohibition on `content_retrieved` data (section 9.1 migration rule) +- Referential integrity within a session document: event ids are distinct; `content_engaged.presentation_id` matches a `content_presented` event id and `citation_id` on `content_presented`/`content_reproduced` matches a `content_cited` event id, in each case identifying the same content; one event-level `ctx_token` binds to one presentation (sections 6.6-6.8, 7.4.1). Standalone envelopes and batch members are exempt - they may reference events delivered elsewhere. +- Manifest rejection rules: duplicate `keys[].id`; `domains` entries that are not the manifest's own host or a subdomain of it; `domains` on a manifest served under a path prefix; `ctx_resolution` on a manifest without the `agent` or `platform` role (sections 8.5-8.7) +- Withdrawn `ip_hash` prohibition on event data (section 9.1 migration rule), in every document shape - Grounding provenance/cache consistency and the prohibition on `preserved_in_output` in `content_fingerprint` (sections 5.7.5, 6.4, 12.1) Valid fixtures must pass both JSON Schema and these checks; `invalid/` fixtures that pass JSON Schema but fail a check are documented in `validate.py`. The `agent_id`-at-Grounding requirement is not fixture-tested: it depends on the emitter's declared conformance level, which the fixtures do not carry. diff --git a/tests/check_examples.py b/tests/check_examples.py index 9c04fca..e36b721 100644 --- a/tests/check_examples.py +++ b/tests/check_examples.py @@ -11,6 +11,8 @@ event batch -> telemetry-event-batch.json manifest -> manifest.json +and against the application-layer conformance rules of validate.py. + A complete bare event object (carrying the required `type` and `timestamp`) is validated against the TelemetryEvent definition in telemetry-session.json. Genuine fragments (a single turn, a one-field snippet, an event elided below @@ -54,6 +56,10 @@ REPO = Path(__file__).resolve().parent.parent SOURCES = [REPO / "SPECIFICATION.md", REPO / "README.md"] + +# The application-layer conformance rules live in validate.py (same directory). +sys.path.insert(0, str(Path(__file__).resolve().parent)) +from validate import check_application_layer, check_manifest_application_layer # noqa: E402 FENCE = re.compile(r"```json\n(.*?)\n```", re.DOTALL) @@ -152,12 +158,22 @@ def main(): checked += 1 by_schema[schema_name] = by_schema.get(schema_name, 0) + 1 errors = sorted(validators[schema_name].iter_errors(doc), key=lambda e: e.path) - if errors: + # Worked examples must also satisfy the application-layer rules; a bare + # event is checked as the sole member of a session. + if schema_name == "manifest.json": + violations = check_manifest_application_layer(doc) + elif schema_name == EVENT_DEF: + violations = check_application_layer({"events": [doc]}) + else: + violations = check_application_layer(doc) + if errors or violations: failed += 1 print(f" FAIL {loc} (against {schema_name})") for e in errors[:3]: path = "/".join(str(p) for p in e.path) or "(root)" print(f" {path}: {e.message}") + for v in violations[:3]: + print(f" application-layer: {v}") else: print(f" PASS {loc} ({schema_name})") diff --git a/tests/invalid/access-context-identifier-missing-value.json b/tests/invalid/access-context-identifier-missing-value.json index f7b38b4..a0e57a6 100644 --- a/tests/invalid/access-context-identifier-missing-value.json +++ b/tests/invalid/access-context-identifier-missing-value.json @@ -1,6 +1,6 @@ { "_test_description": "Session access_context identifier missing its value. The schema requires both scheme and value on every identifier (5.1.3).", - "_expected_error": "access_context/identifiers/0 'value' is a required property", + "_expected_error": "/data/access_context/identifiers/0 'value' is a required property", "document_type": "session", "schema_version": "1.0", "session_id": "880e8400-e29b-41d4-a716-446655440090", diff --git a/tests/invalid/access-context-identifiers-not-array.json b/tests/invalid/access-context-identifiers-not-array.json index 226dc7a..96fc123 100644 --- a/tests/invalid/access-context-identifiers-not-array.json +++ b/tests/invalid/access-context-identifiers-not-array.json @@ -1,6 +1,6 @@ { "_test_description": "Session access_context.identifiers as a bare string. The schema requires an array of {scheme, value} objects (5.1.3).", - "_expected_error": "access_context/identifiers 'https://ror.org/013meh722' is not of type 'array'", + "_expected_error": "/data/access_context/identifiers 'https://ror.org/013meh722' is not of type 'array'", "document_type": "session", "schema_version": "1.0", "session_id": "880e8400-e29b-41d4-a716-446655440091", diff --git a/tests/invalid/batch-ctx-token-bad-pattern.json b/tests/invalid/batch-ctx-token-bad-pattern.json new file mode 100644 index 0000000..9533c15 --- /dev/null +++ b/tests/invalid/batch-ctx-token-bad-pattern.json @@ -0,0 +1,18 @@ +{ + "_test_description": "Event batch envelope whose ctx_token fails the value pattern (section 7.4.1).", + "_expected_error": "/ctx_token 'token123' does not match '^ct_[A-Za-z0-9_-]{16,240}$'", + "document_type": "event_batch", + "schema_version": "1.0", + "ctx_token": "token123", + "events": [ + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:06Z", + "turn_id": "1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + } + } + ] +} diff --git a/tests/invalid/batch-ctx-token-non-engagement.json b/tests/invalid/batch-ctx-token-non-engagement.json new file mode 100644 index 0000000..d5ae302 --- /dev/null +++ b/tests/invalid/batch-ctx-token-non-engagement.json @@ -0,0 +1,28 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: event batch under ctx_token carrying a content_grounded event alongside an engagement (section 7.1).", + "_expected_error": "Envelope ctx_token accompanies a 'content_grounded' event", + "document_type": "event_batch", + "schema_version": "1.0", + "ctx_token": "ct_9f3a1c7e2b8d4a06", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 4200 + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:06Z", + "turn_id": "1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + } + } + ] +} diff --git a/tests/invalid/batch-empty-events.json b/tests/invalid/batch-empty-events.json index 38a17e4..fa13e0c 100644 --- a/tests/invalid/batch-empty-events.json +++ b/tests/invalid/batch-empty-events.json @@ -1,6 +1,6 @@ { "_test_description": "Event batch envelope with an empty events array. The schema requires minItems 1: an empty batch carries no signal and MUST NOT be delivered.", - "_expected_error": "should be non-empty", + "_expected_error": "/events [] should be non-empty", "document_type": "event_batch", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440008", diff --git a/tests/invalid/batch-engaged-missing-presentation-id.json b/tests/invalid/batch-engaged-missing-presentation-id.json new file mode 100644 index 0000000..71621a3 --- /dev/null +++ b/tests/invalid/batch-engaged-missing-presentation-id.json @@ -0,0 +1,20 @@ +{ + "_test_description": "Event batch with session_id (no ctx_token) whose content_engaged lacks presentation_id. The relaxation applies only to envelopes carrying ctx_token (sections 6.8, 7.4).", + "_expected_error": "/events/0 'presentation_id' is a required property", + "document_type": "event_batch", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440617", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:06Z", + "turn_id": "1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + } + } + ] +} diff --git a/tests/invalid/batch-missing-schema-version.json b/tests/invalid/batch-missing-schema-version.json index 094a850..f097746 100644 --- a/tests/invalid/batch-missing-schema-version.json +++ b/tests/invalid/batch-missing-schema-version.json @@ -1,6 +1,6 @@ { "_test_description": "Event batch envelope missing schema_version. The event_batch document_type triggers batch envelope validation, which requires schema_version.", - "_expected_error": "'schema_version' is a required property", + "_expected_error": "/ 'schema_version' is a required property", "document_type": "event_batch", "session_id": "660e8400-e29b-41d4-a716-446655440009", "events": [ diff --git a/tests/invalid/batch-missing-session-mixed-retrieval.json b/tests/invalid/batch-missing-session-mixed-retrieval.json new file mode 100644 index 0000000..c61ec2e --- /dev/null +++ b/tests/invalid/batch-missing-session-mixed-retrieval.json @@ -0,0 +1,24 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: event batch mixing a retrieval with a grounding event, with neither session_id nor ctx_token. The retrieval-only exemption does not extend to a batch that carries other events (sections 5.7.5, 7.1).", + "_expected_error": "neither session_id nor ctx_token", + "document_type": "event_batch", + "schema_version": "1.0", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "agent", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + }, + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 4200 + } + } + ] +} diff --git a/tests/invalid/citation-id-on-grounded.json b/tests/invalid/citation-id-on-grounded.json new file mode 100644 index 0000000..104350f --- /dev/null +++ b/tests/invalid/citation-id-on-grounded.json @@ -0,0 +1,33 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: citation_id on a content_grounded event. Valid only on content_presented and content_reproduced (sections 5.2, 5.7.5).", + "_expected_error": "Field 'citation_id' present on 'content_grounded' event", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440632", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "citation_type": "reference", + "position": "primary" + } + }, + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 4200 + }, + "citation_id": "770e8400-e29b-41d4-a716-446655440601" + } + ] +} diff --git a/tests/invalid/citation-position-invalid.json b/tests/invalid/citation-position-invalid.json new file mode 100644 index 0000000..a918931 --- /dev/null +++ b/tests/invalid/citation-position-invalid.json @@ -0,0 +1,22 @@ +{ + "_test_description": "content_cited with position 'leading'. position is a closed enum: primary, supporting, mentioned, unclassified (section 6.5).", + "_expected_error": "/events/0/data/position 'leading' is not one of ['primary', 'supporting', 'mentioned', 'unclassified']", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440603", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "citation_type": "reference", + "position": "leading" + } + } + ] +} diff --git a/tests/invalid/citation-type-invalid.json b/tests/invalid/citation-type-invalid.json new file mode 100644 index 0000000..77bdcb4 --- /dev/null +++ b/tests/invalid/citation-type-invalid.json @@ -0,0 +1,21 @@ +{ + "_test_description": "content_cited with citation_type 'summary'. citation_type is a closed enum: direct_quote, paraphrase, reference, contradiction, unclassified (section 6.5).", + "_expected_error": "/events/0/data/citation_type 'summary' is not one of ['direct_quote', 'paraphrase', 'reference', 'contradiction', 'unclassified']", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440602", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "citation_type": "summary" + } + } + ] +} diff --git a/tests/invalid/cited-missing-citation-type.json b/tests/invalid/cited-missing-citation-type.json index 9d6427c..e2001f6 100644 --- a/tests/invalid/cited-missing-citation-type.json +++ b/tests/invalid/cited-missing-citation-type.json @@ -1,6 +1,6 @@ { "_test_description": "content_cited must classify how content was used in the response: data.citation_type is required. Emitters that cannot confidently classify use citation_type 'unclassified' rather than omitting the field (section 6.5).", - "_expected_error": "'citation_type' is a required property", + "_expected_error": "/events/0/data 'citation_type' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440204", "started_at": "2026-08-05T09:30:00Z", diff --git a/tests/invalid/cited-missing-data.json b/tests/invalid/cited-missing-data.json new file mode 100644 index 0000000..165af8e --- /dev/null +++ b/tests/invalid/cited-missing-data.json @@ -0,0 +1,18 @@ +{ + "_test_description": "content_cited with no data object. citation_type is required, so data itself is required (section 6.5).", + "_expected_error": "/events/0 'data' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440615", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + } + ] +} diff --git a/tests/invalid/cited-missing-id.json b/tests/invalid/cited-missing-id.json index 633a721..4ed1401 100644 --- a/tests/invalid/cited-missing-id.json +++ b/tests/invalid/cited-missing-id.json @@ -1,6 +1,6 @@ { "_test_description": "content_cited event with output_id and a resolvable source reference but no event id. id is required on citation events so a presentation or reproduction can reference the exact citation via citation_id (section 6.5).", - "_expected_error": "'id' is a required property", + "_expected_error": "/events/0 'id' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440203", "started_at": "2026-08-05T09:20:00Z", diff --git a/tests/invalid/cited-missing-output-id.json b/tests/invalid/cited-missing-output-id.json index 32a3e0c..3644486 100644 --- a/tests/invalid/cited-missing-output-id.json +++ b/tests/invalid/cited-missing-output-id.json @@ -1,6 +1,6 @@ { "_test_description": "content_cited event with id and a resolvable source reference but no output_id. output_id is required on citation events so output construction can be correlated with later presentation.", - "_expected_error": "'output_id' is a required property", + "_expected_error": "/events/0 'output_id' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440160", "started_at": "2026-08-01T14:00:00Z", @@ -10,7 +10,9 @@ "type": "content_cited", "timestamp": "2026-08-01T14:00:01Z", "content_url": "https://www.example-news.com/economy/rate-decision-analysis", - "data": { "citation_type": "reference" } + "data": { + "citation_type": "reference" + } } ] } diff --git a/tests/invalid/cited-missing-source-reference.json b/tests/invalid/cited-missing-source-reference.json index 26e21cf..ad79e50 100644 --- a/tests/invalid/cited-missing-source-reference.json +++ b/tests/invalid/cited-missing-source-reference.json @@ -1,6 +1,6 @@ { "_test_description": "content_cited event carrying neither content_url nor content_id. A source association with no resolvable reference is not a citation; unlike other content events, the JSON Schema enforces the identifier requirement for content_cited (section 6.5).", - "_expected_error": "'content_url' is a required property", + "_expected_error": "/events/0 'content_url' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440140", "agent_id": "research-assistant-v1", diff --git a/tests/invalid/cited-negative-excerpt-chars.json b/tests/invalid/cited-negative-excerpt-chars.json new file mode 100644 index 0000000..07e32b5 --- /dev/null +++ b/tests/invalid/cited-negative-excerpt-chars.json @@ -0,0 +1,22 @@ +{ + "_test_description": "content_cited with a negative excerpt_chars (section 6.5: minimum 0).", + "_expected_error": "/events/0/data/excerpt_chars -5 is less than the minimum of 0", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440622", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "citation_type": "direct_quote", + "excerpt_chars": -5 + } + } + ] +} diff --git a/tests/invalid/cited-null-source-reference.json b/tests/invalid/cited-null-source-reference.json index 92df5f1..fd6b084 100644 --- a/tests/invalid/cited-null-source-reference.json +++ b/tests/invalid/cited-null-source-reference.json @@ -1,6 +1,6 @@ { "_test_description": "content_cited event with content_url explicitly null and no content_id. Presence of a null reference does not satisfy the citation reference requirement: the schema demands a non-null content_url or content_id on content_cited (section 6.5).", - "_expected_error": "is not of type 'string'", + "_expected_error": "/events/0/content_url None is not of type 'string'", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440141", "agent_id": "research-assistant-v1", diff --git a/tests/invalid/ctx-token-bad-pattern.json b/tests/invalid/ctx-token-bad-pattern.json index e1f1560..3470560 100644 --- a/tests/invalid/ctx-token-bad-pattern.json +++ b/tests/invalid/ctx-token-bad-pattern.json @@ -1,6 +1,6 @@ { - "_test_description": "ctx_token failing the value pattern: tokens are opaque but MUST match ^ct_[A-Za-z0-9_-]{8,240}$ so they survive URL carriage and are recognisable in logs (section 7.4). This value has no ct_ prefix and contains reserved characters.", - "_expected_error": "does not match '^ct_[A-Za-z0-9_-]{8,240}$'", + "_test_description": "ctx_token failing the value pattern: tokens are opaque but MUST match ^ct_[A-Za-z0-9_-]{16,240}$ so they survive URL carriage and are recognisable in logs (section 7.4). This value has no ct_ prefix and contains reserved characters.", + "_expected_error": "/ctx_token 'session=660e8400!' does not match '^ct_[A-Za-z0-9_-]{16,240}$'", "document_type": "event", "schema_version": "1.0", "ctx_token": "session=660e8400!", diff --git a/tests/invalid/ctx-token-on-grounded.json b/tests/invalid/ctx-token-on-grounded.json new file mode 100644 index 0000000..aa2e6cc --- /dev/null +++ b/tests/invalid/ctx-token-on-grounded.json @@ -0,0 +1,21 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: event-level ctx_token on a content_grounded event. The field is valid only on content_engaged (sections 5.2, 5.7.5).", + "_expected_error": "Field 'ctx_token' present on 'content_grounded' event", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440631", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 4200 + }, + "ctx_token": "ct_9f3a1c7e2b8d4a06" + } + ] +} diff --git a/tests/invalid/duplicate-event-id.json b/tests/invalid/duplicate-event-id.json new file mode 100644 index 0000000..f055f5d --- /dev/null +++ b/tests/invalid/duplicate-event-id.json @@ -0,0 +1,44 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: two content_presented events share one id, so the engagement cannot name the exact occurrence. Repeated presentations receive distinct ids (section 6.7).", + "_expected_error": "Duplicate event id", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440635", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:05Z", + "turn_id": "1", + "output_id": "response:2", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:06Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440602", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + } + } + ] +} diff --git a/tests/invalid/empty-output-id.json b/tests/invalid/empty-output-id.json new file mode 100644 index 0000000..b61cb42 --- /dev/null +++ b/tests/invalid/empty-output-id.json @@ -0,0 +1,22 @@ +{ + "_test_description": "content_cited with an empty output_id. output_id is an opaque identifier and MUST be non-empty (section 5.2).", + "_expected_error": "/events/0/output_id '' should be non-empty", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440620", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "citation_type": "reference", + "position": "primary" + } + } + ] +} diff --git a/tests/invalid/engaged-missing-presentation-id.json b/tests/invalid/engaged-missing-presentation-id.json index 0e35747..719378f 100644 --- a/tests/invalid/engaged-missing-presentation-id.json +++ b/tests/invalid/engaged-missing-presentation-id.json @@ -1,6 +1,6 @@ { "_test_description": "content_engaged must identify the exact presentation occurrence rather than matching only by URL.", - "_expected_error": "'presentation_id' is a required property", + "_expected_error": "/events/0 'presentation_id' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440092", "started_at": "2026-07-18T13:00:00Z", @@ -9,7 +9,9 @@ "type": "content_engaged", "timestamp": "2026-07-18T13:00:02Z", "content_id": "publisher:article:1", - "data": { "engagement_type": "link_click" } + "data": { + "engagement_type": "link_click" + } } ] } diff --git a/tests/invalid/engaged-presentation-content-mismatch.json b/tests/invalid/engaged-presentation-content-mismatch.json new file mode 100644 index 0000000..c720e69 --- /dev/null +++ b/tests/invalid/engaged-presentation-content-mismatch.json @@ -0,0 +1,32 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: content_engaged references a presentation of different content. The engagement identifies the same content as the presentation it acted on (sections 5.7.5, 6.8).", + "_expected_error": "content_engaged references presentation", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440636", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-review.com/headphones/best-noise-cancelling", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:06Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440602", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + } + } + ] +} diff --git a/tests/invalid/engaged-presentation-id-no-presentations.json b/tests/invalid/engaged-presentation-id-no-presentations.json new file mode 100644 index 0000000..d72098c --- /dev/null +++ b/tests/invalid/engaged-presentation-id-no-presentations.json @@ -0,0 +1,30 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: content_engaged carries a presentation_id in a session with no content_presented events at all (section 6.8).", + "_expected_error": "does not match any content_presented event id", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440639", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 4200 + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:06Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440602", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + } + } + ] +} diff --git a/tests/invalid/event-level-ctx-token-bad-pattern.json b/tests/invalid/event-level-ctx-token-bad-pattern.json new file mode 100644 index 0000000..700329d --- /dev/null +++ b/tests/invalid/event-level-ctx-token-bad-pattern.json @@ -0,0 +1,33 @@ +{ + "_test_description": "Agent-reported content_engaged whose event-level ctx_token is shorter than the 16-character minimum (section 7.4.1). Tokens MUST be unguessable; the pattern enforces the length floor.", + "_expected_error": "/events/1/ctx_token 'ct_short' does not match '^ct_[A-Za-z0-9_-]{16,240}$'", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440618", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:06Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440602", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + }, + "ctx_token": "ct_short" + } + ] +} diff --git a/tests/invalid/grounded-missing-identifier-batch.json b/tests/invalid/grounded-missing-identifier-batch.json new file mode 100644 index 0000000..d8ad604 --- /dev/null +++ b/tests/invalid/grounded-missing-identifier-batch.json @@ -0,0 +1,20 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: batch member content_grounded with neither content_url nor content_id (section 5.7.5), in the event batch shape.", + "_expected_error": "'content_grounded' carries neither content_url nor content_id", + "document_type": "event_batch", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440646", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 4200 + } + } + ] +} diff --git a/tests/invalid/grounded-missing-identifier-standalone.json b/tests/invalid/grounded-missing-identifier-standalone.json new file mode 100644 index 0000000..84d2108 --- /dev/null +++ b/tests/invalid/grounded-missing-identifier-standalone.json @@ -0,0 +1,18 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: standalone content_grounded with neither content_url nor content_id (section 5.7.5), in the standalone envelope shape.", + "_expected_error": "'content_grounded' carries neither content_url nor content_id", + "document_type": "event", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440645", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "event": { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 4200 + } + } +} diff --git a/tests/invalid/grounded-negative-chars-ingested.json b/tests/invalid/grounded-negative-chars-ingested.json index dc44c37..a9352f2 100644 --- a/tests/invalid/grounded-negative-chars-ingested.json +++ b/tests/invalid/grounded-negative-chars-ingested.json @@ -1,6 +1,6 @@ { "_test_description": "content_grounded with a negative chars_ingested. Ingestion counts are Unicode code point counts and cannot be negative (section 6.4); the schema requires a minimum of 0.", - "_expected_error": "chars_ingested -7200 is less than the minimum", + "_expected_error": "/events/0/data/chars_ingested -7200 is less than the minimum of 0", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440210", "started_at": "2026-08-05T10:00:00Z", diff --git a/tests/invalid/grounded-negative-tokens-ingested.json b/tests/invalid/grounded-negative-tokens-ingested.json new file mode 100644 index 0000000..c74aba1 --- /dev/null +++ b/tests/invalid/grounded-negative-tokens-ingested.json @@ -0,0 +1,20 @@ +{ + "_test_description": "content_grounded with a negative tokens_ingested (section 6.4: minimum 0).", + "_expected_error": "/events/0/data/tokens_ingested -10 is less than the minimum of 0", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440621", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "tokens_ingested": -10 + } + } + ] +} diff --git a/tests/invalid/grounding-fingerprint-missing-detected.json b/tests/invalid/grounding-fingerprint-missing-detected.json index 2cff586..a31c603 100644 --- a/tests/invalid/grounding-fingerprint-missing-detected.json +++ b/tests/invalid/grounding-fingerprint-missing-detected.json @@ -1,6 +1,6 @@ { "_test_description": "Grounding event carries content_fingerprint without required detected. Must fail JSON Schema.", - "_expected_error": "content_fingerprint 'detected' is a required property", + "_expected_error": "/events/0/data/content_fingerprint 'detected' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440182", "started_at": "2026-08-12T10:20:00Z", diff --git a/tests/invalid/grounding-fingerprint-missing-scheme.json b/tests/invalid/grounding-fingerprint-missing-scheme.json new file mode 100644 index 0000000..21c5a4c --- /dev/null +++ b/tests/invalid/grounding-fingerprint-missing-scheme.json @@ -0,0 +1,22 @@ +{ + "_test_description": "content_fingerprint without the required scheme (section 6.4).", + "_expected_error": "/events/0/data/content_fingerprint 'scheme' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440625", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "content_fingerprint": { + "detected": true + } + } + } + ] +} diff --git a/tests/invalid/grounding-missing-data.json b/tests/invalid/grounding-missing-data.json new file mode 100644 index 0000000..05c44e0 --- /dev/null +++ b/tests/invalid/grounding-missing-data.json @@ -0,0 +1,15 @@ +{ + "_test_description": "content_grounded with no data object at all. data is required on grounding events because scope is required (section 6.4).", + "_expected_error": "/events/0 'data' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440606", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + } + ] +} diff --git a/tests/invalid/grounding-provenance-cached-conflict-agent-cached.json b/tests/invalid/grounding-provenance-cached-conflict-agent-cached.json new file mode 100644 index 0000000..d6fd7fd --- /dev/null +++ b/tests/invalid/grounding-provenance-cached-conflict-agent-cached.json @@ -0,0 +1,20 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: grounding declares agent_cached with cached false. agent_cached requires cached true (section 6.4).", + "_expected_error": "agent_cached does not carry cached true", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440641", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "provenance": "agent_cached" + } + } + ] +} diff --git a/tests/invalid/grounding-provenance-invalid.json b/tests/invalid/grounding-provenance-invalid.json new file mode 100644 index 0000000..fba1625 --- /dev/null +++ b/tests/invalid/grounding-provenance-invalid.json @@ -0,0 +1,20 @@ +{ + "_test_description": "content_grounded with provenance 'scraped'. provenance is agent_fetched, agent_cached or third_party_sourced (section 6.4).", + "_expected_error": "/events/0/data/provenance 'scraped' is not one of ['agent_fetched', 'agent_cached', 'third_party_sourced']", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440607", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "provenance": "scraped" + } + } + ] +} diff --git a/tests/invalid/grounding-scope-invalid.json b/tests/invalid/grounding-scope-invalid.json new file mode 100644 index 0000000..70a5cbe --- /dev/null +++ b/tests/invalid/grounding-scope-invalid.json @@ -0,0 +1,19 @@ +{ + "_test_description": "content_grounded with scope 'document'. scope is session or turn only (section 6.4).", + "_expected_error": "/events/0/data/scope 'document' is not one of ['session', 'turn']", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440604", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "document", + "cached": false + } + } + ] +} diff --git a/tests/invalid/grounding-scope-missing.json b/tests/invalid/grounding-scope-missing.json new file mode 100644 index 0000000..17cf039 --- /dev/null +++ b/tests/invalid/grounding-scope-missing.json @@ -0,0 +1,19 @@ +{ + "_test_description": "content_grounded without data.scope. scope is required on every grounding event (sections 5.7.2, 6.4): the occurrence boundary and every counting model depend on it.", + "_expected_error": "/events/0/data 'scope' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440605", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "cached": false, + "chars_ingested": 4200 + } + } + ] +} diff --git a/tests/invalid/invalid-event-type.json b/tests/invalid/invalid-event-type.json index 9948b43..09993df 100644 --- a/tests/invalid/invalid-event-type.json +++ b/tests/invalid/invalid-event-type.json @@ -1,6 +1,6 @@ { "_test_description": "Event with type 'content_summarised' which is not in the EventType enum. Fails JSON Schema validation.", - "_expected_error": "'content_summarised' is not one of", + "_expected_error": "/events/0/type 'content_summarised' is not one of ['content_retrieved', 'content_grounded', 'content_reproduced', 'content_cited', 'content_presented', 'content", "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440008", "started_at": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/invalid-privacy-level.json b/tests/invalid/invalid-privacy-level.json index c0e827c..70b313d 100644 --- a/tests/invalid/invalid-privacy-level.json +++ b/tests/invalid/invalid-privacy-level.json @@ -1,6 +1,6 @@ { "_test_description": "Turn with privacy_level 'redacted' which is not in the PrivacyLevel enum. Fails JSON Schema validation.", - "_expected_error": "'redacted' is not one of", + "_expected_error": "/events/0/turn/privacy_level 'redacted' is not one of ['full', 'summary', 'intent', 'minimal']", "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440009", "started_at": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/invalid-schema-version.json b/tests/invalid/invalid-schema-version.json index e851c2c..7c9356a 100644 --- a/tests/invalid/invalid-schema-version.json +++ b/tests/invalid/invalid-schema-version.json @@ -1,6 +1,6 @@ { "_test_description": "schema_version set to '2.0' which does not match the const '1.0'. Fails JSON Schema validation.", - "_expected_error": "'1.0' was expected", + "_expected_error": "/schema_version '1.0' was expected", "schema_version": "2.0", "session_id": "770e8400-e29b-41d4-a716-446655440010", "started_at": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/invalid-source-role.json b/tests/invalid/invalid-source-role.json index 57be3e5..d26336d 100644 --- a/tests/invalid/invalid-source-role.json +++ b/tests/invalid/invalid-source-role.json @@ -1,6 +1,6 @@ { "_test_description": "Event with source_role 'cdn' which is not in the SourceRole enum. Fails JSON Schema validation.", - "_expected_error": "'cdn' is not one of", + "_expected_error": "/events/0/source_role 'cdn' is not one of ['origin', 'edge', 'index', 'agent']", "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440014", "started_at": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/legacy-content-displayed.json b/tests/invalid/legacy-content-displayed.json index 59c6ead..f0c3b36 100644 --- a/tests/invalid/legacy-content-displayed.json +++ b/tests/invalid/legacy-content-displayed.json @@ -1,6 +1,6 @@ { "_test_description": "The v1 presentation event replaces the preview content_displayed name.", - "_expected_error": "'content_displayed' is not one of", + "_expected_error": "/events/0/type 'content_displayed' is not one of ['content_retrieved', 'content_grounded', 'content_reproduced', 'content_cited', 'content_presented', 'content_", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440093", "started_at": "2026-07-18T13:00:00Z", @@ -9,7 +9,9 @@ "type": "content_displayed", "timestamp": "2026-07-18T13:00:03Z", "content_id": "publisher:article:1", - "data": { "display_type": "link" } + "data": { + "display_type": "link" + } } ] } diff --git a/tests/invalid/legacy-schema-version-0-1.json b/tests/invalid/legacy-schema-version-0-1.json new file mode 100644 index 0000000..3e0408b --- /dev/null +++ b/tests/invalid/legacy-schema-version-0-1.json @@ -0,0 +1,16 @@ +{ + "_test_description": "Document declaring the v0.1 wire version. A v1 consumer MUST reject it (sections 5.7.4, 12.1): v0.1 is a different wire version, not a compatible minor.", + "_expected_error": "/schema_version '1.0' was expected", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440601", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "agent", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + } + ] +} diff --git a/tests/invalid/malformed-citation-id.json b/tests/invalid/malformed-citation-id.json new file mode 100644 index 0000000..9bd684c --- /dev/null +++ b/tests/invalid/malformed-citation-id.json @@ -0,0 +1,35 @@ +{ + "_test_description": "content_presented whose citation_id is not a UUID (section 5.2 format: uuid).", + "_expected_error": "/events/1/citation_id 'cite-1' is not a 'uuid'", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440609", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "citation_type": "reference", + "position": "primary" + } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + }, + "citation_id": "cite-1" + } + ] +} diff --git a/tests/invalid/malformed-content-hash.json b/tests/invalid/malformed-content-hash.json new file mode 100644 index 0000000..250da93 --- /dev/null +++ b/tests/invalid/malformed-content-hash.json @@ -0,0 +1,20 @@ +{ + "_test_description": "content_grounded whose content_hash is not sha256:{hex} (section 6.4).", + "_expected_error": "/events/0/data/content_hash 'md5:abc' does not match '^sha256:[a-f0-9]{64}$'", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440619", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "content_hash": "md5:abc" + } + } + ] +} diff --git a/tests/invalid/malformed-content-urls-cited.json b/tests/invalid/malformed-content-urls-cited.json new file mode 100644 index 0000000..c89c9d1 --- /dev/null +++ b/tests/invalid/malformed-content-urls-cited.json @@ -0,0 +1,22 @@ +{ + "_test_description": "Turn whose content_urls_cited carries a value that is not a URI (section 5.4: URI[]).", + "_expected_error": "/events/0/turn/content_urls_cited/0 'not a url' is not a 'uri'", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440613", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "turn_completed", + "timestamp": "2026-08-20T10:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "minimal", + "content_urls_cited": [ + "not a url" + ], + "response_tokens": 10 + } + } + ] +} diff --git a/tests/invalid/malformed-event-id.json b/tests/invalid/malformed-event-id.json new file mode 100644 index 0000000..8fa2726 --- /dev/null +++ b/tests/invalid/malformed-event-id.json @@ -0,0 +1,22 @@ +{ + "_test_description": "content_cited whose id is not a UUID. Event ids carry format: uuid (section 5.2); rejected only when the validator enforces format assertions, which the conformance runners do.", + "_expected_error": "/events/0/id 'cite-1' is not a 'uuid'", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440608", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "cite-1", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "citation_type": "reference", + "position": "primary" + } + } + ] +} diff --git a/tests/invalid/malformed-parent-session-id.json b/tests/invalid/malformed-parent-session-id.json index 3d71576..f7ac83c 100644 --- a/tests/invalid/malformed-parent-session-id.json +++ b/tests/invalid/malformed-parent-session-id.json @@ -1,6 +1,6 @@ { "_test_description": "Session document whose parent_session_id is not a UUID. parent_session_id carries format: uuid (section 5.1); this is only rejected when the validator enforces format assertions, which the conformance runners do.", - "_expected_error": "'orchestrator-main' is not a 'uuid'", + "_expected_error": "/parent_session_id 'orchestrator-main' is not a 'uuid'", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440213", "parent_session_id": "orchestrator-main", diff --git a/tests/invalid/malformed-presentation-id.json b/tests/invalid/malformed-presentation-id.json new file mode 100644 index 0000000..ed23541 --- /dev/null +++ b/tests/invalid/malformed-presentation-id.json @@ -0,0 +1,32 @@ +{ + "_test_description": "content_engaged whose presentation_id is not a UUID (section 5.2 format: uuid).", + "_expected_error": "/events/1/presentation_id 'pres-1' is not a 'uuid'", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440610", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:06Z", + "turn_id": "1", + "presentation_id": "pres-1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + } + } + ] +} diff --git a/tests/invalid/malformed-session-id.json b/tests/invalid/malformed-session-id.json new file mode 100644 index 0000000..6997648 --- /dev/null +++ b/tests/invalid/malformed-session-id.json @@ -0,0 +1,16 @@ +{ + "_test_description": "Session document whose session_id is not a UUID (section 5.1 format: uuid).", + "_expected_error": "/session_id 'session-42' is not a 'uuid'", + "schema_version": "1.0", + "session_id": "session-42", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "agent", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + } + ] +} diff --git a/tests/invalid/malformed-started-at.json b/tests/invalid/malformed-started-at.json new file mode 100644 index 0000000..d7a4452 --- /dev/null +++ b/tests/invalid/malformed-started-at.json @@ -0,0 +1,16 @@ +{ + "_test_description": "Session document whose started_at is a bare date, not an RFC 3339 date-time (section 5.1).", + "_expected_error": "/started_at '2026-08-20' is not a 'date-time'", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440612", + "agent_id": "assistant-v4", + "started_at": "2026-08-20", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "agent", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + } + ] +} diff --git a/tests/invalid/manifest-bad-conformance-level.json b/tests/invalid/manifest-bad-conformance-level.json index 5afadc1..429b796 100644 --- a/tests/invalid/manifest-bad-conformance-level.json +++ b/tests/invalid/manifest-bad-conformance-level.json @@ -1,10 +1,14 @@ { "_test_description": "Manifest with telemetry.conformance_level 'attribution' which is not one of retrieval, grounding, citation. Fails JSON Schema validation.", - "_expected_error": "'attribution' is not one of", + "_expected_error": "/telemetry/conformance_level 'attribution' is not one of ['retrieval', 'grounding', 'citation']", "schema_version": "1.0", "id": "https://searchco.com/agents/web-search/.well-known/content-telemetry.json", - "roles": ["agent"], - "operator": { "name": "SearchCo" }, + "roles": [ + "agent" + ], + "operator": { + "name": "SearchCo" + }, "telemetry": { "endpoint": "https://telemetry.example.com/v1/events", "conformance_level": "attribution" diff --git a/tests/invalid/manifest-bad-key-type.json b/tests/invalid/manifest-bad-key-type.json new file mode 100644 index 0000000..696fd13 --- /dev/null +++ b/tests/invalid/manifest-bad-key-type.json @@ -0,0 +1,19 @@ +{ + "_test_description": "Manifest key of type 'RSA'; v1 defines Ed25519 only (section 8.4).", + "_expected_error": "/keys/0/type 'Ed25519' was expected", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "agent" + ], + "operator": { + "name": "Example Media" + }, + "keys": [ + { + "id": "key-1", + "type": "RSA", + "publicKey": "z6Mk" + } + ] +} diff --git a/tests/invalid/manifest-bad-role.json b/tests/invalid/manifest-bad-role.json index 646d362..65c115c 100644 --- a/tests/invalid/manifest-bad-role.json +++ b/tests/invalid/manifest-bad-role.json @@ -1,8 +1,12 @@ { "_test_description": "Manifest with a role value 'publisher' that is not in the roles enum (content_owner, agent, platform). Fails JSON Schema validation.", - "_expected_error": "'publisher' is not one of", + "_expected_error": "/roles/0 'publisher' is not one of ['content_owner', 'agent', 'platform']", "schema_version": "1.0", "id": "https://example.com/.well-known/content-telemetry.json", - "roles": ["publisher"], - "operator": { "name": "Example Media" } + "roles": [ + "publisher" + ], + "operator": { + "name": "Example Media" + } } diff --git a/tests/invalid/manifest-coverage-bad-mode.json b/tests/invalid/manifest-coverage-bad-mode.json index 653c5ec..6d7e424 100644 --- a/tests/invalid/manifest-coverage-bad-mode.json +++ b/tests/invalid/manifest-coverage-bad-mode.json @@ -1,6 +1,6 @@ { "_test_description": "Manifest coverage entry with a mode outside the enum. The schema requires one of complete, sampled, aggregated, selected (8.5, 5.7.6).", - "_expected_error": "'partial' is not one of ['complete', 'sampled', 'aggregated', 'selected']", + "_expected_error": "/telemetry/coverage/content_grounded/mode 'partial' is not one of ['complete', 'sampled', 'aggregated', 'selected']", "schema_version": "1.0", "id": "https://assistant.example.com/.well-known/content-telemetry.json", "roles": [ diff --git a/tests/invalid/manifest-coverage-missing-mode.json b/tests/invalid/manifest-coverage-missing-mode.json new file mode 100644 index 0000000..78a2614 --- /dev/null +++ b/tests/invalid/manifest-coverage-missing-mode.json @@ -0,0 +1,20 @@ +{ + "_test_description": "Manifest coverage entry without the required mode (sections 8.5, 5.7.6).", + "_expected_error": "/telemetry/coverage/content_grounded 'mode' is a required property", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "agent" + ], + "operator": { + "name": "Example Media" + }, + "telemetry": { + "endpoint": "https://t.example.com/v1/events", + "coverage": { + "content_grounded": { + "terms_ref": "https://terms.example.com/r/1" + } + } + } +} diff --git a/tests/invalid/manifest-ctx-resolution-http.json b/tests/invalid/manifest-ctx-resolution-http.json new file mode 100644 index 0000000..c848a40 --- /dev/null +++ b/tests/invalid/manifest-ctx-resolution-http.json @@ -0,0 +1,16 @@ +{ + "_test_description": "Agent manifest whose ctx_resolution is an http URL; the resolution endpoint MUST be https (sections 7.4, 8.5).", + "_expected_error": "/telemetry/ctx_resolution 'http://t.example.com/v1/ctx/resolve' does not match '^https://'", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "agent" + ], + "operator": { + "name": "Example Media" + }, + "telemetry": { + "endpoint": "https://t.example.com/v1/events", + "ctx_resolution": "http://t.example.com/v1/ctx/resolve" + } +} diff --git a/tests/invalid/manifest-ctx-resolution-on-content-owner.json b/tests/invalid/manifest-ctx-resolution-on-content-owner.json new file mode 100644 index 0000000..108edb7 --- /dev/null +++ b/tests/invalid/manifest-ctx-resolution-on-content-owner.json @@ -0,0 +1,16 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: content_owner manifest declaring telemetry.ctx_resolution, which is valid on agent and platform manifests only (section 8.5).", + "_expected_error": "valid on agent and platform manifests", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "content_owner" + ], + "operator": { + "name": "Example Media" + }, + "telemetry": { + "endpoint": "https://t.example.com/v1/events", + "ctx_resolution": "https://t.example.com/v1/ctx/resolve" + } +} diff --git a/tests/invalid/manifest-domains-on-path-manifest.json b/tests/invalid/manifest-domains-on-path-manifest.json new file mode 100644 index 0000000..cb7b946 --- /dev/null +++ b/tests/invalid/manifest-domains-on-path-manifest.json @@ -0,0 +1,18 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: agent manifest served under a path prefix that carries domains. domains MAY appear only on manifests served from the domain root (section 8.6).", + "_expected_error": "only root manifests carry domains", + "schema_version": "1.0", + "id": "https://example.com/agents/search/.well-known/content-telemetry.json", + "roles": [ + "agent" + ], + "operator": { + "name": "Example Media" + }, + "telemetry": { + "endpoint": "https://t.example.com/v1/events" + }, + "domains": [ + "example.com" + ] +} diff --git a/tests/invalid/manifest-duplicate-roles.json b/tests/invalid/manifest-duplicate-roles.json new file mode 100644 index 0000000..57283a1 --- /dev/null +++ b/tests/invalid/manifest-duplicate-roles.json @@ -0,0 +1,13 @@ +{ + "_test_description": "Manifest repeating a role; roles are a set (section 8.2).", + "_expected_error": "/roles ['agent', 'agent'] has non-unique elements", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "agent", + "agent" + ], + "operator": { + "name": "Example Media" + } +} diff --git a/tests/invalid/manifest-empty-roles.json b/tests/invalid/manifest-empty-roles.json new file mode 100644 index 0000000..7508e9c --- /dev/null +++ b/tests/invalid/manifest-empty-roles.json @@ -0,0 +1,10 @@ +{ + "_test_description": "Manifest with an empty roles array; at least one role is required (section 8.2).", + "_expected_error": "/roles [] should be non-empty", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [], + "operator": { + "name": "Example Media" + } +} diff --git a/tests/invalid/manifest-endpoint-http.json b/tests/invalid/manifest-endpoint-http.json new file mode 100644 index 0000000..813e39f --- /dev/null +++ b/tests/invalid/manifest-endpoint-http.json @@ -0,0 +1,15 @@ +{ + "_test_description": "Manifest whose telemetry endpoint is an http URL; section 8.5 requires an HTTPS URL.", + "_expected_error": "/telemetry/endpoint 'http://t.example.com/v1/events' does not match '^https://'", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "content_owner" + ], + "operator": { + "name": "Example Media" + }, + "telemetry": { + "endpoint": "http://t.example.com/v1/events" + } +} diff --git a/tests/invalid/manifest-id-http.json b/tests/invalid/manifest-id-http.json new file mode 100644 index 0000000..5913ae9 --- /dev/null +++ b/tests/invalid/manifest-id-http.json @@ -0,0 +1,12 @@ +{ + "_test_description": "Manifest whose id is an http URL; manifests are served over https (sections 8.1, 8.2).", + "_expected_error": "/id 'http://example.com/.well-known/content-telemetry.json' does not match '^https://[^?#]+/\\\\.well-known/content-telemetry\\\\.json$'", + "schema_version": "1.0", + "id": "http://example.com/.well-known/content-telemetry.json", + "roles": [ + "content_owner" + ], + "operator": { + "name": "Example Media" + } +} diff --git a/tests/invalid/manifest-id-not-a-uri.json b/tests/invalid/manifest-id-not-a-uri.json new file mode 100644 index 0000000..7f20fb0 --- /dev/null +++ b/tests/invalid/manifest-id-not-a-uri.json @@ -0,0 +1,12 @@ +{ + "_test_description": "Manifest whose id is not a valid URI (a space in the host). format: uri rejects it (section 8.2).", + "_expected_error": "/id 'https://exa mple.com/.well-known/content-telemetry.json' is not a 'uri'", + "schema_version": "1.0", + "id": "https://exa mple.com/.well-known/content-telemetry.json", + "roles": [ + "content_owner" + ], + "operator": { + "name": "Example Media" + } +} diff --git a/tests/invalid/manifest-id-not-well-known.json b/tests/invalid/manifest-id-not-well-known.json new file mode 100644 index 0000000..c85c1ed --- /dev/null +++ b/tests/invalid/manifest-id-not-well-known.json @@ -0,0 +1,12 @@ +{ + "_test_description": "Manifest whose id is not at the well-known location. Every manifest is served at its own /.well-known/content-telemetry.json URL (sections 8.1, 8.2).", + "_expected_error": "/id 'https://example.com/manifest.json' does not match '^https://[^?#]+/\\\\.well-known/content-telemetry\\\\.json$'", + "schema_version": "1.0", + "id": "https://example.com/manifest.json", + "roles": [ + "content_owner" + ], + "operator": { + "name": "Example Media" + } +} diff --git a/tests/invalid/manifest-key-missing-id.json b/tests/invalid/manifest-key-missing-id.json new file mode 100644 index 0000000..7e627c0 --- /dev/null +++ b/tests/invalid/manifest-key-missing-id.json @@ -0,0 +1,18 @@ +{ + "_test_description": "Manifest key without the required id (section 8.4).", + "_expected_error": "/keys/0 'id' is a required property", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "agent" + ], + "operator": { + "name": "Example Media" + }, + "keys": [ + { + "type": "Ed25519", + "publicKey": "z6Mk" + } + ] +} diff --git a/tests/invalid/manifest-lookalike-domain.json b/tests/invalid/manifest-lookalike-domain.json new file mode 100644 index 0000000..c83fbe0 --- /dev/null +++ b/tests/invalid/manifest-lookalike-domain.json @@ -0,0 +1,16 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: manifest at example.com claiming evilexample.com, a lookalike that ends in the host string but is not a subdomain of it (section 8.6).", + "_expected_error": "'evilexample.com' is not the manifest host", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "content_owner" + ], + "operator": { + "name": "Example Media" + }, + "domains": [ + "example.com", + "evilexample.com" + ] +} diff --git a/tests/invalid/manifest-missing-endpoint.json b/tests/invalid/manifest-missing-endpoint.json new file mode 100644 index 0000000..4ff043f --- /dev/null +++ b/tests/invalid/manifest-missing-endpoint.json @@ -0,0 +1,15 @@ +{ + "_test_description": "Manifest whose telemetry block lacks the required endpoint (section 8.5).", + "_expected_error": "/telemetry 'endpoint' is a required property", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "content_owner" + ], + "operator": { + "name": "Example Media" + }, + "telemetry": { + "conformance_level": "retrieval" + } +} diff --git a/tests/invalid/manifest-missing-key-publickey.json b/tests/invalid/manifest-missing-key-publickey.json index 1c49857..33f26df 100644 --- a/tests/invalid/manifest-missing-key-publickey.json +++ b/tests/invalid/manifest-missing-key-publickey.json @@ -1,11 +1,18 @@ { "_test_description": "Manifest with a keys entry missing the required 'publicKey' field. Fails JSON Schema validation.", - "_expected_error": "'publicKey' is a required property", + "_expected_error": "/keys/0 'publicKey' is a required property", "schema_version": "1.0", "id": "https://searchco.com/agents/web-search/.well-known/content-telemetry.json", - "roles": ["agent"], - "operator": { "name": "SearchCo" }, + "roles": [ + "agent" + ], + "operator": { + "name": "SearchCo" + }, "keys": [ - { "id": "key-1", "type": "Ed25519" } + { + "id": "key-1", + "type": "Ed25519" + } ] } diff --git a/tests/invalid/manifest-missing-operator.json b/tests/invalid/manifest-missing-operator.json index 729f251..6d06ee1 100644 --- a/tests/invalid/manifest-missing-operator.json +++ b/tests/invalid/manifest-missing-operator.json @@ -1,7 +1,9 @@ { "_test_description": "Manifest missing the required 'operator' field. Fails JSON Schema validation.", - "_expected_error": "'operator' is a required property", + "_expected_error": "/ 'operator' is a required property", "schema_version": "1.0", "id": "https://example.com/.well-known/content-telemetry.json", - "roles": ["content_owner"] + "roles": [ + "content_owner" + ] } diff --git a/tests/invalid/missing-event-timestamp.json b/tests/invalid/missing-event-timestamp.json index 64437ad..10e4d54 100644 --- a/tests/invalid/missing-event-timestamp.json +++ b/tests/invalid/missing-event-timestamp.json @@ -1,6 +1,6 @@ { "_test_description": "Event without timestamp field. Fails JSON Schema validation: timestamp is required on TelemetryEvent.", - "_expected_error": "'timestamp' is a required property", + "_expected_error": "/events/0 'timestamp' is a required property", "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440005", "started_at": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/missing-event-type.json b/tests/invalid/missing-event-type.json index 07a1c5d..4eb04af 100644 --- a/tests/invalid/missing-event-type.json +++ b/tests/invalid/missing-event-type.json @@ -1,6 +1,6 @@ { "_test_description": "Event without type field. Fails JSON Schema validation: type is required on TelemetryEvent.", - "_expected_error": "'type' is a required property", + "_expected_error": "/events/0 'type' is a required property", "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440004", "started_at": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/missing-privacy-level.json b/tests/invalid/missing-privacy-level.json index fc38f4d..c36d913 100644 --- a/tests/invalid/missing-privacy-level.json +++ b/tests/invalid/missing-privacy-level.json @@ -1,6 +1,6 @@ { "_test_description": "Turn object without privacy_level. Fails JSON Schema validation: privacy_level is required on ConversationTurn.", - "_expected_error": "'privacy_level' is a required property", + "_expected_error": "/events/0/turn 'privacy_level' is a required property", "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440006", "started_at": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/missing-schema-version.json b/tests/invalid/missing-schema-version.json index d275e03..ba6660a 100644 --- a/tests/invalid/missing-schema-version.json +++ b/tests/invalid/missing-schema-version.json @@ -1,6 +1,6 @@ { "_test_description": "Session without schema_version. Fails JSON Schema validation: schema_version is a required field.", - "_expected_error": "'schema_version' is a required property", + "_expected_error": "/ 'schema_version' is a required property", "session_id": "770e8400-e29b-41d4-a716-446655440001", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/missing-session-id.json b/tests/invalid/missing-session-id.json index 20e400a..84cd0ae 100644 --- a/tests/invalid/missing-session-id.json +++ b/tests/invalid/missing-session-id.json @@ -1,6 +1,6 @@ { "_test_description": "Session without session_id. Fails JSON Schema validation: session_id is a required field.", - "_expected_error": "'session_id' is a required property", + "_expected_error": "/ 'session_id' is a required property", "schema_version": "1.0", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/missing-started-at.json b/tests/invalid/missing-started-at.json index 54f9f63..e5d35c4 100644 --- a/tests/invalid/missing-started-at.json +++ b/tests/invalid/missing-started-at.json @@ -1,6 +1,6 @@ { "_test_description": "Session without started_at. Fails JSON Schema validation: started_at is a required field.", - "_expected_error": "'started_at' is a required property", + "_expected_error": "/ 'started_at' is a required property", "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440003", "events": [ diff --git a/tests/invalid/presentation-id-on-cited.json b/tests/invalid/presentation-id-on-cited.json new file mode 100644 index 0000000..c4c93ce --- /dev/null +++ b/tests/invalid/presentation-id-on-cited.json @@ -0,0 +1,35 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: presentation_id on a content_cited event. Valid only on content_engaged (sections 5.2, 5.7.5).", + "_expected_error": "Field 'presentation_id' present on 'content_cited' event", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440633", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "citation_type": "reference", + "position": "primary" + }, + "presentation_id": "770e8400-e29b-41d4-a716-446655440602" + } + ] +} diff --git a/tests/invalid/presentation-kind-invalid.json b/tests/invalid/presentation-kind-invalid.json index 2499eeb..e64cf95 100644 --- a/tests/invalid/presentation-kind-invalid.json +++ b/tests/invalid/presentation-kind-invalid.json @@ -1,6 +1,6 @@ { "_test_description": "content_presented with presentation_kind 'reference'. presentation_kind is a closed two-value distinction between source content and a reference to the source: only 'content' and 'source_reference' are valid (section 6.7).", - "_expected_error": "'reference' is not one of", + "_expected_error": "/events/0/data/presentation_kind 'reference' is not one of ['content', 'source_reference']", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440208", "started_at": "2026-08-05T09:50:00Z", diff --git a/tests/invalid/presented-citation-content-mismatch.json b/tests/invalid/presented-citation-content-mismatch.json new file mode 100644 index 0000000..75d5d0e --- /dev/null +++ b/tests/invalid/presented-citation-content-mismatch.json @@ -0,0 +1,35 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: content_presented.citation_id references a citation of different content (sections 5.7.5, 6.7).", + "_expected_error": "content_presented references citation", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440637", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-review.com/headphones/best-noise-cancelling", + "data": { + "citation_type": "reference", + "position": "primary" + } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "citation_id": "770e8400-e29b-41d4-a716-446655440601", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + } + ] +} diff --git a/tests/invalid/presented-missing-data.json b/tests/invalid/presented-missing-data.json new file mode 100644 index 0000000..049fbff --- /dev/null +++ b/tests/invalid/presented-missing-data.json @@ -0,0 +1,18 @@ +{ + "_test_description": "content_presented with no data object. presentation_kind and presentation_type are required, so data itself is required (section 6.7).", + "_expected_error": "/events/0 'data' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440614", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + } + ] +} diff --git a/tests/invalid/presented-missing-id.json b/tests/invalid/presented-missing-id.json index 6bdd533..8cf702f 100644 --- a/tests/invalid/presented-missing-id.json +++ b/tests/invalid/presented-missing-id.json @@ -1,6 +1,6 @@ { "_test_description": "content_presented event with output_id and presentation data but no event id. id is required on presentation events so a later content_engaged.presentation_id can reference the exact surface occurrence (section 6.7).", - "_expected_error": "'id' is a required property", + "_expected_error": "/events/0 'id' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440200", "started_at": "2026-08-05T09:00:00Z", diff --git a/tests/invalid/presented-missing-kind.json b/tests/invalid/presented-missing-kind.json index 69b4fdf..b0dc287 100644 --- a/tests/invalid/presented-missing-kind.json +++ b/tests/invalid/presented-missing-kind.json @@ -1,6 +1,6 @@ { "_test_description": "content_presented must distinguish source content from a source reference.", - "_expected_error": "'presentation_kind' is a required property", + "_expected_error": "/events/0/data 'presentation_kind' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440090", "started_at": "2026-07-18T13:00:00Z", @@ -11,7 +11,9 @@ "timestamp": "2026-07-18T13:00:01Z", "output_id": "response:1", "content_id": "publisher:article:1", - "data": { "presentation_type": "link" } + "data": { + "presentation_type": "link" + } } ] } diff --git a/tests/invalid/presented-missing-output-id.json b/tests/invalid/presented-missing-output-id.json index ff1ac5c..05c2db5 100644 --- a/tests/invalid/presented-missing-output-id.json +++ b/tests/invalid/presented-missing-output-id.json @@ -1,6 +1,6 @@ { "_test_description": "content_presented event with an event id and presentation data but no output_id. output_id is required on presentation events so presentation can be correlated with the output artifact it delivers (section 6.7).", - "_expected_error": "'output_id' is a required property", + "_expected_error": "/events/0 'output_id' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440201", "started_at": "2026-08-05T09:10:00Z", diff --git a/tests/invalid/presented-missing-presentation-type.json b/tests/invalid/presented-missing-presentation-type.json index 43858c8..15cebb0 100644 --- a/tests/invalid/presented-missing-presentation-type.json +++ b/tests/invalid/presented-missing-presentation-type.json @@ -1,6 +1,6 @@ { "_test_description": "content_presented event carrying presentation_kind but no presentation_type. Both are required: the kind says what crossed the presentation boundary, the type says how it was made perceivable.", - "_expected_error": "'presentation_type' is a required property", + "_expected_error": "/events/0/data 'presentation_type' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440163", "started_at": "2026-08-01T15:00:00Z", @@ -11,7 +11,9 @@ "timestamp": "2026-08-01T15:00:01Z", "output_id": "response:1", "content_id": "publisher:article:4", - "data": { "presentation_kind": "source_reference" } + "data": { + "presentation_kind": "source_reference" + } } ] } diff --git a/tests/invalid/privacy-violation-query-at-minimal-batch.json b/tests/invalid/privacy-violation-query-at-minimal-batch.json new file mode 100644 index 0000000..77b0b70 --- /dev/null +++ b/tests/invalid/privacy-violation-query-at-minimal-batch.json @@ -0,0 +1,21 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: event batch whose turn_completed at minimal privacy carries query_text (section 5.5).", + "_expected_error": "'query_text' present on turn with privacy_level 'minimal'", + "document_type": "event_batch", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440644", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "turn_completed", + "timestamp": "2026-08-20T10:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "minimal", + "query_text": "What is the best coffee grinder under 100 pounds?", + "response_tokens": 120 + } + } + ] +} diff --git a/tests/invalid/privacy-violation-query-at-minimal-standalone.json b/tests/invalid/privacy-violation-query-at-minimal-standalone.json new file mode 100644 index 0000000..4bacc20 --- /dev/null +++ b/tests/invalid/privacy-violation-query-at-minimal-standalone.json @@ -0,0 +1,19 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: standalone envelope carrying a turn_completed at minimal privacy with query_text. The privacy gate applies wherever turns are emitted (section 5.5).", + "_expected_error": "'query_text' present on turn with privacy_level 'minimal'", + "document_type": "event", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440643", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "event": { + "type": "turn_completed", + "timestamp": "2026-08-20T10:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "minimal", + "query_text": "What is the best coffee grinder under 100 pounds?", + "response_tokens": 120 + } + } +} diff --git a/tests/invalid/privacy-violation-response-text-at-intent.json b/tests/invalid/privacy-violation-response-text-at-intent.json new file mode 100644 index 0000000..a02204c --- /dev/null +++ b/tests/invalid/privacy-violation-response-text-at-intent.json @@ -0,0 +1,21 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: turn at intent privacy with response_text present (section 5.5).", + "_expected_error": "'response_text' present on turn with privacy_level 'intent'", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440642", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "turn_completed", + "timestamp": "2026-08-20T10:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "response_text": "The Baratza Encore ESP is the best grinder under 100 pounds.", + "query_intent": "comparison", + "response_tokens": 120 + } + } + ] +} diff --git a/tests/invalid/reproduced-citation-id-unmatched.json b/tests/invalid/reproduced-citation-id-unmatched.json new file mode 100644 index 0000000..5f3a062 --- /dev/null +++ b/tests/invalid/reproduced-citation-id-unmatched.json @@ -0,0 +1,23 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: content_reproduced.citation_id matches no content_cited event id in the session (section 6.6).", + "_expected_error": "content_reproduced citation_id '00000000-0000-0000-0000-000000000000' does not match", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440638", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440603", + "type": "content_reproduced", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "reproduction_type": "verbatim", + "reproduced_chars": 180 + }, + "citation_id": "00000000-0000-0000-0000-000000000000" + } + ] +} diff --git a/tests/invalid/reproduced-missing-data.json b/tests/invalid/reproduced-missing-data.json new file mode 100644 index 0000000..a1c77d1 --- /dev/null +++ b/tests/invalid/reproduced-missing-data.json @@ -0,0 +1,18 @@ +{ + "_test_description": "content_reproduced with no data object. reproduction_type is required, so data itself is required (section 6.6).", + "_expected_error": "/events/0 'data' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440616", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440603", + "type": "content_reproduced", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + } + ] +} diff --git a/tests/invalid/reproduced-missing-id.json b/tests/invalid/reproduced-missing-id.json index cff1bc9..05014a1 100644 --- a/tests/invalid/reproduced-missing-id.json +++ b/tests/invalid/reproduced-missing-id.json @@ -1,6 +1,6 @@ { "_test_description": "content_reproduced event with output_id and a resolvable source reference but no event id. id is required on reproduction events so a crediting citation or verification result can reference the exact reproduction claim.", - "_expected_error": "'id' is a required property", + "_expected_error": "/events/0 'id' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440162", "started_at": "2026-08-01T14:30:00Z", @@ -10,7 +10,10 @@ "timestamp": "2026-08-01T14:30:01Z", "output_id": "response:1", "content_id": "publisher:article:3", - "data": { "reproduction_type": "verbatim", "reproduced_chars": 320 } + "data": { + "reproduction_type": "verbatim", + "reproduced_chars": 320 + } } ] } diff --git a/tests/invalid/reproduced-missing-output-id.json b/tests/invalid/reproduced-missing-output-id.json index 7a2a067..258c4cc 100644 --- a/tests/invalid/reproduced-missing-output-id.json +++ b/tests/invalid/reproduced-missing-output-id.json @@ -1,6 +1,6 @@ { "_test_description": "content_reproduced must identify the output artifact containing the reproduction: id and output_id are required so reproduction can be correlated with citation and presentation.", - "_expected_error": "'output_id' is a required property", + "_expected_error": "/events/0 'output_id' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440130", "started_at": "2026-08-01T11:00:00Z", @@ -10,7 +10,10 @@ "type": "content_reproduced", "timestamp": "2026-08-01T11:00:01Z", "content_id": "publisher:article:2", - "data": { "reproduction_type": "verbatim", "reproduced_chars": 500 } + "data": { + "reproduction_type": "verbatim", + "reproduced_chars": 500 + } } ] } diff --git a/tests/invalid/reproduced-missing-source-reference.json b/tests/invalid/reproduced-missing-source-reference.json index c75a3d7..ac05921 100644 --- a/tests/invalid/reproduced-missing-source-reference.json +++ b/tests/invalid/reproduced-missing-source-reference.json @@ -1,6 +1,6 @@ { "_test_description": "content_reproduced event carrying neither content_url nor content_id. A reproduction claim is only meaningful for an identified source; the JSON Schema requires a non-null content_url or content_id on content_reproduced (section 6.6).", - "_expected_error": "'content_url' is a required property", + "_expected_error": "/events/0 'content_url' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440150", "started_at": "2026-08-01T13:00:00Z", @@ -10,7 +10,10 @@ "type": "content_reproduced", "timestamp": "2026-08-01T13:00:01Z", "output_id": "response:1", - "data": { "reproduction_type": "verbatim", "reproduced_chars": 250 } + "data": { + "reproduction_type": "verbatim", + "reproduced_chars": 250 + } } ] } diff --git a/tests/invalid/reproduced-missing-type.json b/tests/invalid/reproduced-missing-type.json index 25e8e67..09c6fc3 100644 --- a/tests/invalid/reproduced-missing-type.json +++ b/tests/invalid/reproduced-missing-type.json @@ -1,6 +1,6 @@ { "_test_description": "content_reproduced must classify the fidelity of the reproduction: data.reproduction_type is required.", - "_expected_error": "'reproduction_type' is a required property", + "_expected_error": "/events/0/data 'reproduction_type' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440120", "started_at": "2026-08-01T10:00:00Z", @@ -11,7 +11,9 @@ "timestamp": "2026-08-01T10:00:01Z", "output_id": "response:1", "content_id": "publisher:article:1", - "data": { "reproduced_chars": 300 } + "data": { + "reproduced_chars": 300 + } } ] } diff --git a/tests/invalid/reproduced-negative-chars.json b/tests/invalid/reproduced-negative-chars.json index 97374c3..56ae558 100644 --- a/tests/invalid/reproduced-negative-chars.json +++ b/tests/invalid/reproduced-negative-chars.json @@ -1,6 +1,6 @@ { "_test_description": "content_reproduced with a negative reproduced_chars. The reproduced span length is a Unicode code point count and cannot be negative (section 6.6); the schema requires a minimum of 0.", - "_expected_error": "reproduced_chars -412 is less than the minimum", + "_expected_error": "/events/0/data/reproduced_chars -412 is less than the minimum of 0", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440211", "started_at": "2026-08-05T10:10:00Z", diff --git a/tests/invalid/reproduction-type-invalid.json b/tests/invalid/reproduction-type-invalid.json index 04be830..e7e39ac 100644 --- a/tests/invalid/reproduction-type-invalid.json +++ b/tests/invalid/reproduction-type-invalid.json @@ -1,6 +1,6 @@ { "_test_description": "content_reproduced with reproduction_type 'paraphrase'. Paraphrase is not reproduction (section 6.6): a credited paraphrase is a content_cited event with citation_type 'paraphrase', and an uncredited one is silent grounding. reproduction_type accepts only verbatim, near_verbatim, unclassified.", - "_expected_error": "'paraphrase' is not one of", + "_expected_error": "/events/0/data/reproduction_type 'paraphrase' is not one of ['verbatim', 'near_verbatim', 'unclassified']", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440206", "started_at": "2026-08-05T09:40:00Z", diff --git a/tests/invalid/retrieved-bad-country.json b/tests/invalid/retrieved-bad-country.json new file mode 100644 index 0000000..df9ed41 --- /dev/null +++ b/tests/invalid/retrieved-bad-country.json @@ -0,0 +1,19 @@ +{ + "_test_description": "Edge retrieval with country 'usa'; the field is an ISO 3166-1 alpha-2 code (section 6.2).", + "_expected_error": "/events/0/data/country 'usa' does not match '^[A-Z]{2}$'", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440624", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "edge", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "country": "usa" + } + } + ] +} diff --git a/tests/invalid/retrieved-missing-source-role.json b/tests/invalid/retrieved-missing-source-role.json new file mode 100644 index 0000000..055a19c --- /dev/null +++ b/tests/invalid/retrieved-missing-source-role.json @@ -0,0 +1,15 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: content_retrieved without source_role. Every retrieval event MUST say who observed it (sections 5.2.2, 5.7.1).", + "_expected_error": "content_retrieved event carries no source_role", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440630", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + } + ] +} diff --git a/tests/invalid/retrieved-response-status-out-of-range.json b/tests/invalid/retrieved-response-status-out-of-range.json new file mode 100644 index 0000000..cebdceb --- /dev/null +++ b/tests/invalid/retrieved-response-status-out-of-range.json @@ -0,0 +1,19 @@ +{ + "_test_description": "Edge retrieval with response_status 999, outside the HTTP status range 100-599 (section 6.2).", + "_expected_error": "/events/0/data/response_status 999 is greater than the maximum of 599", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440623", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "edge", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "response_status": 999 + } + } + ] +} diff --git a/tests/invalid/shared-ctx-token-two-presentations.json b/tests/invalid/shared-ctx-token-two-presentations.json new file mode 100644 index 0000000..fec7720 --- /dev/null +++ b/tests/invalid/shared-ctx-token-two-presentations.json @@ -0,0 +1,56 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: one event-level ctx_token appears on engagements bound to two different presentations. A token is minted for exactly one presentation occurrence (section 7.4.1).", + "_expected_error": "appears on engagements bound to two presentations", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440640", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440604", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:05Z", + "turn_id": "1", + "output_id": "response:2", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:06Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440602", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + }, + "ctx_token": "ct_5d1c9e4b2a7f8036" + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:07Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440604", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + }, + "ctx_token": "ct_5d1c9e4b2a7f8036" + } + ] +} diff --git a/tests/invalid/standalone-ctx-token-non-engagement.json b/tests/invalid/standalone-ctx-token-non-engagement.json new file mode 100644 index 0000000..8c9f427 --- /dev/null +++ b/tests/invalid/standalone-ctx-token-non-engagement.json @@ -0,0 +1,17 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: standalone envelope whose only session reference is a ctx_token but whose event is a content_grounded. An envelope ctx_token accompanies content_engaged events only (section 7.1).", + "_expected_error": "Envelope ctx_token accompanies a 'content_grounded' event", + "document_type": "event", + "schema_version": "1.0", + "ctx_token": "ct_9f3a1c7e2b8d4a06", + "event": { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 4200 + } + } +} diff --git a/tests/invalid/standalone-malformed-session-id.json b/tests/invalid/standalone-malformed-session-id.json new file mode 100644 index 0000000..03989f4 --- /dev/null +++ b/tests/invalid/standalone-malformed-session-id.json @@ -0,0 +1,13 @@ +{ + "_test_description": "Standalone envelope whose session_id is not a UUID (telemetry-event.json format: uuid).", + "_expected_error": "/session_id 'session-42' is not a 'uuid'", + "document_type": "event", + "schema_version": "1.0", + "session_id": "session-42", + "event": { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "agent", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + } +} diff --git a/tests/invalid/standalone-missing-document-type.json b/tests/invalid/standalone-missing-document-type.json index 497f55a..63508ae 100644 --- a/tests/invalid/standalone-missing-document-type.json +++ b/tests/invalid/standalone-missing-document-type.json @@ -1,6 +1,6 @@ { "_test_description": "Envelope-shaped document with an event key but no document_type. Per section 7.1 a document without document_type is treated as a session, and it fails the session schema.", - "_expected_error": "'session_id' is a required property", + "_expected_error": "/ 'session_id' is a required property", "schema_version": "1.0", "event": { "type": "content_retrieved", diff --git a/tests/invalid/standalone-missing-event.json b/tests/invalid/standalone-missing-event.json index befc66e..22ad0cc 100644 --- a/tests/invalid/standalone-missing-event.json +++ b/tests/invalid/standalone-missing-event.json @@ -1,6 +1,6 @@ { "_test_description": "Standalone event envelope (document_type 'event') with no event field. Fails the envelope schema, which requires event.", - "_expected_error": "'event' is a required property", + "_expected_error": "/ 'event' is a required property", "document_type": "event", "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440010" diff --git a/tests/invalid/standalone-missing-schema-version.json b/tests/invalid/standalone-missing-schema-version.json new file mode 100644 index 0000000..61b474b --- /dev/null +++ b/tests/invalid/standalone-missing-schema-version.json @@ -0,0 +1,12 @@ +{ + "_test_description": "Standalone event envelope missing schema_version, which the envelope schema requires.", + "_expected_error": "/ 'schema_version' is a required property", + "document_type": "event", + "session_id": "660e8400-e29b-41d4-a716-446655440626", + "event": { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "agent", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + } +} diff --git a/tests/invalid/turn-on-content-event.json b/tests/invalid/turn-on-content-event.json new file mode 100644 index 0000000..f6a4301 --- /dev/null +++ b/tests/invalid/turn-on-content-event.json @@ -0,0 +1,23 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: a turn object on a content_grounded event. Turn data is carried on turn_started and turn_completed only (sections 5.2, 5.4).", + "_expected_error": "Field 'turn' present on 'content_grounded' event", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440634", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 4200 + }, + "turn": { + "privacy_level": "minimal" + } + } + ] +} diff --git a/tests/invalid/withdrawn-ip-hash-batch.json b/tests/invalid/withdrawn-ip-hash-batch.json new file mode 100644 index 0000000..60ec597 --- /dev/null +++ b/tests/invalid/withdrawn-ip-hash-batch.json @@ -0,0 +1,18 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: event batch whose retrieval carries the withdrawn ip_hash (section 9.1), in the batch shape.", + "_expected_error": "'ip_hash' in data", + "document_type": "event_batch", + "schema_version": "1.0", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "edge", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "bot_category": "inference", + "ip_hash": "sha256:e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6" + } + } + ] +} diff --git a/tests/invalid/withdrawn-ip-hash-session.json b/tests/invalid/withdrawn-ip-hash-session.json new file mode 100644 index 0000000..ed8a10b --- /dev/null +++ b/tests/invalid/withdrawn-ip-hash-session.json @@ -0,0 +1,20 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: session document whose retrieval carries the withdrawn ip_hash (section 9.1), in the session shape.", + "_expected_error": "'ip_hash' in data", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440647", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "edge", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "bot_category": "inference", + "ip_hash": "sha256:e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6" + } + } + ] +} diff --git a/tests/mutation_smoke.py b/tests/mutation_smoke.py index 9783d5c..0b16cf9 100644 --- a/tests/mutation_smoke.py +++ b/tests/mutation_smoke.py @@ -1,9 +1,9 @@ #!/usr/bin/env python3 """Mutation smoke test for the conformance suite. -Replays the review's key mutations - each of which the suite previously -missed - against a scratch copy of the repository and confirms validate.py -now fails under every one of them. The working tree is never modified. +Replays suite-weakening mutations - each of which the suite once missed - +against a scratch copy of the repository and confirms validate.py fails under +every one of them. The working tree is never modified. A mutation "survives" when the mutated suite still exits 0; any survivor is a real detection gap and fails this script. @@ -88,6 +88,199 @@ def mutate_zero_uuid_engagement(root): ] +# --------------------------------------------------------------------------- +# Review mutations (22 August pre-freeze review). Each weakens one rule the +# suite previously pinned in a single document shape, or not at all, and names +# the fixture that now fails under it. +# --------------------------------------------------------------------------- + +def _edit_text(root, rel, old, new): + path = root / rel + src = path.read_text() + assert old in src, (rel, old) + path.write_text(src.replace(old, new)) + + +def _edit_json(root, rel, fn): + path = root / rel + doc = json.loads(path.read_text()) + fn(doc) + path.write_text(json.dumps(doc, indent=2) + "\n") + + +def _event_branch(schema, etype): + for branch in schema["$defs"]["TelemetryEvent"]["allOf"]: + if branch["if"]["properties"]["type"]["const"] == etype: + return branch["then"] + raise KeyError(etype) + + +def _text(rel, old, new, fixture): + def mutate(root): + _edit_text(root, rel, old, new) + return fixture + return mutate + + +def _json(rel, fn, fixture): + def mutate(root): + _edit_json(root, rel, fn) + return fixture + return mutate + + +_V = "tests/validate.py" +_S = "telemetry-session.json" +_E = "telemetry-event.json" +_B = "telemetry-event-batch.json" +_M = "manifest.json" + +REVIEW_MUTATIONS = [ + # validate.py application-layer weakenings + ("drop content_reproduced from the citation_id integrity check", + _text(_V, 'if etype in ("content_presented", "content_reproduced"):', 'if etype in ("content_presented",):', + "reproduced-citation-id-unmatched.json")), + ("exempt any envelope containing a retrieval from the session/ctx_token rule", + _text(_V, 'if types <= {"content_retrieved"}:', 'if "content_retrieved" in types:', + "batch-missing-session-mixed-retrieval.json")), + ("stop forbidding response_text at intent", + _text(_V, '"intent": {\n "query_text", "response_text",\n },', '"intent": {\n "query_text",\n },', + "privacy-violation-response-text-at-intent.json")), + ("drop the agent_cached -> cached:true half of the provenance rule", + _text(_V, 'if provenance == "agent_cached" and cached is not True:', 'if False:', + "grounding-provenance-cached-conflict-agent-cached.json")), + ("skip referential integrity when the session has no content_presented events", + _text(_V, 'if pid and pid not in presented_ids:', 'if presented_ids and pid and pid not in presented_ids:', + "engaged-presentation-id-no-presentations.json")), + ("privacy gating on session documents only", + _text(_V, ' violations = []\n\n for event in _iter_events(data):\n turn = event.get("turn")', + ' violations = []\n if is_standalone_event(data) or is_event_batch(data):\n return []\n for event in _iter_events(data):\n turn = event.get("turn")', + "privacy-violation-query-at-minimal-standalone.json")), + ("content-identifier rule on session documents only", + _text(_V, ' violations = []\n for event in _iter_events(data):\n if event.get("type") not in CONTENT_EVENT_TYPES:', + ' violations = []\n if is_standalone_event(data) or is_event_batch(data):\n return []\n for event in _iter_events(data):\n if event.get("type") not in CONTENT_EVENT_TYPES:', + "grounded-missing-identifier-standalone.json")), + ("ip_hash prohibition on standalone envelopes only", + _text(_V, ' violations = []\n for event in _iter_events(data):\n event_data = event.get("data")\n if not isinstance(event_data, dict):\n continue\n for field, reason', + ' violations = []\n if not is_standalone_event(data):\n return []\n for event in _iter_events(data):\n event_data = event.get("data")\n if not isinstance(event_data, dict):\n continue\n for field, reason', + "withdrawn-ip-hash-session.json")), + ("domains check accepts lookalike hosts (endswith without the dot)", + _text(_V, 'not bare.endswith("." + host)', 'not bare.endswith(host)', "manifest-lookalike-domain.json")), + ("drop the source_role requirement on retrievals", + _text(_V, 'if etype == "content_retrieved" and not event.get("source_role"):', 'if False:', + "retrieved-missing-source-role.json")), + ("drop the field-placement rules", + _text(_V, ' if event.get(field) is not None and etype not in allowed:', ' if False:', + "ctx-token-on-grounded.json")), + ("drop the duplicate event id rule", + _text(_V, ' if eid in seen_ids:', ' if False:', "duplicate-event-id.json")), + ("drop the same-content rule for engagements", + _text(_V, 'if pid in presented and not _same_content(e, presented[pid]):', 'if False:', + "engaged-presentation-content-mismatch.json")), + ("drop the one-token-one-presentation rule", + _text(_V, ' if bound != pid:', ' if False:', "shared-ctx-token-two-presentations.json")), + ("drop the envelope ctx_token rule", + _text(_V, " if not data.get(\"ctx_token\"):\n return []\n violations = []", " return []\n violations = []", + "standalone-ctx-token-non-engagement.json")), + ("drop the root-manifest rule for domains", + _text(_V, 'if "domains" in data and parsed.path != "/.well-known/content-telemetry.json":', 'if False:', + "manifest-domains-on-path-manifest.json")), + ("drop the agent/platform rule for ctx_resolution", + _text(_V, 'if telemetry.get("ctx_resolution") and not roles & {"agent", "platform"}:', 'if False:', + "manifest-ctx-resolution-on-content-owner.json")), + # telemetry-session.json + ("accept the v0.1 wire version", + _text(_S, '"const": "1.0",\n "description": "Content Telemetry schema version"', + '"enum": ["0.1", "1.0"],\n "description": "Content Telemetry schema version"', + "legacy-schema-version-0-1.json")), + ("remove the event-level ctx_token pattern", + _json(_S, lambda d: d["$defs"]["TelemetryEvent"]["properties"]["ctx_token"].pop("pattern"), + "event-level-ctx-token-bad-pattern.json")), + ("remove format:uuid from event ids", + _json(_S, lambda d: [d["$defs"]["TelemetryEvent"]["properties"][k].pop("format") for k in ("id", "citation_id", "presentation_id")], + "malformed-event-id.json")), + ("open the CitationType enum", + _json(_S, lambda d: d["$defs"]["CitationType"].pop("enum"), "citation-type-invalid.json")), + ("open the CitationPosition enum", + _json(_S, lambda d: d["$defs"]["CitationPosition"].pop("enum"), "citation-position-invalid.json")), + ("open the GroundingScope enum", + _json(_S, lambda d: d["$defs"]["GroundingScope"].pop("enum"), "grounding-scope-invalid.json")), + ("open the SourceProvenance enum", + _json(_S, lambda d: d["$defs"]["SourceProvenance"].pop("enum"), "grounding-provenance-invalid.json")), + ("drop required scope on content_grounded", + _json(_S, lambda d: _event_branch(d, "content_grounded")["properties"]["data"].pop("required"), + "grounding-scope-missing.json")), + ("drop required data on content_grounded", + _json(_S, lambda d: _event_branch(d, "content_grounded")["required"].remove("data"), "grounding-missing-data.json")), + ("drop required data on content_presented", + _json(_S, lambda d: _event_branch(d, "content_presented")["required"].remove("data"), "presented-missing-data.json")), + ("drop required data on content_cited", + _json(_S, lambda d: _event_branch(d, "content_cited")["required"].remove("data"), "cited-missing-data.json")), + ("drop required data on content_reproduced", + _json(_S, lambda d: _event_branch(d, "content_reproduced")["required"].remove("data"), "reproduced-missing-data.json")), + ("remove the sha256 hash patterns", + _text(_S, '"pattern": "^sha256:[a-f0-9]{64}$"', '"type": "string"', "malformed-content-hash.json")), + ("remove minLength on output_id", + _json(_S, lambda d: d["$defs"]["TelemetryEvent"]["properties"]["output_id"].pop("minLength"), "empty-output-id.json")), + ("remove minimum on tokens_ingested", + _json(_S, lambda d: _event_branch(d, "content_grounded")["properties"]["data"]["properties"]["tokens_ingested"].pop("minimum"), + "grounded-negative-tokens-ingested.json")), + ("remove minimum on excerpt_chars", + _json(_S, lambda d: _event_branch(d, "content_cited")["properties"]["data"]["properties"]["excerpt_chars"].pop("minimum"), + "cited-negative-excerpt-chars.json")), + ("remove the response_status bounds", + _json(_S, lambda d: _event_branch(d, "content_retrieved")["properties"]["data"]["properties"]["response_status"].pop("maximum"), + "retrieved-response-status-out-of-range.json")), + ("remove the country pattern", + _json(_S, lambda d: _event_branch(d, "content_retrieved")["properties"]["data"]["properties"]["country"].pop("pattern"), + "retrieved-bad-country.json")), + ("drop required scheme on content_fingerprint", + _json(_S, lambda d: d["$defs"]["ContentFingerprint"]["required"].remove("scheme"), "grounding-fingerprint-missing-scheme.json")), + ("drop format:uri from the turn URL arrays", + _json(_S, lambda d: [d["$defs"]["ConversationTurn"]["properties"][k]["items"].pop("format") for k in ("content_urls_retrieved", "content_urls_cited")], + "malformed-content-urls-cited.json")), + ("drop format:date-time from started_at", + _json(_S, lambda d: d["properties"]["started_at"].pop("format"), "malformed-started-at.json")), + ("drop format:uuid from session_id", + _json(_S, lambda d: d["properties"]["session_id"].pop("format"), "malformed-session-id.json")), + # envelope schemas + ("batch: drop the presentation_id-unless-ctx_token conditional", + _json(_B, lambda d: d.pop("allOf"), "batch-engaged-missing-presentation-id.json")), + ("batch: remove the envelope ctx_token pattern", + _json(_B, lambda d: d["properties"]["ctx_token"].pop("pattern"), "batch-ctx-token-bad-pattern.json")), + ("event envelope: drop schema_version from required", + _json(_E, lambda d: d["required"].remove("schema_version"), "standalone-missing-schema-version.json")), + ("event envelope: drop format:uuid from session_id", + _json(_E, lambda d: d["properties"]["session_id"].pop("format"), "standalone-malformed-session-id.json")), + # manifest.json + ("manifest: drop required mode on coverage entries", + _json(_M, lambda d: d["properties"]["telemetry"]["properties"]["coverage"]["additionalProperties"].pop("required"), + "manifest-coverage-missing-mode.json")), + ("manifest: drop required endpoint", + _json(_M, lambda d: d["properties"]["telemetry"].pop("required"), "manifest-missing-endpoint.json")), + ("manifest: drop minItems on roles", + _json(_M, lambda d: d["properties"]["roles"].pop("minItems"), "manifest-empty-roles.json")), + ("manifest: drop uniqueItems on roles", + _json(_M, lambda d: d["properties"]["roles"].pop("uniqueItems"), "manifest-duplicate-roles.json")), + ("manifest: drop the https pattern on ctx_resolution", + _json(_M, lambda d: d["properties"]["telemetry"]["properties"]["ctx_resolution"].pop("pattern"), "manifest-ctx-resolution-http.json")), + ("manifest: drop the https pattern on endpoint", + _json(_M, lambda d: d["properties"]["telemetry"]["properties"]["endpoint"].pop("pattern"), "manifest-endpoint-http.json")), + ("manifest: drop the https pattern on id", + _json(_M, lambda d: d["properties"]["id"].pop("pattern"), "manifest-id-http.json")), + ("manifest: drop the well-known location pattern on id", + _json(_M, lambda d: d["properties"]["id"].pop("pattern"), "manifest-id-not-well-known.json")), + ("manifest: drop format:uri on id", + _json(_M, lambda d: d["properties"]["id"].pop("format"), "manifest-id-not-a-uri.json")), + ("manifest: drop const Ed25519 on keys[].type", + _json(_M, lambda d: d["properties"]["keys"]["items"]["properties"]["type"].pop("const"), "manifest-bad-key-type.json")), + ("manifest: drop required id on keys[]", + _json(_M, lambda d: d["properties"]["keys"]["items"]["required"].remove("id"), "manifest-key-missing-id.json")), +] + +MUTATIONS += REVIEW_MUTATIONS + + def run_one(name, mutate): with tempfile.TemporaryDirectory(prefix="ct-mutation-") as tmp: root = Path(tmp) / "repo" diff --git a/tests/valid/event-batch-destination-engagements.json b/tests/valid/event-batch-destination-engagements.json new file mode 100644 index 0000000..1fdfc2b --- /dev/null +++ b/tests/valid/event-batch-destination-engagements.json @@ -0,0 +1,25 @@ +{ + "_test_description": "Destination-reported engagements delivered as a batch under one ctx_token: a direct-link surface minted the token per presentation, so two clicks on the same presentation share it and are distinguished by timestamp at resolution (section 7.4.1). No presentation_id and no session_id: the consumer restores both from the token.", + "document_type": "event_batch", + "schema_version": "1.0", + "ctx_token": "ct_77b41f0ac93e5d28", + "manifest_ref": "https://shop.example.com/.well-known/content-telemetry.json", + "events": [ + { + "type": "content_engaged", + "timestamp": "2026-08-20T11:15:42Z", + "content_url": "https://shop.example.com/products/anc-headphones-x9", + "data": { + "engagement_type": "link_click" + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T11:21:09Z", + "content_url": "https://shop.example.com/products/anc-headphones-x9", + "data": { + "engagement_type": "link_click" + } + } + ] +} diff --git a/tests/valid/event-standalone-grounded-provenance.json b/tests/valid/event-standalone-grounded-provenance.json index 769c5ba..0bfff6b 100644 --- a/tests/valid/event-standalone-grounded-provenance.json +++ b/tests/valid/event-standalone-grounded-provenance.json @@ -1,4 +1,5 @@ { + "_test_description": "Standalone grounding of a third-party-sourced representation that the agent had cached (cached true is unconstrained for third_party_sourced, section 6.4), identified by an ISCC content_id alongside the URL, with an iscc fingerprint detection.", "document_type": "event", "schema_version": "1.0", "session_id": "550e8400-e29b-41d4-a716-446655440000", diff --git a/tests/valid/event-standalone-index.json b/tests/valid/event-standalone-index.json new file mode 100644 index 0000000..6441b16 --- /dev/null +++ b/tests/valid/event-standalone-index.json @@ -0,0 +1,18 @@ +{ + "_test_description": "Standalone retrieval reported by a marketplace index (source_role index) that served the content to the agent: identified by the marketplace catalogue id with no canonical URL, resolved by content_id prefix registration (sections 4.4, 7.3).", + "document_type": "event", + "schema_version": "1.0", + "manifest_ref": "https://marketplace.example.com/.well-known/content-telemetry.json", + "event": { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "index", + "content_telemetry_id": "990e8400-e29b-41d4-a716-446655440602", + "content_id": "mkt:gridnews:5520", + "license_ref": "agreement-2026-017:gridnews", + "data": { + "media_type": "text", + "content_depth": "full" + } + } +} diff --git a/tests/valid/event-standalone-origin.json b/tests/valid/event-standalone-origin.json new file mode 100644 index 0000000..6c3f0f6 --- /dev/null +++ b/tests/valid/event-standalone-origin.json @@ -0,0 +1,19 @@ +{ + "_test_description": "Standalone retrieval reported by the content owner's own web server (source_role origin) at Retrieval conformance: no session context, origin enrichment fields (6.3) and content_depth recording that the agent reached the abstract only.", + "document_type": "event", + "schema_version": "1.0", + "event": { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "origin", + "content_telemetry_id": "990e8400-e29b-41d4-a716-446655440601", + "content_url": "https://journals.example.com/article/10.1000/xyz123", + "content_id": "doi:10.1000/xyz123", + "data": { + "user_agent": "ClaudeBot/1.0", + "response_status": 200, + "media_type": "text", + "content_depth": "abstract" + } + } +} diff --git a/tests/valid/manifest-platform.json b/tests/valid/manifest-platform.json new file mode 100644 index 0000000..e658529 --- /dev/null +++ b/tests/valid/manifest-platform.json @@ -0,0 +1,15 @@ +{ + "_test_description": "Platform manifest: an intermediary operating a telemetry consumer declares its inbound endpoint and a click-token resolution endpoint (sections 7.4.3, 8.5).", + "schema_version": "1.0", + "id": "https://telemetry.example.com/.well-known/content-telemetry.json", + "roles": [ + "platform" + ], + "operator": { + "name": "Telemetry Example" + }, + "telemetry": { + "endpoint": "https://telemetry.example.com/v1/events", + "ctx_resolution": "https://telemetry.example.com/v1/ctx/resolve" + } +} diff --git a/tests/valid/session-citation-contradiction.json b/tests/valid/session-citation-contradiction.json new file mode 100644 index 0000000..ede6f51 --- /dev/null +++ b/tests/valid/session-citation-contradiction.json @@ -0,0 +1,79 @@ +{ + "_test_description": "Negative attribution and unclassified values: the response explicitly disagrees with a retrieved source (citation_type contradiction, section 6.5), a second association the agent could not classify carries citation_type and position unclassified, the grounding carries a content_version (ETag), and the turn ran in deep_research mode.", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440651", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "turn_started", + "timestamp": "2026-08-20T10:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "query_intent": "fact_check", + "topics": [ + "battery recycling" + ] + } + }, + { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "agent", + "content_url": "https://www.example-news.com/energy/battery-recycling-claims", + "data": { + "media_type": "text" + } + }, + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/energy/battery-recycling-claims", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 5100, + "content_version": "W/\"a1b2c3\"", + "content_last_modified": "2026-08-19T07:30:00Z" + } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/energy/battery-recycling-claims", + "data": { + "citation_type": "contradiction", + "position": "primary" + }, + "output_element_id": "answer:claim:1" + }, + { + "id": "770e8400-e29b-41d4-a716-446655440604", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/energy/battery-recycling-claims", + "data": { + "citation_type": "unclassified", + "position": "unclassified" + }, + "output_element_id": "answer:aside:1" + }, + { + "type": "turn_completed", + "timestamp": "2026-08-20T10:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "response_type": "fact_check", + "response_mode": "deep_research", + "response_tokens": 900 + } + } + ] +} diff --git a/tests/valid/session-engagement-types.json b/tests/valid/session-engagement-types.json new file mode 100644 index 0000000..560c682 --- /dev/null +++ b/tests/valid/session-engagement-types.json @@ -0,0 +1,113 @@ +{ + "_test_description": "The non-click engagement types: a snippet presentation is expanded, and the detail view it opens is copied from and shared. Each action is one engagement occurrence on one presentation (sections 4.3, 6.8); three actions on two presentations, all in-product and agent-reported.", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440650", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "turn_started", + "timestamp": "2026-08-20T10:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "query_intent": "how_to", + "topics": [ + "sourdough" + ] + } + }, + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.kingarthurbaking.com/recipes/sourdough-starter-maintenance", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 7200 + } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.kingarthurbaking.com/recipes/sourdough-starter-maintenance", + "data": { + "citation_type": "paraphrase", + "position": "primary" + }, + "output_element_id": "answer:steps:1" + }, + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "citation_id": "770e8400-e29b-41d4-a716-446655440601", + "content_url": "https://www.kingarthurbaking.com/recipes/sourdough-starter-maintenance", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "snippet" + }, + "output_element_id": "answer:steps:1" + }, + { + "id": "770e8400-e29b-41d4-a716-446655440603", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.kingarthurbaking.com/recipes/sourdough-starter-maintenance", + "data": { + "presentation_kind": "content", + "presentation_type": "detail_view", + "media_type": "text" + }, + "output_element_id": "panel:source:1" + }, + { + "type": "turn_completed", + "timestamp": "2026-08-20T10:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "response_type": "how_to", + "response_mode": "standard", + "response_tokens": 260 + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:06Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440602", + "content_url": "https://www.kingarthurbaking.com/recipes/sourdough-starter-maintenance", + "data": { + "engagement_type": "expand" + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:07Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440603", + "content_url": "https://www.kingarthurbaking.com/recipes/sourdough-starter-maintenance", + "data": { + "engagement_type": "copy" + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:08Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440603", + "content_url": "https://www.kingarthurbaking.com/recipes/sourdough-starter-maintenance", + "data": { + "engagement_type": "share" + } + } + ] +} diff --git a/tests/valid/session-minimal.json b/tests/valid/session-minimal.json index 39f46f8..a54498b 100644 --- a/tests/valid/session-minimal.json +++ b/tests/valid/session-minimal.json @@ -1,5 +1,5 @@ { - "_test_description": "Bare minimum conforming session: schema_version, session_id, started_at, and one event with type and timestamp.", + "_test_description": "Bare minimum conforming session: schema_version, session_id, started_at, and one content_retrieved event with type, timestamp, source_role and a content_url.", "schema_version": "1.0", "session_id": "550e8400-e29b-41d4-a716-446655440000", "started_at": "2026-03-28T10:00:00Z", @@ -7,6 +7,7 @@ { "type": "content_retrieved", "timestamp": "2026-03-28T10:00:01Z", + "source_role": "agent", "content_url": "https://example.com/article/getting-started" } ] diff --git a/tests/valid/session-multi-owner-catalogue.json b/tests/valid/session-multi-owner-catalogue.json new file mode 100644 index 0000000..bd15288 --- /dev/null +++ b/tests/valid/session-multi-owner-catalogue.json @@ -0,0 +1,142 @@ +{ + "_test_description": "Multi-owner catalogue under one agreement (Annex B.5): content_scope identifies the marketplace agreement, owner resolution is per event - one owner by content_url domain, the other by a registered content_id prefix with no canonical URL at all. Two owners, each event resolving to its own.", + "schema_version": "1.0", + "session_id": "990e8400-e29b-41d4-a716-446655440500", + "agent_id": "research-assistant-v5", + "content_scope": "marketplace-agreement-2026-017", + "manifest_ref": "https://assistant.example.com/.well-known/content-telemetry.json", + "started_at": "2026-08-20T09:00:00Z", + "ended_at": "2026-08-20T09:00:09Z", + "events": [ + { + "type": "turn_started", + "timestamp": "2026-08-20T09:00:00Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "query_intent": "comparison", + "topics": [ + "electric vehicles", + "charging" + ] + } + }, + { + "type": "content_retrieved", + "timestamp": "2026-08-20T09:00:01Z", + "source_role": "index", + "content_telemetry_id": "990e8400-e29b-41d4-a716-446655440501", + "content_url": "https://www.autoreview.example/ev/charging-networks-2026", + "content_id": "mkt:autoreview:88213", + "license_ref": "agreement-2026-017:autoreview", + "data": { + "media_type": "text", + "content_depth": "full" + } + }, + { + "type": "content_retrieved", + "timestamp": "2026-08-20T09:00:01Z", + "source_role": "index", + "content_telemetry_id": "990e8400-e29b-41d4-a716-446655440502", + "content_id": "mkt:gridnews:5520", + "license_ref": "agreement-2026-017:gridnews", + "data": { + "media_type": "text", + "content_depth": "full" + } + }, + { + "type": "content_grounded", + "timestamp": "2026-08-20T09:00:02Z", + "turn_id": "1", + "content_url": "https://www.autoreview.example/ev/charging-networks-2026", + "content_id": "mkt:autoreview:88213", + "data": { + "scope": "turn", + "cached": false, + "provenance": "third_party_sourced", + "chars_ingested": 11200 + } + }, + { + "type": "content_grounded", + "timestamp": "2026-08-20T09:00:02Z", + "turn_id": "1", + "content_id": "mkt:gridnews:5520", + "data": { + "scope": "turn", + "cached": false, + "provenance": "third_party_sourced", + "chars_ingested": 6400 + } + }, + { + "id": "990e8400-e29b-41d4-a716-446655440503", + "type": "content_cited", + "timestamp": "2026-08-20T09:00:06Z", + "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:networks:1", + "content_url": "https://www.autoreview.example/ev/charging-networks-2026", + "content_id": "mkt:autoreview:88213", + "data": { + "citation_type": "paraphrase", + "position": "primary" + } + }, + { + "id": "990e8400-e29b-41d4-a716-446655440504", + "type": "content_cited", + "timestamp": "2026-08-20T09:00:06Z", + "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:tariffs:1", + "content_id": "mkt:gridnews:5520", + "data": { + "citation_type": "reference", + "position": "supporting" + } + }, + { + "id": "990e8400-e29b-41d4-a716-446655440505", + "type": "content_presented", + "timestamp": "2026-08-20T09:00:07Z", + "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:networks:1", + "citation_id": "990e8400-e29b-41d4-a716-446655440503", + "content_url": "https://www.autoreview.example/ev/charging-networks-2026", + "content_id": "mkt:autoreview:88213", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "id": "990e8400-e29b-41d4-a716-446655440506", + "type": "content_presented", + "timestamp": "2026-08-20T09:00:07Z", + "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:tariffs:1", + "citation_id": "990e8400-e29b-41d4-a716-446655440504", + "content_id": "mkt:gridnews:5520", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "card" + } + }, + { + "type": "turn_completed", + "timestamp": "2026-08-20T09:00:09Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "response_type": "comparison", + "response_mode": "standard", + "response_tokens": 410 + } + } + ] +} diff --git a/tests/validate.py b/tests/validate.py index c9a65cb..8dcac1a 100644 --- a/tests/validate.py +++ b/tests/validate.py @@ -82,6 +82,22 @@ # across events. Session documents only: standalone envelopes and batch # members may reference events delivered elsewhere. # +# 7. Field placement and source_role (sections 5.2, 5.2.2, 5.7.1, 5.7.5): +# source_role is required on content_retrieved; presentation_id and the +# event-level ctx_token appear only on content_engaged, citation_id only on +# content_presented/content_reproduced, turn only on turn events. +# +# 8. Session-document integrity beyond rule 6 (sections 6.7, 6.8, 7.4.1): +# event ids are distinct; the engagement and the presentation it references, +# and the presentation/reproduction and the citation it references, identify +# the same content; one event-level ctx_token binds to one presentation. +# +# 9. Envelope ctx_token (section 7.1): an envelope carrying ctx_token carries +# content_engaged events only. +# +# 10. Manifest placement rules (sections 8.5, 8.6): domains only on a manifest +# served at the domain root; ctx_resolution only on agent/platform manifests. +# # Not checked here: agent_id at Grounding/Citation conformance (section # 5.7) depends on the emitter's declared conformance level, which fixtures do # not carry, so it is out of scope for the fixture suite. @@ -182,6 +198,106 @@ "Grounding event declares agent_fetched with cached true. " "Violates section 6.4: agent_fetched requires cached false." ), + "grounding-provenance-cached-conflict-agent-cached.json": ( + "Grounding event declares agent_cached with cached false. " + "Violates section 6.4: agent_cached requires cached true." + ), + "privacy-violation-response-text-at-intent.json": ( + "Turn at intent privacy includes response_text. " + "Violates section 5.5: response_text MUST NOT be present at intent level." + ), + "privacy-violation-query-at-minimal-standalone.json": ( + "Standalone envelope turn at minimal privacy includes query_text. " + "Violates section 5.5: the gate applies wherever turns are emitted." + ), + "privacy-violation-query-at-minimal-batch.json": ( + "Batch member turn at minimal privacy includes query_text. " + "Violates section 5.5: the gate applies wherever turns are emitted." + ), + "grounded-missing-identifier-standalone.json": ( + "Standalone content_grounded event has neither content_url nor content_id. " + "Violates section 5.7.5 in the standalone envelope shape." + ), + "grounded-missing-identifier-batch.json": ( + "Batch member content_grounded event has neither content_url nor content_id. " + "Violates section 5.7.5 in the event batch shape." + ), + "withdrawn-ip-hash-session.json": ( + "Session document retrieval event carries ip_hash in data. " + "Violates section 9.1 in the session document shape." + ), + "withdrawn-ip-hash-batch.json": ( + "Batch member retrieval event carries ip_hash in data. " + "Violates section 9.1 in the event batch shape." + ), + "retrieved-missing-source-role.json": ( + "content_retrieved event carries no source_role. " + "Violates sections 5.2.2 and 5.7.1: source_role MUST be set on every retrieval." + ), + "ctx-token-on-grounded.json": ( + "Event-level ctx_token on a content_grounded event. " + "Violates section 5.2: ctx_token is valid only on content_engaged." + ), + "citation-id-on-grounded.json": ( + "citation_id on a content_grounded event. " + "Violates section 5.2: citation_id is valid only on content_presented and content_reproduced." + ), + "presentation-id-on-cited.json": ( + "presentation_id on a content_cited event. " + "Violates section 5.2: presentation_id is valid only on content_engaged." + ), + "turn-on-content-event.json": ( + "turn object on a content_grounded event. " + "Violates section 5.2: turn data is carried on turn_started and turn_completed only." + ), + "duplicate-event-id.json": ( + "Two content_presented events share one id. " + "Violates section 6.7: repeated presentations receive distinct event IDs." + ), + "engaged-presentation-content-mismatch.json": ( + "content_engaged references a content_presented event of different content. " + "Violates section 6.8: the engagement identifies the same content as the presentation it acted on." + ), + "presented-citation-content-mismatch.json": ( + "content_presented.citation_id references a content_cited event of different content. " + "Violates section 6.7: the presentation and the citation it realises identify the same content." + ), + "reproduced-citation-id-unmatched.json": ( + "content_reproduced.citation_id matches no content_cited event id in the session. " + "Violates section 6.6: citation_id references the crediting content_cited event's id." + ), + "engaged-presentation-id-no-presentations.json": ( + "content_engaged carries a presentation_id in a session with no content_presented events. " + "Violates section 6.8: every engagement references an exact presentation occurrence." + ), + "shared-ctx-token-two-presentations.json": ( + "One event-level ctx_token appears on engagements bound to two different presentations. " + "Violates section 7.4.1: a token is bound to exactly one presentation occurrence." + ), + "standalone-ctx-token-non-engagement.json": ( + "Standalone envelope carries ctx_token with a content_grounded event. " + "Violates section 7.1: an envelope ctx_token accompanies content_engaged events only." + ), + "batch-ctx-token-non-engagement.json": ( + "Event batch under ctx_token carries a content_grounded event. " + "Violates section 7.1: an envelope ctx_token accompanies content_engaged events only." + ), + "batch-missing-session-mixed-retrieval.json": ( + "Event batch mixing a retrieval with a grounding event carries neither session_id nor ctx_token. " + "Violates section 7.1: the retrieval-only exemption does not extend to a batch that carries other events." + ), + "manifest-domains-on-path-manifest.json": ( + "Manifest served under a path prefix carries domains. " + "Violates section 8.6: domains MAY appear only on manifests served from the domain root." + ), + "manifest-ctx-resolution-on-content-owner.json": ( + "content_owner manifest declares telemetry.ctx_resolution. " + "Violates section 8.5: ctx_resolution is valid on agent and platform manifests." + ), + "manifest-lookalike-domain.json": ( + "Manifest at example.com claims evilexample.com in domains. " + "Violates section 8.6: a lookalike host is not a subdomain of the manifest host." + ), } # V0.1 fields prohibited by the v1 migration rule (section 9.1). This is a @@ -497,6 +613,104 @@ def check_grounding_provenance(data): return violations +# Fields that belong to one event type (section 5.2). A key present with a +# non-null value on any other type is a placement violation. +FIELD_PLACEMENT = { + "presentation_id": {"content_engaged"}, + "ctx_token": {"content_engaged"}, + "citation_id": {"content_presented", "content_reproduced"}, + "turn": {"turn_started", "turn_completed"}, +} + + +def check_field_placement(data): + """Check source_role on retrievals (sections 5.2.2, 5.7.1) and the + event-type scoping of presentation_id, ctx_token, citation_id and turn + (section 5.2). Applies to every document shape.""" + violations = [] + for event in _iter_events(data): + etype = event.get("type") + if etype == "content_retrieved" and not event.get("source_role"): + violations.append("content_retrieved event carries no source_role") + for field, allowed in FIELD_PLACEMENT.items(): + if event.get(field) is not None and etype not in allowed: + violations.append( + f"Field '{field}' present on '{etype}' event; valid only on " + + ", ".join(sorted(allowed)) + ) + return violations + + +def _same_content(a, b): + """Two events identify the same content unless a shared identifier field + (content_id or content_url) carries different non-null values.""" + for field in ("content_id", "content_url"): + x, y = a.get(field), b.get(field) + if x is not None and y is not None and x != y: + return False + return True + + +def check_session_integrity(data): + """Session-document rules beyond check_referential_integrity (sections + 6.7, 6.8, 7.4.1): distinct event ids; engagement/presentation and + presentation|reproduction/citation pairs identify the same content; one + event-level ctx_token binds to one presentation. Session documents only.""" + if is_standalone_event(data) or is_event_batch(data): + return [] + events = data.get("events", []) + violations = [] + seen_ids = set() + for e in events: + eid = e.get("id") + if eid: + if eid in seen_ids: + violations.append(f"Duplicate event id '{eid}' in session") + seen_ids.add(eid) + presented = {e.get("id"): e for e in events if e.get("type") == "content_presented" and e.get("id")} + cited = {e.get("id"): e for e in events if e.get("type") == "content_cited" and e.get("id")} + token_binding = {} + for e in events: + etype = e.get("type") + if etype == "content_engaged": + pid = e.get("presentation_id") + if pid in presented and not _same_content(e, presented[pid]): + violations.append( + f"content_engaged references presentation '{pid}' but identifies different content" + ) + token = e.get("ctx_token") + if token and pid: + bound = token_binding.setdefault(token, pid) + if bound != pid: + violations.append( + f"ctx_token '{token}' appears on engagements bound to two presentations" + ) + if etype in ("content_presented", "content_reproduced"): + cid = e.get("citation_id") + if cid in cited and not _same_content(e, cited[cid]): + violations.append( + f"{etype} references citation '{cid}' but identifies different content" + ) + return violations + + +def check_envelope_ctx_token(data): + """An envelope ctx_token accompanies content_engaged events only (section + 7.1); any other event type on such an envelope needs session_id instead.""" + if not (is_standalone_event(data) or is_event_batch(data)): + return [] + if not data.get("ctx_token"): + return [] + violations = [] + for event in _iter_events(data): + etype = event.get("type") + if etype != "content_engaged": + violations.append( + f"Envelope ctx_token accompanies a '{etype}' event; ctx_token is carried only for content_engaged" + ) + return violations + + def check_application_layer(data): """Run every application-layer conformance rule and return all violations.""" return ( @@ -506,15 +720,19 @@ def check_application_layer(data): + check_referential_integrity(data) + check_v1_migration_prohibitions(data) + check_grounding_provenance(data) + + check_field_placement(data) + + check_session_integrity(data) + + check_envelope_ctx_token(data) ) def check_manifest_application_layer(data): """ - Check the manifest rejection rules of section 8.7 that JSON Schema cannot - express: duplicate keys[].id values, and domains entries that are not the - manifest's own host or a subdomain of it (section 8.6). - Returns a list of violation descriptions. + Check the manifest rejection rules that JSON Schema cannot express: + duplicate keys[].id values (section 8.7), domains entries that are not the + manifest's own host or a subdomain of it, domains on a path-prefixed + manifest (section 8.6), and ctx_resolution on a manifest without the agent + or platform role (section 8.5). Returns a list of violation descriptions. """ violations = [] @@ -525,7 +743,8 @@ def check_manifest_application_layer(data): violations.append(f"Duplicate keys[].id '{kid}'") seen.add(kid) - host = urlparse(data.get("id", "")).hostname + parsed = urlparse(data.get("id", "")) + host = parsed.hostname if host: for entry in data.get("domains", []): bare = entry[2:] if entry.startswith("*.") else entry @@ -535,6 +754,22 @@ def check_manifest_application_layer(data): f"'{host}' or a subdomain of it" ) + # domains only on a root manifest (section 8.6) + if "domains" in data and parsed.path != "/.well-known/content-telemetry.json": + violations.append( + f"domains present on a manifest served under a path prefix " + f"('{parsed.path}'); only root manifests carry domains" + ) + + # ctx_resolution only on agent/platform manifests (section 8.5) + telemetry = data.get("telemetry") or {} + roles = set(data.get("roles") or []) + if telemetry.get("ctx_resolution") and not roles & {"agent", "platform"}: + violations.append( + f"telemetry.ctx_resolution on a manifest with roles {sorted(roles)}; " + "valid on agent and platform manifests" + ) + return violations From 4f56b6e590a230ea718aafb1feafcca0a4812bc9 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Wed, 2 Sep 2026 14:24:55 +0100 Subject: [PATCH 34/38] Remove content_reproduced from v1 core The v1 lifecycle keeps the content_displayed -> content_presented rename and drops the separate reproduction event. Output-side reuse reporting (reproduction_type, reproduced_chars/tokens/hash, the reproduced funnel departures and per-reproduction counting) leaves core; a future version or profile can reintroduce it on implementation evidence. The migration note records that the event existed only on the pre-release draft line. Co-Authored-By: Claude Fable 5 --- README.md | 8 +- SCOPE.md | 2 +- SPECIFICATION.md | 132 ++++++------------ telemetry-session.json | 40 +----- tests/README.md | 13 +- ...batch-engaged-missing-presentation-id.json | 2 +- tests/invalid/citation-id-on-grounded.json | 2 +- tests/invalid/cited-missing-id.json | 2 +- ...engaged-presentation-content-mismatch.json | 2 +- ...aged-presentation-id-no-presentations.json | 2 +- .../engaged-presentation-id-unmatched.json | 2 +- ...dalone-missing-presentation-and-token.json | 2 +- ...nding-fingerprint-preserved-in-output.json | 4 +- tests/invalid/invalid-event-type.json | 2 +- tests/invalid/legacy-content-displayed.json | 2 +- .../reproduced-citation-id-unmatched.json | 23 --- tests/invalid/reproduced-missing-data.json | 18 --- tests/invalid/reproduced-missing-id.json | 19 --- .../invalid/reproduced-missing-output-id.json | 19 --- .../reproduced-missing-source-reference.json | 19 --- tests/invalid/reproduced-missing-type.json | 19 --- tests/invalid/reproduced-negative-chars.json | 21 --- tests/invalid/reproduction-type-invalid.json | 21 --- tests/mutation_smoke.py | 5 - tests/valid/session-engagement-types.json | 2 +- .../session-repeated-link-engagement.json | 2 +- .../session-reproduction-credited-quote.json | 115 --------------- .../session-reproduction-no-grounding.json | 50 ------- .../session-reproduction-uncredited.json | 71 ---------- tests/validate.py | 52 ++++--- 30 files changed, 91 insertions(+), 582 deletions(-) delete mode 100644 tests/invalid/reproduced-citation-id-unmatched.json delete mode 100644 tests/invalid/reproduced-missing-data.json delete mode 100644 tests/invalid/reproduced-missing-id.json delete mode 100644 tests/invalid/reproduced-missing-output-id.json delete mode 100644 tests/invalid/reproduced-missing-source-reference.json delete mode 100644 tests/invalid/reproduced-missing-type.json delete mode 100644 tests/invalid/reproduced-negative-chars.json delete mode 100644 tests/invalid/reproduction-type-invalid.json delete mode 100644 tests/valid/session-reproduction-credited-quote.json delete mode 100644 tests/valid/session-reproduction-no-grounding.json delete mode 100644 tests/valid/session-reproduction-uncredited.json diff --git a/README.md b/README.md index cccb7d6..7b42705 100644 --- a/README.md +++ b/README.md @@ -33,12 +33,11 @@ Platforms self-report usage metrics (if they report at all), and content owners ## Telemetry events -Content Telemetry tracks content through six stages: +Content Telemetry tracks content through five stages: ``` Retrieved → content fetched over HTTP (content owner can see this today) Grounded → content loaded into the agent's generation context - Reproduced → content appearing verbatim or near-verbatim in the response Cited → content explicitly referenced in the response Presented → content or a source reference made perceivable on a recipient-facing surface Engaged → user clicked, copied, shared, or directed the agent to act @@ -50,7 +49,6 @@ The gaps between stages show how content was used: - **Retrieval without grounding** - your content was fetched but not used - **Grounding without citation** - your content influenced the answer but you got no credit -- **Reproduction without citation** - your content appeared in the answer without credit - **Citation without engagement** - your content was cited but the user didn't click through The grounding event captures the boundary "this content entered the agent's generation context." It is architecture-neutral and decoupled from retrieval: content cached by the agent for days still produces a grounding event in every session it influences. @@ -61,7 +59,7 @@ Grounding and presentation record different boundary crossings: grounding means **Post-hoc, not pre-declared.** Events report what actually happened, not what the agent said it would do at request time. An agent cannot reliably declare how it will use content before reading it. -**Observable boundaries, not agent internals.** The six content event types mark boundary crossings. What happens between them - the fan-out, relevance evaluation, re-ranking, reasoning chains - is internal to the agent and changes constantly. The spec does not model it. +**Observable boundaries, not agent internals.** The five content event types mark boundary crossings. What happens between them - the fan-out, relevance evaluation, re-ranking, reasoning chains - is internal to the agent and changes constantly. The spec does not model it. **Multiple observers, one event.** A content retrieval can be reported by the content owner's CDN, the content owner's origin server, and the AI agent independently. The `Content-Telemetry-ID` header correlates these into a single corroborated event. Uncorroborated retrievals (no matching agent event) may indicate an agent that does not yet support the telemetry protocol. @@ -208,7 +206,7 @@ This is a preview specification. The following areas are under active discussion **Event volume at scale.** A single deep-research query can produce 100+ retrieval events and dozens of grounding/citation events. The session document format already handles transport - one POST with all events after the session ends, not one request per event. Volume management beyond that (storage, processing, consumer-side aggregation) is an implementation concern, not a protocol gap. Version 1 adds an explicit coverage declaration - `complete`, `sampled`, `aggregated` or `selected` (section 5.7.6) - and a manifest field for it (section 8.5); the standard still sets no default for reporting granularity, leaving it to profiles and deployments. -**Verification of grounding and citation.** Grounding and citation events are reported by the agent, which is also the party that may owe compensation under a licence. In v0.1, manifest signing is informational: consumers may verify signatures but are not required to, and the specification defines no required proof binding an event to its emitter (sections 8.4 and 8.9). The events attribution depends on are therefore self-reported by the reporting party. Verifiable credentials and signed events are deferred (section 8.9). One corroboration mechanism works without signing: the `Content-Telemetry-ID` field correlates an agent-reported retrieval with an origin- or edge-reported one (section 7.2), but it covers retrieval only - grounding, reproduction, citation, presentation, and engagement have no independent observer. Signing, even once required, would prove who reported an event, not that the event is true or that all qualifying events were reported. Input is wanted on what a verification layer should cover and where it belongs. Mechanisms that test truthfulness and completeness rather than origin, such as sampled audits or publisher-seeded canary content, are of particular interest. +**Verification of grounding and citation.** Grounding and citation events are reported by the agent, which is also the party that may owe compensation under a licence. In v0.1, manifest signing is informational: consumers may verify signatures but are not required to, and the specification defines no required proof binding an event to its emitter (sections 8.4 and 8.9). The events attribution depends on are therefore self-reported by the reporting party. Verifiable credentials and signed events are deferred (section 8.9). One corroboration mechanism works without signing: the `Content-Telemetry-ID` field correlates an agent-reported retrieval with an origin- or edge-reported one (section 7.2), but it covers retrieval only - grounding, citation, presentation, and engagement have no independent observer. Signing, even once required, would prove who reported an event, not that the event is true or that all qualifying events were reported. Input is wanted on what a verification layer should cover and where it belongs. Mechanisms that test truthfulness and completeness rather than origin, such as sampled audits or publisher-seeded canary content, are of particular interest. **Reporting granularity.** The standard sets no default for reporting granularity, leaving it to profiles and deployments (see *Event volume* above). The SPUR profile requires event-level delivery and does not permit aggregation. Version 1 answers the first half of the question: coverage modes are defined once, in section 5.7.6, so that profiles reference them rather than each define their own. How event-level delivery scales for the highest-volume case remains open. diff --git a/SCOPE.md b/SCOPE.md index f1fe559..72e3888 100644 --- a/SCOPE.md +++ b/SCOPE.md @@ -18,7 +18,7 @@ Conformance and verification answer five separate questions: 4. **Factual truth and completeness:** did the event happen as claimed, and were all qualifying events reported? 5. **Entitlement:** was the reported use permitted under an applicable grant or agreement? -Events are claims by identified emitters. Evidence applies to a particular assertion. Origin or access evidence can corroborate only what that observer could see; it cannot prove grounding, reproduction, citation, presentation, engagement, truth, completeness or entitlement. +Events are claims by identified emitters. Evidence applies to a particular assertion. Origin or access evidence can corroborate only what that observer could see; it cannot prove grounding, citation, presentation, engagement, truth, completeness or entitlement. Relationship configuration should avoid profile proliferation. A publisher may require `content_grounded` and `content_cited` events, intent-level topics, event delivery to a named endpoint and a set of aggregate reports. Another may require `content_cited`, `content_presented` and `content_engaged` events with a different privacy level. These are deployment choices backed by governing terms, not publisher-specific protocol profiles. A new profile is justified only when a class of relationships introduces semantics or processing rules that multiple implementations must interpret in the same way. diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 7fbdc90..a761431 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -11,7 +11,7 @@ 3. [Terms and definitions](#3-terms-and-definitions) 4. [Concepts](#4-concepts) - roles, sessions, event lifecycle, source roles, content identification 5. [Schema](#5-schema) - session, event, event types, conversation turn, privacy, intent, conformance levels -6. [Data profiles](#6-data-profiles) - retrieval, edge enrichment, origin enrichment, grounding, citation, reproduction, presentation, engagement +6. [Data profiles](#6-data-profiles) - retrieval, edge enrichment, origin enrichment, grounding, citation, presentation, engagement 7. [Transport](#7-transport) - delivery formats, Content-Telemetry-ID header, routing, click context 8. [Manifest](#8-manifest) - discovery, schema, operator, keys, telemetry, domains 9. [Privacy](#9-privacy) - data minimisation, recommended levels, retention @@ -58,9 +58,9 @@ Content Telemetry does not: #### 1.3.1 Inference-time scope -The six-stage lifecycle reports content use observable at inference time: identified content entered a generation context for a particular response, and what the resulting output did with it. +The five-stage lifecycle reports content use observable at inference time: identified content entered a generation context for a particular response, and what the resulting output did with it. -Retrieval is the boundary case. A crawl whose purpose is training or index building can be reported as a `content_retrieved` event, and `bot_category` (section 6.2) distinguishes it, but the event is non-attributable: no grounding, reproduction, citation, presentation or engagement follows it. What the system then does with the content, whether it enters a training corpus, a fine-tuning set, an embedding store or a search index, is outside this specification. Nothing here reports that a model was trained on a work, and a conforming implementation says nothing either way about it. +Retrieval is the boundary case. A crawl whose purpose is training or index building can be reported as a `content_retrieved` event, and `bot_category` (section 6.2) distinguishes it, but the event is non-attributable: no grounding, citation, presentation or engagement follows it. What the system then does with the content, whether it enters a training corpus, a fine-tuning set, an embedding store or a search index, is outside this specification. Nothing here reports that a model was trained on a work, and a conforming implementation says nothing either way about it. Using such a store at inference time is inside scope. When an index built over a content owner's material is queried during a response and returns content that grounds the answer, that is a `content_grounded` event like any other, with `source_role: index` on the retrieval that served it (section 4.4). The line is between constructing a derived artefact and using one to answer a query, not whether an index was involved. @@ -105,7 +105,6 @@ For the purposes of this specification, the following terms apply. | **content owner** | entity that owns or licences content accessed by an AI agent | | **agent operator** | entity running the AI agent that uses content | | **grounding** | content entering the generation model's context, the boundary where content can directly influence output (section 4.3) | -| **reproduction** | verbatim or near-verbatim appearance of identified source content in an output artifact, independent of credit and delivery (section 4.3) | | **presentation** | content or a source reference made perceivable on a recipient-facing surface (section 4.3) | | **source role** | classification of the observer reporting a retrieval event: `origin`, `edge`, `index`, or `agent` (section 4.4) | | **privacy level** | data sharing tier controlling which conversation fields are populated: `full`, `summary`, `intent`, or `minimal` (section 5.5) | @@ -181,7 +180,6 @@ Session │ ├── turn_started │ ├── content_retrieved (HTTP layer) │ ├── content_grounded (influence layer) -│ ├── content_reproduced (response layer) │ ├── content_cited (response layer) │ ├── content_presented (recipient-facing surface) │ ├── turn_completed @@ -194,7 +192,7 @@ These are the event types a session can contain, not a strict ordering: events a ### 4.3 Event lifecycle -Content moves through six stages during an agent interaction: +Content moves through five stages during an agent interaction: 1. **Retrieved** - Content fetched over HTTP from an origin server, CDN, marketplace, or index. This is an infrastructure event observable by the content owner's infrastructure (origin server, edge network) and the agent. A retrieval may be cached by the agent for use across multiple sessions. @@ -208,42 +206,34 @@ Content moves through six stages during an agent interaction: One grounding occurrence is one distinct content item entering a generation context at the declared `data.scope`: at `session` scope, a content item grounds once per session; at `turn` scope, once per turn it enters. A distinct content item is a distinct `content_id`, or its canonical `content_url` where no stable identifier exists (section 4.5). Continued presence within the declared scope is not a further occurrence; re-entry in a later turn is, when the scope is `turn`, and a change of `content_version` is a new occurrence at either scope. An emitter that ingests a content item in chunks MAY emit one grounding event per chunk, preserving the chunk-level hashes of section 6.4; events sharing content identity within one scope describe one occurrence, and consumers count occurrences by deduplicating on content identity and scope, not by counting events. -3. **Reproduced** - The output artifact contains identified source content: a quotation, an excerpt, or a full copy, verbatim or near-verbatim. Reproduction is an output-construction claim by the system that built the output. It is independent of credit and of delivery: reproduced content may or may not also be cited, and the artifact may or may not later be presented. An uncredited excerpt in a response delivered through an API produces a `content_reproduced` event and nothing else - without this event, that use would be unreportable. +3. **Cited** - An output artifact explicitly associates identified source content with a response, claim, passage, quotation, or other output element. Citation is an output-construction relationship, not evidence that the output was delivered. A subset of grounded content is commonly cited, but a citation can also be emitted without a matching grounding event when an agent produces an uncorroborated or hallucinated source association. - One reproduction occurrence is one identified source content item appearing in one output element (or in the output artifact, where no element identity exists), so a credited quotation and its companion citation share an `output_element_id` (section 6.6). Three quoted passages from one source in three elements are three occurrences; repeated appearance within one element is not a further occurrence. - -4. **Cited** - An output artifact explicitly associates identified source content with a response, claim, passage, quotation, or other output element. Citation is an output-construction relationship, not evidence that the output was delivered. A subset of grounded content is commonly cited, but a citation can also be emitted without a matching grounding event when an agent produces an uncorroborated or hallucinated source association. - - A citation MUST carry a resolvable reference to the source it associates: a `content_url` or a `content_id`. A source association with no resolvable reference is not a citation and MUST NOT be emitted as `content_cited`. Unlike other content events, where the identifier requirement is an application-layer rule (section 5.7.5), for `content_cited` and `content_reproduced` it is enforced by the JSON Schema. - - Reproduction and citation are sibling claims about the same artifact: reproduction records the material, citation records the credit. Neither implies the other. A credited quotation produces both events; an uncredited excerpt produces only a reproduction; a reference citation with no quoted material produces only a citation. + A citation MUST carry a resolvable reference to the source it associates: a `content_url` or a `content_id`. A source association with no resolvable reference is not a citation and MUST NOT be emitted as `content_cited`. Unlike other content events, where the identifier requirement is an application-layer rule (section 5.7.5), for `content_cited` it is enforced by the JSON Schema. One citation occurrence is one distinct association between a source and an output element (or the output artifact, where no element identity exists). Associating the same source with three separate output elements produces three citation events; repeating the same association is not a further occurrence. -5. **Presented** - Content or a source reference was rendered, played, spoken, embedded, or otherwise made perceivable on a recipient-facing surface. Presentation does not assert that a person noticed or attended to it. `presentation_kind` distinguishes source content (including a reproduced excerpt or media) from a source reference (such as a link, credit, or card). Not all citations are presented: an output can be stored, suppressed, or passed to another system before delivery. +4. **Presented** - Content or a source reference was rendered, played, spoken, embedded, or otherwise made perceivable on a recipient-facing surface. Presentation does not assert that a person noticed or attended to it. `presentation_kind` distinguishes source content (an excerpt, an embedded page, played media) from a source reference (such as a link, credit, or card). Not all citations are presented: an output can be stored, suppressed, or passed to another system before delivery. Grounding and presentation record different boundary crossings: grounding records entry into a generation context, while presentation records a recipient-facing delivery occurrence. As agent experiences evolve beyond the chat window the two diverge - content can shape an answer whose source is never presented, and an agent can present content that never entered a generation context (see *Departures from the funnel model* below). One presentation occurrence is one rendering of content or a source reference on a recipient-facing surface; the event's `id` names that occurrence. Presenting the same artifact again - on a new surface, or in a new delivery - is a new occurrence. -6. **Engaged** - The recipient or agent performed an observable action on a presentation: clicked a link, expanded a preview, copied text, shared the response, or directed the agent to act on the content. It does not imply attention beyond the reported action. `presentation_id` links the action to the exact presentation occurrence; a click-out can also carry a `ctx_token` that a destination resolves to the click context (section 7.4). +5. **Engaged** - The recipient or agent performed an observable action on a presentation: clicked a link, expanded a preview, copied text, shared the response, or directed the agent to act on the content. It does not imply attention beyond the reported action. `presentation_id` links the action to the exact presentation occurrence; a click-out can also carry a `ctx_token` that a destination resolves to the click context (section 7.4). One engagement occurrence is one observed action on one presentation occurrence. ``` Retrieved (HTTP layer, cacheable) → Grounded (influence layer, per-session or per-turn) - → Reproduced (response layer, per-turn) ⎫ sibling output-construction - → Cited (response layer, per-turn) ⎭ claims; neither implies the other + → Cited (response layer, per-turn) → Presented (recipient-facing surface, per-turn) → Engaged (user action layer) ``` -Each stage after retrieval is typically a progressively narrower subset, except that reproduction and citation are siblings at the response layer rather than steps in the chain. The ratios between stages are meaningful for potential attribution: +Each stage after retrieval is typically a progressively narrower subset. The ratios between stages are meaningful for potential attribution: - **Retrieval-to-grounding** measures content fetched but not used (irrelevant, stale, or a competing source was preferred) - **Grounding-to-citation** measures content that influenced the response without explicit attribution -- **Reproduction-to-citation** measures reproduced material without an explicit source association - the uncredited-reproduction rate - **Citation-to-presentation** measures source associations constructed in output but not made perceivable - **Presentation-to-engagement** measures observable actions on exact presentation occurrences @@ -251,13 +241,11 @@ These ratios are computed over reported events and are comparable across emitter #### Departures from the funnel model -Five cases break the strict subset model: +Three cases break the strict subset model: - **Presented without cited.** An agent may present content references (e.g., a "Sources" sidebar) without semantically associating them with a response element. In this case, a `content_presented` event exists with no corresponding `content_cited` event. - **Cited without grounded.** A hallucinated citation references content the agent never retrieved or loaded into context. The `content_cited` event has no preceding `content_grounded` event. Telemetry consumers SHOULD treat uncorroborated citations (no matching grounding event) as lower-confidence signals. - **Presented without grounded.** An agent can present content without that content entering a generation context: an agentic browser showing a page or an embedded video played on a response surface. A `content_presented` event (typically `presentation_kind: content` and `presentation_type: embed`) exists with no corresponding `content_grounded` event. -- **Reproduced without cited.** An output contains source material with no explicit source association: an uncredited excerpt. The `content_reproduced` event exists with no corresponding `content_cited` event. This is the primary case the reproduction event exists to record. -- **Reproduced without grounded.** An output can reproduce content that never entered this session's generation context, most commonly content the model memorised during training. When the emitter can identify the source, the `content_reproduced` event stands without a grounding event; telemetry consumers SHOULD treat it, like an uncorroborated citation, as a lower-confidence signal. These cases are valid. Emitters SHOULD produce the events that reflect what actually happened, even when the result does not follow the typical funnel ordering. @@ -268,7 +256,7 @@ Conversation turns overlay this lifecycle: 1. **Turn started** - user submits a query 2. **Turn completed** - agent finishes response -A single grounding event with session scope influences all subsequent turns. Reproduction, citation, presentation, and engagement events occur within specific turns. +A single grounding event with session scope influences all subsequent turns. Citation, presentation, and engagement events occur within specific turns. ### 4.4 Source roles @@ -285,7 +273,7 @@ The `origin` and `edge` source roles enable content owners to report AI agent tr A marketplace operating as both emitter and telemetry consumer receives telemetry from platforms (as a consumer), resolves content owner identity from `content_id` or `content_url`, and generates per-content-owner usage reports. The marketplace's own `source_role: index` events provide a corroboration layer - it can cross-reference what it served against what platforms reported using. -`content_grounded`, `content_reproduced`, `content_cited`, and `content_presented` events are reported by the agent (or agent operator) only. These events describe what happened inside the agent, during output construction, or on a recipient-facing surface, which is not observable from the content owner's infrastructure. A third party that detects reproduced content in a delivered output is corroborating or contradicting the emitter's claims, not observing construction; detection results belong to verification tooling, not to these event types. +`content_grounded`, `content_cited`, and `content_presented` events are reported by the agent (or agent operator) only. These events describe what happened inside the agent, during output construction, or on a recipient-facing surface, which is not observable from the content owner's infrastructure. A third party that detects source content in a delivered output is corroborating or contradicting the emitter's claims, not observing construction; detection results belong to verification tooling, not to these event types. `content_engaged` events are usually reported by the agent for in-product interactions. For a click-out to a landing page, a downstream marketplace, affiliate network, or destination site MAY report a corroborating `content_engaged` event using `ctx_token` in place of `session_id` (section 7.1). @@ -409,13 +397,13 @@ Emitters MUST NOT populate `access_context` unless the governing terms of the re | Field | Type | Required | Description | |-------|------|----------|-------------| -| `id` | UUID | For reproduced/cited/presented | Emitter-assigned unique event identifier; optional on other event types | +| `id` | UUID | For cited/presented | Emitter-assigned unique event identifier; optional on other event types | | `type` | EventType | Yes | Event type (see 5.3) | | `timestamp` | datetime | Yes | Event timestamp (UTC) | | `turn_id` | string | No | Associates this event with a conversation turn (see 5.2.1) | -| `output_id` | string | For reproduced/cited/presented | Opaque output-artifact identifier joining construction to later delivery | +| `output_id` | string | For cited/presented | Opaque output-artifact identifier joining construction to later delivery | | `output_element_id` | string | No | Opaque element within `output_id`, such as a passage, media track, caption, link, or card | -| `citation_id` | UUID | No | On `content_presented` or `content_reproduced`, the `id` of the associated citation event; absent for uncited presentations and uncredited reproductions | +| `citation_id` | UUID | No | On `content_presented`, the `id` of the associated citation event; absent for uncited presentations | | `presentation_id` | UUID | For engaged | On agent-reported `content_engaged`, the `id` of the exact presentation occurrence acted upon. Destination-reported events carrying an envelope `ctx_token` omit it (section 7.4) | | `ctx_token` | string | No | On agent-reported `content_engaged`, the click token minted for this engagement's presentation, recorded so destination reports join to it (section 7.4) | | `source_role` | SourceRole | No | Who is reporting: `origin`, `edge`, `index`, `agent` (see 4.4) | @@ -429,7 +417,7 @@ Emitters MUST NOT populate `access_context` unless the governing terms of the re #### 5.2.1 Turn association -The `turn_id` field associates content events with a specific conversation turn. Emitters SHOULD set `turn_id` on `content_reproduced`, `content_cited`, `content_presented`, and `content_engaged` events. Emitters SHOULD also set `turn_id` on `content_grounded` events when `scope` is `turn`. The corresponding `turn_started` and `turn_completed` events SHOULD carry the same `turn_id`. +The `turn_id` field associates content events with a specific conversation turn. Emitters SHOULD set `turn_id` on `content_cited`, `content_presented`, and `content_engaged` events. Emitters SHOULD also set `turn_id` on `content_grounded` events when `scope` is `turn`. The corresponding `turn_started` and `turn_completed` events SHOULD carry the same `turn_id`. `turn_id` is scoped to the session. Format is emitter-defined (sequential integers, UUIDs, or any opaque string). @@ -465,10 +453,9 @@ A processor that stores, forwards or transforms a document MUST preserve `terms_ |------|-------------|-----------------| | `content_retrieved` | Content fetched from source | `content_url`, `source_role`, `data.media_type` | | `content_grounded` | Content loaded into agent context | `content_url` or `content_id`, `data.scope`, `data.cached` | -| `content_reproduced` | Source content appears verbatim or near-verbatim in an output artifact | `id`, `output_id`, `content_url` or `content_id`, `data.reproduction_type` | | `content_cited` | Output explicitly associates source content with an output element | `id`, `output_id`, `content_url` or `content_id`, `data.citation_type` | | `content_presented` | Content or a source reference was made perceivable | `id`, `output_id`, `content_url` or `content_id`, `data.presentation_kind`, `data.presentation_type` | -| `content_engaged` | Observable action on an exact presentation | `presentation_id`, `content_url` or `content_id`, `data.engagement_type` (see 6.8) | +| `content_engaged` | Observable action on an exact presentation | `presentation_id`, `content_url` or `content_id`, `data.engagement_type` (see 6.7) | #### Conversation events @@ -561,7 +548,7 @@ Each level is named for the event it adds: a level proves the emitter produces t | **Grounding** | Above + `content_grounded`, turn events | Content entered the agent's context | Agent with basic instrumentation | | **Citation** | Above + `content_cited` | Content was explicitly referenced in the agent's response | Agent with citation instrumentation | -Reproduction, presentation, and engagement events are optional lifecycle signals outside the Retrieval/Grounding/Citation ladder. A Citation emitter SHOULD emit them when applicable (section 5.7.3), but the Citation level proves citation support, not full retrieval-to-engagement coverage. +Presentation and engagement events are optional lifecycle signals outside the Retrieval/Grounding/Citation ladder. A Citation emitter SHOULD emit them when applicable (section 5.7.3), but the Citation level proves citation support, not full retrieval-to-engagement coverage. #### 5.7.1 Retrieval conformance @@ -600,7 +587,6 @@ The privacy-level field restriction (section 5.5) applies to Citation emitters a A Citation emitter SHOULD: - Emit `content_presented` and `content_engaged` events when applicable -- Emit `content_reproduced` events whenever the response reproduces identified source content, including alongside every `direct_quote` citation (section 6.6) - Include `data.position` on citation events - Include `output_element_id` when the cited or presented element has a stable identity - Include `citation_id` on a presentation of a cited source association @@ -618,15 +604,15 @@ A conforming **telemetry consumer** MUST: The JSON Schema (`telemetry-session.json`) validates structure and types but cannot express every conformance rule. The following are normative requirements verified at the application layer, not by schema validation: -- At least one of `content_url` or `content_id` MUST be present on every content event (section 4.5). For `content_cited` and `content_reproduced` events this requirement is additionally enforced by the JSON Schema, which rejects an event whose reference is absent or null (sections 6.5, 6.6). +- At least one of `content_url` or `content_id` MUST be present on every content event (section 4.5). For `content_cited` events this requirement is additionally enforced by the JSON Schema, which rejects an event whose reference is absent or null (section 6.5). - An event MUST carry either `session_id` or `ctx_token` at Grounding conformance and above (section 7.1). - Conversation-turn fields MUST NOT exceed the turn's declared `privacy_level` (section 5.5). - The conformance-level requirements (sections 5.7.1 to 5.7.3) are cumulative. - When `content_grounded.data.provenance` is `agent_fetched`, `data.cached` MUST be `false`; when it is `agent_cached`, `data.cached` MUST be `true` (section 6.4). -- `content_grounded.data.content_fingerprint` MUST NOT contain `preserved_in_output`; output-side reuse is reported with `content_reproduced` (sections 6.4 and 12.1). +- `content_grounded.data.content_fingerprint` MUST NOT contain `preserved_in_output`; v1 defines no output-side reuse reporting (sections 6.4 and 12.1). - `source_role` MUST be present on every `content_retrieved` event (sections 5.2.2, 5.7.1). -- Fields scoped to an event type MUST NOT appear on other types: `presentation_id` and the event-level `ctx_token` only on `content_engaged`; `citation_id` only on `content_presented` and `content_reproduced`; `turn` only on `turn_started` and `turn_completed` (section 5.2). -- Within a session document, event `id` values MUST be distinct; a `content_engaged.presentation_id` MUST reference a `content_presented` event, and a `citation_id` a `content_cited` event, that identifies the same content - where both events carry `content_id` the values MUST be equal, and likewise for `content_url` (sections 6.7, 6.8); one event-level `ctx_token` MUST NOT appear on engagements bound to two different presentations (section 7.4.1). +- Fields scoped to an event type MUST NOT appear on other types: `presentation_id` and the event-level `ctx_token` only on `content_engaged`; `citation_id` only on `content_presented`; `turn` only on `turn_started` and `turn_completed` (section 5.2). +- Within a session document, event `id` values MUST be distinct; a `content_engaged.presentation_id` MUST reference a `content_presented` event, and a `citation_id` a `content_cited` event, that identifies the same content - where both events carry `content_id` the values MUST be equal, and likewise for `content_url` (sections 6.6, 6.7); one event-level `ctx_token` MUST NOT appear on engagements bound to two different presentations (section 7.4.1). - An envelope `ctx_token` (section 7.1) MUST accompany `content_engaged` events only. - Manifests: `domains` MUST appear only on a manifest served at the domain root, and `telemetry.ctx_resolution` only on a manifest declaring the `agent` or `platform` role (sections 8.5, 8.6). @@ -653,7 +639,7 @@ Whether an emitter's reporting in fact met its declared coverage is the complete ## 6. Data profiles -The `data` field on events carries type-specific metadata. These profiles document the recommended fields by event type and source role. None are required except where a section states otherwise - `scope` in 6.4, `citation_type` in 6.5, `reproduction_type` in 6.6, `presentation_kind` and `presentation_type` in 6.7, each enforced by the JSON Schema - but emitting them enables richer attribution. +The `data` field on events carries type-specific metadata. These profiles document the recommended fields by event type and source role. None are required except where a section states otherwise - `scope` in 6.4, `citation_type` in 6.5, `presentation_kind` and `presentation_type` in 6.6, each enforced by the JSON Schema - but emitting them enables richer attribution. ### 6.1 Retrieved content metadata (`content_retrieved`) @@ -666,7 +652,7 @@ When the reporter is the agent (`source_role: agent`), the following fields are `media_type` on retrieval events allows content owners to see what types of content are being fetched, independent of whether those retrievals result in grounding or citation. Defaults to `text` when absent. -`text`, `image`, `video`, and `audio` are the core values. Emitters MAY use custom string values for media outside the core set (for example `3d` or `dataset`). Telemetry consumers MUST tolerate unknown `media_type` values. This rule applies to `media_type` on every event type that carries it (sections 6.4, 6.5, 6.6, 6.7). +`text`, `image`, `video`, and `audio` are the core values. Emitters MAY use custom string values for media outside the core set (for example `3d` or `dataset`). Telemetry consumers MUST tolerate unknown `media_type` values. This rule applies to `media_type` on every event type that carries it (sections 6.4, 6.5, 6.6). `content_depth` records how much of the content record the retrieval reached: `metadata` for a bibliographic or descriptive record only, `abstract` for an abstract or summary record, `full` for the full content record. These are the core values; emitters MAY use custom values and telemetry consumers MUST tolerate unknown ones. Where entitlement gates depth, a retrieval that reached only an abstract and a retrieval of full text are otherwise indistinguishable at the retrieval layer. Depth records what was reachable at retrieval, independent of what portion later entered a generation context. @@ -756,7 +742,7 @@ Emitters MUST keep `provenance` and `cached` consistent: `agent_fetched` require `detected` reports a grounding-time observation by the emitter. It does not establish that the signal is authentic, identify who applied it, prove that the content was used later in the output, or raise the evidentiary status of the grounding event. Those questions require profile-defined evidence and consumer trust policy outside the core schema. -Output-side reuse is reported with `content_reproduced`, not a fingerprint-preservation field on `content_grounded`. A consumer MAY compare a grounding fingerprint with evidence about a reproduced output, but the two remain separate assertions about separate lifecycle stages. +The fingerprint is a grounding-time claim only: v1 defines no field asserting that a fingerprinted signal was preserved in the output (the pre-release `preserved_in_output` field is withdrawn; section 12.1). A consumer MAY compare a grounding fingerprint with evidence about the output gathered outside this specification, but the two remain separate assertions about separate lifecycle stages. `scheme` is an open identifier. Emitters SHOULD use a globally collision-resistant value. Core does not register schemes, interpret `value`, or assign capabilities or evidence status from a scheme identifier. A profile MAY define scheme-specific processing rules. @@ -835,34 +821,7 @@ The `unclassified` value for `citation_type` indicates the agent did not classif When `content_hash` is absent or does not match any grounding event's hash (for example, because the agent re-chunked content between grounding and citation), consumers SHOULD fall back to matching on `content_url` or `content_id`, accepting that the correlation may be imprecise when the same content appears in multiple grounding events. -A `direct_quote` citation records the credit; the reproduced material itself is recorded by a companion `content_reproduced` event (section 6.6). Emitters SHOULD emit both, sharing `output_element_id`, so that reproduction totals can be computed over reproduction events alone without unioning citation types. - -### 6.6 Reproduction data (`content_reproduced`) - -| Field | Type | Description | -|-------|------|-------------| -| `reproduction_type` | string | Fidelity of the reproduction: `verbatim`, `near_verbatim`, `unclassified` | -| `media_type` | string | Content medium: `text`, `image`, `video`, `audio` (open vocabulary, see 6.1). Defaults to `text` when absent. | -| `reproduced_chars` | integer | Unicode code points in the reproduced span as it appears in the output | -| `reproduced_tokens` | integer | Token count of the reproduced span (supplementary) | -| `reproduced_hash` | string | SHA-256 of the reproduced span as it appears in the output (`sha256:{hex}`) | -| `content_hash` | string | SHA-256 matching the corresponding `content_grounded` event (`sha256:{hex}`) | - -A `content_reproduced` event is the output constructor's claim that the output artifact contains identified source content: a quotation, an excerpt, or a full copy. It records construction only. Whether the reproduction was credited is recorded by a `content_cited` event; whether it reached a person is recorded by a `content_presented` event. The event exists so that reproduction remains reportable when neither occurred - an uncredited excerpt in a response delivered through an API is the canonical case, and produces a `content_reproduced` event with no citation and no presentation. - -Only the system that constructed the output emits this event (section 4.4). - -`verbatim` means the reproduced span matches the source exactly. `near_verbatim` means it differs only by bounded surface edits: truncation, elision marks, whitespace, punctuation, or case. Paraphrase is not reproduction - a credited paraphrase is a `content_cited` event with `citation_type: paraphrase`, and an uncredited one is silent grounding (section 10.2). A translation is likewise not a reproduction of the source text. `unclassified` indicates the emitter identified reproduced content without classifying its fidelity. Reproduction is not limited to excerpts: a full copy is the same event whose span is the entire work. - -`reproduced_chars` counts Unicode code points in the exact reproduced span as it appears in the output, under the same counting rule as `chars_ingested` (section 6.4): no normalisation applied solely for counting. `reproduced_tokens` is the agent-native supplementary measurement, following the same pairing as `excerpt_chars` and `excerpt_tokens` (section 6.5). `reproduced_hash` is the SHA-256 of the reproduced span as produced - for a `verbatim` reproduction it matches a hash of the corresponding source span, and for a credited quotation of the same span it equals the citation's `excerpt_hash`. `content_hash` correlates the reproduction to the grounding event it drew from, with the same fallback rules as citation data (section 6.5). - -The claim is only meaningful for an identified source, so a `content_reproduced` event MUST carry a non-null `content_url` or `content_id`. The JSON Schema enforces this, as it does for `content_cited`; reproduction the emitter cannot attribute to an identified source is not reportable as an event. - -Each `content_reproduced` event MUST have an `id` and `output_id`, and SHOULD carry `output_element_id` identifying the passage or element containing the reproduction. When the reproduction is credited, the event MAY carry `citation_id` referencing the crediting `content_cited` event, and the two SHOULD share `output_element_id`. When reproduced content is later made perceivable, that occurrence is a `content_presented` event with `presentation_kind: content` sharing the same `output_id`; presentation does not re-assert reproduction, and reproduction does not assert presentation. - -For non-text media, `media_type` and `reproduced_hash` identify the reproduced material. Finer-grained portion references for time-based and spatial media (time ranges, regions, segments) are not defined in this version. - -### 6.7 Presentation data (`content_presented`) +### 6.6 Presentation data (`content_presented`) | Field | Type | Description | |-------|------|-------------| @@ -892,7 +851,7 @@ Each presentation event MUST have an `id` and `output_id`. When it presents a ci When a session includes `content_presented` events but no subsequent `content_engaged` events, the telemetry establishes only that content or a reference was made perceivable and no reported interaction followed. It does not establish human attention. Whether this pattern is meaningful depends on the governing terms. Retrieval remains the only lifecycle stage observable from the CDN edge. -### 6.8 Engagement data (`content_engaged`) +### 6.7 Engagement data (`content_engaged`) | Field | Type | Description | |-------|------|-------------| @@ -1048,7 +1007,7 @@ Any party may operate a consumer: an agent operator, a licensing intermediary, o **Content owner resolution.** Telemetry consumers resolve content owner identity from `content_url` domains. Content owners register and verify their domains with the telemetry consumer; the consumer maps incoming event URLs to the owning organisation. This is the primary resolution path and requires `content_url` to be present on events. Events identified only by `content_id` (e.g., cached groundings where the URL was not preserved, or marketplace API content with no canonical URL) cannot be resolved by domain alone. Telemetry consumers SHOULD support `content_id` prefix-based resolution as a secondary path when content owners register their identifier schemes, but this is not yet a normative requirement. -**Cross-consumer correlation.** Origin-side emitters and agent-side emitters MAY use different telemetry consumers. A content owner's CDN sends retrieval events to one telemetry consumer; an agent sends sessions to another. The `content_telemetry_id` field (section 7.2) correlates the same retrieval across consumers - both sides share the same UUID from the HTTP request. This correlation operates at the retrieval level only. Grounding, reproduction, citation, presentation, and engagement events have no independent origin-side counterpart to correlate against. +**Cross-consumer correlation.** Origin-side emitters and agent-side emitters MAY use different telemetry consumers. A content owner's CDN sends retrieval events to one telemetry consumer; an agent sends sessions to another. The `content_telemetry_id` field (section 7.2) correlates the same retrieval across consumers - both sides share the same UUID from the HTTP request. This correlation operates at the retrieval level only. Grounding, citation, presentation, and engagement events have no independent origin-side counterpart to correlate against. ### 7.4 Click context (`ctx_token`) @@ -1079,7 +1038,7 @@ The token stays opaque and carries no routing; the locator travels alongside it. A telemetry consumer that supports resolution exposes, for a presented token, the **click context**: 1. **The engagement.** The `content_engaged` occurrence(s) bound to the token, including the presentation record the token restores: `presentation_id`, `output_id`, `output_element_id` where present, and timestamps. -2. **The lineage of the clicked content, selected by content identity.** The resolved session's `content_retrieved`, `content_grounded`, `content_reproduced`, `content_cited`, and `content_presented` events whose `content_url` or `content_id` identify the same content as the clicked reference - across all turns. The cut is by content identity, not by turn or click timestamp: a click in turn 5 on content grounded in turn 2 resolves that content's full lineage. +2. **The lineage of the clicked content, selected by content identity.** The resolved session's `content_retrieved`, `content_grounded`, `content_cited`, and `content_presented` events whose `content_url` or `content_id` identify the same content as the clicked reference - across all turns. The cut is by content identity, not by turn or click timestamp: a click in turn 5 on content grounded in turn 2 resolves that content's full lineage. 3. **The contributing sources, gated per owner.** The sources that informed the response the click came from: `content_grounded` events in scope for the engaged presentation's turn (including session-scoped groundings) and that turn's `content_cited` and `content_presented` events, for content other than the clicked content. A contributing owner's events appear only when that owner has opted in to contributing-source disclosure with the resolving consumer; owners without a recorded opt-in are visible only through the counts in the session summary. This is the component that supports multi-citation attribution when the clicked content is not the contributing content - a click through to a commerce destination whose recommendation a publisher's review produced - and it restores the consent-gated role of the v0.1 click manifest (section 12.1), scoped to the click's provenance rather than the whole session. 4. **An optional privacy-bounded session summary.** Event counts by type and a distinct-source count. Counts, not events, and no content identifiers of owners not disclosed above. @@ -1364,20 +1323,18 @@ Whether this constitutes one royalty event, three, or ten depends on the commerc | Per-grounding | One event per article entering context per session | Access-based or flat-fee licensing ("you used our content") | | Per-citation | One event per explicit reference in a response | Performance-based licensing ("you cited our content") | | Per-turn-influenced | One event per turn where content was in context | Usage-based licensing ("our content informed N answers") | -| Per-reproduction | One event per reproduced portion appearing in output | Excerpt-based licensing ("N characters of our content appeared in answers") | -The `content_grounded` event with `scope: session` plus the count of subsequent `turn_completed` events provides the inputs for the first three models; the per-reproduction model additionally draws on `content_reproduced` events and their `reproduced_chars`. None of the four requires the schema to embed a commercial opinion. +The `content_grounded` event with `scope: session` plus the count of subsequent `turn_completed` events provides the inputs for all three models without requiring the schema to embed a commercial opinion. ### 10.2 Grounding without citation Content can influence every response in a session without being explicitly cited. A common royalty formula (individual content owner usage / total content owner usage x royalty rate) can be applied at any level of the funnel: - At the **grounding** level: counts all content that was in the agent's context, regardless of citation. This captures the full extent of content influence, including silent grounding. -- At the **reproduction** level: counts source material appearing in the output, credited or not. This captures verbatim reuse that citation-level counting misses, with `reproduced_chars` providing a magnitude. - At the **citation** level: counts only explicitly attributed content. Simpler to verify but undercounts content influence. - At the **presentation** level: counts content or source references made perceivable. It does not prove attention. -Content owners and platforms should agree on which level to count at. The telemetry data supports all four; the choice is commercial, not technical. +Content owners and platforms should agree on which level to count at. The telemetry data supports all three; the choice is commercial, not technical. ## 11. Extensibility @@ -1470,9 +1427,9 @@ on every `content_grounded` event; both are schema-enforced. A v0.1 emitter that omitted them migrates a citation it cannot classify with `citation_type: unclassified`, and a grounding whose scope it did not record with `scope: turn` where the event carries a `turn_id` and `scope: session` otherwise. V1 also -rejects a `content_cited` or `content_reproduced` event whose `content_url` and -`content_id` are both absent or null (sections 6.5, 6.6): a v0.1 citation with no -resolvable reference is not migrated as a citation. +rejects a `content_cited` event whose `content_url` and `content_id` are both +absent or null (section 6.5): a v0.1 citation with no resolvable reference is +not migrated as a citation. V1 withdraws `ip_hash` from the edge and origin retrieval profiles (section 9.1). Emitters remove the field and MUST NOT populate it; `asn`, `asn_org` and `country` @@ -1483,22 +1440,13 @@ remain. V1 also requires `source_role` on every `content_retrieved` event (section 5.2.3): a consumer that read a v0.1 `license_ref` as verification of entitlement now reads it as the emitter's claim about which grant applied. -V1 adds `content_reproduced`. The v0.1 preview has no equivalent: verbatim reuse -was inferable only from `direct_quote` citations, which conflate the credit with -the material and cannot record an uncredited copy. When migrating historical -v0.1 data, consumers MAY treat a `direct_quote` citation as an implied -reproduction of its excerpt. V1 emitters record reproduction explicitly and -SHOULD NOT rely on that inference. - V1 grounding fingerprints report detection only. The published v0.1 preview defined no `content_fingerprint` object; the object, and a `preserved_in_output` field within it, appeared only on the pre-release -`v1-draft` line. An implementation built against that draft removes -`preserved_in_output` during migration. When its value was `true` and the -historical record contains enough information to populate every required -`content_reproduced` field, the implementation MAY also create the corresponding -reproduction event. It MUST NOT synthesize a reproduction event from `false` or -incomplete historical data. A grounding event MAY retain +`v1-draft` line, which also carried a `content_reproduced` event type that is +not part of v1. An implementation built against that draft removes +`preserved_in_output` and any `content_reproduced` events during migration; +v1 defines no output-side reuse reporting. A grounding event MAY retain `content_fingerprint.scheme`, `detected`, and a scheme-defined `value`. V1 narrows `ctx_token` resolution. The v0.1 click manifest returned every @@ -1514,7 +1462,7 @@ from a token to the presentation it was minted for is issuer state: v0.1 defined no `presentation_id`, and v1 never places one in a URL; a destination reports the token, and the consumer restores the binding at resolution. -V1 tightens occurrence boundaries (section 4.3). Each core event now has a stated occurrence and cardinality: retrieval per completed fetch (a cache serve is not a retrieval), grounding per distinct content item per declared scope (chunk-level events deduplicate to one occurrence by content identity), reproduction per source item per output element, presentation per rendering occurrence, engagement per observed action. Preview emitters that emitted per chunk, per passage, or re-emitted `content_retrieved` on cache serves remain schema-valid but SHOULD re-map to the stated boundaries; consumers comparing preview and v1 volumes should expect counts to shift where emitters previously chose finer or coarser units. Coverage becomes an explicit declaration (section 5.7.6) rather than an implication of conformance level. Session-scoped extension metadata belongs in the session-level `data` container (section 5.1.3); custom top-level siblings of `events`, accepted by the preview schema, are undefined. +V1 tightens occurrence boundaries (section 4.3). Each core event now has a stated occurrence and cardinality: retrieval per completed fetch (a cache serve is not a retrieval), grounding per distinct content item per declared scope (chunk-level events deduplicate to one occurrence by content identity), citation per source-to-element association, presentation per rendering occurrence, engagement per observed action. Preview emitters that emitted per chunk, per passage, or re-emitted `content_retrieved` on cache serves remain schema-valid but SHOULD re-map to the stated boundaries; consumers comparing preview and v1 volumes should expect counts to shift where emitters previously chose finer or coarser units. Coverage becomes an explicit declaration (section 5.7.6) rather than an implication of conformance level. Session-scoped extension metadata belongs in the session-level `data` container (section 5.1.3); custom top-level siblings of `events`, accepted by the preview schema, are undefined. Version numbers are `major.minor`, and a document declares the version it was produced under in `schema_version`. From 1.0 onward: diff --git a/telemetry-session.json b/telemetry-session.json index c950987..fe68ba7 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -125,22 +125,22 @@ }, "turn_id": { "type": ["string", "null"], - "description": "Associates this event with a conversation turn. Scoped to the session. SHOULD be set on turn_started, turn_completed, content_reproduced, content_cited, content_presented, content_engaged events, and content_grounded events when scope is turn." + "description": "Associates this event with a conversation turn. Scoped to the session. SHOULD be set on turn_started, turn_completed, content_cited, content_presented, content_engaged events, and content_grounded events when scope is turn." }, "output_id": { "type": "string", "minLength": 1, - "description": "Opaque identifier for the output artifact. REQUIRED on content_reproduced, content_cited, and content_presented events so output construction and later presentation can be correlated across services or times." + "description": "Opaque identifier for the output artifact. REQUIRED on content_cited and content_presented events so output construction and later presentation can be correlated across services or times." }, "output_element_id": { "type": "string", "minLength": 1, - "description": "Opaque identifier for the element within output_id that carries the reproduction, citation, or presentation, such as a passage, media track, caption, link, or card." + "description": "Opaque identifier for the element within output_id that carries the citation or presentation, such as a passage, media track, caption, link, or card." }, "citation_id": { "type": "string", "format": "uuid", - "description": "The id of the content_cited event associated with this presentation or reproduction. Valid only on content_presented and content_reproduced events; absent when the presentation is not a citation or the reproduction is uncredited." + "description": "The id of the content_cited event associated with this presentation. Valid only on content_presented events; absent when the presentation is not a citation." }, "presentation_id": { "type": "string", @@ -243,32 +243,6 @@ } } }, - { - "if": { - "properties": { "type": { "const": "content_reproduced" } }, - "required": ["type"] - }, - "then": { - "required": ["id", "output_id", "data"], - "anyOf": [ - { "required": ["content_url"], "properties": { "content_url": { "type": "string" } } }, - { "required": ["content_id"], "properties": { "content_id": { "type": "string" } } } - ], - "properties": { - "data": { - "required": ["reproduction_type"], - "properties": { - "reproduction_type": { "$ref": "#/$defs/ReproductionType" }, - "media_type": { "$ref": "#/$defs/MediaType" }, - "reproduced_chars": { "type": "integer", "minimum": 0, "description": "Unicode code points in the reproduced span as it appears in the output" }, - "reproduced_tokens": { "type": "integer", "minimum": 0 }, - "reproduced_hash": { "type": "string", "pattern": "^sha256:[a-f0-9]{64}$", "description": "SHA-256 of the reproduced span as it appears in the output (sha256:{hex})" }, - "content_hash": { "type": "string", "pattern": "^sha256:[a-f0-9]{64}$", "description": "SHA-256 matching the corresponding content_grounded event. When the agent chunked the source, this is the chunk hash." } - } - } - } - } - }, { "if": { "properties": { "type": { "const": "content_presented" } }, @@ -338,7 +312,6 @@ "enum": [ "content_retrieved", "content_grounded", - "content_reproduced", "content_cited", "content_presented", "content_engaged", @@ -432,11 +405,6 @@ "description": "How content was used in the response", "enum": ["direct_quote", "paraphrase", "reference", "contradiction", "unclassified"] }, - "ReproductionType": { - "type": "string", - "description": "Fidelity of a reproduction. near_verbatim differs from the source only by bounded surface edits: truncation, elision marks, whitespace, punctuation, or case.", - "enum": ["verbatim", "near_verbatim", "unclassified"] - }, "CitationPosition": { "type": "string", "description": "Prominence of citation in response", diff --git a/tests/README.md b/tests/README.md index 4dde243..25fa583 100644 --- a/tests/README.md +++ b/tests/README.md @@ -30,11 +30,11 @@ Run from the repository root. Without uv: `pip install "jsonschema[format-nongpl - Turn required fields (`privacy_level`) - Enum validation (event types, privacy levels, source roles, schema version) - Citation source-reference requirement (content_cited rejected when content_url/content_id are missing or null) and required citation_type -- Required event and output identifiers on cited, reproduced, and presented events -- Closed enums (citation_type, position, scope, provenance, reproduction_type, presentation_kind) and non-negative counts +- Required event and output identifiers on cited and presented events +- Closed enums (citation_type, position, scope, provenance, presentation_kind) and non-negative counts - Required `data.scope` on grounding events; required `source_role` on retrieval events - Format assertions on every uuid, date-time and uri field (malformed session, event, citation, presentation and parent ids; malformed started_at; malformed turn URL arrays) -- Field placement: presentation_id and event-level ctx_token only on content_engaged, citation_id only on presented/reproduced, turn only on turn events; envelope ctx_token only with engagements +- Field placement: presentation_id and event-level ctx_token only on content_engaged, citation_id only on presented events, turn only on turn events; envelope ctx_token only with engagements - Session integrity: distinct event ids, engagement/presentation and presentation/citation identify the same content, one token per presentation - Rejection of documents declaring the v0.1 wire version - All three conformance levels (Retrieval, Grounding, Citation) and all four source roles (origin, edge, index, agent) @@ -42,8 +42,7 @@ Run from the repository root. Without uv: `pip install "jsonschema[format-nongpl - Manifests for all three roles, including platform; manifest rejection for http endpoints, path-prefixed domains, ctx_resolution on the wrong role, duplicate or empty roles, missing endpoint and coverage mode - Standalone event envelopes (CDN edge, agent with session FK) - Privacy level field gating (application-layer conformance), one fixture per forbidden field at minimal and intent, in session, standalone and batch shapes -- Funnel exceptions (presented-no-cited, cited-no-grounded, presented-no-grounded, reproduced-no-cited, reproduced-no-grounded) -- Reproduction cases: credited quotation (reproduction + direct_quote citation sharing an output element) and uncredited reproduction in unpresented API output +- Funnel exceptions (presented-no-cited, cited-no-grounded, presented-no-grounded) - Text, image, audio, video, suppressed-citation, and repeated-presentation cases - Exact presentation-to-engagement correlation across session and standalone envelopes - Multi-turn sessions and cached grounding @@ -58,9 +57,9 @@ Some rules cannot be expressed in JSON Schema alone. These are tested as applica - Privacy level field gating (e.g. `query_text` MUST NOT be present at `minimal` level), applied to turns wherever they appear: session documents, batches, and standalone envelopes - each shape has its own fixture - `content_url` or `content_id` requirement on every content event, and `source_role` on every `content_retrieved` event (sections 5.2.2, 5.7.5), in every document shape -- Field placement by event type (section 5.7.5): `presentation_id` and event-level `ctx_token` only on `content_engaged`, `citation_id` only on `content_presented`/`content_reproduced`, `turn` only on turn events; an envelope `ctx_token` only with `content_engaged` events +- Field placement by event type (section 5.7.5): `presentation_id` and event-level `ctx_token` only on `content_engaged`, `citation_id` only on `content_presented`, `turn` only on turn events; an envelope `ctx_token` only with `content_engaged` events - `session_id` or `ctx_token` on a standalone event or event batch envelope at Grounding conformance and above (sections 5.7.5, 7.1) -- Referential integrity within a session document: event ids are distinct; `content_engaged.presentation_id` matches a `content_presented` event id and `citation_id` on `content_presented`/`content_reproduced` matches a `content_cited` event id, in each case identifying the same content; one event-level `ctx_token` binds to one presentation (sections 6.6-6.8, 7.4.1). Standalone envelopes and batch members are exempt - they may reference events delivered elsewhere. +- Referential integrity within a session document: event ids are distinct; `content_engaged.presentation_id` matches a `content_presented` event id and `citation_id` on `content_presented` matches a `content_cited` event id, in each case identifying the same content; one event-level `ctx_token` binds to one presentation (sections 6.6-6.7, 7.4.1). Standalone envelopes and batch members are exempt - they may reference events delivered elsewhere. - Manifest rejection rules: duplicate `keys[].id`; `domains` entries that are not the manifest's own host or a subdomain of it; `domains` on a manifest served under a path prefix; `ctx_resolution` on a manifest without the `agent` or `platform` role (sections 8.5-8.7) - Withdrawn `ip_hash` prohibition on event data (section 9.1 migration rule), in every document shape - Grounding provenance/cache consistency and the prohibition on `preserved_in_output` in `content_fingerprint` (sections 5.7.5, 6.4, 12.1) diff --git a/tests/invalid/batch-engaged-missing-presentation-id.json b/tests/invalid/batch-engaged-missing-presentation-id.json index 71621a3..6c51576 100644 --- a/tests/invalid/batch-engaged-missing-presentation-id.json +++ b/tests/invalid/batch-engaged-missing-presentation-id.json @@ -1,5 +1,5 @@ { - "_test_description": "Event batch with session_id (no ctx_token) whose content_engaged lacks presentation_id. The relaxation applies only to envelopes carrying ctx_token (sections 6.8, 7.4).", + "_test_description": "Event batch with session_id (no ctx_token) whose content_engaged lacks presentation_id. The relaxation applies only to envelopes carrying ctx_token (sections 6.7, 7.4).", "_expected_error": "/events/0 'presentation_id' is a required property", "document_type": "event_batch", "schema_version": "1.0", diff --git a/tests/invalid/citation-id-on-grounded.json b/tests/invalid/citation-id-on-grounded.json index 104350f..c6f29dc 100644 --- a/tests/invalid/citation-id-on-grounded.json +++ b/tests/invalid/citation-id-on-grounded.json @@ -1,5 +1,5 @@ { - "_test_description": "APPLICATION-LAYER VIOLATION: citation_id on a content_grounded event. Valid only on content_presented and content_reproduced (sections 5.2, 5.7.5).", + "_test_description": "APPLICATION-LAYER VIOLATION: citation_id on a content_grounded event. Valid only on content_presented (sections 5.2, 5.7.5).", "_expected_error": "Field 'citation_id' present on 'content_grounded' event", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440632", diff --git a/tests/invalid/cited-missing-id.json b/tests/invalid/cited-missing-id.json index 4ed1401..f4cab7e 100644 --- a/tests/invalid/cited-missing-id.json +++ b/tests/invalid/cited-missing-id.json @@ -1,5 +1,5 @@ { - "_test_description": "content_cited event with output_id and a resolvable source reference but no event id. id is required on citation events so a presentation or reproduction can reference the exact citation via citation_id (section 6.5).", + "_test_description": "content_cited event with output_id and a resolvable source reference but no event id. id is required on citation events so a presentation can reference the exact citation via citation_id (section 6.5).", "_expected_error": "/events/0 'id' is a required property", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440203", diff --git a/tests/invalid/engaged-presentation-content-mismatch.json b/tests/invalid/engaged-presentation-content-mismatch.json index c720e69..1298285 100644 --- a/tests/invalid/engaged-presentation-content-mismatch.json +++ b/tests/invalid/engaged-presentation-content-mismatch.json @@ -1,5 +1,5 @@ { - "_test_description": "APPLICATION-LAYER VIOLATION: content_engaged references a presentation of different content. The engagement identifies the same content as the presentation it acted on (sections 5.7.5, 6.8).", + "_test_description": "APPLICATION-LAYER VIOLATION: content_engaged references a presentation of different content. The engagement identifies the same content as the presentation it acted on (sections 5.7.5, 6.7).", "_expected_error": "content_engaged references presentation", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440636", diff --git a/tests/invalid/engaged-presentation-id-no-presentations.json b/tests/invalid/engaged-presentation-id-no-presentations.json index d72098c..e3068f4 100644 --- a/tests/invalid/engaged-presentation-id-no-presentations.json +++ b/tests/invalid/engaged-presentation-id-no-presentations.json @@ -1,5 +1,5 @@ { - "_test_description": "APPLICATION-LAYER VIOLATION: content_engaged carries a presentation_id in a session with no content_presented events at all (section 6.8).", + "_test_description": "APPLICATION-LAYER VIOLATION: content_engaged carries a presentation_id in a session with no content_presented events at all (section 6.7).", "_expected_error": "does not match any content_presented event id", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440639", diff --git a/tests/invalid/engaged-presentation-id-unmatched.json b/tests/invalid/engaged-presentation-id-unmatched.json index 7a0aacd..b2c0302 100644 --- a/tests/invalid/engaged-presentation-id-unmatched.json +++ b/tests/invalid/engaged-presentation-id-unmatched.json @@ -1,5 +1,5 @@ { - "_test_description": "APPLICATION-LAYER VIOLATION: content_engaged whose presentation_id matches no content_presented event id in the session. Section 6.8 requires presentation_id to reference the exact content_presented.id on which the action occurred; an all-zeros UUID satisfies the schema's format assertion but references nothing.", + "_test_description": "APPLICATION-LAYER VIOLATION: content_engaged whose presentation_id matches no content_presented event id in the session. Section 6.7 requires presentation_id to reference the exact content_presented.id on which the action occurred; an all-zeros UUID satisfies the schema's format assertion but references nothing.", "_expected_error": "presentation_id '00000000-0000-0000-0000-000000000000' does not match any content_presented event id", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440236", diff --git a/tests/invalid/engaged-standalone-missing-presentation-and-token.json b/tests/invalid/engaged-standalone-missing-presentation-and-token.json index aab2c10..d61b974 100644 --- a/tests/invalid/engaged-standalone-missing-presentation-and-token.json +++ b/tests/invalid/engaged-standalone-missing-presentation-and-token.json @@ -1,5 +1,5 @@ { - "_test_description": "Agent-reported standalone content_engaged with session_id but no presentation_id. The presentation_id relaxation applies only to envelopes carrying ctx_token; an emitter that knows the session knows the presentation and MUST bind to it (sections 6.8, 7.4).", + "_test_description": "Agent-reported standalone content_engaged with session_id but no presentation_id. The presentation_id relaxation applies only to envelopes carrying ctx_token; an emitter that knows the session knows the presentation and MUST bind to it (sections 6.7, 7.4).", "_expected_error": "/event 'presentation_id' is a required property", "document_type": "event", "schema_version": "1.0", diff --git a/tests/invalid/grounding-fingerprint-preserved-in-output.json b/tests/invalid/grounding-fingerprint-preserved-in-output.json index bd5bfe3..0251418 100644 --- a/tests/invalid/grounding-fingerprint-preserved-in-output.json +++ b/tests/invalid/grounding-fingerprint-preserved-in-output.json @@ -1,6 +1,6 @@ { - "_test_description": "Grounding fingerprint uses withdrawn preserved_in_output instead of a content_reproduced event. Passes the extensible schema but violates the v1 migration rule.", - "_expected_error": "content_fingerprint carries preserved_in_output; use content_reproduced", + "_test_description": "Grounding fingerprint uses withdrawn preserved_in_output. Passes the extensible schema but violates the v1 migration rule (sections 6.4, 12.1).", + "_expected_error": "content_fingerprint carries preserved_in_output; withdrawn in v1", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440183", "started_at": "2026-08-12T10:30:00Z", diff --git a/tests/invalid/invalid-event-type.json b/tests/invalid/invalid-event-type.json index 09993df..9fd1aff 100644 --- a/tests/invalid/invalid-event-type.json +++ b/tests/invalid/invalid-event-type.json @@ -1,6 +1,6 @@ { "_test_description": "Event with type 'content_summarised' which is not in the EventType enum. Fails JSON Schema validation.", - "_expected_error": "/events/0/type 'content_summarised' is not one of ['content_retrieved', 'content_grounded', 'content_reproduced', 'content_cited', 'content_presented', 'content", + "_expected_error": "/events/0/type 'content_summarised' is not one of ['content_retrieved', 'content_grounded', 'content_cited', 'content_presented', 'content", "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440008", "started_at": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/legacy-content-displayed.json b/tests/invalid/legacy-content-displayed.json index f0c3b36..129d17e 100644 --- a/tests/invalid/legacy-content-displayed.json +++ b/tests/invalid/legacy-content-displayed.json @@ -1,6 +1,6 @@ { "_test_description": "The v1 presentation event replaces the preview content_displayed name.", - "_expected_error": "/events/0/type 'content_displayed' is not one of ['content_retrieved', 'content_grounded', 'content_reproduced', 'content_cited', 'content_presented', 'content_", + "_expected_error": "/events/0/type 'content_displayed' is not one of ['content_retrieved', 'content_grounded', 'content_cited', 'content_presented', 'content_", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440093", "started_at": "2026-07-18T13:00:00Z", diff --git a/tests/invalid/reproduced-citation-id-unmatched.json b/tests/invalid/reproduced-citation-id-unmatched.json deleted file mode 100644 index 5f3a062..0000000 --- a/tests/invalid/reproduced-citation-id-unmatched.json +++ /dev/null @@ -1,23 +0,0 @@ -{ - "_test_description": "APPLICATION-LAYER VIOLATION: content_reproduced.citation_id matches no content_cited event id in the session (section 6.6).", - "_expected_error": "content_reproduced citation_id '00000000-0000-0000-0000-000000000000' does not match", - "schema_version": "1.0", - "session_id": "660e8400-e29b-41d4-a716-446655440638", - "agent_id": "assistant-v4", - "started_at": "2026-08-20T10:00:00Z", - "events": [ - { - "id": "770e8400-e29b-41d4-a716-446655440603", - "type": "content_reproduced", - "timestamp": "2026-08-20T10:00:03Z", - "turn_id": "1", - "output_id": "response:1", - "content_url": "https://www.example-news.com/economy/rate-decision-analysis", - "data": { - "reproduction_type": "verbatim", - "reproduced_chars": 180 - }, - "citation_id": "00000000-0000-0000-0000-000000000000" - } - ] -} diff --git a/tests/invalid/reproduced-missing-data.json b/tests/invalid/reproduced-missing-data.json deleted file mode 100644 index a1c77d1..0000000 --- a/tests/invalid/reproduced-missing-data.json +++ /dev/null @@ -1,18 +0,0 @@ -{ - "_test_description": "content_reproduced with no data object. reproduction_type is required, so data itself is required (section 6.6).", - "_expected_error": "/events/0 'data' is a required property", - "schema_version": "1.0", - "session_id": "660e8400-e29b-41d4-a716-446655440616", - "agent_id": "assistant-v4", - "started_at": "2026-08-20T10:00:00Z", - "events": [ - { - "id": "770e8400-e29b-41d4-a716-446655440603", - "type": "content_reproduced", - "timestamp": "2026-08-20T10:00:03Z", - "turn_id": "1", - "output_id": "response:1", - "content_url": "https://www.example-news.com/economy/rate-decision-analysis" - } - ] -} diff --git a/tests/invalid/reproduced-missing-id.json b/tests/invalid/reproduced-missing-id.json deleted file mode 100644 index 05014a1..0000000 --- a/tests/invalid/reproduced-missing-id.json +++ /dev/null @@ -1,19 +0,0 @@ -{ - "_test_description": "content_reproduced event with output_id and a resolvable source reference but no event id. id is required on reproduction events so a crediting citation or verification result can reference the exact reproduction claim.", - "_expected_error": "/events/0 'id' is a required property", - "schema_version": "1.0", - "session_id": "660e8400-e29b-41d4-a716-446655440162", - "started_at": "2026-08-01T14:30:00Z", - "events": [ - { - "type": "content_reproduced", - "timestamp": "2026-08-01T14:30:01Z", - "output_id": "response:1", - "content_id": "publisher:article:3", - "data": { - "reproduction_type": "verbatim", - "reproduced_chars": 320 - } - } - ] -} diff --git a/tests/invalid/reproduced-missing-output-id.json b/tests/invalid/reproduced-missing-output-id.json deleted file mode 100644 index 258c4cc..0000000 --- a/tests/invalid/reproduced-missing-output-id.json +++ /dev/null @@ -1,19 +0,0 @@ -{ - "_test_description": "content_reproduced must identify the output artifact containing the reproduction: id and output_id are required so reproduction can be correlated with citation and presentation.", - "_expected_error": "/events/0 'output_id' is a required property", - "schema_version": "1.0", - "session_id": "660e8400-e29b-41d4-a716-446655440130", - "started_at": "2026-08-01T11:00:00Z", - "events": [ - { - "id": "770e8400-e29b-41d4-a716-446655440131", - "type": "content_reproduced", - "timestamp": "2026-08-01T11:00:01Z", - "content_id": "publisher:article:2", - "data": { - "reproduction_type": "verbatim", - "reproduced_chars": 500 - } - } - ] -} diff --git a/tests/invalid/reproduced-missing-source-reference.json b/tests/invalid/reproduced-missing-source-reference.json deleted file mode 100644 index ac05921..0000000 --- a/tests/invalid/reproduced-missing-source-reference.json +++ /dev/null @@ -1,19 +0,0 @@ -{ - "_test_description": "content_reproduced event carrying neither content_url nor content_id. A reproduction claim is only meaningful for an identified source; the JSON Schema requires a non-null content_url or content_id on content_reproduced (section 6.6).", - "_expected_error": "/events/0 'content_url' is a required property", - "schema_version": "1.0", - "session_id": "660e8400-e29b-41d4-a716-446655440150", - "started_at": "2026-08-01T13:00:00Z", - "events": [ - { - "id": "770e8400-e29b-41d4-a716-446655440151", - "type": "content_reproduced", - "timestamp": "2026-08-01T13:00:01Z", - "output_id": "response:1", - "data": { - "reproduction_type": "verbatim", - "reproduced_chars": 250 - } - } - ] -} diff --git a/tests/invalid/reproduced-missing-type.json b/tests/invalid/reproduced-missing-type.json deleted file mode 100644 index 09c6fc3..0000000 --- a/tests/invalid/reproduced-missing-type.json +++ /dev/null @@ -1,19 +0,0 @@ -{ - "_test_description": "content_reproduced must classify the fidelity of the reproduction: data.reproduction_type is required.", - "_expected_error": "/events/0/data 'reproduction_type' is a required property", - "schema_version": "1.0", - "session_id": "660e8400-e29b-41d4-a716-446655440120", - "started_at": "2026-08-01T10:00:00Z", - "events": [ - { - "id": "770e8400-e29b-41d4-a716-446655440121", - "type": "content_reproduced", - "timestamp": "2026-08-01T10:00:01Z", - "output_id": "response:1", - "content_id": "publisher:article:1", - "data": { - "reproduced_chars": 300 - } - } - ] -} diff --git a/tests/invalid/reproduced-negative-chars.json b/tests/invalid/reproduced-negative-chars.json deleted file mode 100644 index 56ae558..0000000 --- a/tests/invalid/reproduced-negative-chars.json +++ /dev/null @@ -1,21 +0,0 @@ -{ - "_test_description": "content_reproduced with a negative reproduced_chars. The reproduced span length is a Unicode code point count and cannot be negative (section 6.6); the schema requires a minimum of 0.", - "_expected_error": "/events/0/data/reproduced_chars -412 is less than the minimum of 0", - "schema_version": "1.0", - "session_id": "660e8400-e29b-41d4-a716-446655440211", - "started_at": "2026-08-05T10:10:00Z", - "events": [ - { - "id": "770e8400-e29b-41d4-a716-446655440212", - "type": "content_reproduced", - "timestamp": "2026-08-05T10:10:01Z", - "turn_id": "1", - "output_id": "response:1", - "content_id": "publisher:article:9", - "data": { - "reproduction_type": "verbatim", - "reproduced_chars": -412 - } - } - ] -} diff --git a/tests/invalid/reproduction-type-invalid.json b/tests/invalid/reproduction-type-invalid.json deleted file mode 100644 index e7e39ac..0000000 --- a/tests/invalid/reproduction-type-invalid.json +++ /dev/null @@ -1,21 +0,0 @@ -{ - "_test_description": "content_reproduced with reproduction_type 'paraphrase'. Paraphrase is not reproduction (section 6.6): a credited paraphrase is a content_cited event with citation_type 'paraphrase', and an uncredited one is silent grounding. reproduction_type accepts only verbatim, near_verbatim, unclassified.", - "_expected_error": "/events/0/data/reproduction_type 'paraphrase' is not one of ['verbatim', 'near_verbatim', 'unclassified']", - "schema_version": "1.0", - "session_id": "660e8400-e29b-41d4-a716-446655440206", - "started_at": "2026-08-05T09:40:00Z", - "events": [ - { - "id": "770e8400-e29b-41d4-a716-446655440207", - "type": "content_reproduced", - "timestamp": "2026-08-05T09:40:01Z", - "turn_id": "1", - "output_id": "response:1", - "content_id": "publisher:article:7", - "data": { - "reproduction_type": "paraphrase", - "reproduced_chars": 280 - } - } - ] -} diff --git a/tests/mutation_smoke.py b/tests/mutation_smoke.py index 0b16cf9..ebc41ab 100644 --- a/tests/mutation_smoke.py +++ b/tests/mutation_smoke.py @@ -137,9 +137,6 @@ def mutate(root): REVIEW_MUTATIONS = [ # validate.py application-layer weakenings - ("drop content_reproduced from the citation_id integrity check", - _text(_V, 'if etype in ("content_presented", "content_reproduced"):', 'if etype in ("content_presented",):', - "reproduced-citation-id-unmatched.json")), ("exempt any envelope containing a retrieval from the session/ctx_token rule", _text(_V, 'if types <= {"content_retrieved"}:', 'if "content_retrieved" in types:', "batch-missing-session-mixed-retrieval.json")), @@ -216,8 +213,6 @@ def mutate(root): _json(_S, lambda d: _event_branch(d, "content_presented")["required"].remove("data"), "presented-missing-data.json")), ("drop required data on content_cited", _json(_S, lambda d: _event_branch(d, "content_cited")["required"].remove("data"), "cited-missing-data.json")), - ("drop required data on content_reproduced", - _json(_S, lambda d: _event_branch(d, "content_reproduced")["required"].remove("data"), "reproduced-missing-data.json")), ("remove the sha256 hash patterns", _text(_S, '"pattern": "^sha256:[a-f0-9]{64}$"', '"type": "string"', "malformed-content-hash.json")), ("remove minLength on output_id", diff --git a/tests/valid/session-engagement-types.json b/tests/valid/session-engagement-types.json index 560c682..24dcc2e 100644 --- a/tests/valid/session-engagement-types.json +++ b/tests/valid/session-engagement-types.json @@ -1,5 +1,5 @@ { - "_test_description": "The non-click engagement types: a snippet presentation is expanded, and the detail view it opens is copied from and shared. Each action is one engagement occurrence on one presentation (sections 4.3, 6.8); three actions on two presentations, all in-product and agent-reported.", + "_test_description": "The non-click engagement types: a snippet presentation is expanded, and the detail view it opens is copied from and shared. Each action is one engagement occurrence on one presentation (sections 4.3, 6.7); three actions on two presentations, all in-product and agent-reported.", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440650", "agent_id": "assistant-v4", diff --git a/tests/valid/session-repeated-link-engagement.json b/tests/valid/session-repeated-link-engagement.json index 589ee70..ef7fea8 100644 --- a/tests/valid/session-repeated-link-engagement.json +++ b/tests/valid/session-repeated-link-engagement.json @@ -1,5 +1,5 @@ { - "_test_description": "The same URL presented twice in one session: two content_presented events with distinct ids and distinct minted click tokens, and a content_engaged bound by presentation_id to the second occurrence. Matching on URL alone could not tell the two presentations apart (sections 6.8, 7.4).", + "_test_description": "The same URL presented twice in one session: two content_presented events with distinct ids and distinct minted click tokens, and a content_engaged bound by presentation_id to the second occurrence. Matching on URL alone could not tell the two presentations apart (sections 6.7, 7.4).", "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440210", "agent_id": "assistant-v2", diff --git a/tests/valid/session-reproduction-credited-quote.json b/tests/valid/session-reproduction-credited-quote.json deleted file mode 100644 index d9d21c6..0000000 --- a/tests/valid/session-reproduction-credited-quote.json +++ /dev/null @@ -1,115 +0,0 @@ -{ - "_test_description": "Credited quotation: the same span produces a content_reproduced event (the material) and a content_cited direct_quote event (the credit), sharing output_element_id, with reproduced_hash equal to the citation's excerpt_hash. The quote is then presented inline. Demonstrates that reproduction and citation are sibling output-construction claims joined to one presentation.", - "schema_version": "1.0", - "session_id": "660e8400-e29b-41d4-a716-446655440111", - "agent_id": "research-assistant-v5", - "started_at": "2026-08-01T15:00:00Z", - "ended_at": "2026-08-01T15:00:06Z", - "events": [ - { - "type": "turn_started", - "timestamp": "2026-08-01T15:00:00Z", - "turn_id": "1", - "turn": { - "privacy_level": "intent", - "query_intent": "research", - "topics": ["press freedom", "media law"] - } - }, - { - "type": "content_retrieved", - "timestamp": "2026-08-01T15:00:01Z", - "source_role": "agent", - "content_telemetry_id": "880e8400-e29b-41d4-a716-446655440112", - "content_url": "https://www.example-journal.org/analysis/press-freedom-ruling" - }, - { - "type": "content_grounded", - "timestamp": "2026-08-01T15:00:02Z", - "turn_id": "1", - "content_url": "https://www.example-journal.org/analysis/press-freedom-ruling", - "content_id": "examplejournal:2026:press-freedom-ruling", - "data": { - "scope": "turn", - "cached": false, - "chars_ingested": 9800, - "tokens_ingested": 2400, - "content_hash": "sha256:e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6", - "media_type": "text" - } - }, - { - "id": "880e8400-e29b-41d4-a716-446655440113", - "type": "content_cited", - "timestamp": "2026-08-01T15:00:04Z", - "turn_id": "1", - "output_id": "response:1", - "output_element_id": "answer:quote:1", - "content_url": "https://www.example-journal.org/analysis/press-freedom-ruling", - "content_id": "examplejournal:2026:press-freedom-ruling", - "data": { - "citation_type": "direct_quote", - "excerpt_tokens": 54, - "excerpt_chars": 236, - "excerpt_hash": "sha256:f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1", - "position": "primary", - "media_type": "text", - "url_verified": true - } - }, - { - "id": "880e8400-e29b-41d4-a716-446655440114", - "type": "content_reproduced", - "timestamp": "2026-08-01T15:00:04Z", - "turn_id": "1", - "output_id": "response:1", - "output_element_id": "answer:quote:1", - "citation_id": "880e8400-e29b-41d4-a716-446655440113", - "content_url": "https://www.example-journal.org/analysis/press-freedom-ruling", - "content_id": "examplejournal:2026:press-freedom-ruling", - "data": { - "reproduction_type": "verbatim", - "media_type": "text", - "reproduced_chars": 236, - "reproduced_tokens": 54, - "reproduced_hash": "sha256:f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1", - "content_hash": "sha256:e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6" - } - }, - { - "id": "880e8400-e29b-41d4-a716-446655440115", - "type": "content_presented", - "timestamp": "2026-08-01T15:00:05Z", - "turn_id": "1", - "output_id": "response:1", - "output_element_id": "answer:quote:1", - "citation_id": "880e8400-e29b-41d4-a716-446655440113", - "content_url": "https://www.example-journal.org/analysis/press-freedom-ruling", - "content_id": "examplejournal:2026:press-freedom-ruling", - "data": { - "presentation_kind": "content", - "presentation_type": "inline_quote", - "media_type": "text" - } - }, - { - "type": "turn_completed", - "timestamp": "2026-08-01T15:00:06Z", - "turn_id": "1", - "turn": { - "privacy_level": "intent", - "query_intent": "research", - "response_type": "analysis", - "response_mode": "standard", - "topics": ["press freedom", "media law"], - "content_urls_retrieved": [ - "https://www.example-journal.org/analysis/press-freedom-ruling" - ], - "content_urls_cited": [ - "https://www.example-journal.org/analysis/press-freedom-ruling" - ], - "response_tokens": 380 - } - } - ] -} diff --git a/tests/valid/session-reproduction-no-grounding.json b/tests/valid/session-reproduction-no-grounding.json deleted file mode 100644 index 3bc3a78..0000000 --- a/tests/valid/session-reproduction-no-grounding.json +++ /dev/null @@ -1,50 +0,0 @@ -{ - "_test_description": "Funnel exception: content_reproduced with no content_grounded event. The response quotes a passage the model memorised during training - the content never entered this session's generation context, so there is no retrieval and no grounding to report. Valid per section 4.3 (reproduced without grounded); telemetry consumers SHOULD treat it, like an uncorroborated citation, as a lower-confidence signal.", - "schema_version": "1.0", - "session_id": "660e8400-e29b-41d4-a716-446655440241", - "agent_id": "assistant-v4", - "started_at": "2026-08-05T15:00:00Z", - "ended_at": "2026-08-05T15:00:03Z", - "events": [ - { - "type": "turn_started", - "timestamp": "2026-08-05T15:00:00Z", - "turn_id": "1", - "turn": { - "privacy_level": "intent", - "query_intent": "question", - "topics": [ - "english literature", - "opening lines" - ] - } - }, - { - "id": "880e8400-e29b-41d4-a716-446655440242", - "type": "content_reproduced", - "timestamp": "2026-08-05T15:00:02Z", - "turn_id": "1", - "output_id": "response:1", - "output_element_id": "answer:quote:1", - "content_url": "https://www.gutenberg.org/ebooks/1342", - "content_id": "isbn:9780141439518", - "data": { - "reproduction_type": "verbatim", - "media_type": "text", - "reproduced_chars": 122, - "reproduced_hash": "sha256:a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2" - } - }, - { - "type": "turn_completed", - "timestamp": "2026-08-05T15:00:03Z", - "turn_id": "1", - "turn": { - "privacy_level": "intent", - "response_type": "explanation", - "response_mode": "standard", - "response_tokens": 160 - } - } - ] -} diff --git a/tests/valid/session-reproduction-uncredited.json b/tests/valid/session-reproduction-uncredited.json deleted file mode 100644 index b6be589..0000000 --- a/tests/valid/session-reproduction-uncredited.json +++ /dev/null @@ -1,71 +0,0 @@ -{ - "_test_description": "Uncredited reproduction in unpresented output: an API-delivered response contains a verbatim excerpt of grounded content with no citation and no presentation. The content_reproduced event is the only record that source material appears in the output. Valid per section 4.3 (reproduced without cited).", - "schema_version": "1.0", - "session_id": "660e8400-e29b-41d4-a716-446655440101", - "agent_id": "answers-api-v3", - "started_at": "2026-08-01T09:00:00Z", - "ended_at": "2026-08-01T09:00:04Z", - "events": [ - { - "type": "turn_started", - "timestamp": "2026-08-01T09:00:00Z", - "turn_id": "1", - "turn": { - "privacy_level": "minimal", - "query_tokens": 42 - } - }, - { - "type": "content_retrieved", - "timestamp": "2026-08-01T09:00:01Z", - "source_role": "agent", - "content_telemetry_id": "880e8400-e29b-41d4-a716-446655440102", - "content_url": "https://www.example-news.com/economy/rate-decision-analysis" - }, - { - "type": "content_grounded", - "timestamp": "2026-08-01T09:00:02Z", - "turn_id": "1", - "content_url": "https://www.example-news.com/economy/rate-decision-analysis", - "content_id": "examplenews:2026:rate-decision-analysis", - "data": { - "scope": "turn", - "cached": false, - "chars_ingested": 7200, - "tokens_ingested": 1800, - "content_hash": "sha256:c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4", - "media_type": "text" - } - }, - { - "id": "880e8400-e29b-41d4-a716-446655440103", - "type": "content_reproduced", - "timestamp": "2026-08-01T09:00:03Z", - "turn_id": "1", - "output_id": "response:1", - "output_element_id": "answer:para:2", - "content_url": "https://www.example-news.com/economy/rate-decision-analysis", - "content_id": "examplenews:2026:rate-decision-analysis", - "data": { - "reproduction_type": "verbatim", - "media_type": "text", - "reproduced_chars": 412, - "reproduced_tokens": 96, - "reproduced_hash": "sha256:d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5", - "content_hash": "sha256:c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4" - } - }, - { - "type": "turn_completed", - "timestamp": "2026-08-01T09:00:04Z", - "turn_id": "1", - "turn": { - "privacy_level": "minimal", - "response_tokens": 240, - "content_urls_retrieved": [ - "https://www.example-news.com/economy/rate-decision-analysis" - ] - } - } - ] -} diff --git a/tests/validate.py b/tests/validate.py index 8dcac1a..0c725fb 100644 --- a/tests/validate.py +++ b/tests/validate.py @@ -72,12 +72,12 @@ # # 5. Grounding provenance and fingerprint migration (sections 5.7.5, 6.4, 12.1): # agent_fetched requires cached false; agent_cached requires cached true. -# content_fingerprint MUST NOT carry preserved_in_output because output-side -# reuse is represented by content_reproduced. +# content_fingerprint MUST NOT carry preserved_in_output: v1 defines no +# output-side reuse reporting. # -# 6. Referential integrity within a session document (sections 6.6-6.8): +# 6. Referential integrity within a session document (sections 6.6-6.7): # content_engaged.presentation_id references the exact content_presented -# event id, and citation_id on content_presented/content_reproduced +# event id, and citation_id on content_presented # references a content_cited event id. JSON Schema cannot compare values # across events. Session documents only: standalone envelopes and batch # members may reference events delivered elsewhere. @@ -85,11 +85,11 @@ # 7. Field placement and source_role (sections 5.2, 5.2.2, 5.7.1, 5.7.5): # source_role is required on content_retrieved; presentation_id and the # event-level ctx_token appear only on content_engaged, citation_id only on -# content_presented/content_reproduced, turn only on turn events. +# content_presented, turn only on turn events. # -# 8. Session-document integrity beyond rule 6 (sections 6.7, 6.8, 7.4.1): +# 8. Session-document integrity beyond rule 6 (sections 6.6, 6.7, 7.4.1): # event ids are distinct; the engagement and the presentation it references, -# and the presentation/reproduction and the citation it references, identify +# and the presentation and the citation it references, identify # the same content; one event-level ctx_token binds to one presentation. # # 9. Envelope ctx_token (section 7.1): an envelope carrying ctx_token carries @@ -160,7 +160,7 @@ ), "engaged-presentation-id-unmatched.json": ( "content_engaged.presentation_id matches no content_presented event id " - "in the session. Violates section 6.8: every engagement references the " + "in the session. Violates section 6.7: every engagement references the " "exact content_presented.id on which the action occurred." ), "presented-citation-id-unmatched.json": ( @@ -192,7 +192,7 @@ ), "grounding-fingerprint-preserved-in-output.json": ( "Grounding fingerprint carries preserved_in_output. " - "Violates section 6.4: output-side reuse is reported as content_reproduced." + "Violates section 6.4: v1 defines no output-side reuse reporting and the field is withdrawn (section 12.1)." ), "grounding-provenance-cached-conflict.json": ( "Grounding event declares agent_fetched with cached true. " @@ -240,7 +240,7 @@ ), "citation-id-on-grounded.json": ( "citation_id on a content_grounded event. " - "Violates section 5.2: citation_id is valid only on content_presented and content_reproduced." + "Violates section 5.2: citation_id is valid only on content_presented." ), "presentation-id-on-cited.json": ( "presentation_id on a content_cited event. " @@ -252,23 +252,19 @@ ), "duplicate-event-id.json": ( "Two content_presented events share one id. " - "Violates section 6.7: repeated presentations receive distinct event IDs." + "Violates section 6.6: repeated presentations receive distinct event IDs." ), "engaged-presentation-content-mismatch.json": ( "content_engaged references a content_presented event of different content. " - "Violates section 6.8: the engagement identifies the same content as the presentation it acted on." + "Violates section 6.7: the engagement identifies the same content as the presentation it acted on." ), "presented-citation-content-mismatch.json": ( "content_presented.citation_id references a content_cited event of different content. " - "Violates section 6.7: the presentation and the citation it realises identify the same content." - ), - "reproduced-citation-id-unmatched.json": ( - "content_reproduced.citation_id matches no content_cited event id in the session. " - "Violates section 6.6: citation_id references the crediting content_cited event's id." + "Violates section 6.6: the presentation and the citation it realises identify the same content." ), "engaged-presentation-id-no-presentations.json": ( "content_engaged carries a presentation_id in a session with no content_presented events. " - "Violates section 6.8: every engagement references an exact presentation occurrence." + "Violates section 6.7: every engagement references an exact presentation occurrence." ), "shared-ctx-token-two-presentations.json": ( "One event-level ctx_token appears on engagements bound to two different presentations. " @@ -312,7 +308,7 @@ # (content_url or content_id) under section 5.7.5. turn_started and # turn_completed are turn events, not content events, and are exempt. CONTENT_EVENT_TYPES = { - "content_retrieved", "content_grounded", "content_reproduced", + "content_retrieved", "content_grounded", "content_cited", "content_presented", "content_engaged", } @@ -531,9 +527,9 @@ def check_referential_integrity(data): Check the intra-document event references of a session document: - Every content_engaged.presentation_id MUST reference the exact - content_presented.id on which the action occurred (section 6.8). - - Every citation_id on a content_presented or content_reproduced event - references that content_cited event's id (sections 6.6, 6.7). + content_presented.id on which the action occurred (section 6.7). + - Every citation_id on a content_presented event + references that content_cited event's id (section 6.6). Applies only to session documents, where the referenced events live in the same document. Standalone envelopes and batch members legitimately @@ -562,7 +558,7 @@ def check_referential_integrity(data): f"content_engaged presentation_id '{pid}' does not match " "any content_presented event id in the session" ) - if etype in ("content_presented", "content_reproduced"): + if etype == "content_presented": cid = event.get("citation_id") if cid and cid not in cited_ids: violations.append( @@ -608,7 +604,7 @@ def check_grounding_provenance(data): fingerprint = event_data.get("content_fingerprint") if isinstance(fingerprint, dict) and "preserved_in_output" in fingerprint: violations.append( - "content_fingerprint carries preserved_in_output; use content_reproduced" + "content_fingerprint carries preserved_in_output; withdrawn in v1" ) return violations @@ -618,7 +614,7 @@ def check_grounding_provenance(data): FIELD_PLACEMENT = { "presentation_id": {"content_engaged"}, "ctx_token": {"content_engaged"}, - "citation_id": {"content_presented", "content_reproduced"}, + "citation_id": {"content_presented"}, "turn": {"turn_started", "turn_completed"}, } @@ -653,8 +649,8 @@ def _same_content(a, b): def check_session_integrity(data): """Session-document rules beyond check_referential_integrity (sections - 6.7, 6.8, 7.4.1): distinct event ids; engagement/presentation and - presentation|reproduction/citation pairs identify the same content; one + 6.6, 6.7, 7.4.1): distinct event ids; engagement/presentation and + presentation/citation pairs identify the same content; one event-level ctx_token binds to one presentation. Session documents only.""" if is_standalone_event(data) or is_event_batch(data): return [] @@ -685,7 +681,7 @@ def check_session_integrity(data): violations.append( f"ctx_token '{token}' appears on engagements bound to two presentations" ) - if etype in ("content_presented", "content_reproduced"): + if etype == "content_presented": cid = e.get("citation_id") if cid in cited and not _same_content(e, cited[cid]): violations.append( From 3c1e024ddc71479ac998d6b5bc334e6a39b18777 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Wed, 2 Sep 2026 14:30:42 +0100 Subject: [PATCH 35/38] Close the consultation gaps recorded on #12, #17, #31 and #46 - Discovery/outcome-layer clarification in 1.4 and the consumer vs index-emitter role separation in 7.3 (#12) - Optional manifest identifier_schemes block and co-primary URL / content_id owner resolution with defined conflict and failure behaviour (#17, #31), with fixtures and placement checks - The manifest identifies a domain, not a legal person; party-identity declarations recorded as deferred (#46) Co-Authored-By: Claude Fable 5 --- SPECIFICATION.md | 28 +++++++++++++++++-- manifest.json | 21 ++++++++++++++ ...nifest-identifier-scheme-empty-prefix.json | 17 +++++++++++ ...ifest-identifier-schemes-on-non-owner.json | 20 +++++++++++++ tests/mutation_smoke.py | 5 ++++ tests/valid/manifest-identifier-schemes.json | 26 +++++++++++++++++ tests/validate.py | 20 +++++++++++-- 7 files changed, 131 insertions(+), 6 deletions(-) create mode 100644 tests/invalid/manifest-identifier-scheme-empty-prefix.json create mode 100644 tests/invalid/manifest-identifier-schemes-on-non-owner.json create mode 100644 tests/valid/manifest-identifier-schemes.json diff --git a/SPECIFICATION.md b/SPECIFICATION.md index a761431..0e2609f 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -74,6 +74,8 @@ An agent cannot reliably declare how it will use content before reading it - a r Events can reference a licence via the `license_ref` field (section 5.2), connecting telemetry to whatever access protocol issued the licence. The telemetry schema does not depend on any specific access protocol. +Discovery protocols and registries - catalogues that describe where content sources are and what they offer, such as agent resource discovery formats - sit upstream of both layers. A catalogue records where a content source is; telemetry records what happened when an agent used it. Content Telemetry is the outcome layer for discovery in the same sense that it is the reporting counterpart to access: it carries no ranking or discovery metadata of its own, and a discovery service that wants outcome signal consumes telemetry like any other party (section 7.3). + ### 1.5 Conventions The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in [RFC 2119](https://www.rfc-editor.org/rfc/rfc2119) and [RFC 8174](https://www.rfc-editor.org/rfc/rfc8174). @@ -1003,9 +1005,11 @@ Two deployment patterns are common: Any party may operate a consumer: an agent operator, a licensing intermediary, or an independent third party offering it as a service. Both patterns above consume the same session format. The telemetry consumer is responsible for domain resolution, content owner filtering, and access control. The spec does not mandate a specific aggregation topology, nor does it require any particular operator to provide one. +The telemetry-consumer function and the `index` emitter role (section 4.4) are distinct. A discovery registry or ranking service that consumes telemetry to build ranking signal is acting as a telemetry consumer; an operator that also brokers or serves content additionally acts as an `index` emitter. This specification does not restrict combining the two, but where one operator holds both, the telemetry it consumes in the ranking capacity carries the same filtering and access-control responsibilities as any consumer's - operating an index confers no additional visibility into other parties' events. + **Origin-side `.well-known/content-telemetry.json` manifests** declare where origin-emitted retrieval events are sent (CDN → content owner's chosen endpoint). They do not instruct agents where to send session documents. Agent routing is governed by the agent's telemetry configuration, not by content owner manifests. -**Content owner resolution.** Telemetry consumers resolve content owner identity from `content_url` domains. Content owners register and verify their domains with the telemetry consumer; the consumer maps incoming event URLs to the owning organisation. This is the primary resolution path and requires `content_url` to be present on events. Events identified only by `content_id` (e.g., cached groundings where the URL was not preserved, or marketplace API content with no canonical URL) cannot be resolved by domain alone. Telemetry consumers SHOULD support `content_id` prefix-based resolution as a secondary path when content owners register their identifier schemes, but this is not yet a normative requirement. +**Content owner resolution.** Telemetry consumers resolve content owner identity through two co-primary paths: the `content_url` domain, via verified domain registrations, and the registered `content_id` prefix, via identifier-scheme declarations (section 8.6). Content owners register domains and identifier prefixes with the telemetry consumer; the consumer maps each incoming event to the owning organisation by whichever identifier the event carries. An event carrying only one of the two resolves through that path - a cached grounding with no preserved URL, or marketplace API content with no canonical URL, resolves by `content_id` prefix alone. When an event carries both and the two paths resolve to different owners, the consumer MUST NOT deliver the event to either owner's filtered view until its trust policy resolves the conflict, and SHOULD surface the conflict to both registrants. When neither path resolves, the event is unattributed; consumers SHOULD retain unattributed events rather than discard them, so a later registration can claim them. **Cross-consumer correlation.** Origin-side emitters and agent-side emitters MAY use different telemetry consumers. A content owner's CDN sends retrieval events to one telemetry consumer; an agent sends sessions to another. The `content_telemetry_id` field (section 7.2) correlates the same retrieval across consumers - both sides share the same UUID from the HTTP request. This correlation operates at the retrieval level only. Grounding, citation, presentation, and engagement events have no independent origin-side counterpart to correlate against. @@ -1091,6 +1095,8 @@ Token lifetime and requester authentication are likewise not defined in v1: a to Content owners, agents, and platforms publish a manifest declaring their identity and telemetry endpoints. The `manifest_ref` field on session documents (5.1.2) and the routing logic for origin-side emitters (7.3) resolve to manifests defined in this section. +The identity a manifest establishes is control of a domain, or a path under one - not a legal person. `operator.name` is a display name; nothing in a manifest binds the domain to an organisation, and no field carries a jurisdiction or registered legal entity. Where settlement or audit requires a legal counterparty, that binding lives in the governing terms the parties hold (`terms_ref`, section 5.2.4), not in the manifest. A party-identity declaration is deferred (section 8.9). + ### 8.1 Discovery Manifests are served as JSON at: @@ -1123,6 +1129,7 @@ Machine-readable schema: [`./manifest.json`](./manifest.json) (JSON Schema draft | `keys` | object[] | No | Public keys for signing telemetry events (see 8.4). | | `telemetry` | object | No | Telemetry endpoint declaration (see 8.5). | | `domains` | string[] | No | Domains the participant claims authority over (see 8.6). MAY appear only on root manifests. | +| `identifier_schemes` | object[] | No | Identifier prefixes the participant claims and, optionally, where they resolve (see 8.6). MAY appear only on `content_owner` manifests. | Consumers MUST tolerate unknown fields and treat absent optional sections as "not declared" rather than rejecting the manifest. @@ -1159,7 +1166,7 @@ Public keys used to sign telemetry events emitted by this participant. Per-event `conformance_level` is informational. It advertises the level of telemetry the manifest's participant emits. It does **not** constrain what an inbound `endpoint` accepts - an endpoint accepts whatever events it is configured to accept, regardless of any level declared here - and it places **no requirement** on other emitters. On a `content_owner` manifest it describes only the events the owner's own infrastructure emits (typically a CDN edge worker at `retrieval`); it says nothing about what agents or platforms report about the owner's content, which those parties advertise in their own manifests. A `content_owner` manifest SHOULD omit `conformance_level` unless the owner operates its own emitter. There is no field for a content owner to *request* a minimum level from agents; consumers tolerate events from any level (see 5.7), and the protocol does not give a manifest a way to demand more. -### 8.6 Domains +### 8.6 Domains and identifier schemes The `domains` array MAY appear only on manifests served from the domain root (`https:///.well-known/content-telemetry.json`). Manifests under path prefixes MUST NOT include `domains`. @@ -1167,6 +1174,17 @@ In v1, every entry in `domains` MUST be self-validating: either the manifest's o This keeps the v1 protocol fully decentralised: every manifest is a self-contained credential, validated by TLS plus the well-known location, with no dependency on consumer-side validation state or any external registry. Cross-apex claims (one operator unifying several unrelated apex domains in a single manifest) are deferred to a later version. +#### Identifier schemes + +The `identifier_schemes` array declares the `content_id` prefixes a content owner claims, so that events identified only by `content_id` can be routed to their owner (section 7.3). It MAY appear only on manifests declaring the `content_owner` role. Each entry contains: + +| Field | Type | Required | Description | +|-------|------|----------|-------------| +| `prefix` | string | Yes | The identifier prefix claimed, matched against `content_id` values up to and including the prefix (e.g. `ft:`, `iscc:`, `mkt:gridnews:`). | +| `resolution` | string | No | HTTPS URL of an endpoint that maps a `content_id` under this prefix to the owning content record or organisation. | + +Unlike `domains`, a prefix claim is not self-validating: nothing about a manifest's host proves authority over an identifier namespace. The manifest is the carrier of the claim, verified to domain level by TLS and the well-known location; a telemetry consumer verifies prefix ownership at registration, as it verifies domain registrations today, and resolves conflicting claims on the same prefix through its trust policy. A `resolution` endpoint supports owner-identity mapping; it does not verify that grounding or citation happened. + ### 8.7 Consumer behaviour When resolving a manifest from `manifest_ref`, a `content_url` domain, or any other reference: @@ -1193,7 +1211,10 @@ Consumers SHOULD cache resolved manifests respecting the response's `Cache-Contr "telemetry": { "endpoint": "https://telemetry.example.com/v1/events" }, - "domains": ["example.com", "*.example.com"] + "domains": ["example.com", "*.example.com"], + "identifier_schemes": [ + { "prefix": "exm:", "resolution": "https://id.example.com/resolve" } + ] } ``` @@ -1255,6 +1276,7 @@ The two manifests live independently at distinct well-known URLs. The content-ow The following are deferred to later versions: - Content licence declarations (what content the participant is licensed to access) +- Party-identity declarations binding a domain to a legal entity or jurisdiction (the manifest identifies a domain; section 8 opening) - Manifest signing (W3C Verifiable Credentials, JWS proofs) - Training data and model provenance - Deployment context, purpose, brand affiliation diff --git a/manifest.json b/manifest.json index 9ed6b7c..f418b8e 100644 --- a/manifest.json +++ b/manifest.json @@ -119,6 +119,27 @@ "type": "string" }, "description": "Domains the participant claims authority over. MAY appear only on root manifests (enforced at the application layer); each entry MUST be the manifest's own host or a subdomain of it (literal or wildcard). (Section 8.6)" + }, + "identifier_schemes": { + "type": "array", + "items": { + "type": "object", + "required": ["prefix"], + "properties": { + "prefix": { + "type": "string", + "minLength": 1, + "description": "The content_id prefix claimed, matched against content_id values up to and including the prefix, e.g. ft: or mkt:gridnews:. (Section 8.6)" + }, + "resolution": { + "type": "string", + "format": "uri", + "pattern": "^https://", + "description": "HTTPS URL of an endpoint mapping a content_id under this prefix to the owning content record or organisation. (Section 8.6)" + } + } + }, + "description": "Identifier prefixes the participant claims for content_id routing. MAY appear only on content_owner manifests (enforced at the application layer); prefix ownership is verified by the consumer at registration. (Sections 7.3, 8.6)" } } } diff --git a/tests/invalid/manifest-identifier-scheme-empty-prefix.json b/tests/invalid/manifest-identifier-scheme-empty-prefix.json new file mode 100644 index 0000000..4488d17 --- /dev/null +++ b/tests/invalid/manifest-identifier-scheme-empty-prefix.json @@ -0,0 +1,17 @@ +{ + "_test_description": "Identifier-scheme entry with an empty prefix. prefix is required and non-empty: an empty string would match every content_id (section 8.6).", + "_expected_error": "prefix", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "content_owner" + ], + "operator": { + "name": "Example Media" + }, + "identifier_schemes": [ + { + "prefix": "" + } + ] +} diff --git a/tests/invalid/manifest-identifier-schemes-on-non-owner.json b/tests/invalid/manifest-identifier-schemes-on-non-owner.json new file mode 100644 index 0000000..20f759c --- /dev/null +++ b/tests/invalid/manifest-identifier-schemes-on-non-owner.json @@ -0,0 +1,20 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: agent manifest carrying identifier_schemes. The block MAY appear only on manifests declaring the content_owner role (section 8.6).", + "_expected_error": "identifier_schemes only on content_owner manifests", + "schema_version": "1.0", + "id": "https://searchco.com/agents/web-search/.well-known/content-telemetry.json", + "roles": [ + "agent" + ], + "operator": { + "name": "SearchCo" + }, + "telemetry": { + "endpoint": "https://telemetry.searchco.com/v1/events" + }, + "identifier_schemes": [ + { + "prefix": "sc:" + } + ] +} diff --git a/tests/mutation_smoke.py b/tests/mutation_smoke.py index ebc41ab..ec6ba19 100644 --- a/tests/mutation_smoke.py +++ b/tests/mutation_smoke.py @@ -215,6 +215,11 @@ def mutate(root): _json(_S, lambda d: _event_branch(d, "content_cited")["required"].remove("data"), "cited-missing-data.json")), ("remove the sha256 hash patterns", _text(_S, '"pattern": "^sha256:[a-f0-9]{64}$"', '"type": "string"', "malformed-content-hash.json")), + ("remove minLength on identifier-scheme prefix", + _json(_M, lambda d: d["properties"]["identifier_schemes"]["items"]["properties"]["prefix"].pop("minLength"), "manifest-identifier-scheme-empty-prefix.json")), + ("drop the identifier_schemes owner-role placement check", + _text(_V, 'if data.get("identifier_schemes") and "content_owner" not in roles:', 'if False:', + "manifest-identifier-schemes-on-non-owner.json")), ("remove minLength on output_id", _json(_S, lambda d: d["$defs"]["TelemetryEvent"]["properties"]["output_id"].pop("minLength"), "empty-output-id.json")), ("remove minimum on tokens_ingested", diff --git a/tests/valid/manifest-identifier-schemes.json b/tests/valid/manifest-identifier-schemes.json new file mode 100644 index 0000000..b919b3d --- /dev/null +++ b/tests/valid/manifest-identifier-schemes.json @@ -0,0 +1,26 @@ +{ + "_test_description": "Content owner root manifest declaring two identifier-scheme prefixes, one with a resolution endpoint and one without. Supports content_id prefix routing as a co-primary resolution path (sections 7.3, 8.6).", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "content_owner" + ], + "operator": { + "name": "Example Media" + }, + "telemetry": { + "endpoint": "https://telemetry.example.com/v1/events" + }, + "domains": [ + "example.com" + ], + "identifier_schemes": [ + { + "prefix": "exm:", + "resolution": "https://id.example.com/resolve" + }, + { + "prefix": "mkt:example:" + } + ] +} diff --git a/tests/validate.py b/tests/validate.py index 0c725fb..ae3448a 100644 --- a/tests/validate.py +++ b/tests/validate.py @@ -96,7 +96,8 @@ # content_engaged events only. # # 10. Manifest placement rules (sections 8.5, 8.6): domains only on a manifest -# served at the domain root; ctx_resolution only on agent/platform manifests. +# served at the domain root; ctx_resolution only on agent/platform manifests; +# identifier_schemes only on content_owner manifests. # # Not checked here: agent_id at Grounding/Citation conformance (section # 5.7) depends on the emitter's declared conformance level, which fixtures do @@ -290,6 +291,10 @@ "content_owner manifest declares telemetry.ctx_resolution. " "Violates section 8.5: ctx_resolution is valid on agent and platform manifests." ), + "manifest-identifier-schemes-on-non-owner.json": ( + "Agent manifest carries identifier_schemes. " + "Violates section 8.6: identifier_schemes only on content_owner manifests." + ), "manifest-lookalike-domain.json": ( "Manifest at example.com claims evilexample.com in domains. " "Violates section 8.6: a lookalike host is not a subdomain of the manifest host." @@ -727,8 +732,10 @@ def check_manifest_application_layer(data): Check the manifest rejection rules that JSON Schema cannot express: duplicate keys[].id values (section 8.7), domains entries that are not the manifest's own host or a subdomain of it, domains on a path-prefixed - manifest (section 8.6), and ctx_resolution on a manifest without the agent - or platform role (section 8.5). Returns a list of violation descriptions. + manifest (section 8.6), ctx_resolution on a manifest without the agent + or platform role (section 8.5), and identifier_schemes on a manifest + without the content_owner role (section 8.6). Returns a list of violation + descriptions. """ violations = [] @@ -766,6 +773,13 @@ def check_manifest_application_layer(data): "valid on agent and platform manifests" ) + # identifier_schemes only on content_owner manifests (section 8.6) + if data.get("identifier_schemes") and "content_owner" not in roles: + violations.append( + f"identifier_schemes on a manifest with roles {sorted(roles)}; " + "identifier_schemes only on content_owner manifests" + ) + return violations From a6d78b398012f587c4c64adc63da5f17fbe778c6 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Wed, 2 Sep 2026 14:33:40 +0100 Subject: [PATCH 36/38] Publish v1.0: release status, translated-citation guidance, portion caveat - Version 1.0 / current-specification framing in the spec header, README and GOVERNANCE; the consultation section becomes a past-tense record and the checkbox list goes away - GOVERNANCE gains the decision process owed before 1.0 and loses the broken request-for-comment anchor - Cross-language citation guidance in 6.5 (#24 fallback) and the bounded-portion caveat restated in 6.6 (#25) Co-Authored-By: Claude Fable 5 --- GOVERNANCE.md | 18 +++++++------ README.md | 66 +++++++++++++++--------------------------------- SPECIFICATION.md | 10 +++++--- 3 files changed, 38 insertions(+), 56 deletions(-) diff --git a/GOVERNANCE.md b/GOVERNANCE.md index 6e017a1..36cecca 100644 --- a/GOVERNANCE.md +++ b/GOVERNANCE.md @@ -2,7 +2,7 @@ ## Status -Content Telemetry is an open specification stewarded by the SPUR Coalition. It is in preview (v0.x); see [SPECIFICATION.md](./SPECIFICATION.md) section 12 for the versioning policy. +Content Telemetry is an open specification stewarded by the SPUR Coalition. Version 1.0 is the current specification; see [SPECIFICATION.md](./SPECIFICATION.md) section 12 for the versioning policy. ## Stewardship @@ -18,20 +18,22 @@ The SPUR Coalition stewards the specification and holds this repository. The sta The SPUR Coalition is a group of publishers and content owners that maintains the Content Telemetry standard. It holds the intellectual property through the preview period and releases the standard under Apache 2.0 from 12 June 2026. -Contributing to the standard does not require membership. The wire format is developed in the open, and anyone - content owner, agent operator, intermediary, or implementer - can take part through the issue tracker and the comment process described in the [README](./README.md#request-for-comment). +Contributing to the standard does not require membership. The wire format is developed in the open, and anyone - content owner, agent operator, intermediary, or implementer - can take part through the issue tracker and the process described in the [README](./README.md#consultation-record). -The standard is maintained by Alex Springer (alex@spurcoalition.org). The decision-making process will be set out before 1.0. +The standard is maintained by Alex Springer (alex@spurcoalition.org). -## Path to 1.0 +## Version 1.0 and decisions -The specification is at v0.1 (preview). It moves to 1.0 once the open questions are resolved, the conformance suite is stable, and there are independent interoperable implementations from more than one party. There is no fixed date. +The specification reached 1.0 on 2 September 2026, following the public consultation of 12 June to 24 July 2026 and the release-candidate work recorded on the issue tracker. + +Decisions follow the proposal and alignment process in [CONTRIBUTING.md](./CONTRIBUTING.md): anyone may propose a change, a maintainer records the disposition publicly on the issue tracker, and the SPUR Steering Board approves releases. The tracker and pull-request history are the public decision record. Minor versions add optional fields and event types; breaking changes require a major version (SPECIFICATION.md section 12). ## How to participate - File questions and bugs on the [issue tracker](https://github.com/SPUR-Coalition/telemetry/issues) (see the templates). -- Follow the [post-consultation status](./README.md#consultation-status). The - formal v1 comment window is closed, but concrete bugs and implementation - evidence remain welcome on the issue tracker. +- See the [consultation record](./README.md#consultation-record). The formal + v1 comment window is closed, but concrete bugs and implementation evidence + remain welcome on the issue tracker. - Propose new capabilities or changes in behaviour as a short human-written note in [`proposals/`](./proposals/), following [CONTRIBUTING.md](./CONTRIBUTING.md). diff --git a/README.md b/README.md index 7b42705..7673190 100644 --- a/README.md +++ b/README.md @@ -2,16 +2,7 @@ **Signal format for AI content usage reporting.** -This is a preview specification. Field names, event types, and schema structure may change before 1.0. - -> **Consultation status — 12 August 2026:** The public comment period closed on -> **24 July 2026**. The SPUR Steering Board has approved the v1 direction, and a -> final disposition is now recorded on every consultation thread: each carries -> an outcome label, and accepted core changes sit on the -> [v1 release candidate milestone](https://github.com/SPUR-Coalition/telemetry/milestone/1) -> with a schema freeze targeted for **21 August 2026**. Accepted changes are -> merging on the `v1-draft` integration line. Version 0.1 remains the current -> published preview. See [Consultation status](#consultation-status) below. +**Version 1.0** is the current specification, published 2 September 2026. It replaces the v0.1 preview; the changes and migration steps are recorded in [SPECIFICATION.md section 12.1](./SPECIFICATION.md#121-migration-from-the-v01-preview). ## Contents @@ -21,8 +12,8 @@ This is a preview specification. Field names, event types, and schema structure - [Repo contents](#repo-contents) - [Example](#example) - [Relationship to other protocols](#relationship-to-other-protocols) -- [Consultation status](#consultation-status) -- [Open questions in v0.1](#open-questions-in-v01) +- [Consultation record](#consultation-record) +- [Open questions in v1](#open-questions-in-v1) - [Versioning](#versioning) ## Problem @@ -72,7 +63,7 @@ Grounding and presentation record different boundary crossings: grounding means - [telemetry-event-batch.json](./telemetry-event-batch.json) - JSON Schema for event batch envelopes - [manifest.json](./manifest.json) - JSON Schema for the `.well-known/content-telemetry.json` manifest ([section 8](./SPECIFICATION.md#8-manifest)) - [tests/](./tests/) - conformance test suite -- [GOVERNANCE.md](./GOVERNANCE.md) - stewardship, preview status, relationship to profiles +- [GOVERNANCE.md](./GOVERNANCE.md) - stewardship, versioning status, relationship to profiles - [LICENSE](./LICENSE) - Apache License 2.0 This repository is the **standard** - the wire format. Publisher-facing accreditation and the SPUR conformance mark are defined separately in the [SPUR Content Telemetry Profile](https://github.com/SPUR-Coalition/telemetry-profile), which references this specification by version. The standard defines the privacy mechanism (section 5.5); whether a profile makes any privacy level binding is the profile's choice. See [GOVERNANCE.md](./GOVERNANCE.md). @@ -161,52 +152,37 @@ The content owner can derive: FT article `abc123` was in context for the respons Content Telemetry is focussed on **reporting**, while content **access** protocols (Really Simple Licensing, peek-then-pay, IAB CoMP, bilateral APIs) aim to govern how agents discover and license content. The `license_ref` field on events connects telemetry to whatever access protocol issued the licence, but the schemas are independent - telemetry works with any access protocol, or none. -## Consultation status +## Consultation record -The public comment period ran from **12 June to 24 July 2026** and is now closed. -Thank you to everyone who opened an issue, submitted a pull request, joined a -working session or supplied implementation evidence. +The public comment period ran from **12 June to 24 July 2026**. Thank you to +everyone who opened an issue, submitted a pull request, joined a working +session or supplied implementation evidence. The consultation produced 29 specification issue threads, three profile issue -threads and five pull requests. The maintainers are now: - -- [x] reviewing the full consultation record; -- [x] preparing a proposed disposition for every thread; -- [x] recording the approved dispositions on the issue tracker; -- [ ] completing focused v1 changes and migration fixtures on `v1-draft`; -- [ ] publishing a v1 release candidate for implementer testing; and -- [ ] publishing final v1 only after publisher, intermediary and agent/platform - acceptance cases pass. - -Consultation issues remain open while their dispositions are recorded. An open -issue does not mean its proposal has been accepted, and a preparation branch does -not change the published specification. The issue tracker and pull-request -history will remain the public decision record. - -Concrete schema, fixture and documentation bugs may still be filed using the +threads and five pull requests. Every thread carries a recorded outcome, and +the issue tracker and pull-request history remain the public decision record. +The accepted changes were merged on the `v1-draft` integration line and +published as version 1.0 on 2 September 2026. + +Concrete schema, fixture and documentation bugs may be filed using the *Schema or example bug* template, and questions or unclear requirements using *Spec feedback / open question*. For a new capability or change in behaviour, submit a short human-written note to [`proposals/`](./proposals/) and wait for explicit maintainer alignment before beginning implementation (see -[CONTRIBUTING.md](./CONTRIBUTING.md)); new design proposals are not -automatically part of the v1 consultation scope. Pull requests remain welcome -for specific fixes. Feedback on accreditation or the conformance mark belongs -on the [profile +[CONTRIBUTING.md](./CONTRIBUTING.md)). Pull requests remain welcome for +specific fixes. Feedback on accreditation or the conformance mark belongs on +the [profile repository](https://github.com/SPUR-Coalition/telemetry-profile/issues). -Required fields, event types and schema structure may still change before 1.0 -(section 12). Version 0.1 remains the current published preview until a later -version is released. - -## Open questions in v0.1 +## Open questions in v1 -This is a preview specification. The following areas are under active discussion and will be refined with implementer input: +The following areas are expected to develop in 1.x minor versions and profiles, with implementer input: **Grounding boundary.** The spec defines grounding as content entering the generation model's context (sections 4.3 and 6.4). For straightforward RAG pipelines this is clear. For pipelines with multiple processing stages - embedding, re-ranking, summarisation before context insertion - the boundary requires judgement. The spec draws the line at the generation context (not earlier retrieval stages), but edge cases remain. When a re-ranking or summarisation stage is itself a generative model, the multi-step rule in section 6.4 (content entering a sub-agent's generation context is grounded) can pull selection stages back inside the boundary. Input from platform engineering teams building real implementations will sharpen this definition. **Event volume at scale.** A single deep-research query can produce 100+ retrieval events and dozens of grounding/citation events. The session document format already handles transport - one POST with all events after the session ends, not one request per event. Volume management beyond that (storage, processing, consumer-side aggregation) is an implementation concern, not a protocol gap. Version 1 adds an explicit coverage declaration - `complete`, `sampled`, `aggregated` or `selected` (section 5.7.6) - and a manifest field for it (section 8.5); the standard still sets no default for reporting granularity, leaving it to profiles and deployments. -**Verification of grounding and citation.** Grounding and citation events are reported by the agent, which is also the party that may owe compensation under a licence. In v0.1, manifest signing is informational: consumers may verify signatures but are not required to, and the specification defines no required proof binding an event to its emitter (sections 8.4 and 8.9). The events attribution depends on are therefore self-reported by the reporting party. Verifiable credentials and signed events are deferred (section 8.9). One corroboration mechanism works without signing: the `Content-Telemetry-ID` field correlates an agent-reported retrieval with an origin- or edge-reported one (section 7.2), but it covers retrieval only - grounding, citation, presentation, and engagement have no independent observer. Signing, even once required, would prove who reported an event, not that the event is true or that all qualifying events were reported. Input is wanted on what a verification layer should cover and where it belongs. Mechanisms that test truthfulness and completeness rather than origin, such as sampled audits or publisher-seeded canary content, are of particular interest. +**Verification of grounding and citation.** Grounding and citation events are reported by the agent, which is also the party that may owe compensation under a licence. In v1, manifest signing is informational: consumers may verify signatures but are not required to, and the specification defines no required proof binding an event to its emitter (sections 8.4 and 8.9). The events attribution depends on are therefore self-reported by the reporting party. Verifiable credentials and signed events are deferred (section 8.9). One corroboration mechanism works without signing: the `Content-Telemetry-ID` field correlates an agent-reported retrieval with an origin- or edge-reported one (section 7.2), but it covers retrieval only - grounding, citation, presentation, and engagement have no independent observer. Signing, even once required, would prove who reported an event, not that the event is true or that all qualifying events were reported. Input is wanted on what a verification layer should cover and where it belongs. Mechanisms that test truthfulness and completeness rather than origin, such as sampled audits or publisher-seeded canary content, are of particular interest. **Reporting granularity.** The standard sets no default for reporting granularity, leaving it to profiles and deployments (see *Event volume* above). The SPUR profile requires event-level delivery and does not permit aggregation. Version 1 answers the first half of the question: coverage modes are defined once, in section 5.7.6, so that profiles reference them rather than each define their own. How event-level delivery scales for the highest-volume case remains open. @@ -214,4 +190,4 @@ This is a preview specification. The following areas are under active discussion This repo tracks the specification version. SDK repos have their own release cadences and declare which spec version they support. -Current spec version: **0.1** (preview) +Current spec version: **1.0** diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 0e2609f..ffe4131 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -1,8 +1,8 @@ # Content Telemetry Specification -**Version:** 1.0 (release candidate draft) -**Status:** Release candidate in preparation (feature freeze 21 August 2026) -**Last updated:** 2026-08-12 +**Version:** 1.0 +**Status:** Current specification +**Published:** 2026-09-02 ## Contents @@ -819,6 +819,8 @@ The `contradiction` type supports negative attribution: content that was retriev The `unclassified` value for `citation_type` indicates the agent did not classify this citation. The `unclassified` value for `position` indicates the agent did not determine the prominence of the citation. Emitters SHOULD use `unclassified` rather than forcing a classification when the agent cannot confidently determine the citation type or position. +A citation of translated or cross-language content is a `paraphrase` (or `reference`) citation like any other: v1 defines no language fields and no match-confidence claim. `excerpt_hash` identifies the excerpt the agent produced, not the source passage, so a hash that does not match the source is expected for translated and paraphrased citations and does not indicate non-use; establishing which source passage a translated excerpt derives from is verification-layer work outside this specification. + `url_verified` indicates whether the agent confirmed that the cited URL resolves to content matching the citation. When `false` or absent, the citation may reference a hallucinated or outdated URL. `url_verified` MAY be set asynchronously after response generation. Platforms that batch-verify URLs periodically rather than per-request are conforming. A value of `false` indicates the URL was not verified, not that verification failed. When `content_hash` is absent or does not match any grounding event's hash (for example, because the agent re-chunked content between grounding and citation), consumers SHOULD fall back to matching on `content_url` or `content_id`, accepting that the correlation may be imprecise when the same content appears in multiple grounding events. @@ -849,6 +851,8 @@ When `content_hash` is absent or does not match any grounding event's hash (for These are the core values. Platforms with additional presentation surfaces MAY use custom string values. Telemetry consumers MUST tolerate unknown `presentation_type` values. +A presentation identifies whole content items. Finer-grained portion references for time-based and spatial media - time ranges, regions, segments - are not defined in this version. + Each presentation event MUST have an `id` and `output_id`. When it presents a citation, `citation_id` references that `content_cited` event's `id`, and the two events identify the same content; an uncited presentation omits `citation_id`. Repeated presentations of the same source or output element MUST receive distinct event IDs - event `id` values are unique within a session document. This allows a later `content_engaged.presentation_id` to identify the exact surface occurrence rather than matching only by URL. When a session includes `content_presented` events but no subsequent `content_engaged` events, the telemetry establishes only that content or a reference was made perceivable and no reported interaction followed. It does not establish human attention. Whether this pattern is meaningful depends on the governing terms. Retrieval remains the only lifecycle stage observable from the CDN edge. From 11677fdf1358dbb1416a1c1acefd13e0e51e7792 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Wed, 2 Sep 2026 14:37:38 +0100 Subject: [PATCH 37/38] Land the remaining accepted v1 scope from #7 and #8 - bot_category becomes purpose: an open enum classifying the access (training, inference, search, advertising), with the vendor signal mappings moved to informative Annex C and a migration note (#7) - Generic evidence reference slot on content events (6.8): scheme, detached ref plus digest, no status escalation; the status vocabulary and trust policy stay with the evidence profile (#8) Co-Authored-By: Claude Fable 5 --- SPECIFICATION.md | 61 +++++++++++++++----- telemetry-session.json | 2 +- tests/invalid/withdrawn-ip-hash-batch.json | 2 +- tests/invalid/withdrawn-ip-hash-session.json | 2 +- tests/invalid/withdrawn-ip-hash.json | 2 +- tests/valid/event-batch-edge.json | 4 +- tests/valid/event-standalone-edge.json | 2 +- tests/valid/session-evidence-reference.json | 30 ++++++++++ tests/valid/session-retrieval-tier.json | 2 +- 9 files changed, 84 insertions(+), 23 deletions(-) create mode 100644 tests/valid/session-evidence-reference.json diff --git a/SPECIFICATION.md b/SPECIFICATION.md index ffe4131..1442dc2 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -11,7 +11,7 @@ 3. [Terms and definitions](#3-terms-and-definitions) 4. [Concepts](#4-concepts) - roles, sessions, event lifecycle, source roles, content identification 5. [Schema](#5-schema) - session, event, event types, conversation turn, privacy, intent, conformance levels -6. [Data profiles](#6-data-profiles) - retrieval, edge enrichment, origin enrichment, grounding, citation, presentation, engagement +6. [Data profiles](#6-data-profiles) - retrieval, edge enrichment, origin enrichment, grounding, citation, presentation, engagement, evidence references 7. [Transport](#7-transport) - delivery formats, Content-Telemetry-ID header, routing, click context 8. [Manifest](#8-manifest) - discovery, schema, operator, keys, telemetry, domains 9. [Privacy](#9-privacy) - data minimisation, recommended levels, retention @@ -20,6 +20,7 @@ 12. [Versioning](#12-versioning) - [Annex A (normative): JSON Schema](#annex-a-normative-json-schema) - [Annex B (informative): Examples](#annex-b-informative-examples) +- [Annex C (informative): Vendor bot-classification mappings](#annex-c-informative-vendor-bot-classification-mappings) ## 1. Introduction @@ -60,7 +61,7 @@ Content Telemetry does not: The five-stage lifecycle reports content use observable at inference time: identified content entered a generation context for a particular response, and what the resulting output did with it. -Retrieval is the boundary case. A crawl whose purpose is training or index building can be reported as a `content_retrieved` event, and `bot_category` (section 6.2) distinguishes it, but the event is non-attributable: no grounding, citation, presentation or engagement follows it. What the system then does with the content, whether it enters a training corpus, a fine-tuning set, an embedding store or a search index, is outside this specification. Nothing here reports that a model was trained on a work, and a conforming implementation says nothing either way about it. +Retrieval is the boundary case. A crawl whose purpose is training or index building can be reported as a `content_retrieved` event, and `purpose` (section 6.2) distinguishes it, but the event is non-attributable: no grounding, citation, presentation or engagement follows it. What the system then does with the content, whether it enters a training corpus, a fine-tuning set, an embedding store or a search index, is outside this specification. Nothing here reports that a model was trained on a work, and a conforming implementation says nothing either way about it. Using such a store at inference time is inside scope. When an index built over a content owner's material is queried during a response and returns content that grounds the answer, that is a `content_grounded` event like any other, with `source_role: index` on the retrieval that served it (section 4.4). The line is between constructing a derived artefact and using one to answer a query, not whether an index was involved. @@ -562,7 +563,7 @@ A conforming **Retrieval** emitter MUST: This level requires no agent cooperation. Content owners can implement it using CDN edge compute (Cloudflare Workers, Fastly Compute, etc.). -Origin-side emitters operating at the CDN edge SHOULD include `bot_category`, `response_status`, and `response_bytes` alongside the required fields. These fields make retrieval events useful for bot classification and volume analysis. Without them, the event confirms a fetch occurred but cannot support attribution correlation. +Origin-side emitters operating at the CDN edge SHOULD include `purpose`, `response_status`, and `response_bytes` alongside the required fields. These fields make retrieval events useful for bot classification and volume analysis. Without them, the event confirms a fetch occurred but cannot support attribution correlation. #### 5.7.2 Grounding conformance @@ -667,7 +668,7 @@ CDN and edge network integrations SHOULD include these fields: | Field | Type | Description | |-------|------|-------------| | `user_agent` | string | Request User-Agent header | -| `bot_category` | string | Edge platform's bot classification (see below) | +| `purpose` | string | Purpose of the access, as classified by the reporting party (see below) | | `bot_name` | string | Recognised bot family parsed from the User-Agent (e.g., `Claude-User`, `GPTBot`, `Perplexity-User`) | | `verified` | boolean | Whether the bot identity was cryptographically verified | | `cache_status` | string | Edge cache result: `hit`, `miss`, `bypass`, `dynamic` | @@ -678,19 +679,20 @@ CDN and edge network integrations SHOULD include these fields: | `asn_org` | string | Client AS organisation name | | `country` | string | ISO 3166-1 alpha-2 country code | -#### Bot categories +#### Access purpose -The `bot_category` field carries the edge platform's classification of the requesting bot. Recommended values: +The `purpose` field carries the reporting party's classification of what the access was for. It is an open enum: these are the core values, emitters MAY use custom values, and telemetry consumers MUST tolerate unknown ones. It classifies the access, not the organisation - who is reporting stays in `source_role` (section 4.4). -| Value | Description | Fastly signal | Cloudflare signal | -|-------|-------------|---------------|-------------------| -| `training` | Crawling for model training | `AI-CRAWLER` | `AI Crawler` | -| `inference` | Fetching at query time (RAG) | `AI-FETCHER` | `AI Assistant` | -| `search` | AI search indexing | - | `AI Search` | +| Value | Description | +|-------|-------------| +| `training` | Crawling for model training | +| `inference` | Fetching at query time to inform a response | +| `search` | AI search indexing | +| `advertising` | Access to derive advertising signals (contextual classification, brand safety, campaign targeting) | -The `inference` category is where content attribution is most relevant - there is a user, a query, and a session behind the retrieval. `training` crawls have no session context. `bot_category` can distinguish training crawls from inference fetches, but training-specific telemetry is out of scope for this specification (see section 1.3). Edge platforms map their native classification to these values. +The `inference` purpose is where content attribution is most relevant - there is a user, a query, and a session behind the retrieval. `training` crawls have no session context. `purpose` can distinguish training crawls from inference fetches, but training-specific telemetry is out of scope for this specification (see section 1.3). Edge platforms map their native bot classifications to these values; the mappings for common platforms are informative and collected in Annex C. -Emitting a `training`-category `content_retrieved` event is permitted but non-attributable - there is no session, grounding, or citation to follow it. An edge emitter can report these events through its normal pipeline and need not special-case or suppress them. +Emitting a `training`-purpose `content_retrieved` event is permitted but non-attributable - there is no session, grounding, or citation to follow it. An edge emitter can report these events through its normal pipeline and need not special-case or suppress them. ### 6.3 Origin enrichment (`content_retrieved` + `source_role: origin`) @@ -885,6 +887,18 @@ An action that touches several presentations at once - a `share` of a response c A `link_click` or `agent_navigate` engagement reported from the landing page after a click-out crosses a trust boundary. Such events carry a `ctx_token` in place of `session_id`, which the telemetry consumer resolves to the click context (see section 7.4). The agent-authored engagement itself reaches the engaged content's owner through owner-scoped routing whether or not the token survived the redirect chain (section 7.4.5). +### 6.8 Evidence references (any content event) + +The `data.evidence` field MAY appear on any content event: an array of profile-defined evidence references attached to the event's claim. + +| Field | Type | Required | Description | +|-------|------|----------|-------------| +| `scheme` | string | Yes | Open identifier for the evidence scheme, following the same rules as `content_fingerprint.scheme` (section 6.4) | +| `ref` | string | No | URI of a detached evidence artefact, resolvable independently of the event | +| `digest` | string | No | Digest binding the reference to the artefact's bytes (`sha256:{hex}`) | + +Core defines the slot and nothing more. It does not interpret entries, register schemes, or assign evidentiary status: an event remains a claim by its emitter (SCOPE.md), and the presence of evidence entries raises no event's status by itself. Which schemes a consumer accepts, and what a verified entry establishes, is consumer trust policy defined in an evidence profile outside core. Consumers MUST tolerate unknown schemes and unknown fields within entries. A detached reference - a `ref` with a `digest` - is admitted deliberately, so evidence can remain independently verifiable after the fact without travelling inline. + ## 7. Transport Content Telemetry defines a signal format, not a wire protocol. Common delivery patterns include HTTP postback, bulk upload after session end, MCP tool calls, message queues (Kafka, SQS), and direct database writes. The choice of transport is left to implementers. @@ -913,7 +927,7 @@ A standalone event carries `document_type`, `schema_version`, and optionally `se "content_telemetry_id": "770e8400-e29b-41d4-a716-446655440300", "content_url": "https://www.ft.com/content/abc123", "data": { - "bot_category": "inference", + "purpose": "inference", "cache_status": "miss", "response_status": 200 } @@ -1462,6 +1476,13 @@ Emitters remove the field and MUST NOT populate it; `asn`, `asn_org` and `countr remain. V1 also requires `source_role` on every `content_retrieved` event (section 5.7.1); preview emitters that omitted it add the role they report under. +V1 renames the retrieval profile's `bot_category` field to `purpose` and defines +it as an open enum classifying the access rather than the bot: the v0.1 values +`training`, `inference` and `search` keep their meaning, `advertising` is added, +and the vendor signal mappings move to informative Annex C. Emitters rename the +field; consumers MAY read a v0.1 `bot_category` value as `purpose` when +migrating historical data. + `license_ref` keeps its wire form but no longer asserts that the use was licensed (section 5.2.3): a consumer that read a v0.1 `license_ref` as verification of entitlement now reads it as the emitter's claim about which grant applied. @@ -1641,7 +1662,7 @@ A content owner's CDN detects an AI agent fetching content. The agent also repor "content_url": "https://www.wirecutter.com/reviews/best-wireless-headphones", "data": { "user_agent": "ClaudeBot/1.0", - "bot_category": "inference", + "purpose": "inference", "bot_name": "ClaudeBot", "verified": true, "cache_status": "miss", @@ -1938,3 +1959,13 @@ A marketplace intermediary delivers content from many publishers under a single ``` The telemetry consumer resolves `autoreview.example` by domain registration and the `mkt:gridnews:` prefix by identifier registration, and produces two owner-filtered views: AutoReview sees its retrieval, grounding, citation and link presentation; GridNews sees its own four events and nothing of AutoReview's. Neither view carries the other owner's identifiers, and both carry the shared `content_scope` so the marketplace can reconcile the session against the agreement. The second retrieval has no `content_url` at all - marketplace API content with no canonical URL - and resolves by `content_id` alone. + +## Annex C (informative): Vendor bot-classification mappings + +Edge platforms classify AI bot traffic in their own vocabularies. These mappings to the `purpose` values of section 6.2 are informative, reflect the platforms' published categories at the time of writing, and change on the platforms' own cadence. + +| `purpose` | Fastly signal | Cloudflare signal | +|-----------|---------------|-------------------| +| `training` | `AI-CRAWLER` | `AI Crawler` | +| `inference` | `AI-FETCHER` | `AI Assistant` | +| `search` | - | `AI Search` | diff --git a/telemetry-session.json b/telemetry-session.json index fe68ba7..c3aed4f 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -289,7 +289,7 @@ "media_type": { "$ref": "#/$defs/MediaType" }, "content_depth": { "type": "string", "description": "Depth of the content record reached: metadata, abstract, full. Open vocabulary; consumers MUST tolerate unknown values (section 6.1)." }, "user_agent": { "type": "string" }, - "bot_category": { "type": "string", "description": "Edge platform's bot classification. Recommended values: training, inference, search" }, + "purpose": { "type": "string", "description": "Purpose of the access as classified by the reporting party. Open enum; core values: training, inference, search, advertising" }, "bot_name": { "type": "string", "description": "Recognised bot family parsed from the User-Agent (e.g., Claude-User, GPTBot, Perplexity-User). Stable across product variants within a vendor." }, "verified": { "type": "boolean" }, "cache_status": { "type": "string", "description": "Edge cache result. Recommended values: hit, miss, bypass, dynamic" }, diff --git a/tests/invalid/withdrawn-ip-hash-batch.json b/tests/invalid/withdrawn-ip-hash-batch.json index 60ec597..a79b8d5 100644 --- a/tests/invalid/withdrawn-ip-hash-batch.json +++ b/tests/invalid/withdrawn-ip-hash-batch.json @@ -10,7 +10,7 @@ "source_role": "edge", "content_url": "https://www.example-news.com/economy/rate-decision-analysis", "data": { - "bot_category": "inference", + "purpose": "inference", "ip_hash": "sha256:e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6" } } diff --git a/tests/invalid/withdrawn-ip-hash-session.json b/tests/invalid/withdrawn-ip-hash-session.json index ed8a10b..fc09f04 100644 --- a/tests/invalid/withdrawn-ip-hash-session.json +++ b/tests/invalid/withdrawn-ip-hash-session.json @@ -12,7 +12,7 @@ "source_role": "edge", "content_url": "https://www.example-news.com/economy/rate-decision-analysis", "data": { - "bot_category": "inference", + "purpose": "inference", "ip_hash": "sha256:e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6" } } diff --git a/tests/invalid/withdrawn-ip-hash.json b/tests/invalid/withdrawn-ip-hash.json index c634676..9ff7f2a 100644 --- a/tests/invalid/withdrawn-ip-hash.json +++ b/tests/invalid/withdrawn-ip-hash.json @@ -11,7 +11,7 @@ "content_url": "https://www.telegraph.co.uk/business/2026/03/28/ftse-100-markets-live", "data": { "user_agent": "PerplexityBot/1.0", - "bot_category": "inference", + "purpose": "inference", "response_status": 200, "asn": 396982, "asn_org": "Perplexity AI", diff --git a/tests/valid/event-batch-edge.json b/tests/valid/event-batch-edge.json index 03ee1c5..8611ff4 100644 --- a/tests/valid/event-batch-edge.json +++ b/tests/valid/event-batch-edge.json @@ -11,7 +11,7 @@ "content_url": "https://www.telegraph.co.uk/business/2026/03/28/ftse-100-markets-live", "data": { "user_agent": "PerplexityBot/1.0", - "bot_category": "inference", + "purpose": "inference", "verified": false, "cache_status": "miss", "response_status": 200 @@ -24,7 +24,7 @@ "content_url": "https://www.telegraph.co.uk/business/2026/03/28/bank-of-england-rates", "data": { "user_agent": "GPTBot/1.2", - "bot_category": "training", + "purpose": "training", "verified": true, "cache_status": "hit", "response_status": 200 diff --git a/tests/valid/event-standalone-edge.json b/tests/valid/event-standalone-edge.json index 086bb35..5a7c415 100644 --- a/tests/valid/event-standalone-edge.json +++ b/tests/valid/event-standalone-edge.json @@ -10,7 +10,7 @@ "content_url": "https://www.telegraph.co.uk/business/2026/03/28/ftse-100-markets-live", "data": { "user_agent": "PerplexityBot/1.0", - "bot_category": "inference", + "purpose": "inference", "verified": false, "cache_status": "miss", "response_status": 200, diff --git a/tests/valid/session-evidence-reference.json b/tests/valid/session-evidence-reference.json new file mode 100644 index 0000000..ae086db --- /dev/null +++ b/tests/valid/session-evidence-reference.json @@ -0,0 +1,30 @@ +{ + "_test_description": "Grounding event carrying a data.evidence array (section 6.8): one detached reference (ref plus digest) and one entry in an unknown scheme, which consumers tolerate. The slot attaches profile-defined evidence to an event's claim without raising its status.", + "schema_version": "1.0", + "session_id": "aa0e8400-e29b-41d4-a716-446655440600", + "agent_id": "assistant-v1", + "started_at": "2026-08-30T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-30T10:00:01Z", + "content_url": "https://example.com/articles/energy-prices", + "data": { + "scope": "session", + "cached": false, + "chars_ingested": 5400, + "evidence": [ + { + "scheme": "org.example.retrieval-receipt", + "ref": "https://evidence.example.com/receipts/8f3a", + "digest": "sha256:c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4" + }, + { + "scheme": "com.vendor.unknown-scheme", + "inline_note": "consumers tolerate unknown schemes and fields" + } + ] + } + } + ] +} diff --git a/tests/valid/session-retrieval-tier.json b/tests/valid/session-retrieval-tier.json index 8e0a803..3047baf 100644 --- a/tests/valid/session-retrieval-tier.json +++ b/tests/valid/session-retrieval-tier.json @@ -12,7 +12,7 @@ "content_url": "https://www.theguardian.com/technology/2026/mar/28/ai-agents-roundup", "data": { "user_agent": "ClaudeBot/1.0", - "bot_category": "inference", + "purpose": "inference", "verified": true, "cache_status": "miss", "response_status": 200, From 7703807d47687e0e6f36ce5697ffb22d2c0b51b8 Mon Sep 17 00:00:00 2001 From: NarrativAI Agent Date: Wed, 2 Sep 2026 14:39:19 +0100 Subject: [PATCH 38/38] Add the multi-agent worked example and extension-event namespacing B.6 shows a delegated child session with parent_session_id, closing the example promised on #1; 5.3 states the namespaced-name rule that 5.7.6 already relied on (#9). Co-Authored-By: Claude Fable 5 --- SPECIFICATION.md | 40 +++++++++++++++++++++++++++++++++++++++- 1 file changed, 39 insertions(+), 1 deletion(-) diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 1442dc2..1c9855d 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -469,7 +469,7 @@ A processor that stores, forwards or transforms a document MUST preserve `terms_ #### Extension events -The core schema defines content and conversation events. Implementations MAY define additional event types using the `data` field for type-specific metadata. Commerce-specific fields (product identifiers, checkout events) are a planned extension. +The core schema defines content and conversation events. Implementations MAY define additional event types using the `data` field for type-specific metadata. Extension event types SHOULD use namespaced names (for example `com.example.checkout_completed`) so they cannot collide with each other or with future core types. Commerce-specific fields (product identifiers, checkout events) are a planned extension. ### 5.4 Conversation turn @@ -1960,6 +1960,44 @@ A marketplace intermediary delivers content from many publishers under a single The telemetry consumer resolves `autoreview.example` by domain registration and the `mkt:gridnews:` prefix by identifier registration, and produces two owner-filtered views: AutoReview sees its retrieval, grounding, citation and link presentation; GridNews sees its own four events and nothing of AutoReview's. Neither view carries the other owner's identifiers, and both carry the shared `content_scope` so the marketplace can reconcile the session against the agreement. The second retrieval has no `content_url` at all - marketplace API content with no canonical URL - and resolves by `content_id` alone. +### B.6 Delegated sub-agent session + +An orchestrating agent delegates a research step to a sub-agent. The child session links to its parent with `parent_session_id`, keeps its own `session_id` and events, and grounds the source in its own generation context (section 5.1). No citation or presentation is emitted merely because the child returns an internal result to its orchestrator - those events belong to the session whose output reaches the end user. + +```json +{ + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440021", + "parent_session_id": "660e8400-e29b-41d4-a716-446655440020", + "agent_id": "research-subagent-v1", + "started_at": "2026-07-29T09:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-07-29T09:00:02Z", + "turn_id": "child-turn-1", + "content_url": "https://example.org/research/source", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 3600 + } + }, + { + "type": "turn_completed", + "timestamp": "2026-07-29T09:00:05Z", + "turn_id": "child-turn-1", + "turn": { + "privacy_level": "minimal", + "response_tokens": 180 + } + } + ] +} +``` + +If the orchestrator later cites and presents this source to the user, those `content_cited` and `content_presented` events appear in the parent session, and a consumer holding both sessions joins them through `parent_session_id`. Emitters MAY omit the link when the relationship is unavailable or its disclosure is not appropriate; consumers MUST NOT infer that an unlinked session had no parent. + ## Annex C (informative): Vendor bot-classification mappings Edge platforms classify AI bot traffic in their own vocabularies. These mappings to the `purpose` values of section 6.2 are informative, reflect the platforms' published categories at the time of writing, and change on the platforms' own cadence.