diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 9c523ad..2dfce92 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -17,6 +17,8 @@ jobs: - name: Install uv uses: astral-sh/setup-uv@v5 - name: Run conformance test suite - run: uv run --with jsonschema python tests/validate.py + run: uv run --with "jsonschema[format-nongpl]" python tests/validate.py - name: Validate worked examples in the spec - run: uv run --with jsonschema python tests/check_examples.py + run: uv run --with "jsonschema[format-nongpl]" python tests/check_examples.py + - name: Replay suite-weakening mutations + run: uv run --with "jsonschema[format-nongpl]" python tests/mutation_smoke.py diff --git a/GOVERNANCE.md b/GOVERNANCE.md index 6e017a1..36cecca 100644 --- a/GOVERNANCE.md +++ b/GOVERNANCE.md @@ -2,7 +2,7 @@ ## Status -Content Telemetry is an open specification stewarded by the SPUR Coalition. It is in preview (v0.x); see [SPECIFICATION.md](./SPECIFICATION.md) section 12 for the versioning policy. +Content Telemetry is an open specification stewarded by the SPUR Coalition. Version 1.0 is the current specification; see [SPECIFICATION.md](./SPECIFICATION.md) section 12 for the versioning policy. ## Stewardship @@ -18,20 +18,22 @@ The SPUR Coalition stewards the specification and holds this repository. The sta The SPUR Coalition is a group of publishers and content owners that maintains the Content Telemetry standard. It holds the intellectual property through the preview period and releases the standard under Apache 2.0 from 12 June 2026. -Contributing to the standard does not require membership. The wire format is developed in the open, and anyone - content owner, agent operator, intermediary, or implementer - can take part through the issue tracker and the comment process described in the [README](./README.md#request-for-comment). +Contributing to the standard does not require membership. The wire format is developed in the open, and anyone - content owner, agent operator, intermediary, or implementer - can take part through the issue tracker and the process described in the [README](./README.md#consultation-record). -The standard is maintained by Alex Springer (alex@spurcoalition.org). The decision-making process will be set out before 1.0. +The standard is maintained by Alex Springer (alex@spurcoalition.org). -## Path to 1.0 +## Version 1.0 and decisions -The specification is at v0.1 (preview). It moves to 1.0 once the open questions are resolved, the conformance suite is stable, and there are independent interoperable implementations from more than one party. There is no fixed date. +The specification reached 1.0 on 2 September 2026, following the public consultation of 12 June to 24 July 2026 and the release-candidate work recorded on the issue tracker. + +Decisions follow the proposal and alignment process in [CONTRIBUTING.md](./CONTRIBUTING.md): anyone may propose a change, a maintainer records the disposition publicly on the issue tracker, and the SPUR Steering Board approves releases. The tracker and pull-request history are the public decision record. Minor versions add optional fields and event types; breaking changes require a major version (SPECIFICATION.md section 12). ## How to participate - File questions and bugs on the [issue tracker](https://github.com/SPUR-Coalition/telemetry/issues) (see the templates). -- Follow the [post-consultation status](./README.md#consultation-status). The - formal v1 comment window is closed, but concrete bugs and implementation - evidence remain welcome on the issue tracker. +- See the [consultation record](./README.md#consultation-record). The formal + v1 comment window is closed, but concrete bugs and implementation evidence + remain welcome on the issue tracker. - Propose new capabilities or changes in behaviour as a short human-written note in [`proposals/`](./proposals/), following [CONTRIBUTING.md](./CONTRIBUTING.md). diff --git a/README.md b/README.md index c82cc9c..3ddad71 100644 --- a/README.md +++ b/README.md @@ -2,16 +2,7 @@ **Signal format for AI content usage reporting.** -This is a preview specification. Field names, event types, and schema structure may change before 1.0. - -> **Consultation status — 12 August 2026:** The public comment period closed on -> **24 July 2026**. The SPUR Steering Board has approved the v1 direction, and a -> final disposition is now recorded on every consultation thread: each carries -> an outcome label, and accepted core changes sit on the -> [v1 release candidate milestone](https://github.com/SPUR-Coalition/telemetry/milestone/1) -> with a schema freeze targeted for **21 August 2026**. Accepted changes are -> merging on the `v1-draft` integration line. Version 0.1 remains the current -> published preview. See [Consultation status](#consultation-status) below. +**Version 1.0** is the current specification, published 2 September 2026. It replaces the v0.1 preview; the changes and migration steps are recorded in [SPECIFICATION.md section 12.1](./SPECIFICATION.md#121-migration-from-the-v01-preview). ## Contents @@ -21,13 +12,13 @@ This is a preview specification. Field names, event types, and schema structure - [Repo contents](#repo-contents) - [Example](#example) - [Relationship to other protocols](#relationship-to-other-protocols) -- [Consultation status](#consultation-status) -- [Open questions in v0.1](#open-questions-in-v01) +- [Consultation record](#consultation-record) +- [Open questions in v1](#open-questions-in-v1) - [Versioning](#versioning) ## Problem -AI agents retrieve a content owner's content, use it to generate responses, and sometimes cite it. Content owners currently see an initial retrieval event - HTTP requests hitting their servers or access logs from content repositories. Whether the content actually influenced the response, whether it was cited, whether a user saw the citation, whether they clicked through - is not reported back to content owners. +AI agents retrieve a content owner's content, use it to generate responses, and sometimes cite it. Content owners currently see an initial retrieval event - HTTP requests hitting their servers or access logs from content repositories. Whether the content actually entered a generation context, whether it was cited, whether content or a source reference was made perceivable, and whether anyone interacted with that presentation is not reported back to content owners. Platforms self-report usage metrics (if they report at all), and content owners have no way to verify the numbers or compare across platforms. @@ -39,7 +30,7 @@ Content Telemetry tracks content through five stages: Retrieved → content fetched over HTTP (content owner can see this today) Grounded → content loaded into the agent's generation context Cited → content explicitly referenced in the response - Displayed → user saw it - a reference, or the content embedded in the answer + Presented → content or a source reference made perceivable on a recipient-facing surface Engaged → user clicked, copied, shared, or directed the agent to act ``` @@ -53,13 +44,13 @@ The gaps between stages show how content was used: The grounding event captures the boundary "this content entered the agent's generation context." It is architecture-neutral and decoupled from retrieval: content cached by the agent for days still produces a grounding event in every session it influences. -Grounding and display record two different kinds of influence: grounding means the content influenced the agent, display means it reached the user. The two diverge as agent experiences move beyond the chat window - an agentic browser can render a page to the user that never entered a generation context, reported as a `content_displayed` event with `display_type: embed` and no grounding event. +Grounding and presentation record different boundary crossings: grounding means the content entered a generation context; presentation means content or a source reference was made perceivable on a recipient-facing surface. Presentation does not prove attention. The two diverge as agent experiences move beyond the chat window - an agentic browser can render a page that never entered a generation context, reported as a `content_presented` event with `presentation_kind: content` and `presentation_type: embed` and no grounding event. ## Design principles **Post-hoc, not pre-declared.** Events report what actually happened, not what the agent said it would do at request time. An agent cannot reliably declare how it will use content before reading it. -**Observable boundaries, not agent internals.** The five event types mark boundary crossings. What happens between them - the fan-out, relevance evaluation, re-ranking, reasoning chains - is internal to the agent and changes constantly. The spec does not model it. +**Observable boundaries, not agent internals.** The five content event types mark boundary crossings. What happens between them - the fan-out, relevance evaluation, re-ranking, reasoning chains - is internal to the agent and changes constantly. The spec does not model it. **Multiple observers, one event.** A content retrieval can be reported by the content owner's CDN, the content owner's origin server, and the AI agent independently. The `Content-Telemetry-ID` header correlates these into a single corroborated event. Uncorroborated retrievals (no matching agent event) may indicate an agent that does not yet support the telemetry protocol. @@ -72,7 +63,7 @@ Grounding and display record two different kinds of influence: grounding means t - [telemetry-event-batch.json](./telemetry-event-batch.json) - JSON Schema for event batch envelopes - [manifest.json](./manifest.json) - JSON Schema for the `.well-known/content-telemetry.json` manifest ([section 8](./SPECIFICATION.md#8-manifest)) - [tests/](./tests/) - conformance test suite -- [GOVERNANCE.md](./GOVERNANCE.md) - stewardship, preview status, relationship to profiles +- [GOVERNANCE.md](./GOVERNANCE.md) - stewardship, versioning status, relationship to profiles - [LICENSE](./LICENSE) - Apache License 2.0 This repository is the **standard** - the wire format. Publisher-facing accreditation and the SPUR conformance mark are defined separately in the [SPUR Content Telemetry Profile](https://github.com/SPUR-Coalition/telemetry-profile), which references this specification by version. The standard defines the privacy mechanism (section 5.5); whether a profile makes any privacy level binding is the profile's choice. See [GOVERNANCE.md](./GOVERNANCE.md). @@ -83,7 +74,7 @@ A user asks an AI agent about UK interest rates. The agent grounds its response ```json { - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440000", "agent_id": "copilot-v3", "started_at": "2026-03-28T09:00:00Z", @@ -96,6 +87,7 @@ A user asks an AI agent about UK interest rates. The agent grounds its response "data": { "scope": "session", "cached": true, + "chars_ingested": 12800, "tokens_ingested": 3200, "content_last_modified": "2026-03-27T18:30:00Z" } @@ -111,9 +103,12 @@ A user asks an AI agent about UK interest rates. The agent grounds its response } }, { + "id": "550e8400-e29b-41d4-a716-446655440001", "type": "content_cited", "timestamp": "2026-03-28T09:00:05Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:paragraph:2", "content_url": "https://www.ft.com/content/abc123", "content_id": "ft:abc123", "data": { @@ -122,12 +117,19 @@ A user asks an AI agent about UK interest rates. The agent grounds its response } }, { - "type": "content_displayed", + "id": "550e8400-e29b-41d4-a716-446655440002", + "type": "content_presented", "timestamp": "2026-03-28T09:00:05Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:paragraph:2", + "citation_id": "550e8400-e29b-41d4-a716-446655440001", "content_url": "https://www.ft.com/content/abc123", "content_id": "ft:abc123", - "data": { "display_type": "link" } + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } }, { "type": "turn_completed", @@ -144,67 +146,50 @@ A user asks an AI agent about UK interest rates. The agent grounds its response } ``` -The content owner can derive: FT article `abc123` was in context for the response, cited as a paraphrase, link was displayed, user never clicked, ads were shown alongside. +The content owner can derive: FT article `abc123` was in context for the response, cited as a paraphrase, its link was made perceivable, and no engagement was reported; ads were shown alongside. ## Relationship to other protocols Content Telemetry is focussed on **reporting**, while content **access** protocols (Really Simple Licensing, peek-then-pay, IAB CoMP, bilateral APIs) aim to govern how agents discover and license content. The `license_ref` field on events connects telemetry to whatever access protocol issued the licence, but the schemas are independent - telemetry works with any access protocol, or none. -## Consultation status +## Consultation record -The public comment period ran from **12 June to 24 July 2026** and is now closed. -Thank you to everyone who opened an issue, submitted a pull request, joined a -working session or supplied implementation evidence. +The public comment period ran from **12 June to 24 July 2026**. Thank you to +everyone who opened an issue, submitted a pull request, joined a working +session or supplied implementation evidence. The consultation produced 29 specification issue threads, three profile issue -threads and five pull requests. The maintainers are now: - -- [x] reviewing the full consultation record; -- [x] preparing a proposed disposition for every thread; -- [x] recording the approved dispositions on the issue tracker; -- [ ] completing focused v1 changes and migration fixtures on `v1-draft`; -- [ ] publishing a v1 release candidate for implementer testing; and -- [ ] publishing final v1 only after publisher, intermediary and agent/platform - acceptance cases pass. - -A final disposition is now recorded on every consultation thread, and accepted -core changes are tracked on the -[v1 release candidate milestone](https://github.com/SPUR-Coalition/telemetry/milestone/1) -towards the 21 August 2026 schema freeze. Each thread's outcome label and -disposition comment, not its open or closed state, record the decision: threads -that remain open do so pending their recorded follow-ups. A preparation branch -does not change the published specification. The issue tracker and pull-request -history remain the public decision record. - -Concrete schema, fixture and documentation bugs may still be filed using the +threads and five pull requests. Every thread carries a recorded outcome, and +the issue tracker and pull-request history remain the public decision record. +The accepted core changes were tracked on the +[v1 release candidate milestone](https://github.com/SPUR-Coalition/telemetry/milestone/1), +merged on the `v1-draft` integration line, and published as version 1.0 on +2 September 2026. + +Concrete schema, fixture and documentation bugs may be filed using the *Schema or example bug* template, and questions or unclear requirements using *Spec feedback / open question*. For a new capability or change in behaviour, submit a short human-written note to [`proposals/`](./proposals/) and wait for explicit maintainer alignment before beginning implementation (see -[CONTRIBUTING.md](./CONTRIBUTING.md)); new design proposals are not -automatically part of the v1 consultation scope. Pull requests remain welcome -for specific fixes. Feedback on accreditation or the conformance mark belongs -on the [profile +[CONTRIBUTING.md](./CONTRIBUTING.md)). Pull requests remain welcome for +specific fixes. Feedback on accreditation or the conformance mark belongs on +the [profile repository](https://github.com/SPUR-Coalition/telemetry-profile/issues). -Required fields, event types and schema structure may still change before 1.0 -(section 12). Version 0.1 remains the current published preview until a later -version is released. - -## Open questions in v0.1 +## Open questions in v1 -This is a preview specification. The following areas are under active discussion and will be refined with implementer input: +The following areas are expected to develop in 1.x minor versions and profiles, with implementer input: **Grounding boundary.** The spec defines grounding as content entering the generation model's context (sections 4.3 and 6.4). For straightforward RAG pipelines this is clear. For pipelines with multiple processing stages - embedding, re-ranking, summarisation before context insertion - the boundary requires judgement. The spec draws the line at the generation context (not earlier retrieval stages), but edge cases remain. When a re-ranking or summarisation stage is itself a generative model, the multi-step rule in section 6.4 (content entering a sub-agent's generation context is grounded) can pull selection stages back inside the boundary. Input from platform engineering teams building real implementations will sharpen this definition. -**Event volume at scale.** A single deep-research query can produce 100+ retrieval events and dozens of grounding/citation events. The session document format already handles transport - one POST with all events after the session ends, not one request per event. Volume management beyond that (storage, processing, consumer-side aggregation) is an implementation concern, not a protocol gap. Sampling and aggregation are options for future versions but are not in v0.1; the standard sets no default for reporting granularity, leaving it to profiles and deployments. +**Event volume at scale.** A single deep-research query can produce 100+ retrieval events and dozens of grounding/citation events. The session document format already handles transport - one POST with all events after the session ends, not one request per event. Volume management beyond that (storage, processing, consumer-side aggregation) is an implementation concern, not a protocol gap. Version 1 adds an explicit coverage declaration - `complete`, `sampled`, `aggregated` or `selected` (section 5.7.6) - and a manifest field for it (section 8.5); the standard still sets no default for reporting granularity, leaving it to profiles and deployments. -**Verification of grounding and citation.** Grounding and citation events are reported by the agent, which is also the party that may owe compensation under a licence. In v0.1, manifest signing is informational: consumers may verify signatures but are not required to, and the specification defines no required proof binding an event to its emitter (sections 8.4 and 8.9). The events attribution depends on are therefore self-reported by the reporting party. Verifiable credentials and signed events are deferred (section 8.9). One corroboration mechanism works without signing: the `Content-Telemetry-ID` field correlates an agent-reported retrieval with an origin- or edge-reported one (section 7.2), but it covers retrieval only - grounding, citation, display, and engagement have no independent observer. Signing, even once required, would prove who reported an event, not that the event is true or that all qualifying events were reported. Input is wanted on what a verification layer should cover and where it belongs. Mechanisms that test truthfulness and completeness rather than origin, such as sampled audits or publisher-seeded canary content, are of particular interest. +**Verification of grounding and citation.** Grounding and citation events are reported by the agent, which is also the party that may owe compensation under a licence. In v1, manifest signing is informational: consumers may verify signatures but are not required to, and the specification defines no required proof binding an event to its emitter (sections 8.4 and 8.9). The events attribution depends on are therefore self-reported by the reporting party. Verifiable credentials and signed events are deferred (section 8.9). One corroboration mechanism works without signing: the `Content-Telemetry-ID` field correlates an agent-reported retrieval with an origin- or edge-reported one (section 7.2), but it covers retrieval only - grounding, citation, presentation, and engagement have no independent observer. Signing, even once required, would prove who reported an event, not that the event is true or that all qualifying events were reported. Input is wanted on what a verification layer should cover and where it belongs. Mechanisms that test truthfulness and completeness rather than origin, such as sampled audits or publisher-seeded canary content, are of particular interest. -**Reporting granularity.** The standard sets no default for reporting granularity, leaving it to profiles and deployments (see *Event volume* above). The SPUR profile requires event-level delivery and does not permit aggregation. The open question is whether the standard should say more about sampling and aggregation so that profiles do not each define it separately, and how event-level delivery scales for the highest-volume case. No mechanism is selected in v0.1. +**Reporting granularity.** The standard sets no default for reporting granularity, leaving it to profiles and deployments (see *Event volume* above). The SPUR profile requires event-level delivery and does not permit aggregation. Version 1 answers the first half of the question: coverage modes are defined once, in section 5.7.6, so that profiles reference them rather than each define their own. How event-level delivery scales for the highest-volume case remains open. ## Versioning This repo tracks the specification version. SDK repos have their own release cadences and declare which spec version they support. -Current spec version: **0.1** (preview) +Current spec version: **1.0** diff --git a/SCOPE.md b/SCOPE.md index f1fe559..72e3888 100644 --- a/SCOPE.md +++ b/SCOPE.md @@ -18,7 +18,7 @@ Conformance and verification answer five separate questions: 4. **Factual truth and completeness:** did the event happen as claimed, and were all qualifying events reported? 5. **Entitlement:** was the reported use permitted under an applicable grant or agreement? -Events are claims by identified emitters. Evidence applies to a particular assertion. Origin or access evidence can corroborate only what that observer could see; it cannot prove grounding, reproduction, citation, presentation, engagement, truth, completeness or entitlement. +Events are claims by identified emitters. Evidence applies to a particular assertion. Origin or access evidence can corroborate only what that observer could see; it cannot prove grounding, citation, presentation, engagement, truth, completeness or entitlement. Relationship configuration should avoid profile proliferation. A publisher may require `content_grounded` and `content_cited` events, intent-level topics, event delivery to a named endpoint and a set of aggregate reports. Another may require `content_cited`, `content_presented` and `content_engaged` events with a different privacy level. These are deployment choices backed by governing terms, not publisher-specific protocol profiles. A new profile is justified only when a class of relationships introduces semantics or processing rules that multiple implementations must interpret in the same way. diff --git a/SPECIFICATION.md b/SPECIFICATION.md index 746e4a4..1c9855d 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -1,18 +1,18 @@ # Content Telemetry Specification -**Version:** 0.1 -**Status:** Preview -**Last updated:** 2026-06-11 +**Version:** 1.0 +**Status:** Current specification +**Published:** 2026-09-02 ## Contents -1. [Introduction](#1-introduction) - problem, goals, non-goals, relationship to access protocols, conventions +1. [Introduction](#1-introduction) - problem, goals, non-goals, inference-time scope, relationship to access protocols, conventions 2. [Normative references](#2-normative-references) 3. [Terms and definitions](#3-terms-and-definitions) 4. [Concepts](#4-concepts) - roles, sessions, event lifecycle, source roles, content identification 5. [Schema](#5-schema) - session, event, event types, conversation turn, privacy, intent, conformance levels -6. [Data profiles](#6-data-profiles) - retrieval, edge enrichment, origin enrichment, grounding, citation, display, engagement -7. [Transport](#7-transport) - delivery formats, Content-Telemetry-ID header +6. [Data profiles](#6-data-profiles) - retrieval, edge enrichment, origin enrichment, grounding, citation, presentation, engagement, evidence references +7. [Transport](#7-transport) - delivery formats, Content-Telemetry-ID header, routing, click context 8. [Manifest](#8-manifest) - discovery, schema, operator, keys, telemetry, domains 9. [Privacy](#9-privacy) - data minimisation, recommended levels, retention 10. [Attribution](#10-attribution) - counting semantics, grounding without citation @@ -20,6 +20,7 @@ 12. [Versioning](#12-versioning) - [Annex A (normative): JSON Schema](#annex-a-normative-json-schema) - [Annex B (informative): Examples](#annex-b-informative-examples) +- [Annex C (informative): Vendor bot-classification mappings](#annex-c-informative-vendor-bot-classification-mappings) ## 1. Introduction @@ -53,9 +54,17 @@ Content Telemetry does not: - Mandate specific privacy policies (left to agreements between parties) - Require specific transport protocols (HTTP, gRPC, etc. all valid) - Define content access or licensing protocols (see 1.4) -- Model content usage for model training. The five-stage lifecycle covers inference-time usage only. The `bot_category` field on retrieval events (section 6.2) can distinguish training crawls from inference fetches, but training-specific telemetry is out of scope. +- Model what a system does with content other than at inference time. Assembling a training corpus, training or fine-tuning a model, computing embeddings and constructing a retrieval index are all outside scope (see 1.3.1). - Define accreditation tiers, conformance marks, or community-specific conformance requirements. These belong in profiles layered on this specification (see [GOVERNANCE.md](./GOVERNANCE.md)). +#### 1.3.1 Inference-time scope + +The five-stage lifecycle reports content use observable at inference time: identified content entered a generation context for a particular response, and what the resulting output did with it. + +Retrieval is the boundary case. A crawl whose purpose is training or index building can be reported as a `content_retrieved` event, and `purpose` (section 6.2) distinguishes it, but the event is non-attributable: no grounding, citation, presentation or engagement follows it. What the system then does with the content, whether it enters a training corpus, a fine-tuning set, an embedding store or a search index, is outside this specification. Nothing here reports that a model was trained on a work, and a conforming implementation says nothing either way about it. + +Using such a store at inference time is inside scope. When an index built over a content owner's material is queried during a response and returns content that grounds the answer, that is a `content_grounded` event like any other, with `source_role: index` on the retrieval that served it (section 4.4). The line is between constructing a derived artefact and using one to answer a query, not whether an index was involved. + ### 1.4 Relationship to content access protocols Content access protocols govern how AI agents discover and license content. Examples include peek-then-pay (HTTP 203 previews with JWT licensing), IAB CoMP (content package negotiation), and bilateral API agreements. @@ -66,6 +75,8 @@ An agent cannot reliably declare how it will use content before reading it - a r Events can reference a licence via the `license_ref` field (section 5.2), connecting telemetry to whatever access protocol issued the licence. The telemetry schema does not depend on any specific access protocol. +Discovery protocols and registries - catalogues that describe where content sources are and what they offer, such as agent resource discovery formats - sit upstream of both layers. A catalogue records where a content source is; telemetry records what happened when an agent used it. Content Telemetry is the outcome layer for discovery in the same sense that it is the reporting counterpart to access: it carries no ranking or discovery metadata of its own, and a discovery service that wants outcome signal consumes telemetry like any other party (section 7.3). + ### 1.5 Conventions The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in [RFC 2119](https://www.rfc-editor.org/rfc/rfc2119) and [RFC 8174](https://www.rfc-editor.org/rfc/rfc8174). @@ -97,10 +108,14 @@ For the purposes of this specification, the following terms apply. | **content owner** | entity that owns or licences content accessed by an AI agent | | **agent operator** | entity running the AI agent that uses content | | **grounding** | content entering the generation model's context, the boundary where content can directly influence output (section 4.3) | +| **presentation** | content or a source reference made perceivable on a recipient-facing surface (section 4.3) | | **source role** | classification of the observer reporting a retrieval event: `origin`, `edge`, `index`, or `agent` (section 4.4) | | **privacy level** | data sharing tier controlling which conversation fields are populated: `full`, `summary`, `intent`, or `minimal` (section 5.5) | | **conformance level** | emitter capability tier: Retrieval, Grounding, or Citation (section 5.7) | | **content scope** | opaque identifier grouping sessions by their content access context (section 5.1.1) | +| **governing terms** | licence, contract or other terms selecting which events a relationship requires and the coverage, cadence, delivery, privacy and reports owed (sections 5.2.4, 5.7.6) | +| **qualifying occurrence** | occurrence satisfying an event type's core definition and occurrence boundary within the relationship scope reported under (section 5.7.6) | +| **coverage** | declared relationship between qualifying occurrences and emitted events: `complete`, `sampled`, `aggregated` or `selected` (section 5.7.6) | ## 4. Concepts @@ -169,7 +184,7 @@ Session │ ├── content_retrieved (HTTP layer) │ ├── content_grounded (influence layer) │ ├── content_cited (response layer) -│ ├── content_displayed (UI layer) +│ ├── content_presented (recipient-facing surface) │ ├── turn_completed │ ├── content_engaged (user action layer) │ └── ... @@ -184,42 +199,56 @@ Content moves through five stages during an agent interaction: 1. **Retrieved** - Content fetched over HTTP from an origin server, CDN, marketplace, or index. This is an infrastructure event observable by the content owner's infrastructure (origin server, edge network) and the agent. A retrieval may be cached by the agent for use across multiple sessions. + One retrieval occurrence is one completed fetch of a content representation as observed by the reporting party: a redirect chain resolving to one representation is one occurrence, and a revalidation returning no new representation (an HTTP 304) is not a new occurrence. Serving content from the agent's own cache is not a new retrieval; the reuse surfaces as grounding (stage 2), not as a repeated `content_retrieved` event. + 2. **Grounded** - Content used in the agent's generation context for this session or turn. The boundary is "this content entered the generation model's context" - the point where content can directly influence the model's output. Content used only for retrieval selection (embedding similarity search, re-ranking scores, routing decisions) without entering the generation context is not grounded. Grounding is architecture-neutral: same event whether the agent uses RAG, chain-of-thought reasoning, embeddings, or multi-step delegation (see section 6.4 for architecture-specific guidance). Grounding is decoupled from retrieval: content may be grounded from a live fetch, from agent-side cache, or from a pre-loaded index. Only the agent can report grounding events. -3. **Cited** - Content explicitly referenced in the agent's response: quoted, paraphrased, or linked. A subset of grounded content. Content can influence every response in a session without being cited once. + One grounding occurrence is one distinct content item entering a generation context at the declared `data.scope`: at `session` scope, a content item grounds once per session; at `turn` scope, once per turn it enters. A distinct content item is a distinct `content_id`, or its canonical `content_url` where no stable identifier exists (section 4.5). Continued presence within the declared scope is not a further occurrence; re-entry in a later turn is, when the scope is `turn`, and a change of `content_version` is a new occurrence at either scope. An emitter that ingests a content item in chunks MAY emit one grounding event per chunk, preserving the chunk-level hashes of section 6.4; events sharing content identity within one scope describe one occurrence, and consumers count occurrences by deduplicating on content identity and scope, not by counting events. + +3. **Cited** - An output artifact explicitly associates identified source content with a response, claim, passage, quotation, or other output element. Citation is an output-construction relationship, not evidence that the output was delivered. A subset of grounded content is commonly cited, but a citation can also be emitted without a matching grounding event when an agent produces an uncorroborated or hallucinated source association. -4. **Displayed** - Content presented to the end user: a reference (a link, snippet, inline quote, or preview card) or the content itself embedded in the response surface (an iframe, a page rendered by an agentic browser, an embedded media player). Not all citations result in display (e.g., when the agent uses content internally without surfacing the source). + A citation MUST carry a resolvable reference to the source it associates: a `content_url` or a `content_id`. A source association with no resolvable reference is not a citation and MUST NOT be emitted as `content_cited`. Unlike other content events, where the identifier requirement is an application-layer rule (section 5.7.5), for `content_cited` it is enforced by the JSON Schema. - Grounding and display record two different kinds of influence: grounding records that content influenced the agent, display records that it reached the user. As agent experiences evolve beyond the chat window the two diverge - content can shape an answer whose source the user never sees, and an agent can render content to the user that never entered a generation context (see *Departures from the funnel model* below). + One citation occurrence is one distinct association between a source and an output element (or the output artifact, where no element identity exists). Associating the same source with three separate output elements produces three citation events; repeating the same association is not a further occurrence. -5. **Engaged** - The user acted on displayed or cited content: clicked a link, expanded a preview, copied text, shared the response, or directed the agent to act on the content on their behalf (opening the linked page, retrieving more from the source). Engagement connects the preceding events to down-funnel activity: a click-out carries a `ctx_token` that a destination can resolve to the session's click manifest - the content that influenced the response (section 7.1). +4. **Presented** - Content or a source reference was rendered, played, spoken, embedded, or otherwise made perceivable on a recipient-facing surface. Presentation does not assert that a person noticed or attended to it. `presentation_kind` distinguishes source content (an excerpt, an embedded page, played media) from a source reference (such as a link, credit, or card). Not all citations are presented: an output can be stored, suppressed, or passed to another system before delivery. + + Grounding and presentation record different boundary crossings: grounding records entry into a generation context, while presentation records a recipient-facing delivery occurrence. As agent experiences evolve beyond the chat window the two diverge - content can shape an answer whose source is never presented, and an agent can present content that never entered a generation context (see *Departures from the funnel model* below). + + One presentation occurrence is one rendering of content or a source reference on a recipient-facing surface; the event's `id` names that occurrence. Presenting the same artifact again - on a new surface, or in a new delivery - is a new occurrence. + +5. **Engaged** - The recipient or agent performed an observable action on a presentation: clicked a link, expanded a preview, copied text, shared the response, or directed the agent to act on the content. It does not imply attention beyond the reported action. `presentation_id` links the action to the exact presentation occurrence; a click-out can also carry a `ctx_token` that a destination resolves to the click context (section 7.4). + + One engagement occurrence is one observed action on one presentation occurrence. ``` Retrieved (HTTP layer, cacheable) → Grounded (influence layer, per-session or per-turn) → Cited (response layer, per-turn) - → Displayed (UI layer, per-turn) + → Presented (recipient-facing surface, per-turn) → Engaged (user action layer) ``` -Each stage is typically a progressively narrower subset. The ratios between stages are meaningful for potential attribution: +Each stage after retrieval is typically a progressively narrower subset. The ratios between stages are meaningful for potential attribution: - **Retrieval-to-grounding** measures content fetched but not used (irrelevant, stale, or a competing source was preferred) - **Grounding-to-citation** measures content that influenced the response without explicit attribution -- **Citation-to-display** measures content attributed internally but not shown to the user -- **Display-to-engagement** measures interactions where the user did or did not visit the source +- **Citation-to-presentation** measures source associations constructed in output but not made perceivable +- **Presentation-to-engagement** measures observable actions on exact presentation occurrences + +These ratios are computed over reported events and are comparable across emitters only at known coverage (section 5.7.6). #### Departures from the funnel model Three cases break the strict subset model: -- **Displayed without cited.** An agent may display content references (e.g., a "Sources" sidebar) without citing the content in the response text. In this case, a `content_displayed` event exists with no corresponding `content_cited` event. +- **Presented without cited.** An agent may present content references (e.g., a "Sources" sidebar) without semantically associating them with a response element. In this case, a `content_presented` event exists with no corresponding `content_cited` event. - **Cited without grounded.** A hallucinated citation references content the agent never retrieved or loaded into context. The `content_cited` event has no preceding `content_grounded` event. Telemetry consumers SHOULD treat uncorroborated citations (no matching grounding event) as lower-confidence signals. -- **Displayed without grounded.** An agent can render content to the user without that content entering a generation context: an agentic browser showing a page, an embedded video played in the response surface. A `content_displayed` event (typically `display_type: embed`) exists with no corresponding `content_grounded` event. The content influenced the user directly rather than through the model, and the engagement that follows it is consumption the content owner cannot otherwise observe. +- **Presented without grounded.** An agent can present content without that content entering a generation context: an agentic browser showing a page or an embedded video played on a response surface. A `content_presented` event (typically `presentation_kind: content` and `presentation_type: embed`) exists with no corresponding `content_grounded` event. These cases are valid. Emitters SHOULD produce the events that reflect what actually happened, even when the result does not follow the typical funnel ordering. @@ -230,7 +259,7 @@ Conversation turns overlay this lifecycle: 1. **Turn started** - user submits a query 2. **Turn completed** - agent finishes response -A single grounding event with session scope influences all subsequent turns. Citation, display, and engagement events occur within specific turns. +A single grounding event with session scope influences all subsequent turns. Citation, presentation, and engagement events occur within specific turns. ### 4.4 Source roles @@ -247,12 +276,18 @@ The `origin` and `edge` source roles enable content owners to report AI agent tr A marketplace operating as both emitter and telemetry consumer receives telemetry from platforms (as a consumer), resolves content owner identity from `content_id` or `content_url`, and generates per-content-owner usage reports. The marketplace's own `source_role: index` events provide a corroboration layer - it can cross-reference what it served against what platforms reported using. -`content_grounded`, `content_cited`, and `content_displayed` events are reported by the agent (or agent operator) only. These events describe what happened inside the agent or in the agent's user interface, which is not observable from the content owner's infrastructure. +`content_grounded`, `content_cited`, and `content_presented` events are reported by the agent (or agent operator) only. These events describe what happened inside the agent, during output construction, or on a recipient-facing surface, which is not observable from the content owner's infrastructure. A third party that detects source content in a delivered output is corroborating or contradicting the emitter's claims, not observing construction; detection results belong to verification tooling, not to these event types. `content_engaged` events are usually reported by the agent for in-product interactions. For a click-out to a landing page, a downstream marketplace, affiliate network, or destination site MAY report a corroborating `content_engaged` event using `ctx_token` in place of `session_id` (section 7.1). When multiple observers report the same retrieval, events are correlated using the `Content-Telemetry-ID` header (see section 7.2). A retrieval corroborated by multiple sources is a stronger signal than either alone. An uncorroborated origin- or edge-reported retrieval (no matching agent event) may indicate a scraper that does not support the telemetry protocol, or missing header propagation. +#### Supply paths with a non-emitting intermediary + +Telemetry cannot describe a supply path whose middle does not emit. Where an agent obtains content from an intermediary that is not a telemetry participant, the agent's events are the only record, and they carry what the agent was told: usually a URL or identifier supplied by that intermediary. Core provides no way to establish from telemetry alone that the intermediary held the content lawfully, or that the content owner ever served it. + +An origin or edge event correlated by `Content-Telemetry-ID` is what closes the gap, and it exists only where the content owner observed the original request. Where it is absent, a consumer SHOULD treat the path back to the content owner as unestablished rather than infer it from the agent's report. This is a limit of the observation model, not a defect in the emitter: an agent reporting honestly cannot supply evidence about a party it did not observe. + ### 4.5 Content identification Events identify content using at least one of two fields: @@ -286,8 +321,9 @@ Additional content metadata - version, last-modified timestamp, content hash, me | Field | Type | Required | Description | |-------|------|----------|-------------| -| `schema_version` | string | Yes | Schema version (e.g., "0.1") | +| `schema_version` | string | Yes | Schema version as `major.minor` (v1 documents declare "1.0"; see section 12) | | `session_id` | UUID | Yes | Unique session identifier | +| `parent_session_id` | UUID | No | Immediate parent session that delegated work to this session | | `agent_id` | string | No | Responding agent identifier | | `content_scope` | string | No | Opaque content collection identifier (see 5.1.1) | | `manifest_ref` | string | No | Manifest reference (see 5.1.2 and section 8) | @@ -295,8 +331,22 @@ Additional content metadata - version, last-modified timestamp, content hash, me | `ended_at` | datetime | No | Session end (UTC) | | `conformance_level` | string | No | Informational conformance level advertised by the emitter (see section 5.7). Values: `retrieval`, `grounding`, `citation` | | `document_type` | string | No | `"session"` for session documents (see section 7.1 for the standalone event and event batch formats) | +| `data` | object | No | Session-level extension container, including access context (see 5.1.3) | | `events` | Event[] | No | Ordered list of events | +`parent_session_id` links a delegated session to its immediate parent without +requiring an emitter to disclose the agent system's full internal topology. A +child session retains its own `session_id`, events and conformance obligations. +Emitters MAY omit the link when the relationship is unavailable or its disclosure +is not appropriate. Consumers MUST NOT infer that an unlinked session had no +parent. + +The event boundary does not change in a multi-agent system. Content entering a +sub-agent's generation context is grounded in the child session. A source +reference that appears only in the sub-agent's response to its orchestrator is +not thereby a citation or presentation to the end user; those events require the +corresponding relationship or presentation in the recipient-facing output. + #### 5.1.1 Content scope The `content_scope` field is an opaque identifier that groups sessions by their content access context. Implementers define its meaning: @@ -316,25 +366,61 @@ The `manifest_ref` field optionally references a manifest (section 8), identifyi Format: the URL of a manifest served at `/.well-known/content-telemetry.json` under a path the participant controls. +#### 5.1.3 Session data and access context + +Sessions carry an optional `data` object mirroring the event-level `data` field (section 11.1): an extension container for session-scoped metadata. Extensions SHOULD namespace custom fields or use containers documented in this specification, and consumers MUST tolerate unknown fields within it. The session root itself is not an extension point: custom top-level siblings of `events` are not defined by this specification, and consumers are not required to preserve or interpret them. + +One container is defined in core. `access_context` records the context from which the session's access rights derive - the institution, not the individual: + +```json +{ + "schema_version": "1.0", + "session_id": "770e8400-e29b-41d4-a716-446655440000", + "content_scope": "consortium-agreement-4471", + "started_at": "2026-08-13T14:02:10Z", + "data": { + "access_context": { + "identifiers": [ + { "scheme": "ror", "value": "https://ror.org/013meh722" }, + { "scheme": "saml_entity_id", "value": "https://idp.example.ac.uk/shibboleth" } + ] + } + }, + "events": [] +} +``` + +`identifiers` is an array of typed identifiers, each a `scheme` and a `value`. `ror`, `saml_entity_id` and `isni` are the core scheme values; emitters MAY use other schemes and telemetry consumers MUST tolerate unknown ones, as with `media_type` (section 6.1). Access rights can derive through consortia, federated identity and proxies at once, so a session may carry both a SAML entity ID and the ROR ID it maps to. + +`access_context` identifies an institution, never an individual, and like every session field it is a claim by the emitter. The field serves the third-party agent that holds the entitlement and asserts the affiliation to the content owner; where the owner authenticated the session itself (`source_role` of `origin` or `edge`), it already knows the institution. Corroborating an asserted affiliation is verification-layer work, outside core. + +Emitters MUST NOT populate `access_context` unless the governing terms of the relationship require it, and SHOULD pair it with `intent` or `minimal` conversation-turn data (section 5.5): an identified institution combined with query text can come close to identifying an individual at a small subscriber. + ### 5.2 Event | Field | Type | Required | Description | |-------|------|----------|-------------| -| `id` | UUID | No | Unique event identifier (generated by server if not provided) | +| `id` | UUID | For cited/presented | Emitter-assigned unique event identifier; optional on other event types | | `type` | EventType | Yes | Event type (see 5.3) | | `timestamp` | datetime | Yes | Event timestamp (UTC) | | `turn_id` | string | No | Associates this event with a conversation turn (see 5.2.1) | +| `output_id` | string | For cited/presented | Opaque output-artifact identifier joining construction to later delivery | +| `output_element_id` | string | No | Opaque element within `output_id`, such as a passage, media track, caption, link, or card | +| `citation_id` | UUID | No | On `content_presented`, the `id` of the associated citation event; absent for uncited presentations | +| `presentation_id` | UUID | For engaged | On agent-reported `content_engaged`, the `id` of the exact presentation occurrence acted upon. Destination-reported events carrying an envelope `ctx_token` omit it (section 7.4) | +| `ctx_token` | string | No | On agent-reported `content_engaged`, the click token minted for this engagement's presentation, recorded so destination reports join to it (section 7.4) | | `source_role` | SourceRole | No | Who is reporting: `origin`, `edge`, `index`, `agent` (see 4.4) | | `content_telemetry_id` | UUID | No | Correlation ID for cross-observer deduplication (see 7.2) | | `content_url` | string | No | Content URL as fetched or canonical URL | | `content_id` | string | No | Content owner's stable content identifier (see 4.5) | -| `license_ref` | string | No | Reference to the licence under which content was accessed | +| `license_ref` | string | No | Reference to a licence or grant the emitter associates with this event (see 5.2.3) | +| `terms_ref` | string | No | Reference to the governing terms the emitter associates with this event (see 5.2.4) | | `turn` | ConversationTurn | No | Conversation data (for turn events) | | `data` | object | No | Type-specific metadata (see section 6) | #### 5.2.1 Turn association -The `turn_id` field associates content events with a specific conversation turn. Emitters SHOULD set `turn_id` on `content_cited`, `content_displayed`, and `content_engaged` events. Emitters SHOULD also set `turn_id` on `content_grounded` events when `scope` is `turn`. The corresponding `turn_started` and `turn_completed` events SHOULD carry the same `turn_id`. +The `turn_id` field associates content events with a specific conversation turn. Emitters SHOULD set `turn_id` on `content_cited`, `content_presented`, and `content_engaged` events. Emitters SHOULD also set `turn_id` on `content_grounded` events when `scope` is `turn`. The corresponding `turn_started` and `turn_completed` events SHOULD carry the same `turn_id`. `turn_id` is scoped to the session. Format is emitter-defined (sequential integers, UUIDs, or any opaque string). @@ -342,11 +428,25 @@ Content events without a `turn_id` (e.g., `content_grounded` with `scope: sessio #### 5.2.2 Source role -The `source_role` field SHOULD be set on `content_retrieved` events. When multiple systems observe the same retrieval, the `content_telemetry_id` field correlates their events for deduplication. +The `source_role` field MUST be set on `content_retrieved` events (section 5.7.1): without it a consumer cannot tell an agent-reported fetch from an origin- or edge-reported one, or correlate the observers of one retrieval. When multiple systems observe the same retrieval, the `content_telemetry_id` field correlates their events for deduplication. #### 5.2.3 Licence reference -The `license_ref` field connects a telemetry event to the content access licence that authorised it. The format depends on the access protocol: a JWT `jti` claim, a CoMP package ID, or any opaque identifier that both parties can resolve. When present, telemetry consumers can verify that content usage was licensed. +The `license_ref` field associates a telemetry event with a licence or grant the emitter references. The format depends on the access protocol: a JWT `jti` claim, a CoMP package ID, or any opaque identifier that both parties can resolve. + +Core does not resolve, validate or interpret the reference. `license_ref` is part of the emitter's claim about the event: it records which grant the emitter says applied. It does not establish that the grant existed, that it covered this content, that it was valid at the time of use, or that the use was permitted. A consumer that needs any of those has to check the issuer's own records, or use evidence defined outside this specification. + +`license_ref` also does not identify the party whose entitlement was used. Where a publisher issues one grant per subscriber the value may work as a proxy for that subscriber, but only within the issuing publisher's namespace: nothing here requires the value to be typed, stable across sessions, or comparable between emitters. + +#### 5.2.4 Terms reference + +The `terms_ref` field associates a telemetry event with the governing terms under which it is reported: a licence agreement, a standard-form contract, a tariff, a profile's terms, or any other terms document. Like `license_ref`, the value MAY be a public URL or an opaque identifier that both parties can resolve. Nothing requires the terms to be published: per-relationship terms are often confidential, and an opaque identifier resolved privately is a conforming reference. + +`license_ref` records which grant the emitter says applied (5.2.3). `terms_ref` records which terms govern the event's commercial consequences and the emitter's reporting obligations. Either may appear without the other: an access outside any grant carries no `license_ref`, and can still carry the `terms_ref` of the terms that attach consequences to that access. + +Core does not resolve, validate or interpret the reference, and `terms_ref` does not redefine core event semantics. Governing terms select which events a relationship requires and at what coverage (section 5.7.6), together with the cadence, delivery, privacy and reports owed (SCOPE.md); the meaning and occurrence boundary of each event remain those defined in sections 4.3 and 6, whatever `terms_ref` points to. + +A processor that stores, forwards or transforms a document MUST preserve `terms_ref` unchanged and MUST NOT remove or rewrite it. A `terms_ref` value MUST always refer to the same terms: when terms change, the emitter references them with a new value, so events emitted under earlier terms remain resolvable to them. A consumer MUST NOT read the absence of `terms_ref` as a statement that no terms governed the event (the same rule as event absence under coverage, section 5.7.6). ### 5.3 Event types @@ -356,9 +456,9 @@ The `license_ref` field connects a telemetry event to the content access licence |------|-------------|-----------------| | `content_retrieved` | Content fetched from source | `content_url`, `source_role`, `data.media_type` | | `content_grounded` | Content loaded into agent context | `content_url` or `content_id`, `data.scope`, `data.cached` | -| `content_cited` | Content referenced in response | `content_url`, `data.citation_type`, `data.position` | -| `content_displayed` | Content or a reference to it shown to user | `content_url`, `data.display_type` | -| `content_engaged` | User acted on content | `content_url`, `data.engagement_type` (see 6.7) | +| `content_cited` | Output explicitly associates source content with an output element | `id`, `output_id`, `content_url` or `content_id`, `data.citation_type` | +| `content_presented` | Content or a source reference was made perceivable | `id`, `output_id`, `content_url` or `content_id`, `data.presentation_kind`, `data.presentation_type` | +| `content_engaged` | Observable action on an exact presentation | `presentation_id`, `content_url` or `content_id`, `data.engagement_type` (see 6.7) | #### Conversation events @@ -369,7 +469,7 @@ The `license_ref` field connects a telemetry event to the content access licence #### Extension events -The core schema defines content and conversation events. Implementations MAY define additional event types using the `data` field for type-specific metadata. Commerce-specific fields (product identifiers, checkout events) are a planned extension. +The core schema defines content and conversation events. Implementations MAY define additional event types using the `data` field for type-specific metadata. Extension event types SHOULD use namespaced names (for example `com.example.checkout_completed`) so they cannot collide with each other or with future core types. Commerce-specific fields (product identifiers, checkout events) are a planned extension. ### 5.4 Conversation turn @@ -389,7 +489,7 @@ A conversation turn represents one query-response exchange. Turn data is carried | `query_tokens` | integer | No | Query token count | | `response_tokens` | integer | No | Response token count | | `model_id` | string | No | Model identifier | -| `ad_rendered` | boolean | No | Whether advertising was displayed alongside the response | +| `ad_rendered` | boolean | No | Whether advertising was rendered alongside the response | #### 5.4.1 Response modes @@ -415,7 +515,9 @@ These are the recommended values. Platforms with additional product surfaces (co An emitter that populates a conversation turn MUST NOT include a field above that turn's declared `privacy_level` - for example, `query_text` MUST NOT be present when `privacy_level` is `intent` or `minimal`. This restriction is a property of `privacy_level` itself: it applies wherever conversation turns are emitted, independent of the emitter's conformance level. -**Token counts** includes `query_tokens` and `response_tokens`. These are available at all levels because they are needed for token-based counting models and do not reveal user intent or platform strategy. +The session-level `access_context` container (section 5.1.3) does not appear in this table but is subject to the privacy model. It is populated only where governing terms require it, and pairing it with `full` or `summary` turn data is discouraged (section 5.1.3): `privacy_level` controls how much of the query and response is visible, `access_context` identifies whose access rights the session used, and populating both makes re-identification easier. + +**Token counts** includes `query_tokens` and `response_tokens`. These are available at all levels because they are needed for token-based counting models and do not reveal user intent or platform strategy. They carry the same portability limit as `tokens_ingested` (section 6.4): both are measured in the emitter's own tokeniser and are not comparable between agents. Version 1 does not define corresponding turn-level character counts; `chars_ingested` measures source content placed in a generation context, not query or response length. **Response classification** includes `response_type` (e.g., `"recommendation"`, `"explanation"`). Available at `intent` level and above, as it can reveal the nature of the user's query. @@ -441,7 +543,7 @@ These are the core values. Extensions MAY define additional intent category valu Emitters that advertise a standard capability tier use one of three conformance levels. The authoritative declaration lives in the emitter's manifest (section 8). Emitters MAY also include an optional `conformance_level` field on individual session documents; when present it is informational and consumers MUST NOT treat it as a substitute for verifying the manifest's declaration. -Each level is named for the event it adds: a level proves the emitter produces that event and everything below it. These levels describe what an emitter reports, not what a consumer computes from it; attribution - the apportioning of credit across content - is performed by a telemetry consumer at whatever funnel level the parties agree (section 10), and can be computed from grounding alone, without citation. An emitter does not need to reach the Citation level for its telemetry to support attribution. +Each level is named for the event it adds: a level proves the emitter produces that event and everything below it. A level does not assert that every qualifying occurrence was reported (section 5.7.6). These levels describe what an emitter reports, not what a consumer computes from it; attribution - the apportioning of credit across content - is performed by a telemetry consumer at whatever funnel level the parties agree (section 10), and can be computed from grounding alone, without citation. An emitter does not need to reach the Citation level for its telemetry to support attribution. | Level | Events | What it proves | Typical emitter | |-------|--------|----------------|-----------------| @@ -449,31 +551,31 @@ Each level is named for the event it adds: a level proves the emitter produces t | **Grounding** | Above + `content_grounded`, turn events | Content entered the agent's context | Agent with basic instrumentation | | **Citation** | Above + `content_cited` | Content was explicitly referenced in the agent's response | Agent with citation instrumentation | -Display and engagement events are optional lifecycle signals. A Citation emitter SHOULD emit them when applicable (section 5.7.3), but the Citation level proves citation support, not full retrieval-to-engagement coverage. +Presentation and engagement events are optional lifecycle signals outside the Retrieval/Grounding/Citation ladder. A Citation emitter SHOULD emit them when applicable (section 5.7.3), but the Citation level proves citation support, not full retrieval-to-engagement coverage. #### 5.7.1 Retrieval conformance A conforming **Retrieval** emitter MUST: - Set `source_role` on `content_retrieved` events -- Include at least one of `content_url` or `content_id` on every event +- Include at least one of `content_url` or `content_id` on every content event - Set `type` and `timestamp` on every event This level requires no agent cooperation. Content owners can implement it using CDN edge compute (Cloudflare Workers, Fastly Compute, etc.). -Origin-side emitters operating at the CDN edge SHOULD include `bot_category`, `response_status`, and `response_bytes` alongside the required fields. These fields make retrieval events useful for bot classification and volume analysis. Without them, the event confirms a fetch occurred but cannot support attribution correlation. +Origin-side emitters operating at the CDN edge SHOULD include `purpose`, `response_status`, and `response_bytes` alongside the required fields. These fields make retrieval events useful for bot classification and volume analysis. Without them, the event confirms a fetch occurred but cannot support attribution correlation. #### 5.7.2 Grounding conformance A conforming **Grounding** emitter MUST satisfy Retrieval requirements and also: - Produce sessions with `schema_version`, `session_id`, `agent_id`, and `started_at` -- Emit `content_grounded` events with `data.scope` +- Emit `content_grounded` events with `data.scope` (schema-enforced; section 6.4) - Include at least one of `content_url` or `content_id` on every content event - Emit `turn_started` and `turn_completed` events with `privacy_level` - Restrict conversation turn fields to the declared `privacy_level` (section 5.5) -A Grounding emitter SHOULD include `data.tokens_ingested` and `data.cached` on grounding events. +A Grounding emitter SHOULD include `data.chars_ingested` and `data.cached` on grounding events, and MAY add `data.tokens_ingested` alongside them (section 6.4). Emitters using standalone event delivery (section 7.1) MUST include `agent_id`, `started_at`, and either `session_id` or, for click-out engagement events, `ctx_token` on the standalone event envelope to satisfy Grounding conformance. @@ -481,21 +583,22 @@ Emitters using standalone event delivery (section 7.1) MUST include `agent_id`, A conforming **Citation** emitter MUST satisfy Grounding requirements and also: -- Emit `content_cited` events with `data.citation_type` +- Emit `content_cited` events with `id`, `output_id`, `data.citation_type`, and a non-null `content_url` or `content_id` (schema-enforced; section 6.5) The privacy-level field restriction (section 5.5) applies to Citation emitters as it does to any emitter producing conversation turns; it is inherited through the Grounding requirements above. A Citation emitter SHOULD: -- Emit `content_displayed` and `content_engaged` events when applicable +- Emit `content_presented` and `content_engaged` events when applicable - Include `data.position` on citation events -- Include `data.display_type` on display events +- Include `output_element_id` when the cited or presented element has a stable identity +- Include `citation_id` on a presentation of a cited source association #### 5.7.4 Telemetry consumers A conforming **telemetry consumer** MUST: -- Accept sessions with any `schema_version` that shares the same major version. During the preview period (0.x), consumers MUST accept sessions with the exact same minor version (e.g., a 0.1 consumer accepts 0.1 only). The major-version compatibility rule takes effect from 1.0.0 onward. +- Accept documents declaring any `schema_version` with the same major version as the one the consumer implements. `schema_version` is `major.minor` (section 12): v1.0 documents declare `"1.0"`, and the v1.0 schemas accept that value only. Each minor version publishes its own schemas. A consumer implementing 1.y validates a document declaring 1.x, x ≤ y, against the 1.x schemas, and a document declaring a later minor against the latest schemas it implements, tolerating the optional fields that minor added. A v1 consumer MUST reject documents declaring `"0.1"`: v0.1 is a different wire version, not a compatible minor. Conversely, a v0.1 consumer following the preview rule (a 0.x consumer accepts only the exact same minor version, so a 0.1 consumer accepts 0.1 only) rejects documents declaring `"1.0"`. - Tolerate unknown fields without error - Tolerate events from any conformance level - Accept the session-document, standalone-event, and event-batch delivery formats, reconstructing sessions from standalone events and event batches where needed (see section 7.1) @@ -504,16 +607,42 @@ A conforming **telemetry consumer** MUST: The JSON Schema (`telemetry-session.json`) validates structure and types but cannot express every conformance rule. The following are normative requirements verified at the application layer, not by schema validation: -- At least one of `content_url` or `content_id` MUST be present on every content event (section 4.5). +- At least one of `content_url` or `content_id` MUST be present on every content event (section 4.5). For `content_cited` events this requirement is additionally enforced by the JSON Schema, which rejects an event whose reference is absent or null (section 6.5). - An event MUST carry either `session_id` or `ctx_token` at Grounding conformance and above (section 7.1). - Conversation-turn fields MUST NOT exceed the turn's declared `privacy_level` (section 5.5). - The conformance-level requirements (sections 5.7.1 to 5.7.3) are cumulative. +- When `content_grounded.data.provenance` is `agent_fetched`, `data.cached` MUST be `false`; when it is `agent_cached`, `data.cached` MUST be `true` (section 6.4). +- `content_grounded.data.content_fingerprint` MUST NOT contain `preserved_in_output`; v1 defines no output-side reuse reporting (sections 6.4 and 12.1). +- `source_role` MUST be present on every `content_retrieved` event (sections 5.2.2, 5.7.1). +- Fields scoped to an event type MUST NOT appear on other types: `presentation_id` and the event-level `ctx_token` only on `content_engaged`; `citation_id` only on `content_presented`; `turn` only on `turn_started` and `turn_completed` (section 5.2). +- Within a session document, event `id` values MUST be distinct; a `content_engaged.presentation_id` MUST reference a `content_presented` event, and a `citation_id` a `content_cited` event, that identifies the same content - where both events carry `content_id` the values MUST be equal, and likewise for `content_url` (sections 6.6, 6.7); one event-level `ctx_token` MUST NOT appear on engagements bound to two different presentations (section 7.4.1). +- An envelope `ctx_token` (section 7.1) MUST accompany `content_engaged` events only. +- Manifests: `domains` MUST appear only on a manifest served at the domain root, and `telemetry.ctx_resolution` only on a manifest declaring the `agent` or `platform` role (sections 8.5, 8.6). The `tests/` directory provides an informative reference suite for these rules. A consumer that receives a privacy-violating turn (e.g., `query_text` present at `minimal` level) SHOULD strip the offending fields rather than reject the document carrying them. +#### 5.7.6 Occurrence, qualifying events and coverage + +Each core event type has the meaning and occurrence boundary defined in sections 4.3 and 6. Profiles, deployment configurations and governing terms MUST NOT redefine them. A relationship that needs a different assertion defines a namespaced extension event (sections 5.3 and 11.1); it does not reuse a core type with altered semantics. + +An occurrence is **qualifying** for an emitter when it satisfies the core definition and occurrence boundary of its event type and falls within the relationship scope the emitter reports under - the content, domains or relationships selected by the applicable governing terms or deployment configuration. + +A conformance level (sections 5.7.1 to 5.7.3) does not assert that every qualifying occurrence was reported. Reporting coverage is a separate, explicit declaration, stated as one of four modes: + +- **complete** - every qualifying occurrence is emitted +- **sampled** - qualifying occurrences are emitted under a stated sampling rule +- **aggregated** - qualifying occurrences are reported only through a stated aggregation rule +- **selected** - only qualifying occurrences satisfying a further stated condition are emitted + +A coverage declaration states its mode together with the relationship scope it applies over; both MUST be disclosed to the receiving party. The rule or condition for `sampled`, `aggregated` and `selected` MUST be objectively decidable from information available at emission time and MUST NOT depend on the emitter's discretion at the moment of emission. An emitter reporting under governing terms that state a coverage mode MUST report at that mode, and an emitter MUST NOT declare or describe its reporting as `complete` for an event type unless every qualifying occurrence is emitted. A consumer MUST NOT treat the absence of an event as evidence that no occurrence happened except where complete coverage applies. + +An emitter MAY declare its coverage modes machine-readably in its manifest (`telemetry.coverage`, section 8.5); a manifest declaration is subject to the same rules, and where governing terms and a manifest declaration conflict, the governing terms take precedence for the relationships they cover. + +Whether an emitter's reporting in fact met its declared coverage is the completeness question of SCOPE.md's conformance list: it is answered by verification and audit mechanisms outside core, not by the declaration itself. + ## 6. Data profiles -The `data` field on events carries type-specific metadata. These profiles document the recommended fields by event type and source role, in lifecycle order. None are required, but emitting them enables richer attribution. +The `data` field on events carries type-specific metadata. These profiles document the recommended fields by event type and source role. None are required except where a section states otherwise - `scope` in 6.4, `citation_type` in 6.5, `presentation_kind` and `presentation_type` in 6.6, each enforced by the JSON Schema - but emitting them enables richer attribution. ### 6.1 Retrieved content metadata (`content_retrieved`) @@ -522,11 +651,16 @@ When the reporter is the agent (`source_role: agent`), the following fields are | Field | Type | Description | |-------|------|-------------| | `media_type` | string | Content medium: `text`, `image`, `video`, `audio` (see below) | +| `content_depth` | string | Depth of the content record reached: `metadata`, `abstract`, `full` (see below) | `media_type` on retrieval events allows content owners to see what types of content are being fetched, independent of whether those retrievals result in grounding or citation. Defaults to `text` when absent. `text`, `image`, `video`, and `audio` are the core values. Emitters MAY use custom string values for media outside the core set (for example `3d` or `dataset`). Telemetry consumers MUST tolerate unknown `media_type` values. This rule applies to `media_type` on every event type that carries it (sections 6.4, 6.5, 6.6). +`content_depth` records how much of the content record the retrieval reached: `metadata` for a bibliographic or descriptive record only, `abstract` for an abstract or summary record, `full` for the full content record. These are the core values; emitters MAY use custom values and telemetry consumers MUST tolerate unknown ones. Where entitlement gates depth, a retrieval that reached only an abstract and a retrieval of full text are otherwise indistinguishable at the retrieval layer. Depth records what was reachable at retrieval, independent of what portion later entered a generation context. + +Although listed in the agent profile, `content_depth` applies to `content_retrieved` events from any `source_role`. The origin that served the response knows the depth authoritatively, and origin and edge reporters SHOULD include it alongside their fields in sections 6.2 and 6.3 where entitlement gates depth. + ### 6.2 Edge enrichment (`content_retrieved` + `source_role: edge`) CDN and edge network integrations SHOULD include these fields: @@ -534,7 +668,7 @@ CDN and edge network integrations SHOULD include these fields: | Field | Type | Description | |-------|------|-------------| | `user_agent` | string | Request User-Agent header | -| `bot_category` | string | Edge platform's bot classification (see below) | +| `purpose` | string | Purpose of the access, as classified by the reporting party (see below) | | `bot_name` | string | Recognised bot family parsed from the User-Agent (e.g., `Claude-User`, `GPTBot`, `Perplexity-User`) | | `verified` | boolean | Whether the bot identity was cryptographically verified | | `cache_status` | string | Edge cache result: `hit`, `miss`, `bypass`, `dynamic` | @@ -544,43 +678,77 @@ CDN and edge network integrations SHOULD include these fields: | `asn` | integer | Client AS number | | `asn_org` | string | Client AS organisation name | | `country` | string | ISO 3166-1 alpha-2 country code | -| `ip_hash` | string | SHA-256 of client IP (`sha256:{hex}`) | -#### Bot categories +#### Access purpose -The `bot_category` field carries the edge platform's classification of the requesting bot. Recommended values: +The `purpose` field carries the reporting party's classification of what the access was for. It is an open enum: these are the core values, emitters MAY use custom values, and telemetry consumers MUST tolerate unknown ones. It classifies the access, not the organisation - who is reporting stays in `source_role` (section 4.4). -| Value | Description | Fastly signal | Cloudflare signal | -|-------|-------------|---------------|-------------------| -| `training` | Crawling for model training | `AI-CRAWLER` | `AI Crawler` | -| `inference` | Fetching at query time (RAG) | `AI-FETCHER` | `AI Assistant` | -| `search` | AI search indexing | - | `AI Search` | +| Value | Description | +|-------|-------------| +| `training` | Crawling for model training | +| `inference` | Fetching at query time to inform a response | +| `search` | AI search indexing | +| `advertising` | Access to derive advertising signals (contextual classification, brand safety, campaign targeting) | -The `inference` category is where content attribution is most relevant - there is a user, a query, and a session behind the retrieval. `training` crawls have no session context. `bot_category` can distinguish training crawls from inference fetches, but training-specific telemetry is out of scope for this specification (see section 1.3). Edge platforms map their native classification to these values. +The `inference` purpose is where content attribution is most relevant - there is a user, a query, and a session behind the retrieval. `training` crawls have no session context. `purpose` can distinguish training crawls from inference fetches, but training-specific telemetry is out of scope for this specification (see section 1.3). Edge platforms map their native bot classifications to these values; the mappings for common platforms are informative and collected in Annex C. -Emitting a `training`-category `content_retrieved` event is permitted but non-attributable - there is no session, grounding, or citation to follow it. An edge emitter can report these events through its normal pipeline and need not special-case or suppress them. +Emitting a `training`-purpose `content_retrieved` event is permitted but non-attributable - there is no session, grounding, or citation to follow it. An edge emitter can report these events through its normal pipeline and need not special-case or suppress them. ### 6.3 Origin enrichment (`content_retrieved` + `source_role: origin`) | Field | Type | Description | |-------|------|-------------| | `user_agent` | string | Request User-Agent header | -| `ip_hash` | string | SHA-256 of client IP | | `response_status` | integer | HTTP response status code | ### 6.4 Grounding data (`content_grounded`) | Field | Type | Description | |-------|------|-------------| -| `scope` | string | Influence scope: `session` or `turn` (see below) | +| `scope` | string | Required. Influence scope: `session` or `turn` (see below) | | `cached` | boolean | Content served from agent-side cache rather than a live fetch | -| `tokens_ingested` | integer | Token count of content placed in the generation context (see below) | +| `provenance` | string | How content reached the context: `agent_fetched`, `agent_cached`, or `third_party_sourced` (see below) | +| `chars_ingested` | integer | Character count of content placed in the generation context (see below) | +| `tokens_ingested` | integer | Token count of the same content, supplementary (see below) | | `content_version` | string | Content version identifier (ETag, revision ID, CMS version) | | `content_last_modified` | datetime | When the content was last modified at source | | `content_hash` | string | SHA-256 of the content as ingested (`sha256:{hex}`) | | `media_type` | string | Content medium: `text`, `image`, `video`, `audio` (open vocabulary, see 6.1) | +| `content_fingerprint` | object | Agent-reported detection of a fingerprint or provenance signal in the grounded content (see below) | -`tokens_ingested` counts tokens actually placed in the generation model's context. For chunked retrieval, count only the tokens used, not the full source document. The token count uses the generation model's tokeniser (the model identified in `model_id` on the corresponding `turn_completed` event), not the retrieval or embedding model's tokeniser. +Both fields measure the content actually placed in the generation model's context. For chunked retrieval, count only the portion used, not the full source document. + +`chars_ingested` counts Unicode code points in the exact text placed in context. Count the string as ingested: an emitter MUST NOT apply Unicode normalisation solely to calculate this field. It is the portable measure: two emitters that ingest the same code-point sequence agree, so a content owner can compare volumes across agents and over time without knowing which model produced the number. Different normalised representations remain different ingested sequences and may therefore produce different counts. + +`tokens_ingested` counts the same content in the generation model's tokeniser (the model identified in `model_id` on the corresponding `turn_completed` event), not the retrieval or embedding model's tokeniser. It is supplementary. Token counts are model-specific, change when a vendor revises a tokeniser, and are not comparable between agents, so a consumer cannot aggregate them across emitters or treat a difference as a difference in volume. Emitters SHOULD send `chars_ingested` where they send `tokens_ingested`, and consumers that receive only token counts SHOULD record which model produced them. At `minimal` privacy the turn carries no `model_id` (section 5.5); token counts reported at that level name no tokeniser, and consumers SHOULD NOT compare them with counts from any other emitter or model. + +#### Provenance and content fingerprints + +`provenance` describes the delivery path by which the grounded representation reached the agent: + +| Value | Description | +|-------|-------------| +| `agent_fetched` | The agent obtained the representation directly from the publisher or from publisher-authorised origin or edge infrastructure for this session | +| `agent_cached` | The agent reused a representation it had obtained before this session | +| `third_party_sourced` | The representation reached the agent through an intermediary rather than through a direct publisher-authorised retrieval by the agent in this session | + +The field describes delivery path, not evidence quality. An emitter declaring Grounding or Citation conformance SHOULD include it when the path is known. It remains optional because an agent may not be able to distinguish its own earlier fetch from intermediary delivery. Consumers MUST NOT infer a value when it is absent. + +Emitters MUST keep `provenance` and `cached` consistent: `agent_fetched` requires `cached: false`, and `agent_cached` requires `cached: true`. `third_party_sourced` leaves `cached` unconstrained because an intermediary-sourced representation may be used immediately or cached by the agent before grounding. + +`content_fingerprint` contains: + +| Field | Type | Required | Description | +|-------|------|----------|-------------| +| `scheme` | string | Yes | Open identifier for the fingerprint or provenance scheme checked | +| `detected` | boolean | Yes | Emitter claim that the scheme's signal was found in the exact grounded representation | +| `value` | string | No | Scheme-defined fingerprint or identifier value, when the scheme produces one | + +`detected` reports a grounding-time observation by the emitter. It does not establish that the signal is authentic, identify who applied it, prove that the content was used later in the output, or raise the evidentiary status of the grounding event. Those questions require profile-defined evidence and consumer trust policy outside the core schema. + +The fingerprint is a grounding-time claim only: v1 defines no field asserting that a fingerprinted signal was preserved in the output (the pre-release `preserved_in_output` field is withdrawn; section 12.1). A consumer MAY compare a grounding fingerprint with evidence about the output gathered outside this specification, but the two remain separate assertions about separate lifecycle stages. + +`scheme` is an open identifier. Emitters SHOULD use a globally collision-resistant value. Core does not register schemes, interpret `value`, or assign capabilities or evidence status from a scheme identifier. A profile MAY define scheme-specific processing rules. #### Grounding scope @@ -591,6 +759,8 @@ Emitting a `training`-category `content_retrieved` event is permitted but non-at For session-scoped grounding, the number of turns influenced is derivable from the session's `turn_started` events following the grounding event. This avoids redundant per-turn grounding events for content that persists across responses. +`scope` is required on every `content_grounded` event and the JSON Schema enforces it: the occurrence boundary of section 4.3 and every counting model in section 10 depend on knowing whether a grounding informed one response or the rest of the session. + #### Agent architecture and the grounding boundary The grounding event marks the point where content enters the generation model's context - the boundary where content can directly influence the model's output text. Content used only for retrieval selection (embedding similarity search, re-ranking, query routing) without entering the generation context is not grounded. @@ -618,7 +788,7 @@ The `cached` field distinguishes live fetches from cached reuse. A live fetch pr Telemetry consumers may weight cached and live groundings differently. An agent may cache an article for days or weeks, grounding it in multiple sessions from a single retrieval. A single retrieval produces one `content_retrieved` event but potentially many `content_grounded` events across subsequent sessions. -Agents SHOULD preserve the `license_ref` from the original retrieval when emitting cached grounding events. Without this, telemetry consumers cannot link cached usage to the licence that authorised the original access. +Agents SHOULD preserve the `license_ref` from the original retrieval when emitting cached grounding events. Without this, telemetry consumers cannot link cached usage to the grant referenced at the original access. #### Freshness and verification @@ -630,7 +800,7 @@ Agents SHOULD preserve the `license_ref` from the original retrieval when emitti | Field | Type | Description | |-------|------|-------------| -| `citation_type` | string | How content was used: `direct_quote`, `paraphrase`, `reference`, `contradiction`, `unclassified` | +| `citation_type` | string | Required. How content was used: `direct_quote`, `paraphrase`, `reference`, `contradiction`, `unclassified` | | `media_type` | string | Content medium: `text`, `image`, `video`, `audio` (open vocabulary, see 6.1) | | `excerpt_tokens` | integer | Token count of the excerpt used | | `excerpt_chars` | integer | Character count of the excerpt used | @@ -639,9 +809,11 @@ Agents SHOULD preserve the `license_ref` from the original retrieval when emitti | `content_hash` | string | SHA-256 matching the corresponding `content_grounded` event (`sha256:{hex}`). When the agent chunked the source, this is the chunk hash, not the full document hash. | | `url_verified` | boolean | Whether the cited URL was verified to resolve to matching content | +A citation MUST carry a resolvable source reference: a non-null `content_url` or `content_id` at the event level. This is what distinguishes a citation from vague attribution - the credit names a source that owner routing (section 7.3) can resolve. The JSON Schema enforces this for `content_cited` events; an association the emitter cannot resolve to a URL or identifier is not reportable as a citation. This is stricter than the application-layer identifier rule that applies to content events generally (section 5.7.5). `citation_type` is likewise required and schema-enforced; an emitter that cannot classify a citation uses `unclassified` rather than omitting the field. + `media_type` identifies the content medium. Defaults to `text` when absent. -`excerpt_tokens` is the agent-native measurement. `excerpt_chars` provides the same information in a unit familiar to content owners and licensors. Emitters SHOULD include both when available. +`excerpt_chars` counts Unicode code points in the cited excerpt under the same counting rule as `chars_ingested` (section 6.4): no normalisation applied solely for counting. It is the portable primary measurement, comparable across emitters and stated in a unit familiar to content owners and licensors. `excerpt_tokens` counts the same excerpt in the generation model's tokeniser; it is the agent-native supplementary measurement, carrying the same portability limits as `tokens_ingested`. Emitters SHOULD send `excerpt_chars` where they send `excerpt_tokens`. `excerpt_hash` is the SHA-256 of the excerpt text as it appears in the agent's response - the exact string the agent produced, not the source text it was derived from. For `direct_quote` citations, a matching hash against the source content confirms verbatim fidelity. For `paraphrase` citations, a non-matching hash is expected; verification tooling can use the hash to confirm which specific excerpt was cited and compare it against known source passages. Emitters SHOULD include `excerpt_hash` when `excerpt_tokens` or `excerpt_chars` is present. The hash uses the same `sha256:{hex}` format as `content_hash`. @@ -649,18 +821,23 @@ The `contradiction` type supports negative attribution: content that was retriev The `unclassified` value for `citation_type` indicates the agent did not classify this citation. The `unclassified` value for `position` indicates the agent did not determine the prominence of the citation. Emitters SHOULD use `unclassified` rather than forcing a classification when the agent cannot confidently determine the citation type or position. +A citation of translated or cross-language content is a `paraphrase` (or `reference`) citation like any other: v1 defines no language fields and no match-confidence claim. `excerpt_hash` identifies the excerpt the agent produced, not the source passage, so a hash that does not match the source is expected for translated and paraphrased citations and does not indicate non-use; establishing which source passage a translated excerpt derives from is verification-layer work outside this specification. + `url_verified` indicates whether the agent confirmed that the cited URL resolves to content matching the citation. When `false` or absent, the citation may reference a hallucinated or outdated URL. `url_verified` MAY be set asynchronously after response generation. Platforms that batch-verify URLs periodically rather than per-request are conforming. A value of `false` indicates the URL was not verified, not that verification failed. When `content_hash` is absent or does not match any grounding event's hash (for example, because the agent re-chunked content between grounding and citation), consumers SHOULD fall back to matching on `content_url` or `content_id`, accepting that the correlation may be imprecise when the same content appears in multiple grounding events. -### 6.6 Display data (`content_displayed`) +### 6.6 Presentation data (`content_presented`) | Field | Type | Description | |-------|------|-------------| -| `display_type` | string | How the content or a reference to it was presented (see below) | -| `media_type` | string | Content medium: `text`, `image`, `video`, `audio` (open vocabulary, see 6.1). Defaults to `text` when absent. Most useful on `embed` displays (an embedded video reports `media_type: video`). | +| `presentation_kind` | string | What was made perceivable: `content` or `source_reference` | +| `presentation_type` | string | How it was made perceivable (see below) | +| `media_type` | string | Medium made perceivable: `text`, `image`, `video`, `audio` (open vocabulary, see 6.1). Defaults to `text` when absent. | + +`presentation_kind: content` means source content itself, a bounded excerpt, or a derived representation was made perceivable. It does not claim that the whole source was reproduced. `presentation_kind: source_reference` means a credit, identifier, link, card, or other reference to the source was made perceivable. This distinction is independent of modality: a spoken credit is a source reference; played source audio is content. -#### Display types +#### Presentation types | Value | Description | |-------|-------------| @@ -670,12 +847,17 @@ When `content_hash` is absent or does not match any grounding event's hash (for | `card` | Rich preview card (title, description, image) | | `detail_view` | Expanded or full-content presentation within the agent's own interface | | `embed` | Source content rendered in the response surface: an iframe, a page rendered by an agentic browser, an embedded media player | +| `spoken_credit` | Source reference spoken in an audio output or assistive surface | -The first five values present a reference or excerpt within the agent's interface; `embed` presents the source content itself. An `embed` display can occur without a grounding event when the content never entered a generation context (section 4.3, *Departures from the funnel model*). +`presentation_kind`, rather than `presentation_type`, determines whether the occurrence carries source content or a source reference. For example, a snippet may be an attributed source reference or an uncredited content excerpt. An embed can occur without a grounding event when the content never entered a generation context (section 4.3, *Departures from the funnel model*). -These are the core values. Platforms with additional presentation surfaces MAY use custom string values. Telemetry consumers MUST tolerate unknown `display_type` values. +These are the core values. Platforms with additional presentation surfaces MAY use custom string values. Telemetry consumers MUST tolerate unknown `presentation_type` values. -When a session includes `content_displayed` events but no subsequent `content_engaged` events, the user saw a content reference but did not interact with it. Whether this pattern is meaningful depends on the commercial agreement - a per-citation deal may not care about clickthrough, while a traffic-based deal will. This pattern is only detectable from platform-reported `content_displayed` and `content_engaged` events. Retrieval is the only event stage observable from the CDN edge. +A presentation identifies whole content items. Finer-grained portion references for time-based and spatial media - time ranges, regions, segments - are not defined in this version. + +Each presentation event MUST have an `id` and `output_id`. When it presents a citation, `citation_id` references that `content_cited` event's `id`, and the two events identify the same content; an uncited presentation omits `citation_id`. Repeated presentations of the same source or output element MUST receive distinct event IDs - event `id` values are unique within a session document. This allows a later `content_engaged.presentation_id` to identify the exact surface occurrence rather than matching only by URL. + +When a session includes `content_presented` events but no subsequent `content_engaged` events, the telemetry establishes only that content or a reference was made perceivable and no reported interaction followed. It does not establish human attention. Whether this pattern is meaningful depends on the governing terms. Retrieval remains the only lifecycle stage observable from the CDN edge. ### 6.7 Engagement data (`content_engaged`) @@ -683,7 +865,7 @@ When a session includes `content_displayed` events but no subsequent `content_en |-------|------|-------------| | `engagement_type` | string | Type of interaction (see below) | -The content URL is identified by the event-level `content_url` field (section 5.2), not duplicated in `data`. +The content URL is identified by the event-level `content_url` field (section 5.2), not duplicated in `data`. Every agent-reported engagement MUST carry `presentation_id`, referencing the exact `content_presented.id` on which the action occurred, and identifies the same content as that presentation (section 5.7.5). Matching on URL alone is insufficient because the same source reference can be presented more than once. A destination-reported engagement carries `ctx_token` on its envelope instead: the destination cannot know the presentation UUID, and the telemetry consumer restores the binding from the token at resolution (section 7.4). #### Engagement types @@ -699,9 +881,23 @@ These are the core values. Extensions MAY define additional engagement actions - `agent_navigate` is the agent-mediated counterpart of a click: the user reached the source through the agent rather than through a browser. Consumers measuring traffic SHOULD count it alongside `link_click`, distinguishing the two where the commercial agreement does. -`link_click` is the primary signal for clickthrough rate calculation. Telemetry consumers can derive per-content-owner and aggregate clickthrough rates from the ratio of `link_click` engagements to `content_displayed` events for the same `content_url`. +An action that touches several presentations at once - a `share` of a response containing three source cards - is one engagement occurrence per presentation shared (section 4.3), each bound to its own `presentation_id`; a surface that shares a single card reports one. An `agent_navigate` to a URL the recipient supplied, which no presentation made perceivable, is not an engagement: it begins with a `content_retrieved` event like any other fetch. + +`link_click` is the primary signal for clickthrough rate calculation. Telemetry consumers can derive per-content-owner and aggregate clickthrough rates from `link_click` engagements and link presentations, joining each engagement through `presentation_id` rather than URL alone. + +A `link_click` or `agent_navigate` engagement reported from the landing page after a click-out crosses a trust boundary. Such events carry a `ctx_token` in place of `session_id`, which the telemetry consumer resolves to the click context (see section 7.4). The agent-authored engagement itself reaches the engaged content's owner through owner-scoped routing whether or not the token survived the redirect chain (section 7.4.5). + +### 6.8 Evidence references (any content event) + +The `data.evidence` field MAY appear on any content event: an array of profile-defined evidence references attached to the event's claim. -A `link_click` or `agent_navigate` engagement reported from the landing page after a click-out crosses a trust boundary. Such events carry a `ctx_token` in place of `session_id`, which the telemetry consumer resolves to the originating session's click manifest (see section 7.1). +| Field | Type | Required | Description | +|-------|------|----------|-------------| +| `scheme` | string | Yes | Open identifier for the evidence scheme, following the same rules as `content_fingerprint.scheme` (section 6.4) | +| `ref` | string | No | URI of a detached evidence artefact, resolvable independently of the event | +| `digest` | string | No | Digest binding the reference to the artefact's bytes (`sha256:{hex}`) | + +Core defines the slot and nothing more. It does not interpret entries, register schemes, or assign evidentiary status: an event remains a claim by its emitter (SCOPE.md), and the presence of evidence entries raises no event's status by itself. Which schemes a consumer accepts, and what a verified entry establishes, is consumer trust policy defined in an evidence profile outside core. Consumers MUST tolerate unknown schemes and unknown fields within entries. A detached reference - a `ref` with a `digest` - is admitted deliberately, so evidence can remain independently verifiable after the fact without travelling inline. ## 7. Transport @@ -717,12 +913,12 @@ The schema supports three delivery formats: **Event batch.** Multiple events sharing one session context, delivered together. The envelope carries the same fields as a standalone event, with an `events` array in place of the single `event`. Suitable for emitters that buffer events and flush periodically: edge platforms aggregating detections across requests, or SDKs batching events within a session. -A standalone event carries `document_type`, `schema_version`, and optionally `session_id` alongside the event fields. The `document_type` field distinguishes standalone events from session documents: +A standalone event carries `document_type`, `schema_version`, and optionally `session_id` and `parent_session_id` alongside the event fields. The `document_type` field distinguishes standalone events from session documents: ```json { "document_type": "event", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "550e8400-e29b-41d4-a716-446655440000", "event": { "type": "content_retrieved", @@ -731,7 +927,7 @@ A standalone event carries `document_type`, `schema_version`, and optionally `se "content_telemetry_id": "770e8400-e29b-41d4-a716-446655440300", "content_url": "https://www.ft.com/content/abc123", "data": { - "bot_category": "inference", + "purpose": "inference", "cache_status": "miss", "response_status": 200 } @@ -739,12 +935,12 @@ A standalone event carries `document_type`, `schema_version`, and optionally `se } ``` -An event batch carries the same envelope fields with `"document_type": "event_batch"` and an `events` array. Envelope-level fields (`session_id`, `ctx_token`, `agent_id`, `started_at`) apply to every event in the batch; events belonging to different sessions MUST be delivered in separate batches or as session documents. +An event batch carries the same envelope fields with `"document_type": "event_batch"` and an `events` array. Envelope-level fields (`session_id`, `parent_session_id`, `ctx_token`, `agent_id`, `started_at`, `manifest_ref`) apply to every event in the batch; events belonging to different sessions MUST be delivered in separate batches or as session documents. ```json { "document_type": "event_batch", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "550e8400-e29b-41d4-a716-446655440000", "events": [ { @@ -754,10 +950,15 @@ An event batch carries the same envelope fields with `"document_type": "event_ba "content_url": "https://www.ft.com/content/abc123" }, { + "id": "770e8400-e29b-41d4-a716-446655440301", "type": "content_cited", "timestamp": "2026-01-15T10:30:04Z", + "output_id": "response:1", "source_role": "agent", - "content_url": "https://www.ft.com/content/abc123" + "content_url": "https://www.ft.com/content/abc123", + "data": { + "citation_type": "reference" + } } ] } @@ -767,9 +968,7 @@ Session documents use `"document_type": "session"`. When `document_type` is abse For origin-side emitters at Retrieval conformance level, `session_id` MAY be omitted when the content owner has no session context. Telemetry consumers correlate these events with agent-reported sessions using the `content_telemetry_id` field. -For `content_engaged` events emitted from a landing page after a click-out (typically by a content marketplace, affiliate network, or destination site), `session_id` MAY be replaced by a `ctx_token` field that carries an opaque click-token issued by the originating agent. Telemetry consumers resolve the token to the owning session. This lets a downstream observer report a corroborating engagement event without sharing the session UUID across trust boundaries. An event MUST carry either `session_id` or `ctx_token` at Grounding conformance and above. - -**ctx_token resolution.** A telemetry consumer that supports `ctx_token` resolution exposes, for a resolved token, the **click manifest**: the set of `content_grounded`, `content_cited`, and `content_displayed` events belonging to the resolved session, identifying every source that informed the response that produced the click. The manifest is gated by the resolved session's `privacy_level` and by consent. A consumer MUST return the manifest only when the issuing agent has opted in to sharing sessions via click tokens; when the agent opt-in is absent, the consumer MUST NOT disclose the manifest. Within a returned manifest, a source MUST appear only when its content owner has opted in to being visible in click-token lookups; the consumer MUST withhold the events of any content owner whose opt-in is absent while returning the remainder of the manifest. A resolution response MUST NOT include the resolved session's raw `session_id`: the token exists so that the session UUID never crosses the trust boundary, and a resolution response that returned it would undo that. The mechanism by which an agent and a content owner record these opt-ins is operator-defined; the consent gate is normative. +For `content_engaged` events emitted from a landing page after a click-out (typically by a content marketplace, affiliate network, or destination site), `session_id` MAY be replaced by a `ctx_token` field that carries an opaque click token issued by the originating agent. This lets a downstream observer report a corroborating engagement event without sharing the session UUID across trust boundaries. An event MUST carry either `session_id` or `ctx_token` at Grounding conformance and above, and an envelope `ctx_token` accompanies `content_engaged` events only: an envelope carrying any other event type carries `session_id`. Token issuance, carriage, binding, resolver discovery, and the resolution response are defined in section 7.4. The primary schema (`telemetry-session.json`) validates session documents. A standalone event envelope schema (`telemetry-event.json`) validates the event delivery format, and a batch envelope schema (`telemetry-event-batch.json`) validates the event batch format. All three schemas share the `TelemetryEvent` definition. @@ -777,6 +976,8 @@ The primary schema (`telemetry-session.json`) validates session documents. A sta An agent emitter that uses standalone events or event batches for streaming delivery and wants to achieve Grounding or Citation conformance MUST include the optional `agent_id` and `started_at` fields on the envelope. Each envelope MUST also carry `session_id`, except for click-out engagement events where `ctx_token` is used instead. Consumers reconstruct the session from the stream of envelopes sharing the same `session_id`, or resolve the owning session from `ctx_token`. +The optional `manifest_ref` field is available on standalone event and event batch envelopes, mirroring the session-level field (section 5.1.2). It identifies the emitter's manifest where no session document carries one. An event delivered standalone in support of settlement, audit or other obligations under governing terms (section 5.2.4) SHOULD carry `manifest_ref`, since it is the only envelope field that names the manifest - and so the domain - under which the emitter claims to report; verifying that claim uses the manifest mechanisms of section 8. + Origin-side emitters (source role `origin` or `edge`) are not expected to achieve Grounding conformance and do not need these fields. Telemetry consumers MUST accept all three delivery formats, reconstructing sessions from standalone events and event batches where needed. @@ -822,16 +1023,98 @@ Two deployment patterns are common: Any party may operate a consumer: an agent operator, a licensing intermediary, or an independent third party offering it as a service. Both patterns above consume the same session format. The telemetry consumer is responsible for domain resolution, content owner filtering, and access control. The spec does not mandate a specific aggregation topology, nor does it require any particular operator to provide one. +The telemetry-consumer function and the `index` emitter role (section 4.4) are distinct. A discovery registry or ranking service that consumes telemetry to build ranking signal is acting as a telemetry consumer; an operator that also brokers or serves content additionally acts as an `index` emitter. This specification does not restrict combining the two, but where one operator holds both, the telemetry it consumes in the ranking capacity carries the same filtering and access-control responsibilities as any consumer's - operating an index confers no additional visibility into other parties' events. + **Origin-side `.well-known/content-telemetry.json` manifests** declare where origin-emitted retrieval events are sent (CDN → content owner's chosen endpoint). They do not instruct agents where to send session documents. Agent routing is governed by the agent's telemetry configuration, not by content owner manifests. -**Content owner resolution.** Telemetry consumers resolve content owner identity from `content_url` domains. Content owners register and verify their domains with the telemetry consumer; the consumer maps incoming event URLs to the owning organisation. This is the primary resolution path and requires `content_url` to be present on events. Events identified only by `content_id` (e.g., cached groundings where the URL was not preserved, or marketplace API content with no canonical URL) cannot be resolved by domain alone. Telemetry consumers SHOULD support `content_id` prefix-based resolution as a secondary path when content owners register their identifier schemes, but this is not yet a normative requirement. +**Content owner resolution.** Telemetry consumers resolve content owner identity through two co-primary paths: the `content_url` domain, via verified domain registrations, and the registered `content_id` prefix, via identifier-scheme declarations (section 8.6). Content owners register domains and identifier prefixes with the telemetry consumer; the consumer maps each incoming event to the owning organisation by whichever identifier the event carries. An event carrying only one of the two resolves through that path - a cached grounding with no preserved URL, or marketplace API content with no canonical URL, resolves by `content_id` prefix alone. When an event carries both and the two paths resolve to different owners, the consumer MUST NOT deliver the event to either owner's filtered view until its trust policy resolves the conflict, and SHOULD surface the conflict to both registrants. When neither path resolves, the event is unattributed; consumers SHOULD retain unattributed events rather than discard them, so a later registration can claim them. + +**Cross-consumer correlation.** Origin-side emitters and agent-side emitters MAY use different telemetry consumers. A content owner's CDN sends retrieval events to one telemetry consumer; an agent sends sessions to another. The `content_telemetry_id` field (section 7.2) correlates the same retrieval across consumers - both sides share the same UUID from the HTTP request. This correlation operates at the retrieval level only. Grounding, citation, presentation, and engagement events have no independent origin-side counterpart to correlate against. + +### 7.4 Click context (`ctx_token`) + +A click-out is the moment content usage becomes traffic the destination can observe. The click token lets the destination corroborate that moment and learn what produced it, without receiving the session UUID or any other publisher's activity. + +#### 7.4.1 Token issuance + +A `ctx_token` is an opaque token minted by the originating agent. Its value MUST match `^ct_[A-Za-z0-9_-]{16,240}$`, MUST be unguessable - the suffix is drawn from at least 96 bits of cryptographically secure randomness, or is a keyed construction of equivalent strength, so that holding one token gives no way to derive or enumerate another - and MUST NOT encode content, session, or user identifiers recoverable without the issuer's state. + +A token MUST be bound to exactly one `content_presented` occurrence at mint time. The same URL presented twice receives two tokens; a token observed on two presentations is malformed issuance and consumers MUST NOT resolve it. Surfaces that route outbound navigation through the agent SHOULD mint per click, additionally binding the token to the resulting `content_engaged` event. Direct-link surfaces mint per presentation; repeated clicks on one presentation then share a token, and are distinguished at resolution by the destination's event timestamps. + +The token-to-presentation binding is issuer state. It never travels in the URL: destinations do not receive `presentation_id`, and the consumer restores the binding at resolution. The agent SHOULD record the minted token on its own `content_engaged` event (the event-level `ctx_token` field) so the consumer can join destination reports to it. + +#### 7.4.2 Carriage and redirects + +Agents that decorate outbound link URLs MUST use the reserved query parameters `ctx_token` (the token) and `ctx_iss` (the issuer locator, section 7.4.3), and MUST NOT place other telemetry data in the URL. + +Parties operating redirects SHOULD propagate both parameters through same-domain redirect hops, mirroring the `Content-Telemetry-ID` redirect guidance in section 7.2. Destinations relying on redirect-based routing SHOULD capture the parameters at the earliest point in the chain. This is a transport and correlation convention: it does not claim that every intermediary preserved the value, and it does not enforce downstream behaviour. + +#### 7.4.3 Resolver discovery + +`ctx_iss` carries the issuer manifest locator: a host, optionally with a path prefix, identifying a well-known manifest location (section 8.1). `ctx_iss=example.com/agents/search` resolves to `https://example.com/agents/search/.well-known/content-telemetry.json`. That manifest declares the resolution endpoint in `telemetry.ctx_resolution` (section 8.5). + +The token stays opaque and carries no routing; the locator travels alongside it. No central registry is required or defined. + +#### 7.4.4 Resolution response - the click context + +A telemetry consumer that supports resolution exposes, for a presented token, the **click context**: + +1. **The engagement.** The `content_engaged` occurrence(s) bound to the token, including the presentation record the token restores: `presentation_id`, `output_id`, `output_element_id` where present, and timestamps. +2. **The lineage of the clicked content, selected by content identity.** The resolved session's `content_retrieved`, `content_grounded`, `content_cited`, and `content_presented` events whose `content_url` or `content_id` identify the same content as the clicked reference - across all turns. The cut is by content identity, not by turn or click timestamp: a click in turn 5 on content grounded in turn 2 resolves that content's full lineage. +3. **The contributing sources, gated per owner.** The sources that informed the response the click came from: `content_grounded` events in scope for the engaged presentation's turn (including session-scoped groundings) and that turn's `content_cited` and `content_presented` events, for content other than the clicked content. A contributing owner's events appear only when that owner has opted in to contributing-source disclosure with the resolving consumer; owners without a recorded opt-in are visible only through the counts in the session summary. This is the component that supports multi-citation attribution when the clicked content is not the contributing content - a click through to a commerce destination whose recommendation a publisher's review produced - and it restores the consent-gated role of the v0.1 click manifest (section 12.1), scoped to the click's provenance rather than the whole session. +4. **An optional privacy-bounded session summary.** Event counts by type and a distinct-source count. Counts, not events, and no content identifiers of owners not disclosed above. -**Cross-consumer correlation.** Origin-side emitters and agent-side emitters MAY use different telemetry consumers. A content owner's CDN sends retrieval events to one telemetry consumer; an agent sends sessions to another. The `content_telemetry_id` field (section 7.2) correlates the same retrieval across consumers - both sides share the same UUID from the HTTP request. This correlation operates at the retrieval level only. Grounding, citation, and engagement events have no independent origin-side counterpart to correlate against. +A resolution response MUST NOT include the resolved session's raw `session_id`: the token exists so the session UUID never crosses the trust boundary. Outside the contributing-source component, a resolution response MUST NOT include events for content other than the clicked content; within it, a response MUST NOT include events for an owner without a recorded contributing-source opt-in. Whole-session cross-content detail remains a reporting concern, delivered through publisher-filtered views (section 7.3), not through per-click resolution. A consumer MUST resolve a token only when the issuing agent has opted in to click-token resolution. The response is further gated by the `privacy_level` of the turn the engaged presentation belongs to - the turn named by the presentation's `turn_id`, or the turn in progress at its timestamp: at `minimal` the consumer returns the engagement and the lineage (components 1 and 2) only, withholding the contributing-source component and the session summary; at `intent` and above all four components are available. A resolution response never carries conversation-turn fields at any level. The mechanisms by which the issuer and contributing-owner opt-ins are recorded are operator-defined; all three gates are normative. + +A worked resolution response (informative): + +``` +{ + "engagement": { + "engagement_type": "link_click", + "timestamp": "2026-08-10T09:02:31Z", + "presentation_id": "880e8400-e29b-41d4-a716-446655440213", + "output_id": "response:2" + }, + "lineage": [ + { "type": "content_grounded", "timestamp": "2026-08-10T09:00:01Z", "turn_id": "1", "data": { "chars_ingested": 9400 } }, + { "type": "content_cited", "timestamp": "2026-08-10T09:00:04Z", "turn_id": "1", "data": { "citation_type": "paraphrase", "position": "primary" } }, + { "type": "content_presented", "timestamp": "2026-08-10T09:00:04Z", "turn_id": "1", "data": { "presentation_kind": "source_reference", "presentation_type": "link" } }, + { "type": "content_presented", "timestamp": "2026-08-10T09:02:10Z", "turn_id": "2", "data": { "presentation_kind": "source_reference", "presentation_type": "link" } } + ], + "contributing_sources": [ + { + "content_url": "https://publisher-a.example/heaters/space-heater-review", + "events": [ + { "type": "content_grounded", "timestamp": "2026-08-10T09:02:08Z", "turn_id": "2", "data": { "chars_ingested": 7200 } }, + { "type": "content_cited", "timestamp": "2026-08-10T09:02:10Z", "turn_id": "2", "data": { "citation_type": "paraphrase", "position": "primary" } } + ] + } + ], + "session_summary": { "turns": 2, "distinct_sources": 3, "events_by_type": { "content_grounded": 4, "content_cited": 3, "content_presented": 5 } } +} +``` + +The response shape above is informative in v1; the constraints in this section are normative. A response schema can follow implementation evidence during the release-candidate window. + +#### 7.4.5 Owner-scoped delivery of the engagement + +The agent-authored click `content_engaged` is a session event like any other: it reaches content owners through routing and aggregation (section 7.3), independent of whether the token in the URL survived the redirect chain. A telemetry consumer that provides owner-filtered views MUST include the agent-authored `content_engaged` event in the filtered view of the engaged content's owner, on the same terms as `content_grounded` and `content_cited` events - including the session identifier that owner-scoped delivery carries. The `session_id` prohibition in section 7.4.4 binds token resolution, where the requesting party is authenticated by nothing more than possession of a URL-carried value; it does not bind section 7.3 delivery to an owner whose domain registration the consumer has verified. + +This delivery is deliberately redundant with the token path. It notifies the destination owner of the click even when `ctx_token` was stripped in transit; it lets that owner join the click to their own grounded and cited events on `session_id` without calling a resolver; and it lets a party processing owner-scoped streams for both a contributing publisher and a click destination match its clients' events on `session_id` for attribution. The owner's filtered view SHOULD carry the event-level `ctx_token`, so a destination that captured the query parameters at landing can join the URL-channel observation to the server-side event directly. This is not token distribution to contributing owners (the limit recorded in section 7.4.6): only the owner of the clicked content receives the event, and that owner already saw the token in the URL. Where the clicked content's owner is also the destination, that owner therefore holds both the token from the URL and the session identifier from this delivery: the boundary of section 7.4.4 protects sessions from unregistered holders of a URL, not from the verified owner of the content that was clicked. + +#### 7.4.6 Recorded limit: consumer custody + +Resolution depends on the telemetry consumer the agent chose, because that consumer holds the session. Grounding and citation events precede the click and cannot carry its later token, and distributing tokens to every contributing content owner after the fact would weaken the privacy boundary this section maintains. Publisher-derived tokens would require a federation and key-management design; that belongs in a later attribution or evidence profile. Core v1 mitigates the dependency with resolver discoverability (7.4.3) and exact click binding (7.4.1). + +Token lifetime and requester authentication are likewise not defined in v1: a token resolves for as long as the consumer retains the session, and the resolver authenticates the requester by possession of the token alone (section 7.4.5). A resolution window after issuance, and requester credentials - for example authenticating a destination against the manifest its domain serves (section 8) - belong to the evidence profile, together with the federation design above. ## 8. Manifest Content owners, agents, and platforms publish a manifest declaring their identity and telemetry endpoints. The `manifest_ref` field on session documents (5.1.2) and the routing logic for origin-side emitters (7.3) resolve to manifests defined in this section. +The identity a manifest establishes is control of a domain, or a path under one - not a legal person. `operator.name` is a display name; nothing in a manifest binds the domain to an organisation, and no field carries a jurisdiction or registered legal entity. Where settlement or audit requires a legal counterparty, that binding lives in the governing terms the parties hold (`terms_ref`, section 5.2.4), not in the manifest. A party-identity declaration is deferred (section 8.9). + ### 8.1 Discovery Manifests are served as JSON at: @@ -849,7 +1132,7 @@ https://example.com/agents/search/.well-known/content-telemetry.json # operated Each manifest is self-contained at its own well-known URL. -Trust derives from TLS and DNS control of the domain. Manifests are unsigned in v0.1. +Trust derives from TLS and DNS control of the domain. Manifests are unsigned in v1 (section 8.9). ### 8.2 Schema @@ -857,13 +1140,14 @@ Machine-readable schema: [`./manifest.json`](./manifest.json) (JSON Schema draft | Field | Type | Required | Description | |-------|------|----------|-------------| -| `schema_version` | string | Yes | Manifest schema version. v0.1 emitters MUST use `"0.1"`. | -| `id` | string | Yes | The manifest's canonical URL (e.g. `https://example.com/.well-known/content-telemetry.json`). | +| `schema_version` | string | Yes | Manifest schema version. v1 emitters MUST use `"1.0"`. | +| `id` | string | Yes | The manifest's canonical `https://` URL, ending in `/.well-known/content-telemetry.json` (e.g. `https://example.com/.well-known/content-telemetry.json`); the schema rejects other schemes and locations (section 8.1). | | `roles` | string[] | Yes | One or more of `content_owner`, `agent`, `platform`. | | `operator` | object | Yes | Operating organisation (see 8.3). | | `keys` | object[] | No | Public keys for signing telemetry events (see 8.4). | | `telemetry` | object | No | Telemetry endpoint declaration (see 8.5). | | `domains` | string[] | No | Domains the participant claims authority over (see 8.6). MAY appear only on root manifests. | +| `identifier_schemes` | object[] | No | Identifier prefixes the participant claims and, optionally, where they resolve (see 8.6). MAY appear only on `content_owner` manifests. | Consumers MUST tolerate unknown fields and treat absent optional sections as "not declared" rather than rejecting the manifest. @@ -878,12 +1162,12 @@ A manifest MAY declare multiple roles (e.g. `["content_owner", "agent"]`). A mor ### 8.4 Keys -Public keys used to sign telemetry events emitted by this participant. Per-event signing is informational in v0.1; consumers MAY verify signatures but are not required to. +Public keys used to sign telemetry events emitted by this participant. Per-event signing remains informational in v1; consumers MAY verify signatures but are not required to. | Field | Type | Required | Description | |-------|------|----------|-------------| | `id` | string | Yes | Key identifier, unique within the manifest. | -| `type` | string | Yes | Key type. v0.1: `Ed25519`. | +| `type` | string | Yes | Key type. v1 defines `Ed25519` only. | | `publicKey` | string | Yes | Multibase-encoded public key (multicodec prefix, base58btc - the same format as `did:key`). | | `expires` | datetime | No | ISO 8601 expiry. | @@ -891,18 +1175,33 @@ Public keys used to sign telemetry events emitted by this participant. Per-event | Field | Type | Required | Description | |-------|------|----------|-------------| -| `endpoint` | string | Yes | HTTPS URL. For agents and platforms, the outbound submission endpoint. For content owners, the inbound destination for events about the content owner's content. | +| `endpoint` | string | Yes | HTTPS URL (schema-enforced). For agents and platforms, the outbound submission endpoint. For content owners, the inbound destination for events about the content owner's content. | | `conformance_level` | string | No | Conformance level advertised by this participant's own emitter(s). One of `retrieval`, `grounding`, `citation` (see 5.7). | +| `ctx_resolution` | string | No | HTTPS URL of the click-token resolution endpoint operated by or for this participant (see 7.4). Valid on `agent` and `platform` manifests. | +| `coverage` | object | No | Per-event-type coverage declaration: a map from event type to `{ "mode": …, "terms_ref": … }`, where `mode` is one of `complete`, `sampled`, `aggregated`, `selected` (see 5.7.6) and `terms_ref` optionally names the terms stating the rule or condition. | + +`coverage` makes the emitter's declared coverage machine-visible. It is a claim like the rest of the manifest, subject to the rules of section 5.7.6: a `complete` entry asserts that every qualifying occurrence of that event type within the declared relationship scope is emitted, and the other modes are meaningful only with their rule or condition reachable through `terms_ref` or otherwise disclosed to the receiving party. `conformance_level` is informational. It advertises the level of telemetry the manifest's participant emits. It does **not** constrain what an inbound `endpoint` accepts - an endpoint accepts whatever events it is configured to accept, regardless of any level declared here - and it places **no requirement** on other emitters. On a `content_owner` manifest it describes only the events the owner's own infrastructure emits (typically a CDN edge worker at `retrieval`); it says nothing about what agents or platforms report about the owner's content, which those parties advertise in their own manifests. A `content_owner` manifest SHOULD omit `conformance_level` unless the owner operates its own emitter. There is no field for a content owner to *request* a minimum level from agents; consumers tolerate events from any level (see 5.7), and the protocol does not give a manifest a way to demand more. -### 8.6 Domains +### 8.6 Domains and identifier schemes The `domains` array MAY appear only on manifests served from the domain root (`https:///.well-known/content-telemetry.json`). Manifests under path prefixes MUST NOT include `domains`. -In v0.1, every entry in `domains` MUST be self-validating: either the manifest's own host, or a subdomain of it (literal `news.example.com` or wildcard `*.example.com`). Control of the apex - proven by serving the manifest at the apex over TLS - implies DNS control of subdomains, so no further validation is needed. A manifest containing entries that are not subdomains of its own host is malformed. +In v1, every entry in `domains` MUST be self-validating: either the manifest's own host, or a subdomain of it (literal `news.example.com` or wildcard `*.example.com`). Control of the apex - proven by serving the manifest at the apex over TLS - implies DNS control of subdomains, so no further validation is needed. A manifest containing entries that are not subdomains of its own host is malformed. + +This keeps the v1 protocol fully decentralised: every manifest is a self-contained credential, validated by TLS plus the well-known location, with no dependency on consumer-side validation state or any external registry. Cross-apex claims (one operator unifying several unrelated apex domains in a single manifest) are deferred to a later version. + +#### Identifier schemes + +The `identifier_schemes` array declares the `content_id` prefixes a content owner claims, so that events identified only by `content_id` can be routed to their owner (section 7.3). It MAY appear only on manifests declaring the `content_owner` role. Each entry contains: + +| Field | Type | Required | Description | +|-------|------|----------|-------------| +| `prefix` | string | Yes | The identifier prefix claimed, matched against `content_id` values up to and including the prefix (e.g. `ft:`, `iscc:`, `mkt:gridnews:`). | +| `resolution` | string | No | HTTPS URL of an endpoint that maps a `content_id` under this prefix to the owning content record or organisation. | -This keeps the v0.1 protocol fully decentralised: every manifest is a self-contained credential, validated by TLS plus the well-known location, with no dependency on consumer-side validation state or any external registry. Cross-apex claims (one operator unifying several unrelated apex domains in a single manifest) are deferred to a later version. +Unlike `domains`, a prefix claim is not self-validating: nothing about a manifest's host proves authority over an identifier namespace. The manifest is the carrier of the claim, verified to domain level by TLS and the well-known location; a telemetry consumer verifies prefix ownership at registration, as it verifies domain registrations today, and resolves conflicting claims on the same prefix through its trust policy. A `resolution` endpoint supports owner-identity mapping; it does not verify that grounding or citation happened. ### 8.7 Consumer behaviour @@ -910,10 +1209,10 @@ When resolving a manifest from `manifest_ref`, a `content_url` domain, or any ot - **404 or network error.** Treat the participant as unverified. Do not reject telemetry events on this basis alone. - **Invalid JSON or schema validation failure.** Reject the manifest. Treat the participant as unverified. -- **Unknown `schema_version`.** During the v0.x preview period, consumers MUST accept only the exact same minor version. The semver-major compatibility rule applies from 1.0.0 onward (see section 12). +- **Unknown `schema_version`.** Manifests follow the same rule as telemetry documents (section 5.7.4): accept any `1.x` the consumer implements, validating against that minor's schema, and reject `0.x` manifests. During the v0.x preview period consumers accepted only the exact same minor version. - **Duplicate `keys[].id`.** Reject the manifest. - **`domains` entry that is not the manifest's host or a subdomain of it.** Reject the manifest as malformed (see 8.6). -- **Missing `keys` on a manifest referenced by `manifest_ref`.** Not an error in v0.1, since signing is informational. +- **Missing `keys` on a manifest referenced by `manifest_ref`.** Not an error in v1, since signing is informational. Consumers SHOULD cache resolved manifests respecting the response's `Cache-Control` headers. Manifest hosts SHOULD set `Cache-Control: max-age=3600` during onboarding and `max-age=86400` steady-state. @@ -923,14 +1222,17 @@ Consumers SHOULD cache resolved manifests respecting the response's `Cache-Contr ```json { - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://example.com/.well-known/content-telemetry.json", "roles": ["content_owner"], "operator": { "name": "Example Media" }, "telemetry": { "endpoint": "https://telemetry.example.com/v1/events" }, - "domains": ["example.com", "*.example.com"] + "domains": ["example.com", "*.example.com"], + "identifier_schemes": [ + { "prefix": "exm:", "resolution": "https://id.example.com/resolve" } + ] } ``` @@ -938,7 +1240,7 @@ Consumers SHOULD cache resolved manifests respecting the response's `Cache-Contr ```json { - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://searchco.com/agents/web-search/.well-known/content-telemetry.json", "roles": ["agent"], "operator": { "name": "SearchCo" }, @@ -957,7 +1259,7 @@ Consumers SHOULD cache resolved manifests respecting the response's `Cache-Contr ```json // https://publisher.com/.well-known/content-telemetry.json { - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://publisher.com/.well-known/content-telemetry.json", "roles": ["content_owner"], "operator": { "name": "Publisher Co" }, @@ -971,7 +1273,7 @@ Consumers SHOULD cache resolved manifests respecting the response's `Cache-Contr ```json // https://publisher.com/agents/assistant/.well-known/content-telemetry.json { - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://publisher.com/agents/assistant/.well-known/content-telemetry.json", "roles": ["agent"], "operator": { "name": "Publisher Co" }, @@ -987,17 +1289,18 @@ Consumers SHOULD cache resolved manifests respecting the response's `Cache-Contr The two manifests live independently at distinct well-known URLs. The content-owner manifest's `domains` and `telemetry` apply to publisher.com's content; the agent manifest's `keys` and `telemetry` apply to events emitted by the assistant. -### 8.9 Out of scope for v0.1 +### 8.9 Out of scope for v1 The following are deferred to later versions: - Content licence declarations (what content the participant is licensed to access) +- Party-identity declarations binding a domain to a legal entity or jurisdiction (the manifest identifies a domain; section 8 opening) - Manifest signing (W3C Verifiable Credentials, JWS proofs) - Training data and model provenance - Deployment context, purpose, brand affiliation - Revocation registries - Key rotation procedures beyond the `expires` field -- `did:web` compatibility (the `id` field uses the manifest URL in v0.1) +- `did:web` compatibility (the `id` field uses the manifest URL in v1) - Cross-apex claims (one operator unifying several unrelated apex domains in a single manifest) --- @@ -1009,8 +1312,14 @@ The following are deferred to later versions: Emitters SHOULD: - Use the minimum `privacy_level` necessary -- Hash or anonymise identifiers where possible - Use coarse `topics` values that do not identify sensitive categories (health, political or religious affiliation, sexuality) +- Carry network-level context about the request rather than about the client: `asn`, `asn_org` and `country` describe the network path and do not identify the individual behind it + +Hashing does not anonymise a value drawn from a space small enough to enumerate. The entire IPv4 address space can be hashed and compared against a candidate digest on commodity hardware, so a hashed IP address is a pseudonym rather than an anonymous value and should be treated as personal data. Version 0.1 defined an `ip_hash` field in the edge and origin data profiles (sections 6.2 and 6.3). Version 1 withdraws it, and emitters MUST NOT populate it. + +The schemas cannot enforce this v1 migration rule: event `data` accepts additional properties by design, so `ip_hash` would otherwise validate as an ordinary extension. The conformance suite therefore checks this specific prohibition at the application layer. This does not establish a general registry of withdrawn extension names. + +`privacy_level` gates the named conversation-turn fields of section 5.5 and nothing else. Extension fields in `turn` or `data`, content identifiers and URLs, `license_ref`, `terms_ref`, `output_id`, `turn_id` and every other opaque string pass through at every level, so their contents are the emitter's responsibility: an emitter MUST NOT use them to carry the end user's identity, or the query and response text that the declared level withholds, and SHOULD apply the minimisation guidance above to them as it does to the named fields. ### 9.2 Recommended levels @@ -1063,7 +1372,7 @@ Content can influence every response in a session without being explicitly cited - At the **grounding** level: counts all content that was in the agent's context, regardless of citation. This captures the full extent of content influence, including silent grounding. - At the **citation** level: counts only explicitly attributed content. Simpler to verify but undercounts content influence. -- At the **display** level: counts only content references shown to users. Narrowest scope, highest confidence. +- At the **presentation** level: counts content or source references made perceivable. It does not prove attention. Content owners and platforms should agree on which level to count at. The telemetry data supports all three; the choice is commercial, not technical. @@ -1086,6 +1395,8 @@ Implementations MAY extend core event types with custom fields in the `data` obj New event types (e.g., a commerce extension's `checkout_completed`) require a schema extension. The core schema validates only the event types listed in section 5.3. +Sessions carry a parallel extension container: the session-level `data` object (section 5.1.3). Session-scoped extension metadata belongs there, not in custom top-level fields on the session document. + ### 11.2 Custom intent categories `query_intent` accepts custom string values beyond the core set. Extensions SHOULD namespace their values to avoid collisions (e.g., `price_check` for a commerce extension). For ad-hoc categories that don't warrant a formal extension, use `other` with details in `topics`. @@ -1126,13 +1437,87 @@ Telemetry consumers MUST tolerate unknown `response_mode` values. ## 12. Versioning -Preview versions (0.x) use two-component version numbers. From 1.0.0 onward, versions follow [semantic versioning](https://semver.org/): - -- **Major** (1.0.0 → 2.0.0) - breaking changes to required fields -- **Minor** (1.0.0 → 1.1.0) - new optional fields, new event types -- **Patch** (1.0.0 → 1.0.1) - clarifications - -From 1.0.0 onward, consumers SHOULD accept sessions with any compatible minor version (same major version). During the preview period (0.x) the stricter rule in section 5.7.4 applies: a consumer accepts only the exact same minor version (a 0.1 consumer accepts 0.1 only). +### 12.1 Migration from the v0.1 preview + +V1 documents declare `schema_version` `"1.0"`, and the schemas' `$id` URLs move +from `/schema/v0.1/` to `/schema/v1/`. The two versions are distinguishable on +the wire and do not interoperate: a v0.1 consumer, applying the preview rule of +section 5.7.4, rejects a document declaring `"1.0"`, and a v1 consumer rejects a +document declaring `"0.1"`. An emitter moves to v1 by declaring `"1.0"` on +documents that satisfy this section; it MUST NOT declare `"0.1"` on a document +using v1 event types or fields. + +V1 replaces `content_displayed` with `content_presented`; emitters MUST NOT send +the old event name on the v1 integration line. Rename `data.display_type` to +`data.presentation_type` and add `data.presentation_kind` with either `content` +or `source_reference`. This is an intentional pre-1.0 breaking change: merely +renaming the event would preserve the visual-only ambiguity and would not say +what crossed the presentation boundary. + +For every `content_cited` event, assign an event `id` and `output_id`. For every +`content_presented` event, assign a distinct event `id` and the `output_id` of the +artifact made perceivable; add `output_element_id` when a stable element identity +exists and `citation_id` when the presentation carries a citation. For every +`content_engaged` event, add `presentation_id` referencing the exact presentation +event. Do not migrate clicks by matching URL alone: repeated presentations of the +same URL are distinct occurrences. + +V1 requires `data.citation_type` on every `content_cited` event and `data.scope` +on every `content_grounded` event; both are schema-enforced. A v0.1 emitter that +omitted them migrates a citation it cannot classify with `citation_type: +unclassified`, and a grounding whose scope it did not record with `scope: turn` +where the event carries a `turn_id` and `scope: session` otherwise. V1 also +rejects a `content_cited` event whose `content_url` and `content_id` are both +absent or null (section 6.5): a v0.1 citation with no resolvable reference is +not migrated as a citation. + +V1 withdraws `ip_hash` from the edge and origin retrieval profiles (section 9.1). +Emitters remove the field and MUST NOT populate it; `asn`, `asn_org` and `country` +remain. V1 also requires `source_role` on every `content_retrieved` event +(section 5.7.1); preview emitters that omitted it add the role they report under. + +V1 renames the retrieval profile's `bot_category` field to `purpose` and defines +it as an open enum classifying the access rather than the bot: the v0.1 values +`training`, `inference` and `search` keep their meaning, `advertising` is added, +and the vendor signal mappings move to informative Annex C. Emitters rename the +field; consumers MAY read a v0.1 `bot_category` value as `purpose` when +migrating historical data. + +`license_ref` keeps its wire form but no longer asserts that the use was licensed +(section 5.2.3): a consumer that read a v0.1 `license_ref` as verification of +entitlement now reads it as the emitter's claim about which grant applied. + +V1 grounding fingerprints report detection only. The published v0.1 preview +defined no `content_fingerprint` object; the object, and a +`preserved_in_output` field within it, appeared only on the pre-release +`v1-draft` line, which also carried a `content_reproduced` event type that is +not part of v1. An implementation built against that draft removes +`preserved_in_output` and any `content_reproduced` events during migration; +v1 defines no output-side reuse reporting. A grounding event MAY retain +`content_fingerprint.scheme`, `detected`, and a scheme-defined `value`. + +V1 narrows `ctx_token` resolution. The v0.1 click manifest returned every +source that informed the resolved session, gated by per-owner opt-in; the v1 +click context (section 7.4) returns the engagement, the clicked content's +lineage by content identity, a contributing-source set under the same +per-owner opt-in gate - scoped to the turn the click came from rather than +the whole session - and at most a count-based session summary. Consumers +implementing v0.1 resolution narrow the contributing-source scope accordingly +and MUST NOT return events for owners without a recorded opt-in. Token values +gain the `ct_` pattern and the unguessability rule of section 7.4.1. The binding +from a token to the presentation it was minted for is issuer state: v0.1 defined +no `presentation_id`, and v1 never places one in a URL; a destination reports +the token, and the consumer restores the binding at resolution. + +V1 tightens occurrence boundaries (section 4.3). Each core event now has a stated occurrence and cardinality: retrieval per completed fetch (a cache serve is not a retrieval), grounding per distinct content item per declared scope (chunk-level events deduplicate to one occurrence by content identity), citation per source-to-element association, presentation per rendering occurrence, engagement per observed action. Preview emitters that emitted per chunk, per passage, or re-emitted `content_retrieved` on cache serves remain schema-valid but SHOULD re-map to the stated boundaries; consumers comparing preview and v1 volumes should expect counts to shift where emitters previously chose finer or coarser units. Coverage becomes an explicit declaration (section 5.7.6) rather than an implication of conformance level. Session-scoped extension metadata belongs in the session-level `data` container (section 5.1.3); custom top-level siblings of `events`, accepted by the preview schema, are undefined. + +Version numbers are `major.minor`, and a document declares the version it was produced under in `schema_version`. From 1.0 onward: + +- **Major** (1.0 → 2.0) - breaking changes to required fields or to the meaning of an event +- **Minor** (1.0 → 1.1) - new optional fields and new event types; each minor version publishes its own schemas +- **Patch** - clarifications and errata to prose, examples and fixtures that change no field or constraint; a patch does not change `schema_version` + +From 1.0 onward, consumers accept documents with any compatible minor version (same major version) as described in section 5.7.4. During the preview period (0.x) the stricter rule applied: a consumer accepted only the exact same minor version (a 0.1 consumer accepts 0.1 only). ## Annex A (normative): JSON Schema @@ -1146,7 +1531,7 @@ A user asks a shopping assistant to compare noise-cancelling headphones. The age ```json { - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "550e8400-e29b-41d4-a716-446655440000", "agent_id": "shopping-assistant-v2", "content_scope": "electronics-reviews", @@ -1180,15 +1565,19 @@ A user asks a shopping assistant to compare noise-cancelling headphones. The age "data": { "scope": "session", "cached": false, + "chars_ingested": 16800, "tokens_ingested": 4200, "content_last_modified": "2026-01-10T14:00:00Z", "media_type": "text" } }, { + "id": "770e8400-e29b-41d4-a716-446655440302", "type": "content_cited", "timestamp": "2026-01-15T10:30:05Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:recommendation:1", "content_url": "https://www.wirecutter.com/reviews/best-wireless-headphones", "content_id": "wirecutter:best-wireless-headphones-2026", "data": { @@ -1198,13 +1587,18 @@ A user asks a shopping assistant to compare noise-cancelling headphones. The age } }, { - "type": "content_displayed", + "id": "770e8400-e29b-41d4-a716-446655440303", + "type": "content_presented", "timestamp": "2026-01-15T10:30:05Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:recommendation:1", + "citation_id": "770e8400-e29b-41d4-a716-446655440302", "content_url": "https://www.wirecutter.com/reviews/best-wireless-headphones", "content_id": "wirecutter:best-wireless-headphones-2026", "data": { - "display_type": "link" + "presentation_kind": "source_reference", + "presentation_type": "link" } }, { @@ -1230,6 +1624,7 @@ A user asks a shopping assistant to compare noise-cancelling headphones. The age "type": "content_engaged", "timestamp": "2026-01-15T10:32:00Z", "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440303", "content_url": "https://www.wirecutter.com/reviews/best-wireless-headphones", "content_id": "wirecutter:best-wireless-headphones-2026", "data": { @@ -1267,7 +1662,7 @@ A content owner's CDN detects an AI agent fetching content. The agent also repor "content_url": "https://www.wirecutter.com/reviews/best-wireless-headphones", "data": { "user_agent": "ClaudeBot/1.0", - "bot_category": "inference", + "purpose": "inference", "bot_name": "ClaudeBot", "verified": true, "cache_status": "miss", @@ -1276,8 +1671,7 @@ A content owner's CDN detects an AI agent fetching content. The agent also repor "ja4": "t13d1517h2_8daaf6152771_02e4c6ae3e16", "asn": 14618, "asn_org": "Anthropic", - "country": "US", - "ip_hash": "sha256:d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5" + "country": "US" } } ``` @@ -1290,7 +1684,7 @@ An AI agent previously fetched an FT article and cached it. In a new session, th ```json { - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440000", "agent_id": "copilot-v3", "started_at": "2026-03-28T09:00:00Z", @@ -1304,6 +1698,7 @@ An AI agent previously fetched an FT article and cached it. In a new session, th "data": { "scope": "session", "cached": true, + "chars_ingested": 12800, "tokens_ingested": 3200, "content_last_modified": "2026-03-27T18:30:00Z", "content_hash": "sha256:a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2", @@ -1321,9 +1716,12 @@ An AI agent previously fetched an FT article and cached it. In a new session, th } }, { + "id": "770e8400-e29b-41d4-a716-446655440304", "type": "content_cited", "timestamp": "2026-03-28T09:00:05Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:paragraph:2", "content_url": "https://www.ft.com/content/abc123", "content_id": "ft:abc123", "data": { @@ -1335,13 +1733,18 @@ An AI agent previously fetched an FT article and cached it. In a new session, th } }, { - "type": "content_displayed", + "id": "770e8400-e29b-41d4-a716-446655440305", + "type": "content_presented", "timestamp": "2026-03-28T09:00:05Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:paragraph:2", + "citation_id": "770e8400-e29b-41d4-a716-446655440304", "content_url": "https://www.ft.com/content/abc123", "content_id": "ft:abc123", "data": { - "display_type": "link" + "presentation_kind": "source_reference", + "presentation_type": "link" } }, { @@ -1368,9 +1771,11 @@ An AI agent previously fetched an FT article and cached it. In a new session, th } }, { + "id": "770e8400-e29b-41d4-a716-446655440306", "type": "content_cited", "timestamp": "2026-03-28T09:01:08Z", "turn_id": "2", + "output_id": "response:2", "content_url": "https://www.ft.com/content/abc123", "content_id": "ft:abc123", "data": { @@ -1420,11 +1825,11 @@ In this session: - 1 article grounded from cache (no `content_retrieved` event - the CDN saw nothing) - 3 turns of conversation - 2 explicit citations (turns 1 and 2) -- 1 display event (link shown in turn 1) -- 0 engagement events (user did not click through to ft.com) +- 1 presentation event (link made perceivable in turn 1) +- 0 engagement events (no click-through was reported) - Advertising was rendered alongside the first response -The content owner can derive: article `ft:abc123` was in context for all turns, cited twice, displayed once, never clicked. The content was 14.5 hours old (cached from previous day). The response was monetised with advertising. +The content owner can derive: article `ft:abc123` was in context for all turns, cited twice, presented once, and had no reported engagement. The content was 14.5 hours old (cached from previous day). The response was monetised with advertising. ### B.4 Minimal privacy level @@ -1444,3 +1849,161 @@ The same turn from B.3 at `minimal` privacy. No intent, no topics, no platform m ``` Compare with the `intent` version in B.3: `query_intent`, `topics`, `response_type`, `response_mode`, and `ad_rendered` are all absent. + +### B.5 Multi-owner catalogue under one agreement + +A marketplace intermediary delivers content from many publishers under a single agreement. `content_scope` identifies the agreement, and is the same for every session reported under it; content owner resolution is per event, from each event's `content_url` domain or registered `content_id` prefix (section 7.3). One session, two owners, each event resolving to its own owner. + +```json +{ + "schema_version": "1.0", + "session_id": "990e8400-e29b-41d4-a716-446655440500", + "agent_id": "research-assistant-v5", + "content_scope": "marketplace-agreement-2026-017", + "manifest_ref": "https://assistant.example.com/.well-known/content-telemetry.json", + "started_at": "2026-08-20T09:00:00Z", + "ended_at": "2026-08-20T09:00:09Z", + "events": [ + { + "type": "turn_started", + "timestamp": "2026-08-20T09:00:00Z", + "turn_id": "1", + "turn": { "privacy_level": "intent", "query_intent": "comparison", "topics": ["electric vehicles", "charging"] } + }, + { + "type": "content_retrieved", + "timestamp": "2026-08-20T09:00:01Z", + "source_role": "index", + "content_telemetry_id": "990e8400-e29b-41d4-a716-446655440501", + "content_url": "https://www.autoreview.example/ev/charging-networks-2026", + "content_id": "mkt:autoreview:88213", + "license_ref": "agreement-2026-017:autoreview", + "data": { "media_type": "text", "content_depth": "full" } + }, + { + "type": "content_retrieved", + "timestamp": "2026-08-20T09:00:01Z", + "source_role": "index", + "content_telemetry_id": "990e8400-e29b-41d4-a716-446655440502", + "content_id": "mkt:gridnews:5520", + "license_ref": "agreement-2026-017:gridnews", + "data": { "media_type": "text", "content_depth": "full" } + }, + { + "type": "content_grounded", + "timestamp": "2026-08-20T09:00:02Z", + "turn_id": "1", + "content_url": "https://www.autoreview.example/ev/charging-networks-2026", + "content_id": "mkt:autoreview:88213", + "data": { "scope": "turn", "cached": false, "provenance": "third_party_sourced", "chars_ingested": 11200 } + }, + { + "type": "content_grounded", + "timestamp": "2026-08-20T09:00:02Z", + "turn_id": "1", + "content_id": "mkt:gridnews:5520", + "data": { "scope": "turn", "cached": false, "provenance": "third_party_sourced", "chars_ingested": 6400 } + }, + { + "id": "990e8400-e29b-41d4-a716-446655440503", + "type": "content_cited", + "timestamp": "2026-08-20T09:00:06Z", + "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:networks:1", + "content_url": "https://www.autoreview.example/ev/charging-networks-2026", + "content_id": "mkt:autoreview:88213", + "data": { "citation_type": "paraphrase", "position": "primary" } + }, + { + "id": "990e8400-e29b-41d4-a716-446655440504", + "type": "content_cited", + "timestamp": "2026-08-20T09:00:06Z", + "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:tariffs:1", + "content_id": "mkt:gridnews:5520", + "data": { "citation_type": "reference", "position": "supporting" } + }, + { + "id": "990e8400-e29b-41d4-a716-446655440505", + "type": "content_presented", + "timestamp": "2026-08-20T09:00:07Z", + "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:networks:1", + "citation_id": "990e8400-e29b-41d4-a716-446655440503", + "content_url": "https://www.autoreview.example/ev/charging-networks-2026", + "content_id": "mkt:autoreview:88213", + "data": { "presentation_kind": "source_reference", "presentation_type": "link" } + }, + { + "id": "990e8400-e29b-41d4-a716-446655440506", + "type": "content_presented", + "timestamp": "2026-08-20T09:00:07Z", + "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:tariffs:1", + "citation_id": "990e8400-e29b-41d4-a716-446655440504", + "content_id": "mkt:gridnews:5520", + "data": { "presentation_kind": "source_reference", "presentation_type": "card" } + }, + { + "type": "turn_completed", + "timestamp": "2026-08-20T09:00:09Z", + "turn_id": "1", + "turn": { "privacy_level": "intent", "response_type": "comparison", "response_mode": "standard", "response_tokens": 410 } + } + ] +} +``` + +The telemetry consumer resolves `autoreview.example` by domain registration and the `mkt:gridnews:` prefix by identifier registration, and produces two owner-filtered views: AutoReview sees its retrieval, grounding, citation and link presentation; GridNews sees its own four events and nothing of AutoReview's. Neither view carries the other owner's identifiers, and both carry the shared `content_scope` so the marketplace can reconcile the session against the agreement. The second retrieval has no `content_url` at all - marketplace API content with no canonical URL - and resolves by `content_id` alone. + +### B.6 Delegated sub-agent session + +An orchestrating agent delegates a research step to a sub-agent. The child session links to its parent with `parent_session_id`, keeps its own `session_id` and events, and grounds the source in its own generation context (section 5.1). No citation or presentation is emitted merely because the child returns an internal result to its orchestrator - those events belong to the session whose output reaches the end user. + +```json +{ + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440021", + "parent_session_id": "660e8400-e29b-41d4-a716-446655440020", + "agent_id": "research-subagent-v1", + "started_at": "2026-07-29T09:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-07-29T09:00:02Z", + "turn_id": "child-turn-1", + "content_url": "https://example.org/research/source", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 3600 + } + }, + { + "type": "turn_completed", + "timestamp": "2026-07-29T09:00:05Z", + "turn_id": "child-turn-1", + "turn": { + "privacy_level": "minimal", + "response_tokens": 180 + } + } + ] +} +``` + +If the orchestrator later cites and presents this source to the user, those `content_cited` and `content_presented` events appear in the parent session, and a consumer holding both sessions joins them through `parent_session_id`. Emitters MAY omit the link when the relationship is unavailable or its disclosure is not appropriate; consumers MUST NOT infer that an unlinked session had no parent. + +## Annex C (informative): Vendor bot-classification mappings + +Edge platforms classify AI bot traffic in their own vocabularies. These mappings to the `purpose` values of section 6.2 are informative, reflect the platforms' published categories at the time of writing, and change on the platforms' own cadence. + +| `purpose` | Fastly signal | Cloudflare signal | +|-----------|---------------|-------------------| +| `training` | `AI-CRAWLER` | `AI Crawler` | +| `inference` | `AI-FETCHER` | `AI Assistant` | +| `search` | - | `AI Search` | diff --git a/manifest.json b/manifest.json index 435a5e2..f418b8e 100644 --- a/manifest.json +++ b/manifest.json @@ -1,6 +1,6 @@ { "$schema": "https://json-schema.org/draft/2020-12/schema", - "$id": "https://contenttelemetry.org/schema/v0.1/manifest.json", + "$id": "https://contenttelemetry.org/schema/v1/manifest.json", "title": "Content Telemetry Manifest", "description": "Schema for the .well-known/content-telemetry.json manifest defined in section 8 of the Content Telemetry specification. A manifest declares a participant's identity, roles, telemetry endpoint, signing keys, and claimed domains.", "type": "object", @@ -8,13 +8,14 @@ "properties": { "schema_version": { "type": "string", - "const": "0.1", - "description": "Manifest schema version. v0.1 emitters MUST use '0.1'. (Section 8.2)" + "const": "1.0", + "description": "Manifest schema version. v1 emitters MUST use '1.0'. (Section 8.2)" }, "id": { "type": "string", "format": "uri", - "description": "The manifest's canonical URL, e.g. https://example.com/.well-known/content-telemetry.json. (Section 8.2)" + "pattern": "^https://[^?#]+/\\.well-known/content-telemetry\\.json$", + "description": "The manifest's canonical https URL, ending in /.well-known/content-telemetry.json, e.g. https://example.com/.well-known/content-telemetry.json. (Section 8.2)" }, "roles": { "type": "array", @@ -76,12 +77,38 @@ "endpoint": { "type": "string", "format": "uri", - "description": "HTTPS URL. For agents and platforms, the outbound submission endpoint. For content owners, the inbound destination for events about the owner's content. (Section 8.5)" + "pattern": "^https://", + "description": "HTTPS URL (schema-enforced). For agents and platforms, the outbound submission endpoint. For content owners, the inbound destination for events about the owner's content. (Section 8.5)" }, "conformance_level": { "type": "string", "enum": ["retrieval", "grounding", "citation"], "description": "Conformance level advertised by this participant's own emitter(s). Informational. (Sections 8.5, 5.7)" + }, + "ctx_resolution": { + "type": "string", + "format": "uri", + "pattern": "^https://", + "description": "HTTPS URL of the click-token resolution endpoint operated by or for this participant. Valid on agent and platform manifests (enforced at the application layer). Destinations resolve ctx_iss to this manifest and present the token here. (Section 7.4)" + }, + "coverage": { + "type": "object", + "description": "Per-event-type coverage declaration, a map from event type to a mode object. A claim subject to the rules of section 5.7.6: complete asserts every qualifying occurrence within the declared relationship scope is emitted. (Sections 8.5, 5.7.6)", + "additionalProperties": { + "type": "object", + "required": ["mode"], + "properties": { + "mode": { + "type": "string", + "enum": ["complete", "sampled", "aggregated", "selected"], + "description": "Coverage mode (section 5.7.6)" + }, + "terms_ref": { + "type": ["string", "null"], + "description": "Reference to the terms stating the sampling, aggregation or selection rule and the relationship scope (section 5.2.4)" + } + } + } } }, "description": "Telemetry endpoint declaration. (Section 8.5)" @@ -91,7 +118,28 @@ "items": { "type": "string" }, - "description": "Domains the participant claims authority over. MAY appear only on root manifests; each entry MUST be the manifest's own host or a subdomain of it (literal or wildcard). (Section 8.6)" + "description": "Domains the participant claims authority over. MAY appear only on root manifests (enforced at the application layer); each entry MUST be the manifest's own host or a subdomain of it (literal or wildcard). (Section 8.6)" + }, + "identifier_schemes": { + "type": "array", + "items": { + "type": "object", + "required": ["prefix"], + "properties": { + "prefix": { + "type": "string", + "minLength": 1, + "description": "The content_id prefix claimed, matched against content_id values up to and including the prefix, e.g. ft: or mkt:gridnews:. (Section 8.6)" + }, + "resolution": { + "type": "string", + "format": "uri", + "pattern": "^https://", + "description": "HTTPS URL of an endpoint mapping a content_id under this prefix to the owning content record or organisation. (Section 8.6)" + } + } + }, + "description": "Identifier prefixes the participant claims for content_id routing. MAY appear only on content_owner manifests (enforced at the application layer); prefix ownership is verified by the consumer at registration. (Sections 7.3, 8.6)" } } } diff --git a/telemetry-event-batch.json b/telemetry-event-batch.json index 3738c14..4e0b0ea 100644 --- a/telemetry-event-batch.json +++ b/telemetry-event-batch.json @@ -1,6 +1,6 @@ { "$schema": "https://json-schema.org/draft/2020-12/schema", - "$id": "https://contenttelemetry.org/schema/v0.1/telemetry-event-batch.json", + "$id": "https://contenttelemetry.org/schema/v1/telemetry-event-batch.json", "title": "Content Telemetry Event Batch", "description": "Schema for batches of Content Telemetry events sharing one session context, delivered outside a session document", "type": "object", @@ -13,7 +13,7 @@ }, "schema_version": { "type": "string", - "const": "0.1", + "const": "1.0", "description": "Content Telemetry schema version" }, "session_id": { @@ -21,12 +21,18 @@ "format": "uuid", "description": "Session identifier, applying to every event in the batch. MAY be omitted by origin-side emitters with no session context. REQUIRED for emitters at Grounding conformance or above (unless ctx_token is carried instead - see section 7.1). Events belonging to different sessions MUST be delivered in separate batches." }, + "parent_session_id": { + "type": "string", + "format": "uuid", + "description": "Optional identifier of the immediate parent session that delegated work to this session. Applies to every event in the batch. See section 5.1." + }, "ctx_token": { "type": "string", - "description": "Opaque click-token issued by the originating agent, carried in place of session_id on batches of content_engaged events emitted from a landing page after a click-out. Applies to every event in the batch and is resolved by the telemetry consumer to the owning session. See section 7.1. An event MUST carry either session_id or ctx_token at Grounding conformance and above." + "pattern": "^ct_[A-Za-z0-9_-]{16,240}$", + "description": "Opaque click token issued by the originating agent, carried in place of session_id on batches of content_engaged events emitted from a landing page after a click-out. Applies to every event in the batch and is resolved by the telemetry consumer to the click context (section 7.4). An event MUST carry either session_id or ctx_token at Grounding conformance and above." }, "agent_id": { - "type": "string", + "type": ["string", "null"], "description": "Responding agent identifier. REQUIRED for emitters at Grounding conformance or above when using event batch delivery. Mirrors the session-level agent_id field." }, "started_at": { @@ -34,6 +40,11 @@ "format": "date-time", "description": "Session start timestamp (UTC). Mirrors the session-level started_at field. REQUIRED for emitters at Grounding conformance or above when using event batch delivery." }, + "manifest_ref": { + "type": ["string", "null"], + "format": "uri", + "description": "Manifest reference - the URL of a manifest at /.well-known/content-telemetry.json identifying the emitter. Applies to every event in the batch and mirrors the session-level manifest_ref field. See sections 7.1 and 8." + }, "events": { "type": "array", "minItems": 1, @@ -42,5 +53,25 @@ }, "description": "The telemetry events in the batch" } - } + }, + "allOf": [ + { + "if": { + "not": { "required": ["ctx_token"] } + }, + "then": { + "properties": { + "events": { + "items": { + "if": { + "properties": { "type": { "const": "content_engaged" } }, + "required": ["type"] + }, + "then": { "required": ["presentation_id"] } + } + } + } + } + } + ] } diff --git a/telemetry-event.json b/telemetry-event.json index a7c80a5..5477047 100644 --- a/telemetry-event.json +++ b/telemetry-event.json @@ -1,6 +1,6 @@ { "$schema": "https://json-schema.org/draft/2020-12/schema", - "$id": "https://contenttelemetry.org/schema/v0.1/telemetry-event.json", + "$id": "https://contenttelemetry.org/schema/v1/telemetry-event.json", "title": "Content Telemetry Standalone Event", "description": "Schema for standalone Content Telemetry events delivered outside a session document", "type": "object", @@ -13,7 +13,7 @@ }, "schema_version": { "type": "string", - "const": "0.1", + "const": "1.0", "description": "Content Telemetry schema version" }, "session_id": { @@ -21,12 +21,18 @@ "format": "uuid", "description": "Session identifier. MAY be omitted by origin-side emitters with no session context. REQUIRED for emitters at Grounding conformance or above (unless ctx_token is carried instead - see section 7.1)." }, + "parent_session_id": { + "type": "string", + "format": "uuid", + "description": "Optional identifier of the immediate parent session that delegated work to this session. See section 5.1." + }, "ctx_token": { "type": "string", - "description": "Opaque click-token issued by the originating agent, carried on content_engaged events emitted from a landing page after a click-out in place of session_id. Resolved by the telemetry consumer to the owning session (the click manifest). See section 7.1. An event MUST carry either session_id or ctx_token at Grounding conformance and above." + "pattern": "^ct_[A-Za-z0-9_-]{16,240}$", + "description": "Opaque click token issued by the originating agent, carried on content_engaged events emitted from a landing page after a click-out in place of session_id. Resolved by the telemetry consumer to the click context (section 7.4). An event MUST carry either session_id or ctx_token at Grounding conformance and above." }, "agent_id": { - "type": "string", + "type": ["string", "null"], "description": "Responding agent identifier. REQUIRED for emitters at Grounding conformance or above when using standalone event delivery. Mirrors the session-level agent_id field." }, "started_at": { @@ -34,9 +40,33 @@ "format": "date-time", "description": "Session start timestamp (UTC). Mirrors the session-level started_at field. REQUIRED for emitters at Grounding conformance or above when using standalone event delivery." }, + "manifest_ref": { + "type": ["string", "null"], + "format": "uri", + "description": "Manifest reference - the URL of a manifest at /.well-known/content-telemetry.json identifying the emitter. Mirrors the session-level manifest_ref field. RECOMMENDED on standalone events supporting settlement or audit obligations under governing terms. See sections 7.1 and 8." + }, "event": { "$ref": "telemetry-session.json#/$defs/TelemetryEvent", "description": "The telemetry event" } - } + }, + "allOf": [ + { + "if": { + "required": ["event"], + "properties": { + "event": { + "properties": { "type": { "const": "content_engaged" } }, + "required": ["type"] + } + }, + "not": { "required": ["ctx_token"] } + }, + "then": { + "properties": { + "event": { "required": ["presentation_id"] } + } + } + } + ] } diff --git a/telemetry-session.json b/telemetry-session.json index 9804802..c3aed4f 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -1,6 +1,6 @@ { "$schema": "https://json-schema.org/draft/2020-12/schema", - "$id": "https://contenttelemetry.org/schema/v0.1/telemetry-session.json", + "$id": "https://contenttelemetry.org/schema/v1/telemetry-session.json", "title": "Content Telemetry Session", "description": "Schema for Content Telemetry sessions - tracking content usage in AI agent interactions", "type": "object", @@ -14,7 +14,7 @@ }, "schema_version": { "type": "string", - "const": "0.1", + "const": "1.0", "description": "Content Telemetry schema version" }, "conformance_level": { @@ -27,6 +27,11 @@ "format": "uuid", "description": "Unique session identifier" }, + "parent_session_id": { + "type": "string", + "format": "uuid", + "description": "Optional identifier of the immediate parent session that delegated work to this session. See section 5.1." + }, "agent_id": { "type": ["string", "null"], "description": "Responding agent identifier" @@ -37,6 +42,7 @@ }, "manifest_ref": { "type": ["string", "null"], + "format": "uri", "description": "Manifest reference - the URL of a manifest at /.well-known/content-telemetry.json. See section 8." }, "started_at": { @@ -49,12 +55,53 @@ "format": "date-time", "description": "Session end timestamp (UTC)" }, + "data": { + "type": ["object", "null"], + "additionalProperties": true, + "description": "Session-level extension container mirroring the event-level data field (section 5.1.3). Custom fields SHOULD be namespaced; consumers MUST tolerate unknown fields.", + "properties": { + "access_context": { + "type": "object", + "additionalProperties": true, + "description": "Context from which the session's access rights derive - an institution, never an individual. Populated only where the governing terms of the relationship require it. See section 5.1.3.", + "properties": { + "identifiers": { + "type": "array", + "description": "Typed institutional identifiers. Access rights can derive through consortia, federated identity and proxies at once.", + "items": { + "type": "object", + "required": ["scheme", "value"], + "properties": { + "scheme": { + "type": "string", + "description": "Identifier scheme. Core values: ror, saml_entity_id, isni. Consumers MUST tolerate unknown schemes." + }, + "value": { + "type": "string", + "description": "Identifier value in the scheme's own format" + } + } + } + } + } + } + } + }, "events": { "type": "array", "items": { - "$ref": "#/$defs/TelemetryEvent" + "allOf": [ + { "$ref": "#/$defs/TelemetryEvent" }, + { + "if": { + "properties": { "type": { "const": "content_engaged" } }, + "required": ["type"] + }, + "then": { "required": ["presentation_id"] } + } + ] }, - "description": "Ordered list of events in the session (chronological by timestamp)" + "description": "Ordered list of events in the session (chronological by timestamp). Session documents are agent-reported, so content_engaged events here always carry presentation_id; the ctx_token relaxation applies only to standalone/batch envelopes (section 7.4)." } }, "$defs": { @@ -66,7 +113,7 @@ "id": { "type": "string", "format": "uuid", - "description": "Unique event identifier" + "description": "Unique event identifier; distinct within a session document (section 6.7)" }, "type": { "$ref": "#/$defs/EventType" @@ -78,11 +125,36 @@ }, "turn_id": { "type": ["string", "null"], - "description": "Associates this event with a conversation turn. Scoped to the session. SHOULD be set on turn_started, turn_completed, content_cited, content_displayed, content_engaged events, and content_grounded events when scope is turn." + "description": "Associates this event with a conversation turn. Scoped to the session. SHOULD be set on turn_started, turn_completed, content_cited, content_presented, content_engaged events, and content_grounded events when scope is turn." + }, + "output_id": { + "type": "string", + "minLength": 1, + "description": "Opaque identifier for the output artifact. REQUIRED on content_cited and content_presented events so output construction and later presentation can be correlated across services or times." + }, + "output_element_id": { + "type": "string", + "minLength": 1, + "description": "Opaque identifier for the element within output_id that carries the citation or presentation, such as a passage, media track, caption, link, or card." + }, + "citation_id": { + "type": "string", + "format": "uuid", + "description": "The id of the content_cited event associated with this presentation. Valid only on content_presented events; absent when the presentation is not a citation." + }, + "presentation_id": { + "type": "string", + "format": "uuid", + "description": "The id of the exact content_presented event on which the engagement occurred. REQUIRED on agent-reported content_engaged events. Destination-reported events carrying ctx_token omit it; the consumer restores the binding at resolution (section 7.4)." + }, + "ctx_token": { + "type": "string", + "pattern": "^ct_[A-Za-z0-9_-]{16,240}$", + "description": "The click token the agent minted for this engagement's presentation, recorded so the consumer can join destination-reported events to it. Valid only on content_engaged events (enforced at the application layer, section 5.7.5). Unguessable, at least 16 characters after the ct_ prefix. See section 7.4." }, "source_role": { "$ref": "#/$defs/SourceRole", - "description": "Who is reporting this event (see section 4.4). SHOULD be set on content_retrieved events." + "description": "Who is reporting this event (see section 4.4). MUST be set on content_retrieved events (sections 5.2.2, 5.7.1; enforced at the application layer)." }, "content_telemetry_id": { "type": ["string", "null"], @@ -100,7 +172,11 @@ }, "license_ref": { "type": ["string", "null"], - "description": "Reference to the content access licence (JWT jti, CoMP package ID, or opaque identifier)" + "description": "Reference to a licence or grant the emitter associates with this event (JWT jti, CoMP package ID, or opaque identifier). Core does not resolve or interpret it. See section 5.2.3." + }, + "terms_ref": { + "type": ["string", "null"], + "description": "Reference to the governing terms the emitter associates with this event (public URL or opaque identifier both parties can resolve). Distinct from license_ref, which records the grant that applied. Core does not resolve or interpret it, and it does not redefine core event semantics. See section 5.2.4." }, "turn": { "oneOf": [{ "$ref": "#/$defs/ConversationTurn" }, { "type": "null" }], @@ -119,16 +195,21 @@ "required": ["type"] }, "then": { + "required": ["data"], "properties": { "data": { + "required": ["scope"], "properties": { - "scope": { "$ref": "#/$defs/GroundingScope" }, + "scope": { "$ref": "#/$defs/GroundingScope", "description": "REQUIRED on every content_grounded event: the occurrence boundary and every counting model depend on it (sections 4.3, 6.4, 10)." }, "cached": { "type": "boolean" }, - "tokens_ingested": { "type": "integer", "minimum": 0 }, + "provenance": { "$ref": "#/$defs/SourceProvenance" }, + "chars_ingested": { "type": "integer", "minimum": 0, "description": "Unicode code points in the exact text placed in the generation context, without normalising solely for counting. Portable across emitters; preferred over tokens_ingested (section 6.4)." }, + "tokens_ingested": { "type": "integer", "minimum": 0, "description": "Token count of the same content in the emitter's own tokeniser. Supplementary: not comparable between emitters (section 6.4)." }, "content_version": { "type": "string" }, "content_last_modified": { "type": "string", "format": "date-time" }, "content_hash": { "type": "string", "pattern": "^sha256:[a-f0-9]{64}$" }, - "media_type": { "$ref": "#/$defs/MediaType" } + "media_type": { "$ref": "#/$defs/MediaType" }, + "content_fingerprint": { "$ref": "#/$defs/ContentFingerprint" } } } } @@ -140,8 +221,14 @@ "required": ["type"] }, "then": { + "required": ["id", "output_id", "data"], + "anyOf": [ + { "required": ["content_url"], "properties": { "content_url": { "type": "string" } } }, + { "required": ["content_id"], "properties": { "content_id": { "type": "string" } } } + ], "properties": { "data": { + "required": ["citation_type"], "properties": { "citation_type": { "$ref": "#/$defs/CitationType" }, "media_type": { "$ref": "#/$defs/MediaType" }, @@ -158,14 +245,17 @@ }, { "if": { - "properties": { "type": { "const": "content_displayed" } }, + "properties": { "type": { "const": "content_presented" } }, "required": ["type"] }, "then": { + "required": ["id", "output_id", "data"], "properties": { "data": { + "required": ["presentation_kind", "presentation_type"], "properties": { - "display_type": { "$ref": "#/$defs/DisplayType" }, + "presentation_kind": { "$ref": "#/$defs/PresentationKind" }, + "presentation_type": { "$ref": "#/$defs/PresentationType" }, "media_type": { "$ref": "#/$defs/MediaType" } } } @@ -197,8 +287,9 @@ "data": { "properties": { "media_type": { "$ref": "#/$defs/MediaType" }, + "content_depth": { "type": "string", "description": "Depth of the content record reached: metadata, abstract, full. Open vocabulary; consumers MUST tolerate unknown values (section 6.1)." }, "user_agent": { "type": "string" }, - "bot_category": { "type": "string", "description": "Edge platform's bot classification. Recommended values: training, inference, search" }, + "purpose": { "type": "string", "description": "Purpose of the access as classified by the reporting party. Open enum; core values: training, inference, search, advertising" }, "bot_name": { "type": "string", "description": "Recognised bot family parsed from the User-Agent (e.g., Claude-User, GPTBot, Perplexity-User). Stable across product variants within a vendor." }, "verified": { "type": "boolean" }, "cache_status": { "type": "string", "description": "Edge cache result. Recommended values: hit, miss, bypass, dynamic" }, @@ -207,8 +298,7 @@ "ja4": { "type": "string", "description": "JA4 TLS client fingerprint" }, "asn": { "type": "integer", "description": "Client AS number" }, "asn_org": { "type": "string", "description": "Client AS organisation name" }, - "country": { "type": "string", "pattern": "^[A-Z]{2}$", "description": "ISO 3166-1 alpha-2 country code" }, - "ip_hash": { "type": "string", "pattern": "^sha256:[a-f0-9]{64}$", "description": "SHA-256 of client IP (sha256:{hex})" } + "country": { "type": "string", "pattern": "^[A-Z]{2}$", "description": "ISO 3166-1 alpha-2 country code" } } } } @@ -223,7 +313,7 @@ "content_retrieved", "content_grounded", "content_cited", - "content_displayed", + "content_presented", "content_engaged", "turn_started", "turn_completed" @@ -288,7 +378,7 @@ }, "ad_rendered": { "type": ["boolean", "null"], - "description": "Whether advertising was displayed alongside the response" + "description": "Whether advertising was rendered alongside the response" } } }, @@ -334,9 +424,41 @@ "description": "Whether content informed all subsequent responses in the session or a specific turn only", "enum": ["session", "turn"] }, - "DisplayType": { + "SourceProvenance": { + "type": "string", + "description": "How the grounded representation reached the agent. Describes delivery path, not evidence quality.", + "enum": ["agent_fetched", "agent_cached", "third_party_sourced"] + }, + "ContentFingerprint": { + "type": "object", + "description": "Emitter-reported detection of a fingerprint or provenance signal in the exact grounded representation. Core does not interpret schemes or assign evidence status from detection.", + "required": ["scheme", "detected"], + "properties": { + "scheme": { + "type": "string", + "minLength": 1, + "description": "Open identifier for the fingerprint or provenance scheme checked. Globally collision-resistant values are recommended." + }, + "detected": { + "type": "boolean", + "description": "Emitter claim that the scheme's signal was found in the grounded representation" + }, + "value": { + "type": "string", + "minLength": 1, + "description": "Optional scheme-defined fingerprint or identifier value" + } + }, + "additionalProperties": true + }, + "PresentationKind": { + "type": "string", + "description": "What was made perceivable: source content itself, including a bounded excerpt or derived representation, or a reference to the source.", + "enum": ["content", "source_reference"] + }, + "PresentationType": { "type": "string", - "description": "How content or a content reference was presented to the user. Core values: link, snippet, inline_quote, card, detail_view, embed. Platforms MAY use custom values for additional presentation surfaces. Consumers MUST tolerate unknown values." + "description": "How content or a source reference was made perceivable. Core values: link, snippet, inline_quote, card, detail_view, embed, spoken_credit. Platforms MAY use custom values for additional presentation surfaces. Consumers MUST tolerate unknown values." } } } diff --git a/tests/README.md b/tests/README.md index b22d382..25fa583 100644 --- a/tests/README.md +++ b/tests/README.md @@ -1,6 +1,6 @@ # Conformance test suite -Tests for the Content Telemetry Specification v0.1. +Tests for the Content Telemetry Specification v1. ## Structure @@ -8,19 +8,20 @@ Tests for the Content Telemetry Specification v0.1. - `invalid/` - JSON files that MUST fail validation (either JSON Schema or application-layer conformance) - `validate.py` - Conformance test runner (requires `jsonschema`) - `check_examples.py` - Validates the worked examples in SPECIFICATION.md and README.md against the schemas +- `mutation_smoke.py` - Replays known suite-weakening mutations against a scratch copy and confirms the suite fails under each one ## Running From a clean checkout, with no setup beyond [uv](https://docs.astral.sh/uv/): ```sh -uv run --with jsonschema python tests/validate.py -uv run --with jsonschema python tests/check_examples.py +uv run --with "jsonschema[format-nongpl]" python tests/validate.py +uv run --with "jsonschema[format-nongpl]" python tests/check_examples.py ``` -Run from the repository root. Without uv: `pip install jsonschema`, then `python3 tests/validate.py`. Both commands run in CI on every pull request. +Run from the repository root. Without uv: `pip install "jsonschema[format-nongpl]"`, then `python3 tests/validate.py`. `tests/mutation_smoke.py` runs the same way, and in CI. The `format-nongpl` extra pulls in the format validators (`rfc3339-validator` and friends) that make `format: uuid` / `date-time` / `uri` assertions enforce rather than annotate; both scripts hard-error at startup if they are missing. Both commands run in CI on every pull request. -`check_examples.py` extracts every fenced `json` block from the spec and README, validates the complete top-level documents (sessions, standalone events, manifests) against the matching schema, and reports the number of fragments it skipped. A worked example that no longer matches its schema fails the build. +`check_examples.py` extracts every fenced `json` block from the spec and README, validates the complete top-level documents (sessions, standalone events, manifests) against the matching schema and against the application-layer rules of `validate.py`, and reports the number of fragments it skipped. A worked example that no longer matches its schema or violates a conformance rule fails the build. ## What it covers @@ -28,23 +29,39 @@ Run from the repository root. Without uv: `pip install jsonschema`, then `python - Event required fields (`type`, `timestamp`) - Turn required fields (`privacy_level`) - Enum validation (event types, privacy levels, source roles, schema version) -- All three conformance levels (Retrieval, Grounding, Citation) +- Citation source-reference requirement (content_cited rejected when content_url/content_id are missing or null) and required citation_type +- Required event and output identifiers on cited and presented events +- Closed enums (citation_type, position, scope, provenance, presentation_kind) and non-negative counts +- Required `data.scope` on grounding events; required `source_role` on retrieval events +- Format assertions on every uuid, date-time and uri field (malformed session, event, citation, presentation and parent ids; malformed started_at; malformed turn URL arrays) +- Field placement: presentation_id and event-level ctx_token only on content_engaged, citation_id only on presented events, turn only on turn events; envelope ctx_token only with engagements +- Session integrity: distinct event ids, engagement/presentation and presentation/citation identify the same content, one token per presentation +- Rejection of documents declaring the v0.1 wire version +- All three conformance levels (Retrieval, Grounding, Citation) and all four source roles (origin, edge, index, agent) +- Every core engagement type, citation type, position and presentation type; a destination-reported engagement batch; a multi-owner catalogue session under one agreement +- Manifests for all three roles, including platform; manifest rejection for http endpoints, path-prefixed domains, ctx_resolution on the wrong role, duplicate or empty roles, missing endpoint and coverage mode - Standalone event envelopes (CDN edge, agent with session FK) -- Privacy level field gating (application-layer conformance) -- Funnel exceptions (displayed-no-cited, cited-no-grounded, displayed-no-grounded) -- Embedded display (`display_type: embed`) and agent-mediated engagement (`agent_navigate`) -- Multi-turn sessions, cached grounding +- Privacy level field gating (application-layer conformance), one fixture per forbidden field at minimal and intent, in session, standalone and batch shapes +- Funnel exceptions (presented-no-cited, cited-no-grounded, presented-no-grounded) +- Text, image, audio, video, suppressed-citation, and repeated-presentation cases +- Exact presentation-to-engagement correlation across session and standalone envelopes +- Multi-turn sessions and cached grounding +- Grounding provenance paths and generic fingerprint detection across session, standalone-event, and event-batch envelopes - Custom response_mode values -Each test file has a `_test_description` field explaining what it demonstrates. +Each test file has a `_test_description` field explaining what it demonstrates. Every `invalid/` fixture also has an `_expected_error` field: a substring that must appear in the actual error (the first schema error's JSON pointer and message, or the application-layer violation text). Schema pins carry the pointer (`/events/0 'id' is a required property`) so that the same message at a different location does not satisfy them. The runner fails a fixture that fails for a different reason than the one it pins, and fails any invalid fixture missing the field. ## Application-layer conformance Some rules cannot be expressed in JSON Schema alone. These are tested as application-layer conformance checks in `validate.py`: -- Privacy level field gating (e.g. `query_text` MUST NOT be present at `minimal` level) -- `content_url` or `content_id` requirement on every content event (section 5.7.5) +- Privacy level field gating (e.g. `query_text` MUST NOT be present at `minimal` level), applied to turns wherever they appear: session documents, batches, and standalone envelopes - each shape has its own fixture +- `content_url` or `content_id` requirement on every content event, and `source_role` on every `content_retrieved` event (sections 5.2.2, 5.7.5), in every document shape +- Field placement by event type (section 5.7.5): `presentation_id` and event-level `ctx_token` only on `content_engaged`, `citation_id` only on `content_presented`, `turn` only on turn events; an envelope `ctx_token` only with `content_engaged` events - `session_id` or `ctx_token` on a standalone event or event batch envelope at Grounding conformance and above (sections 5.7.5, 7.1) -- Manifest rejection rules: duplicate `keys[].id`, and `domains` entries that are not the manifest's own host or a subdomain of it (sections 8.6, 8.7) +- Referential integrity within a session document: event ids are distinct; `content_engaged.presentation_id` matches a `content_presented` event id and `citation_id` on `content_presented` matches a `content_cited` event id, in each case identifying the same content; one event-level `ctx_token` binds to one presentation (sections 6.6-6.7, 7.4.1). Standalone envelopes and batch members are exempt - they may reference events delivered elsewhere. +- Manifest rejection rules: duplicate `keys[].id`; `domains` entries that are not the manifest's own host or a subdomain of it; `domains` on a manifest served under a path prefix; `ctx_resolution` on a manifest without the `agent` or `platform` role (sections 8.5-8.7) +- Withdrawn `ip_hash` prohibition on event data (section 9.1 migration rule), in every document shape +- Grounding provenance/cache consistency and the prohibition on `preserved_in_output` in `content_fingerprint` (sections 5.7.5, 6.4, 12.1) Valid fixtures must pass both JSON Schema and these checks; `invalid/` fixtures that pass JSON Schema but fail a check are documented in `validate.py`. The `agent_id`-at-Grounding requirement is not fixture-tested: it depends on the emitter's declared conformance level, which the fixtures do not carry. diff --git a/tests/check_examples.py b/tests/check_examples.py index 3a0e8a4..e36b721 100644 --- a/tests/check_examples.py +++ b/tests/check_examples.py @@ -8,15 +8,20 @@ session document -> telemetry-session.json standalone event -> telemetry-event.json + event batch -> telemetry-event-batch.json manifest -> manifest.json -Fragments (a bare event object, a single turn, a one-field snippet) are not -top-level documents and cannot be validated against a top-level schema. They are -counted and listed rather than validated. +and against the application-layer conformance rules of validate.py. + +A complete bare event object (carrying the required `type` and `timestamp`) is +validated against the TelemetryEvent definition in telemetry-session.json. +Genuine fragments (a single turn, a one-field snippet, an event elided below +its required fields) are not validatable and are counted and listed rather +than validated. Usage: - uv run --with jsonschema python tests/check_examples.py - # or: pip install jsonschema && python tests/check_examples.py + uv run --with "jsonschema[format-nongpl]" python tests/check_examples.py + # or: pip install "jsonschema[format-nongpl]" && python tests/check_examples.py """ import json @@ -28,11 +33,33 @@ from jsonschema import Draft202012Validator from referencing import Registry, Resource except ImportError: - print("ERROR: jsonschema package required. Install with: pip install jsonschema") + print('ERROR: jsonschema package required. Install with: pip install "jsonschema[format-nongpl]"') + sys.exit(1) + +# Format assertions (uuid, date-time, uri) are annotation-only unless a +# FormatChecker is attached to the validator. Guard at startup so a missing +# optional dependency (rfc3339-validator) hard-errors instead of silently +# downgrading every format assertion to a no-op. +FORMAT_CHECKER = Draft202012Validator.FORMAT_CHECKER +_missing_formats = {"uuid", "date-time"} - set(FORMAT_CHECKER.checkers) +if _missing_formats: + print( + "ERROR: format checker cannot enforce " + + ", ".join(sorted(_missing_formats)) + + '. Install with: pip install "jsonschema[format-nongpl]"' + ) sys.exit(1) +# Pseudo-schema name for complete bare event objects in the prose, validated +# against the TelemetryEvent definition rather than a top-level document schema. +EVENT_DEF = "telemetry-session.json#/$defs/TelemetryEvent" + REPO = Path(__file__).resolve().parent.parent SOURCES = [REPO / "SPECIFICATION.md", REPO / "README.md"] + +# The application-layer conformance rules live in validate.py (same directory). +sys.path.insert(0, str(Path(__file__).resolve().parent)) +from validate import check_application_layer, check_manifest_application_layer # noqa: E402 FENCE = re.compile(r"```json\n(.*?)\n```", re.DOTALL) @@ -51,10 +78,21 @@ def load_validators(): registry = registry.with_resource( schema.get("$id", name), Resource.from_contents(schema) ) - return { - name: Draft202012Validator(schema, registry=registry) + validators = { + name: Draft202012Validator( + schema, registry=registry, format_checker=FORMAT_CHECKER + ) for name, schema in schemas.items() } + # Bare event objects validate against the TelemetryEvent definition, + # referenced through the session schema's $id so $refs resolve. + session_id = schemas["telemetry-session.json"].get("$id", "") + validators[EVENT_DEF] = Draft202012Validator( + {"$ref": f"{session_id}#/$defs/TelemetryEvent"}, + registry=registry, + format_checker=FORMAT_CHECKER, + ) + return validators def strip_comments(block): @@ -66,7 +104,7 @@ def strip_comments(block): def classify(doc): - """Return the schema a complete document validates against, or None for a fragment.""" + """Return the schema a complete example validates against, or None for a fragment.""" if not isinstance(doc, dict): return None if doc.get("document_type") == "session": @@ -79,6 +117,10 @@ def classify(doc): return "manifest.json" if "session_id" in doc and "started_at" in doc and "events" in doc: return "telemetry-session.json" + if {"type", "timestamp"} <= doc.keys(): + # A bare event object carrying the definition's required fields is a + # complete event, validatable against the TelemetryEvent definition. + return EVENT_DEF return None @@ -116,12 +158,22 @@ def main(): checked += 1 by_schema[schema_name] = by_schema.get(schema_name, 0) + 1 errors = sorted(validators[schema_name].iter_errors(doc), key=lambda e: e.path) - if errors: + # Worked examples must also satisfy the application-layer rules; a bare + # event is checked as the sole member of a session. + if schema_name == "manifest.json": + violations = check_manifest_application_layer(doc) + elif schema_name == EVENT_DEF: + violations = check_application_layer({"events": [doc]}) + else: + violations = check_application_layer(doc) + if errors or violations: failed += 1 print(f" FAIL {loc} (against {schema_name})") for e in errors[:3]: path = "/".join(str(p) for p in e.path) or "(root)" print(f" {path}: {e.message}") + for v in violations[:3]: + print(f" application-layer: {v}") else: print(f" PASS {loc} ({schema_name})") diff --git a/tests/invalid/access-context-identifier-missing-value.json b/tests/invalid/access-context-identifier-missing-value.json new file mode 100644 index 0000000..a0e57a6 --- /dev/null +++ b/tests/invalid/access-context-identifier-missing-value.json @@ -0,0 +1,18 @@ +{ + "_test_description": "Session access_context identifier missing its value. The schema requires both scheme and value on every identifier (5.1.3).", + "_expected_error": "/data/access_context/identifiers/0 'value' is a required property", + "document_type": "session", + "schema_version": "1.0", + "session_id": "880e8400-e29b-41d4-a716-446655440090", + "started_at": "2026-08-13T14:02:10Z", + "data": { + "access_context": { + "identifiers": [ + { + "scheme": "ror" + } + ] + } + }, + "events": [] +} diff --git a/tests/invalid/access-context-identifiers-not-array.json b/tests/invalid/access-context-identifiers-not-array.json new file mode 100644 index 0000000..96fc123 --- /dev/null +++ b/tests/invalid/access-context-identifiers-not-array.json @@ -0,0 +1,14 @@ +{ + "_test_description": "Session access_context.identifiers as a bare string. The schema requires an array of {scheme, value} objects (5.1.3).", + "_expected_error": "/data/access_context/identifiers 'https://ror.org/013meh722' is not of type 'array'", + "document_type": "session", + "schema_version": "1.0", + "session_id": "880e8400-e29b-41d4-a716-446655440091", + "started_at": "2026-08-13T14:02:10Z", + "data": { + "access_context": { + "identifiers": "https://ror.org/013meh722" + } + }, + "events": [] +} diff --git a/tests/invalid/batch-ctx-token-bad-pattern.json b/tests/invalid/batch-ctx-token-bad-pattern.json new file mode 100644 index 0000000..9533c15 --- /dev/null +++ b/tests/invalid/batch-ctx-token-bad-pattern.json @@ -0,0 +1,18 @@ +{ + "_test_description": "Event batch envelope whose ctx_token fails the value pattern (section 7.4.1).", + "_expected_error": "/ctx_token 'token123' does not match '^ct_[A-Za-z0-9_-]{16,240}$'", + "document_type": "event_batch", + "schema_version": "1.0", + "ctx_token": "token123", + "events": [ + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:06Z", + "turn_id": "1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + } + } + ] +} diff --git a/tests/invalid/batch-ctx-token-non-engagement.json b/tests/invalid/batch-ctx-token-non-engagement.json new file mode 100644 index 0000000..d5ae302 --- /dev/null +++ b/tests/invalid/batch-ctx-token-non-engagement.json @@ -0,0 +1,28 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: event batch under ctx_token carrying a content_grounded event alongside an engagement (section 7.1).", + "_expected_error": "Envelope ctx_token accompanies a 'content_grounded' event", + "document_type": "event_batch", + "schema_version": "1.0", + "ctx_token": "ct_9f3a1c7e2b8d4a06", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 4200 + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:06Z", + "turn_id": "1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + } + } + ] +} diff --git a/tests/invalid/batch-empty-events.json b/tests/invalid/batch-empty-events.json index f070acb..fa13e0c 100644 --- a/tests/invalid/batch-empty-events.json +++ b/tests/invalid/batch-empty-events.json @@ -1,7 +1,8 @@ { "_test_description": "Event batch envelope with an empty events array. The schema requires minItems 1: an empty batch carries no signal and MUST NOT be delivered.", + "_expected_error": "/events [] should be non-empty", "document_type": "event_batch", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440008", "events": [] } diff --git a/tests/invalid/batch-engaged-missing-presentation-id.json b/tests/invalid/batch-engaged-missing-presentation-id.json new file mode 100644 index 0000000..6c51576 --- /dev/null +++ b/tests/invalid/batch-engaged-missing-presentation-id.json @@ -0,0 +1,20 @@ +{ + "_test_description": "Event batch with session_id (no ctx_token) whose content_engaged lacks presentation_id. The relaxation applies only to envelopes carrying ctx_token (sections 6.7, 7.4).", + "_expected_error": "/events/0 'presentation_id' is a required property", + "document_type": "event_batch", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440617", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:06Z", + "turn_id": "1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + } + } + ] +} diff --git a/tests/invalid/batch-missing-schema-version.json b/tests/invalid/batch-missing-schema-version.json index 0b872db..f097746 100644 --- a/tests/invalid/batch-missing-schema-version.json +++ b/tests/invalid/batch-missing-schema-version.json @@ -1,5 +1,6 @@ { "_test_description": "Event batch envelope missing schema_version. The event_batch document_type triggers batch envelope validation, which requires schema_version.", + "_expected_error": "/ 'schema_version' is a required property", "document_type": "event_batch", "session_id": "660e8400-e29b-41d4-a716-446655440009", "events": [ diff --git a/tests/invalid/batch-missing-session-and-ctx-token.json b/tests/invalid/batch-missing-session-and-ctx-token.json index 6c70711..0c60b89 100644 --- a/tests/invalid/batch-missing-session-and-ctx-token.json +++ b/tests/invalid/batch-missing-session-and-ctx-token.json @@ -1,13 +1,19 @@ { "_test_description": "Event batch envelope carrying a content_cited event with neither session_id nor ctx_token on the envelope. Passes JSON Schema (both fields are optional) but violates section 7.1: an event MUST carry one at Grounding conformance and above.", + "_expected_error": "neither session_id nor ctx_token", "document_type": "event_batch", - "schema_version": "0.1", + "schema_version": "1.0", "events": [ { + "id": "990e8400-e29b-41d4-a716-446655440063", "type": "content_cited", "timestamp": "2026-03-28T10:05:00Z", + "output_id": "response:1", "source_role": "agent", - "content_url": "https://example.com/article/test" + "content_url": "https://example.com/article/test", + "data": { + "citation_type": "reference" + } } ] } diff --git a/tests/invalid/batch-missing-session-mixed-retrieval.json b/tests/invalid/batch-missing-session-mixed-retrieval.json new file mode 100644 index 0000000..c61ec2e --- /dev/null +++ b/tests/invalid/batch-missing-session-mixed-retrieval.json @@ -0,0 +1,24 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: event batch mixing a retrieval with a grounding event, with neither session_id nor ctx_token. The retrieval-only exemption does not extend to a batch that carries other events (sections 5.7.5, 7.1).", + "_expected_error": "neither session_id nor ctx_token", + "document_type": "event_batch", + "schema_version": "1.0", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "agent", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + }, + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 4200 + } + } + ] +} diff --git a/tests/invalid/citation-id-on-grounded.json b/tests/invalid/citation-id-on-grounded.json new file mode 100644 index 0000000..c6f29dc --- /dev/null +++ b/tests/invalid/citation-id-on-grounded.json @@ -0,0 +1,33 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: citation_id on a content_grounded event. Valid only on content_presented (sections 5.2, 5.7.5).", + "_expected_error": "Field 'citation_id' present on 'content_grounded' event", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440632", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "citation_type": "reference", + "position": "primary" + } + }, + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 4200 + }, + "citation_id": "770e8400-e29b-41d4-a716-446655440601" + } + ] +} diff --git a/tests/invalid/citation-position-invalid.json b/tests/invalid/citation-position-invalid.json new file mode 100644 index 0000000..a918931 --- /dev/null +++ b/tests/invalid/citation-position-invalid.json @@ -0,0 +1,22 @@ +{ + "_test_description": "content_cited with position 'leading'. position is a closed enum: primary, supporting, mentioned, unclassified (section 6.5).", + "_expected_error": "/events/0/data/position 'leading' is not one of ['primary', 'supporting', 'mentioned', 'unclassified']", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440603", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "citation_type": "reference", + "position": "leading" + } + } + ] +} diff --git a/tests/invalid/citation-type-invalid.json b/tests/invalid/citation-type-invalid.json new file mode 100644 index 0000000..77bdcb4 --- /dev/null +++ b/tests/invalid/citation-type-invalid.json @@ -0,0 +1,21 @@ +{ + "_test_description": "content_cited with citation_type 'summary'. citation_type is a closed enum: direct_quote, paraphrase, reference, contradiction, unclassified (section 6.5).", + "_expected_error": "/events/0/data/citation_type 'summary' is not one of ['direct_quote', 'paraphrase', 'reference', 'contradiction', 'unclassified']", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440602", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "citation_type": "summary" + } + } + ] +} diff --git a/tests/invalid/cited-missing-citation-type.json b/tests/invalid/cited-missing-citation-type.json new file mode 100644 index 0000000..e2001f6 --- /dev/null +++ b/tests/invalid/cited-missing-citation-type.json @@ -0,0 +1,21 @@ +{ + "_test_description": "content_cited must classify how content was used in the response: data.citation_type is required. Emitters that cannot confidently classify use citation_type 'unclassified' rather than omitting the field (section 6.5).", + "_expected_error": "/events/0/data 'citation_type' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440204", + "started_at": "2026-08-05T09:30:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440205", + "type": "content_cited", + "timestamp": "2026-08-05T09:30:01Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "position": "primary", + "url_verified": true + } + } + ] +} diff --git a/tests/invalid/cited-missing-data.json b/tests/invalid/cited-missing-data.json new file mode 100644 index 0000000..165af8e --- /dev/null +++ b/tests/invalid/cited-missing-data.json @@ -0,0 +1,18 @@ +{ + "_test_description": "content_cited with no data object. citation_type is required, so data itself is required (section 6.5).", + "_expected_error": "/events/0 'data' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440615", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + } + ] +} diff --git a/tests/invalid/cited-missing-id.json b/tests/invalid/cited-missing-id.json new file mode 100644 index 0000000..f4cab7e --- /dev/null +++ b/tests/invalid/cited-missing-id.json @@ -0,0 +1,20 @@ +{ + "_test_description": "content_cited event with output_id and a resolvable source reference but no event id. id is required on citation events so a presentation can reference the exact citation via citation_id (section 6.5).", + "_expected_error": "/events/0 'id' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440203", + "started_at": "2026-08-05T09:20:00Z", + "events": [ + { + "type": "content_cited", + "timestamp": "2026-08-05T09:20:01Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "citation_type": "reference", + "position": "primary" + } + } + ] +} diff --git a/tests/invalid/cited-missing-output-id.json b/tests/invalid/cited-missing-output-id.json new file mode 100644 index 0000000..3644486 --- /dev/null +++ b/tests/invalid/cited-missing-output-id.json @@ -0,0 +1,18 @@ +{ + "_test_description": "content_cited event with id and a resolvable source reference but no output_id. output_id is required on citation events so output construction can be correlated with later presentation.", + "_expected_error": "/events/0 'output_id' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440160", + "started_at": "2026-08-01T14:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440161", + "type": "content_cited", + "timestamp": "2026-08-01T14:00:01Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "citation_type": "reference" + } + } + ] +} diff --git a/tests/invalid/cited-missing-source-reference.json b/tests/invalid/cited-missing-source-reference.json new file mode 100644 index 0000000..ad79e50 --- /dev/null +++ b/tests/invalid/cited-missing-source-reference.json @@ -0,0 +1,21 @@ +{ + "_test_description": "content_cited event carrying neither content_url nor content_id. A source association with no resolvable reference is not a citation; unlike other content events, the JSON Schema enforces the identifier requirement for content_cited (section 6.5).", + "_expected_error": "/events/0 'content_url' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440140", + "agent_id": "research-assistant-v1", + "started_at": "2026-08-01T12:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440142", + "type": "content_cited", + "timestamp": "2026-08-01T12:00:01Z", + "turn_id": "1", + "output_id": "response:1", + "data": { + "citation_type": "reference", + "position": "primary" + } + } + ] +} diff --git a/tests/invalid/cited-negative-excerpt-chars.json b/tests/invalid/cited-negative-excerpt-chars.json new file mode 100644 index 0000000..07e32b5 --- /dev/null +++ b/tests/invalid/cited-negative-excerpt-chars.json @@ -0,0 +1,22 @@ +{ + "_test_description": "content_cited with a negative excerpt_chars (section 6.5: minimum 0).", + "_expected_error": "/events/0/data/excerpt_chars -5 is less than the minimum of 0", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440622", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "citation_type": "direct_quote", + "excerpt_chars": -5 + } + } + ] +} diff --git a/tests/invalid/cited-null-source-reference.json b/tests/invalid/cited-null-source-reference.json new file mode 100644 index 0000000..fd6b084 --- /dev/null +++ b/tests/invalid/cited-null-source-reference.json @@ -0,0 +1,22 @@ +{ + "_test_description": "content_cited event with content_url explicitly null and no content_id. Presence of a null reference does not satisfy the citation reference requirement: the schema demands a non-null content_url or content_id on content_cited (section 6.5).", + "_expected_error": "/events/0/content_url None is not of type 'string'", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440141", + "agent_id": "research-assistant-v1", + "started_at": "2026-08-01T12:30:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440143", + "type": "content_cited", + "timestamp": "2026-08-01T12:30:01Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": null, + "data": { + "citation_type": "direct_quote", + "excerpt_chars": 180 + } + } + ] +} diff --git a/tests/invalid/content-event-missing-identifier.json b/tests/invalid/content-event-missing-identifier.json index 9f58702..2ff18a2 100644 --- a/tests/invalid/content-event-missing-identifier.json +++ b/tests/invalid/content-event-missing-identifier.json @@ -1,6 +1,7 @@ { "_test_description": "content_grounded event carrying neither content_url nor content_id. Passes JSON Schema (both are individually optional) but violates section 5.7.5: at least one of content_url or content_id MUST be present on every content event.", - "schema_version": "0.1", + "_expected_error": "'content_grounded' carries neither content_url nor content_id", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440099", "agent_id": "research-assistant-v1", "started_at": "2026-03-28T12:00:00Z", diff --git a/tests/invalid/ctx-token-bad-pattern.json b/tests/invalid/ctx-token-bad-pattern.json new file mode 100644 index 0000000..3470560 --- /dev/null +++ b/tests/invalid/ctx-token-bad-pattern.json @@ -0,0 +1,15 @@ +{ + "_test_description": "ctx_token failing the value pattern: tokens are opaque but MUST match ^ct_[A-Za-z0-9_-]{16,240}$ so they survive URL carriage and are recognisable in logs (section 7.4). This value has no ct_ prefix and contains reserved characters.", + "_expected_error": "/ctx_token 'session=660e8400!' does not match '^ct_[A-Za-z0-9_-]{16,240}$'", + "document_type": "event", + "schema_version": "1.0", + "ctx_token": "session=660e8400!", + "event": { + "type": "content_engaged", + "timestamp": "2026-08-10T11:20:00Z", + "content_url": "https://news.example.com/markets/rate-decision", + "data": { + "engagement_type": "link_click" + } + } +} diff --git a/tests/invalid/ctx-token-on-grounded.json b/tests/invalid/ctx-token-on-grounded.json new file mode 100644 index 0000000..aa2e6cc --- /dev/null +++ b/tests/invalid/ctx-token-on-grounded.json @@ -0,0 +1,21 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: event-level ctx_token on a content_grounded event. The field is valid only on content_engaged (sections 5.2, 5.7.5).", + "_expected_error": "Field 'ctx_token' present on 'content_grounded' event", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440631", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 4200 + }, + "ctx_token": "ct_9f3a1c7e2b8d4a06" + } + ] +} diff --git a/tests/invalid/duplicate-event-id.json b/tests/invalid/duplicate-event-id.json new file mode 100644 index 0000000..f055f5d --- /dev/null +++ b/tests/invalid/duplicate-event-id.json @@ -0,0 +1,44 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: two content_presented events share one id, so the engagement cannot name the exact occurrence. Repeated presentations receive distinct ids (section 6.7).", + "_expected_error": "Duplicate event id", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440635", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:05Z", + "turn_id": "1", + "output_id": "response:2", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:06Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440602", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + } + } + ] +} diff --git a/tests/invalid/empty-output-id.json b/tests/invalid/empty-output-id.json new file mode 100644 index 0000000..b61cb42 --- /dev/null +++ b/tests/invalid/empty-output-id.json @@ -0,0 +1,22 @@ +{ + "_test_description": "content_cited with an empty output_id. output_id is an opaque identifier and MUST be non-empty (section 5.2).", + "_expected_error": "/events/0/output_id '' should be non-empty", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440620", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "citation_type": "reference", + "position": "primary" + } + } + ] +} diff --git a/tests/invalid/engaged-missing-identifier.json b/tests/invalid/engaged-missing-identifier.json new file mode 100644 index 0000000..aa22303 --- /dev/null +++ b/tests/invalid/engaged-missing-identifier.json @@ -0,0 +1,30 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: content_engaged event carrying neither content_url nor content_id. The engagement references its presentation via presentation_id but still MUST identify the content acted on (section 5.7.5).", + "_expected_error": "'content_engaged' carries neither content_url nor content_id", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440234", + "started_at": "2026-08-05T12:20:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440235", + "type": "content_presented", + "timestamp": "2026-08-05T12:20:01Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-05T12:20:04Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440235", + "data": { + "engagement_type": "link_click" + } + } + ] +} diff --git a/tests/invalid/engaged-missing-presentation-id.json b/tests/invalid/engaged-missing-presentation-id.json new file mode 100644 index 0000000..719378f --- /dev/null +++ b/tests/invalid/engaged-missing-presentation-id.json @@ -0,0 +1,17 @@ +{ + "_test_description": "content_engaged must identify the exact presentation occurrence rather than matching only by URL.", + "_expected_error": "/events/0 'presentation_id' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440092", + "started_at": "2026-07-18T13:00:00Z", + "events": [ + { + "type": "content_engaged", + "timestamp": "2026-07-18T13:00:02Z", + "content_id": "publisher:article:1", + "data": { + "engagement_type": "link_click" + } + } + ] +} diff --git a/tests/invalid/engaged-presentation-content-mismatch.json b/tests/invalid/engaged-presentation-content-mismatch.json new file mode 100644 index 0000000..1298285 --- /dev/null +++ b/tests/invalid/engaged-presentation-content-mismatch.json @@ -0,0 +1,32 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: content_engaged references a presentation of different content. The engagement identifies the same content as the presentation it acted on (sections 5.7.5, 6.7).", + "_expected_error": "content_engaged references presentation", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440636", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-review.com/headphones/best-noise-cancelling", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:06Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440602", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + } + } + ] +} diff --git a/tests/invalid/engaged-presentation-id-no-presentations.json b/tests/invalid/engaged-presentation-id-no-presentations.json new file mode 100644 index 0000000..e3068f4 --- /dev/null +++ b/tests/invalid/engaged-presentation-id-no-presentations.json @@ -0,0 +1,30 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: content_engaged carries a presentation_id in a session with no content_presented events at all (section 6.7).", + "_expected_error": "does not match any content_presented event id", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440639", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 4200 + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:06Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440602", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + } + } + ] +} diff --git a/tests/invalid/engaged-presentation-id-unmatched.json b/tests/invalid/engaged-presentation-id-unmatched.json new file mode 100644 index 0000000..b2c0302 --- /dev/null +++ b/tests/invalid/engaged-presentation-id-unmatched.json @@ -0,0 +1,31 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: content_engaged whose presentation_id matches no content_presented event id in the session. Section 6.7 requires presentation_id to reference the exact content_presented.id on which the action occurred; an all-zeros UUID satisfies the schema's format assertion but references nothing.", + "_expected_error": "presentation_id '00000000-0000-0000-0000-000000000000' does not match any content_presented event id", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440236", + "started_at": "2026-08-05T12:30:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440237", + "type": "content_presented", + "timestamp": "2026-08-05T12:30:01Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-05T12:30:04Z", + "turn_id": "1", + "presentation_id": "00000000-0000-0000-0000-000000000000", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + } + } + ] +} diff --git a/tests/invalid/engaged-standalone-missing-presentation-and-token.json b/tests/invalid/engaged-standalone-missing-presentation-and-token.json new file mode 100644 index 0000000..d61b974 --- /dev/null +++ b/tests/invalid/engaged-standalone-missing-presentation-and-token.json @@ -0,0 +1,17 @@ +{ + "_test_description": "Agent-reported standalone content_engaged with session_id but no presentation_id. The presentation_id relaxation applies only to envelopes carrying ctx_token; an emitter that knows the session knows the presentation and MUST bind to it (sections 6.7, 7.4).", + "_expected_error": "/event 'presentation_id' is a required property", + "document_type": "event", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440230", + "agent_id": "assistant-v2", + "started_at": "2026-08-10T12:00:00Z", + "event": { + "type": "content_engaged", + "timestamp": "2026-08-10T12:00:30Z", + "content_url": "https://news.example.com/markets/rate-decision", + "data": { + "engagement_type": "link_click" + } + } +} diff --git a/tests/invalid/event-level-ctx-token-bad-pattern.json b/tests/invalid/event-level-ctx-token-bad-pattern.json new file mode 100644 index 0000000..700329d --- /dev/null +++ b/tests/invalid/event-level-ctx-token-bad-pattern.json @@ -0,0 +1,33 @@ +{ + "_test_description": "Agent-reported content_engaged whose event-level ctx_token is shorter than the 16-character minimum (section 7.4.1). Tokens MUST be unguessable; the pattern enforces the length floor.", + "_expected_error": "/events/1/ctx_token 'ct_short' does not match '^ct_[A-Za-z0-9_-]{16,240}$'", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440618", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:06Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440602", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + }, + "ctx_token": "ct_short" + } + ] +} diff --git a/tests/invalid/grounded-missing-identifier-batch.json b/tests/invalid/grounded-missing-identifier-batch.json new file mode 100644 index 0000000..d8ad604 --- /dev/null +++ b/tests/invalid/grounded-missing-identifier-batch.json @@ -0,0 +1,20 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: batch member content_grounded with neither content_url nor content_id (section 5.7.5), in the event batch shape.", + "_expected_error": "'content_grounded' carries neither content_url nor content_id", + "document_type": "event_batch", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440646", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 4200 + } + } + ] +} diff --git a/tests/invalid/grounded-missing-identifier-standalone.json b/tests/invalid/grounded-missing-identifier-standalone.json new file mode 100644 index 0000000..84d2108 --- /dev/null +++ b/tests/invalid/grounded-missing-identifier-standalone.json @@ -0,0 +1,18 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: standalone content_grounded with neither content_url nor content_id (section 5.7.5), in the standalone envelope shape.", + "_expected_error": "'content_grounded' carries neither content_url nor content_id", + "document_type": "event", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440645", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "event": { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 4200 + } + } +} diff --git a/tests/invalid/grounded-negative-chars-ingested.json b/tests/invalid/grounded-negative-chars-ingested.json new file mode 100644 index 0000000..a9352f2 --- /dev/null +++ b/tests/invalid/grounded-negative-chars-ingested.json @@ -0,0 +1,19 @@ +{ + "_test_description": "content_grounded with a negative chars_ingested. Ingestion counts are Unicode code point counts and cannot be negative (section 6.4); the schema requires a minimum of 0.", + "_expected_error": "/events/0/data/chars_ingested -7200 is less than the minimum of 0", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440210", + "started_at": "2026-08-05T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-05T10:00:01Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "session", + "cached": false, + "chars_ingested": -7200 + } + } + ] +} diff --git a/tests/invalid/grounded-negative-tokens-ingested.json b/tests/invalid/grounded-negative-tokens-ingested.json new file mode 100644 index 0000000..c74aba1 --- /dev/null +++ b/tests/invalid/grounded-negative-tokens-ingested.json @@ -0,0 +1,20 @@ +{ + "_test_description": "content_grounded with a negative tokens_ingested (section 6.4: minimum 0).", + "_expected_error": "/events/0/data/tokens_ingested -10 is less than the minimum of 0", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440621", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "tokens_ingested": -10 + } + } + ] +} diff --git a/tests/invalid/grounding-fingerprint-missing-detected.json b/tests/invalid/grounding-fingerprint-missing-detected.json new file mode 100644 index 0000000..a31c603 --- /dev/null +++ b/tests/invalid/grounding-fingerprint-missing-detected.json @@ -0,0 +1,22 @@ +{ + "_test_description": "Grounding event carries content_fingerprint without required detected. Must fail JSON Schema.", + "_expected_error": "/events/0/data/content_fingerprint 'detected' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440182", + "started_at": "2026-08-12T10:20:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-12T10:20:01Z", + "content_id": "publisher:article:99", + "data": { + "scope": "turn", + "cached": false, + "provenance": "agent_fetched", + "content_fingerprint": { + "scheme": "https://example.org/fingerprint/v1" + } + } + } + ] +} diff --git a/tests/invalid/grounding-fingerprint-missing-scheme.json b/tests/invalid/grounding-fingerprint-missing-scheme.json new file mode 100644 index 0000000..21c5a4c --- /dev/null +++ b/tests/invalid/grounding-fingerprint-missing-scheme.json @@ -0,0 +1,22 @@ +{ + "_test_description": "content_fingerprint without the required scheme (section 6.4).", + "_expected_error": "/events/0/data/content_fingerprint 'scheme' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440625", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "content_fingerprint": { + "detected": true + } + } + } + ] +} diff --git a/tests/invalid/grounding-fingerprint-preserved-in-output.json b/tests/invalid/grounding-fingerprint-preserved-in-output.json new file mode 100644 index 0000000..0251418 --- /dev/null +++ b/tests/invalid/grounding-fingerprint-preserved-in-output.json @@ -0,0 +1,24 @@ +{ + "_test_description": "Grounding fingerprint uses withdrawn preserved_in_output. Passes the extensible schema but violates the v1 migration rule (sections 6.4, 12.1).", + "_expected_error": "content_fingerprint carries preserved_in_output; withdrawn in v1", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440183", + "started_at": "2026-08-12T10:30:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-12T10:30:01Z", + "content_id": "publisher:article:100", + "data": { + "scope": "turn", + "cached": false, + "provenance": "agent_fetched", + "content_fingerprint": { + "scheme": "https://example.org/fingerprint/v1", + "detected": true, + "preserved_in_output": true + } + } + } + ] +} diff --git a/tests/invalid/grounding-missing-data.json b/tests/invalid/grounding-missing-data.json new file mode 100644 index 0000000..05c44e0 --- /dev/null +++ b/tests/invalid/grounding-missing-data.json @@ -0,0 +1,15 @@ +{ + "_test_description": "content_grounded with no data object at all. data is required on grounding events because scope is required (section 6.4).", + "_expected_error": "/events/0 'data' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440606", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + } + ] +} diff --git a/tests/invalid/grounding-provenance-cached-conflict-agent-cached.json b/tests/invalid/grounding-provenance-cached-conflict-agent-cached.json new file mode 100644 index 0000000..d6fd7fd --- /dev/null +++ b/tests/invalid/grounding-provenance-cached-conflict-agent-cached.json @@ -0,0 +1,20 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: grounding declares agent_cached with cached false. agent_cached requires cached true (section 6.4).", + "_expected_error": "agent_cached does not carry cached true", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440641", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "provenance": "agent_cached" + } + } + ] +} diff --git a/tests/invalid/grounding-provenance-cached-conflict.json b/tests/invalid/grounding-provenance-cached-conflict.json new file mode 100644 index 0000000..638310c --- /dev/null +++ b/tests/invalid/grounding-provenance-cached-conflict.json @@ -0,0 +1,19 @@ +{ + "_test_description": "Grounding event declares agent_fetched with cached true. Passes JSON Schema but violates section 6.4 provenance consistency.", + "_expected_error": "content_grounded with agent_fetched does not carry cached false", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440184", + "started_at": "2026-08-12T10:40:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-12T10:40:01Z", + "content_id": "publisher:article:101", + "data": { + "scope": "turn", + "cached": true, + "provenance": "agent_fetched" + } + } + ] +} diff --git a/tests/invalid/grounding-provenance-invalid.json b/tests/invalid/grounding-provenance-invalid.json new file mode 100644 index 0000000..fba1625 --- /dev/null +++ b/tests/invalid/grounding-provenance-invalid.json @@ -0,0 +1,20 @@ +{ + "_test_description": "content_grounded with provenance 'scraped'. provenance is agent_fetched, agent_cached or third_party_sourced (section 6.4).", + "_expected_error": "/events/0/data/provenance 'scraped' is not one of ['agent_fetched', 'agent_cached', 'third_party_sourced']", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440607", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "provenance": "scraped" + } + } + ] +} diff --git a/tests/invalid/grounding-scope-invalid.json b/tests/invalid/grounding-scope-invalid.json new file mode 100644 index 0000000..70a5cbe --- /dev/null +++ b/tests/invalid/grounding-scope-invalid.json @@ -0,0 +1,19 @@ +{ + "_test_description": "content_grounded with scope 'document'. scope is session or turn only (section 6.4).", + "_expected_error": "/events/0/data/scope 'document' is not one of ['session', 'turn']", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440604", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "document", + "cached": false + } + } + ] +} diff --git a/tests/invalid/grounding-scope-missing.json b/tests/invalid/grounding-scope-missing.json new file mode 100644 index 0000000..17cf039 --- /dev/null +++ b/tests/invalid/grounding-scope-missing.json @@ -0,0 +1,19 @@ +{ + "_test_description": "content_grounded without data.scope. scope is required on every grounding event (sections 5.7.2, 6.4): the occurrence boundary and every counting model depend on it.", + "_expected_error": "/events/0/data 'scope' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440605", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "cached": false, + "chars_ingested": 4200 + } + } + ] +} diff --git a/tests/invalid/invalid-event-type.json b/tests/invalid/invalid-event-type.json index af0a4be..9fd1aff 100644 --- a/tests/invalid/invalid-event-type.json +++ b/tests/invalid/invalid-event-type.json @@ -1,6 +1,7 @@ { "_test_description": "Event with type 'content_summarised' which is not in the EventType enum. Fails JSON Schema validation.", - "schema_version": "0.1", + "_expected_error": "/events/0/type 'content_summarised' is not one of ['content_retrieved', 'content_grounded', 'content_cited', 'content_presented', 'content", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440008", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/invalid-privacy-level.json b/tests/invalid/invalid-privacy-level.json index 6449c79..70b313d 100644 --- a/tests/invalid/invalid-privacy-level.json +++ b/tests/invalid/invalid-privacy-level.json @@ -1,6 +1,7 @@ { "_test_description": "Turn with privacy_level 'redacted' which is not in the PrivacyLevel enum. Fails JSON Schema validation.", - "schema_version": "0.1", + "_expected_error": "/events/0/turn/privacy_level 'redacted' is not one of ['full', 'summary', 'intent', 'minimal']", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440009", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/invalid-schema-version.json b/tests/invalid/invalid-schema-version.json index 529831f..7c9356a 100644 --- a/tests/invalid/invalid-schema-version.json +++ b/tests/invalid/invalid-schema-version.json @@ -1,5 +1,6 @@ { - "_test_description": "schema_version set to '2.0' which does not match the const '0.1'. Fails JSON Schema validation.", + "_test_description": "schema_version set to '2.0' which does not match the const '1.0'. Fails JSON Schema validation.", + "_expected_error": "/schema_version '1.0' was expected", "schema_version": "2.0", "session_id": "770e8400-e29b-41d4-a716-446655440010", "started_at": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/invalid-source-role.json b/tests/invalid/invalid-source-role.json index bb8453e..d26336d 100644 --- a/tests/invalid/invalid-source-role.json +++ b/tests/invalid/invalid-source-role.json @@ -1,6 +1,7 @@ { "_test_description": "Event with source_role 'cdn' which is not in the SourceRole enum. Fails JSON Schema validation.", - "schema_version": "0.1", + "_expected_error": "/events/0/source_role 'cdn' is not one of ['origin', 'edge', 'index', 'agent']", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440014", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/legacy-content-displayed.json b/tests/invalid/legacy-content-displayed.json new file mode 100644 index 0000000..129d17e --- /dev/null +++ b/tests/invalid/legacy-content-displayed.json @@ -0,0 +1,17 @@ +{ + "_test_description": "The v1 presentation event replaces the preview content_displayed name.", + "_expected_error": "/events/0/type 'content_displayed' is not one of ['content_retrieved', 'content_grounded', 'content_cited', 'content_presented', 'content_", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440093", + "started_at": "2026-07-18T13:00:00Z", + "events": [ + { + "type": "content_displayed", + "timestamp": "2026-07-18T13:00:03Z", + "content_id": "publisher:article:1", + "data": { + "display_type": "link" + } + } + ] +} diff --git a/tests/invalid/legacy-schema-version-0-1.json b/tests/invalid/legacy-schema-version-0-1.json new file mode 100644 index 0000000..3e0408b --- /dev/null +++ b/tests/invalid/legacy-schema-version-0-1.json @@ -0,0 +1,16 @@ +{ + "_test_description": "Document declaring the v0.1 wire version. A v1 consumer MUST reject it (sections 5.7.4, 12.1): v0.1 is a different wire version, not a compatible minor.", + "_expected_error": "/schema_version '1.0' was expected", + "schema_version": "0.1", + "session_id": "660e8400-e29b-41d4-a716-446655440601", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "agent", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + } + ] +} diff --git a/tests/invalid/malformed-citation-id.json b/tests/invalid/malformed-citation-id.json new file mode 100644 index 0000000..9bd684c --- /dev/null +++ b/tests/invalid/malformed-citation-id.json @@ -0,0 +1,35 @@ +{ + "_test_description": "content_presented whose citation_id is not a UUID (section 5.2 format: uuid).", + "_expected_error": "/events/1/citation_id 'cite-1' is not a 'uuid'", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440609", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "citation_type": "reference", + "position": "primary" + } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + }, + "citation_id": "cite-1" + } + ] +} diff --git a/tests/invalid/malformed-content-hash.json b/tests/invalid/malformed-content-hash.json new file mode 100644 index 0000000..250da93 --- /dev/null +++ b/tests/invalid/malformed-content-hash.json @@ -0,0 +1,20 @@ +{ + "_test_description": "content_grounded whose content_hash is not sha256:{hex} (section 6.4).", + "_expected_error": "/events/0/data/content_hash 'md5:abc' does not match '^sha256:[a-f0-9]{64}$'", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440619", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "content_hash": "md5:abc" + } + } + ] +} diff --git a/tests/invalid/malformed-content-urls-cited.json b/tests/invalid/malformed-content-urls-cited.json new file mode 100644 index 0000000..c89c9d1 --- /dev/null +++ b/tests/invalid/malformed-content-urls-cited.json @@ -0,0 +1,22 @@ +{ + "_test_description": "Turn whose content_urls_cited carries a value that is not a URI (section 5.4: URI[]).", + "_expected_error": "/events/0/turn/content_urls_cited/0 'not a url' is not a 'uri'", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440613", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "turn_completed", + "timestamp": "2026-08-20T10:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "minimal", + "content_urls_cited": [ + "not a url" + ], + "response_tokens": 10 + } + } + ] +} diff --git a/tests/invalid/malformed-event-id.json b/tests/invalid/malformed-event-id.json new file mode 100644 index 0000000..8fa2726 --- /dev/null +++ b/tests/invalid/malformed-event-id.json @@ -0,0 +1,22 @@ +{ + "_test_description": "content_cited whose id is not a UUID. Event ids carry format: uuid (section 5.2); rejected only when the validator enforces format assertions, which the conformance runners do.", + "_expected_error": "/events/0/id 'cite-1' is not a 'uuid'", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440608", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "cite-1", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "citation_type": "reference", + "position": "primary" + } + } + ] +} diff --git a/tests/invalid/malformed-parent-session-id.json b/tests/invalid/malformed-parent-session-id.json new file mode 100644 index 0000000..f7ac83c --- /dev/null +++ b/tests/invalid/malformed-parent-session-id.json @@ -0,0 +1,17 @@ +{ + "_test_description": "Session document whose parent_session_id is not a UUID. parent_session_id carries format: uuid (section 5.1); this is only rejected when the validator enforces format assertions, which the conformance runners do.", + "_expected_error": "/parent_session_id 'orchestrator-main' is not a 'uuid'", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440213", + "parent_session_id": "orchestrator-main", + "agent_id": "research-subagent-v2", + "started_at": "2026-08-05T10:20:00Z", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-05T10:20:01Z", + "source_role": "agent", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + } + ] +} diff --git a/tests/invalid/malformed-presentation-id.json b/tests/invalid/malformed-presentation-id.json new file mode 100644 index 0000000..ed23541 --- /dev/null +++ b/tests/invalid/malformed-presentation-id.json @@ -0,0 +1,32 @@ +{ + "_test_description": "content_engaged whose presentation_id is not a UUID (section 5.2 format: uuid).", + "_expected_error": "/events/1/presentation_id 'pres-1' is not a 'uuid'", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440610", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:06Z", + "turn_id": "1", + "presentation_id": "pres-1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + } + } + ] +} diff --git a/tests/invalid/malformed-session-id.json b/tests/invalid/malformed-session-id.json new file mode 100644 index 0000000..6997648 --- /dev/null +++ b/tests/invalid/malformed-session-id.json @@ -0,0 +1,16 @@ +{ + "_test_description": "Session document whose session_id is not a UUID (section 5.1 format: uuid).", + "_expected_error": "/session_id 'session-42' is not a 'uuid'", + "schema_version": "1.0", + "session_id": "session-42", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "agent", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + } + ] +} diff --git a/tests/invalid/malformed-started-at.json b/tests/invalid/malformed-started-at.json new file mode 100644 index 0000000..d7a4452 --- /dev/null +++ b/tests/invalid/malformed-started-at.json @@ -0,0 +1,16 @@ +{ + "_test_description": "Session document whose started_at is a bare date, not an RFC 3339 date-time (section 5.1).", + "_expected_error": "/started_at '2026-08-20' is not a 'date-time'", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440612", + "agent_id": "assistant-v4", + "started_at": "2026-08-20", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "agent", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + } + ] +} diff --git a/tests/invalid/manifest-bad-conformance-level.json b/tests/invalid/manifest-bad-conformance-level.json index 85cb563..429b796 100644 --- a/tests/invalid/manifest-bad-conformance-level.json +++ b/tests/invalid/manifest-bad-conformance-level.json @@ -1,9 +1,14 @@ { "_test_description": "Manifest with telemetry.conformance_level 'attribution' which is not one of retrieval, grounding, citation. Fails JSON Schema validation.", - "schema_version": "0.1", + "_expected_error": "/telemetry/conformance_level 'attribution' is not one of ['retrieval', 'grounding', 'citation']", + "schema_version": "1.0", "id": "https://searchco.com/agents/web-search/.well-known/content-telemetry.json", - "roles": ["agent"], - "operator": { "name": "SearchCo" }, + "roles": [ + "agent" + ], + "operator": { + "name": "SearchCo" + }, "telemetry": { "endpoint": "https://telemetry.example.com/v1/events", "conformance_level": "attribution" diff --git a/tests/invalid/manifest-bad-key-type.json b/tests/invalid/manifest-bad-key-type.json new file mode 100644 index 0000000..696fd13 --- /dev/null +++ b/tests/invalid/manifest-bad-key-type.json @@ -0,0 +1,19 @@ +{ + "_test_description": "Manifest key of type 'RSA'; v1 defines Ed25519 only (section 8.4).", + "_expected_error": "/keys/0/type 'Ed25519' was expected", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "agent" + ], + "operator": { + "name": "Example Media" + }, + "keys": [ + { + "id": "key-1", + "type": "RSA", + "publicKey": "z6Mk" + } + ] +} diff --git a/tests/invalid/manifest-bad-role.json b/tests/invalid/manifest-bad-role.json index d7208b0..65c115c 100644 --- a/tests/invalid/manifest-bad-role.json +++ b/tests/invalid/manifest-bad-role.json @@ -1,7 +1,12 @@ { "_test_description": "Manifest with a role value 'publisher' that is not in the roles enum (content_owner, agent, platform). Fails JSON Schema validation.", - "schema_version": "0.1", + "_expected_error": "/roles/0 'publisher' is not one of ['content_owner', 'agent', 'platform']", + "schema_version": "1.0", "id": "https://example.com/.well-known/content-telemetry.json", - "roles": ["publisher"], - "operator": { "name": "Example Media" } + "roles": [ + "publisher" + ], + "operator": { + "name": "Example Media" + } } diff --git a/tests/invalid/manifest-coverage-bad-mode.json b/tests/invalid/manifest-coverage-bad-mode.json new file mode 100644 index 0000000..6d7e424 --- /dev/null +++ b/tests/invalid/manifest-coverage-bad-mode.json @@ -0,0 +1,20 @@ +{ + "_test_description": "Manifest coverage entry with a mode outside the enum. The schema requires one of complete, sampled, aggregated, selected (8.5, 5.7.6).", + "_expected_error": "/telemetry/coverage/content_grounded/mode 'partial' is not one of ['complete', 'sampled', 'aggregated', 'selected']", + "schema_version": "1.0", + "id": "https://assistant.example.com/.well-known/content-telemetry.json", + "roles": [ + "agent" + ], + "operator": { + "name": "Assistant Example" + }, + "telemetry": { + "endpoint": "https://telemetry.assistant.example.com/v1/events", + "coverage": { + "content_grounded": { + "mode": "partial" + } + } + } +} diff --git a/tests/invalid/manifest-coverage-missing-mode.json b/tests/invalid/manifest-coverage-missing-mode.json new file mode 100644 index 0000000..78a2614 --- /dev/null +++ b/tests/invalid/manifest-coverage-missing-mode.json @@ -0,0 +1,20 @@ +{ + "_test_description": "Manifest coverage entry without the required mode (sections 8.5, 5.7.6).", + "_expected_error": "/telemetry/coverage/content_grounded 'mode' is a required property", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "agent" + ], + "operator": { + "name": "Example Media" + }, + "telemetry": { + "endpoint": "https://t.example.com/v1/events", + "coverage": { + "content_grounded": { + "terms_ref": "https://terms.example.com/r/1" + } + } + } +} diff --git a/tests/invalid/manifest-ctx-resolution-http.json b/tests/invalid/manifest-ctx-resolution-http.json new file mode 100644 index 0000000..c848a40 --- /dev/null +++ b/tests/invalid/manifest-ctx-resolution-http.json @@ -0,0 +1,16 @@ +{ + "_test_description": "Agent manifest whose ctx_resolution is an http URL; the resolution endpoint MUST be https (sections 7.4, 8.5).", + "_expected_error": "/telemetry/ctx_resolution 'http://t.example.com/v1/ctx/resolve' does not match '^https://'", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "agent" + ], + "operator": { + "name": "Example Media" + }, + "telemetry": { + "endpoint": "https://t.example.com/v1/events", + "ctx_resolution": "http://t.example.com/v1/ctx/resolve" + } +} diff --git a/tests/invalid/manifest-ctx-resolution-on-content-owner.json b/tests/invalid/manifest-ctx-resolution-on-content-owner.json new file mode 100644 index 0000000..108edb7 --- /dev/null +++ b/tests/invalid/manifest-ctx-resolution-on-content-owner.json @@ -0,0 +1,16 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: content_owner manifest declaring telemetry.ctx_resolution, which is valid on agent and platform manifests only (section 8.5).", + "_expected_error": "valid on agent and platform manifests", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "content_owner" + ], + "operator": { + "name": "Example Media" + }, + "telemetry": { + "endpoint": "https://t.example.com/v1/events", + "ctx_resolution": "https://t.example.com/v1/ctx/resolve" + } +} diff --git a/tests/invalid/manifest-domains-on-path-manifest.json b/tests/invalid/manifest-domains-on-path-manifest.json new file mode 100644 index 0000000..cb7b946 --- /dev/null +++ b/tests/invalid/manifest-domains-on-path-manifest.json @@ -0,0 +1,18 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: agent manifest served under a path prefix that carries domains. domains MAY appear only on manifests served from the domain root (section 8.6).", + "_expected_error": "only root manifests carry domains", + "schema_version": "1.0", + "id": "https://example.com/agents/search/.well-known/content-telemetry.json", + "roles": [ + "agent" + ], + "operator": { + "name": "Example Media" + }, + "telemetry": { + "endpoint": "https://t.example.com/v1/events" + }, + "domains": [ + "example.com" + ] +} diff --git a/tests/invalid/manifest-duplicate-key-id.json b/tests/invalid/manifest-duplicate-key-id.json index 16a2d71..2040a5e 100644 --- a/tests/invalid/manifest-duplicate-key-id.json +++ b/tests/invalid/manifest-duplicate-key-id.json @@ -1,6 +1,7 @@ { "_test_description": "Agent manifest with two keys sharing the id 'key-1'. Passes JSON Schema (the entries differ in publicKey) but violates section 8.7: consumers reject a manifest with duplicate keys[].id.", - "schema_version": "0.1", + "_expected_error": "Duplicate keys[].id", + "schema_version": "1.0", "id": "https://searchco.com/agents/web-search/.well-known/content-telemetry.json", "roles": ["agent"], "operator": { "name": "SearchCo" }, diff --git a/tests/invalid/manifest-duplicate-roles.json b/tests/invalid/manifest-duplicate-roles.json new file mode 100644 index 0000000..57283a1 --- /dev/null +++ b/tests/invalid/manifest-duplicate-roles.json @@ -0,0 +1,13 @@ +{ + "_test_description": "Manifest repeating a role; roles are a set (section 8.2).", + "_expected_error": "/roles ['agent', 'agent'] has non-unique elements", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "agent", + "agent" + ], + "operator": { + "name": "Example Media" + } +} diff --git a/tests/invalid/manifest-empty-roles.json b/tests/invalid/manifest-empty-roles.json new file mode 100644 index 0000000..7508e9c --- /dev/null +++ b/tests/invalid/manifest-empty-roles.json @@ -0,0 +1,10 @@ +{ + "_test_description": "Manifest with an empty roles array; at least one role is required (section 8.2).", + "_expected_error": "/roles [] should be non-empty", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [], + "operator": { + "name": "Example Media" + } +} diff --git a/tests/invalid/manifest-endpoint-http.json b/tests/invalid/manifest-endpoint-http.json new file mode 100644 index 0000000..813e39f --- /dev/null +++ b/tests/invalid/manifest-endpoint-http.json @@ -0,0 +1,15 @@ +{ + "_test_description": "Manifest whose telemetry endpoint is an http URL; section 8.5 requires an HTTPS URL.", + "_expected_error": "/telemetry/endpoint 'http://t.example.com/v1/events' does not match '^https://'", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "content_owner" + ], + "operator": { + "name": "Example Media" + }, + "telemetry": { + "endpoint": "http://t.example.com/v1/events" + } +} diff --git a/tests/invalid/manifest-foreign-domain.json b/tests/invalid/manifest-foreign-domain.json index 23968f9..db2c2c3 100644 --- a/tests/invalid/manifest-foreign-domain.json +++ b/tests/invalid/manifest-foreign-domain.json @@ -1,6 +1,7 @@ { "_test_description": "Content-owner manifest at example.com whose domains array claims othersite.com. Passes JSON Schema but violates section 8.6: every domains entry MUST be the manifest's own host or a subdomain of it. Consumers reject the manifest as malformed (section 8.7).", - "schema_version": "0.1", + "_expected_error": "'othersite.com' is not the manifest host", + "schema_version": "1.0", "id": "https://example.com/.well-known/content-telemetry.json", "roles": ["content_owner"], "operator": { "name": "Example Media" }, diff --git a/tests/invalid/manifest-id-http.json b/tests/invalid/manifest-id-http.json new file mode 100644 index 0000000..5913ae9 --- /dev/null +++ b/tests/invalid/manifest-id-http.json @@ -0,0 +1,12 @@ +{ + "_test_description": "Manifest whose id is an http URL; manifests are served over https (sections 8.1, 8.2).", + "_expected_error": "/id 'http://example.com/.well-known/content-telemetry.json' does not match '^https://[^?#]+/\\\\.well-known/content-telemetry\\\\.json$'", + "schema_version": "1.0", + "id": "http://example.com/.well-known/content-telemetry.json", + "roles": [ + "content_owner" + ], + "operator": { + "name": "Example Media" + } +} diff --git a/tests/invalid/manifest-id-not-a-uri.json b/tests/invalid/manifest-id-not-a-uri.json new file mode 100644 index 0000000..7f20fb0 --- /dev/null +++ b/tests/invalid/manifest-id-not-a-uri.json @@ -0,0 +1,12 @@ +{ + "_test_description": "Manifest whose id is not a valid URI (a space in the host). format: uri rejects it (section 8.2).", + "_expected_error": "/id 'https://exa mple.com/.well-known/content-telemetry.json' is not a 'uri'", + "schema_version": "1.0", + "id": "https://exa mple.com/.well-known/content-telemetry.json", + "roles": [ + "content_owner" + ], + "operator": { + "name": "Example Media" + } +} diff --git a/tests/invalid/manifest-id-not-well-known.json b/tests/invalid/manifest-id-not-well-known.json new file mode 100644 index 0000000..c85c1ed --- /dev/null +++ b/tests/invalid/manifest-id-not-well-known.json @@ -0,0 +1,12 @@ +{ + "_test_description": "Manifest whose id is not at the well-known location. Every manifest is served at its own /.well-known/content-telemetry.json URL (sections 8.1, 8.2).", + "_expected_error": "/id 'https://example.com/manifest.json' does not match '^https://[^?#]+/\\\\.well-known/content-telemetry\\\\.json$'", + "schema_version": "1.0", + "id": "https://example.com/manifest.json", + "roles": [ + "content_owner" + ], + "operator": { + "name": "Example Media" + } +} diff --git a/tests/invalid/manifest-identifier-scheme-empty-prefix.json b/tests/invalid/manifest-identifier-scheme-empty-prefix.json new file mode 100644 index 0000000..4488d17 --- /dev/null +++ b/tests/invalid/manifest-identifier-scheme-empty-prefix.json @@ -0,0 +1,17 @@ +{ + "_test_description": "Identifier-scheme entry with an empty prefix. prefix is required and non-empty: an empty string would match every content_id (section 8.6).", + "_expected_error": "prefix", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "content_owner" + ], + "operator": { + "name": "Example Media" + }, + "identifier_schemes": [ + { + "prefix": "" + } + ] +} diff --git a/tests/invalid/manifest-identifier-schemes-on-non-owner.json b/tests/invalid/manifest-identifier-schemes-on-non-owner.json new file mode 100644 index 0000000..20f759c --- /dev/null +++ b/tests/invalid/manifest-identifier-schemes-on-non-owner.json @@ -0,0 +1,20 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: agent manifest carrying identifier_schemes. The block MAY appear only on manifests declaring the content_owner role (section 8.6).", + "_expected_error": "identifier_schemes only on content_owner manifests", + "schema_version": "1.0", + "id": "https://searchco.com/agents/web-search/.well-known/content-telemetry.json", + "roles": [ + "agent" + ], + "operator": { + "name": "SearchCo" + }, + "telemetry": { + "endpoint": "https://telemetry.searchco.com/v1/events" + }, + "identifier_schemes": [ + { + "prefix": "sc:" + } + ] +} diff --git a/tests/invalid/manifest-key-missing-id.json b/tests/invalid/manifest-key-missing-id.json new file mode 100644 index 0000000..7e627c0 --- /dev/null +++ b/tests/invalid/manifest-key-missing-id.json @@ -0,0 +1,18 @@ +{ + "_test_description": "Manifest key without the required id (section 8.4).", + "_expected_error": "/keys/0 'id' is a required property", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "agent" + ], + "operator": { + "name": "Example Media" + }, + "keys": [ + { + "type": "Ed25519", + "publicKey": "z6Mk" + } + ] +} diff --git a/tests/invalid/manifest-lookalike-domain.json b/tests/invalid/manifest-lookalike-domain.json new file mode 100644 index 0000000..c83fbe0 --- /dev/null +++ b/tests/invalid/manifest-lookalike-domain.json @@ -0,0 +1,16 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: manifest at example.com claiming evilexample.com, a lookalike that ends in the host string but is not a subdomain of it (section 8.6).", + "_expected_error": "'evilexample.com' is not the manifest host", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "content_owner" + ], + "operator": { + "name": "Example Media" + }, + "domains": [ + "example.com", + "evilexample.com" + ] +} diff --git a/tests/invalid/manifest-missing-endpoint.json b/tests/invalid/manifest-missing-endpoint.json new file mode 100644 index 0000000..4ff043f --- /dev/null +++ b/tests/invalid/manifest-missing-endpoint.json @@ -0,0 +1,15 @@ +{ + "_test_description": "Manifest whose telemetry block lacks the required endpoint (section 8.5).", + "_expected_error": "/telemetry 'endpoint' is a required property", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "content_owner" + ], + "operator": { + "name": "Example Media" + }, + "telemetry": { + "conformance_level": "retrieval" + } +} diff --git a/tests/invalid/manifest-missing-key-publickey.json b/tests/invalid/manifest-missing-key-publickey.json index dec6aa9..33f26df 100644 --- a/tests/invalid/manifest-missing-key-publickey.json +++ b/tests/invalid/manifest-missing-key-publickey.json @@ -1,10 +1,18 @@ { "_test_description": "Manifest with a keys entry missing the required 'publicKey' field. Fails JSON Schema validation.", - "schema_version": "0.1", + "_expected_error": "/keys/0 'publicKey' is a required property", + "schema_version": "1.0", "id": "https://searchco.com/agents/web-search/.well-known/content-telemetry.json", - "roles": ["agent"], - "operator": { "name": "SearchCo" }, + "roles": [ + "agent" + ], + "operator": { + "name": "SearchCo" + }, "keys": [ - { "id": "key-1", "type": "Ed25519" } + { + "id": "key-1", + "type": "Ed25519" + } ] } diff --git a/tests/invalid/manifest-missing-operator.json b/tests/invalid/manifest-missing-operator.json index ceb5df7..6d06ee1 100644 --- a/tests/invalid/manifest-missing-operator.json +++ b/tests/invalid/manifest-missing-operator.json @@ -1,6 +1,9 @@ { "_test_description": "Manifest missing the required 'operator' field. Fails JSON Schema validation.", - "schema_version": "0.1", + "_expected_error": "/ 'operator' is a required property", + "schema_version": "1.0", "id": "https://example.com/.well-known/content-telemetry.json", - "roles": ["content_owner"] + "roles": [ + "content_owner" + ] } diff --git a/tests/invalid/missing-event-timestamp.json b/tests/invalid/missing-event-timestamp.json index f626857..10e4d54 100644 --- a/tests/invalid/missing-event-timestamp.json +++ b/tests/invalid/missing-event-timestamp.json @@ -1,6 +1,7 @@ { "_test_description": "Event without timestamp field. Fails JSON Schema validation: timestamp is required on TelemetryEvent.", - "schema_version": "0.1", + "_expected_error": "/events/0 'timestamp' is a required property", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440005", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/missing-event-type.json b/tests/invalid/missing-event-type.json index d0ea6dc..4eb04af 100644 --- a/tests/invalid/missing-event-type.json +++ b/tests/invalid/missing-event-type.json @@ -1,6 +1,7 @@ { "_test_description": "Event without type field. Fails JSON Schema validation: type is required on TelemetryEvent.", - "schema_version": "0.1", + "_expected_error": "/events/0 'type' is a required property", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440004", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/missing-privacy-level.json b/tests/invalid/missing-privacy-level.json index 28cab62..c36d913 100644 --- a/tests/invalid/missing-privacy-level.json +++ b/tests/invalid/missing-privacy-level.json @@ -1,6 +1,7 @@ { "_test_description": "Turn object without privacy_level. Fails JSON Schema validation: privacy_level is required on ConversationTurn.", - "schema_version": "0.1", + "_expected_error": "/events/0/turn 'privacy_level' is a required property", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440006", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/missing-schema-version.json b/tests/invalid/missing-schema-version.json index 0294251..ba6660a 100644 --- a/tests/invalid/missing-schema-version.json +++ b/tests/invalid/missing-schema-version.json @@ -1,5 +1,6 @@ { "_test_description": "Session without schema_version. Fails JSON Schema validation: schema_version is a required field.", + "_expected_error": "/ 'schema_version' is a required property", "session_id": "770e8400-e29b-41d4-a716-446655440001", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/missing-session-id.json b/tests/invalid/missing-session-id.json index 4474fa6..84cd0ae 100644 --- a/tests/invalid/missing-session-id.json +++ b/tests/invalid/missing-session-id.json @@ -1,6 +1,7 @@ { "_test_description": "Session without session_id. Fails JSON Schema validation: session_id is a required field.", - "schema_version": "0.1", + "_expected_error": "/ 'session_id' is a required property", + "schema_version": "1.0", "started_at": "2026-03-28T10:00:00Z", "events": [ { diff --git a/tests/invalid/missing-started-at.json b/tests/invalid/missing-started-at.json index caa221f..e5d35c4 100644 --- a/tests/invalid/missing-started-at.json +++ b/tests/invalid/missing-started-at.json @@ -1,6 +1,7 @@ { "_test_description": "Session without started_at. Fails JSON Schema validation: started_at is a required field.", - "schema_version": "0.1", + "_expected_error": "/ 'started_at' is a required property", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440003", "events": [ { diff --git a/tests/invalid/presentation-id-on-cited.json b/tests/invalid/presentation-id-on-cited.json new file mode 100644 index 0000000..c4c93ce --- /dev/null +++ b/tests/invalid/presentation-id-on-cited.json @@ -0,0 +1,35 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: presentation_id on a content_cited event. Valid only on content_engaged (sections 5.2, 5.7.5).", + "_expected_error": "Field 'presentation_id' present on 'content_cited' event", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440633", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "citation_type": "reference", + "position": "primary" + }, + "presentation_id": "770e8400-e29b-41d4-a716-446655440602" + } + ] +} diff --git a/tests/invalid/presentation-kind-invalid.json b/tests/invalid/presentation-kind-invalid.json new file mode 100644 index 0000000..e64cf95 --- /dev/null +++ b/tests/invalid/presentation-kind-invalid.json @@ -0,0 +1,21 @@ +{ + "_test_description": "content_presented with presentation_kind 'reference'. presentation_kind is a closed two-value distinction between source content and a reference to the source: only 'content' and 'source_reference' are valid (section 6.7).", + "_expected_error": "/events/0/data/presentation_kind 'reference' is not one of ['content', 'source_reference']", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440208", + "started_at": "2026-08-05T09:50:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440209", + "type": "content_presented", + "timestamp": "2026-08-05T09:50:01Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "reference", + "presentation_type": "link" + } + } + ] +} diff --git a/tests/invalid/presented-citation-content-mismatch.json b/tests/invalid/presented-citation-content-mismatch.json new file mode 100644 index 0000000..75d5d0e --- /dev/null +++ b/tests/invalid/presented-citation-content-mismatch.json @@ -0,0 +1,35 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: content_presented.citation_id references a citation of different content (sections 5.7.5, 6.7).", + "_expected_error": "content_presented references citation", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440637", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-review.com/headphones/best-noise-cancelling", + "data": { + "citation_type": "reference", + "position": "primary" + } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "citation_id": "770e8400-e29b-41d4-a716-446655440601", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + } + ] +} diff --git a/tests/invalid/presented-citation-id-unmatched.json b/tests/invalid/presented-citation-id-unmatched.json new file mode 100644 index 0000000..4b604d2 --- /dev/null +++ b/tests/invalid/presented-citation-id-unmatched.json @@ -0,0 +1,34 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: content_presented whose citation_id matches no content_cited event id in the session. Section 6.7 requires citation_id to reference the presented content_cited event's id; here it points at an all-zeros UUID while the session's actual citation has a different id.", + "_expected_error": "citation_id '00000000-0000-0000-0000-000000000000' does not match any content_cited event id", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440238", + "started_at": "2026-08-05T12:40:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440239", + "type": "content_cited", + "timestamp": "2026-08-05T12:40:01Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "citation_type": "reference", + "position": "primary" + } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440240", + "type": "content_presented", + "timestamp": "2026-08-05T12:40:02Z", + "turn_id": "1", + "output_id": "response:1", + "citation_id": "00000000-0000-0000-0000-000000000000", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + } + ] +} diff --git a/tests/invalid/presented-missing-data.json b/tests/invalid/presented-missing-data.json new file mode 100644 index 0000000..049fbff --- /dev/null +++ b/tests/invalid/presented-missing-data.json @@ -0,0 +1,18 @@ +{ + "_test_description": "content_presented with no data object. presentation_kind and presentation_type are required, so data itself is required (section 6.7).", + "_expected_error": "/events/0 'data' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440614", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + } + ] +} diff --git a/tests/invalid/presented-missing-id.json b/tests/invalid/presented-missing-id.json new file mode 100644 index 0000000..8cf702f --- /dev/null +++ b/tests/invalid/presented-missing-id.json @@ -0,0 +1,20 @@ +{ + "_test_description": "content_presented event with output_id and presentation data but no event id. id is required on presentation events so a later content_engaged.presentation_id can reference the exact surface occurrence (section 6.7).", + "_expected_error": "/events/0 'id' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440200", + "started_at": "2026-08-05T09:00:00Z", + "events": [ + { + "type": "content_presented", + "timestamp": "2026-08-05T09:00:01Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + } + ] +} diff --git a/tests/invalid/presented-missing-identifier.json b/tests/invalid/presented-missing-identifier.json new file mode 100644 index 0000000..441d056 --- /dev/null +++ b/tests/invalid/presented-missing-identifier.json @@ -0,0 +1,20 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: content_presented event carrying neither content_url nor content_id. Passes JSON Schema (both are individually optional) but violates section 5.7.5: at least one MUST be present on every content event.", + "_expected_error": "'content_presented' carries neither content_url nor content_id", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440230", + "started_at": "2026-08-05T12:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440231", + "type": "content_presented", + "timestamp": "2026-08-05T12:00:01Z", + "turn_id": "1", + "output_id": "response:1", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + } + ] +} diff --git a/tests/invalid/presented-missing-kind.json b/tests/invalid/presented-missing-kind.json new file mode 100644 index 0000000..b0dc287 --- /dev/null +++ b/tests/invalid/presented-missing-kind.json @@ -0,0 +1,19 @@ +{ + "_test_description": "content_presented must distinguish source content from a source reference.", + "_expected_error": "/events/0/data 'presentation_kind' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440090", + "started_at": "2026-07-18T13:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440091", + "type": "content_presented", + "timestamp": "2026-07-18T13:00:01Z", + "output_id": "response:1", + "content_id": "publisher:article:1", + "data": { + "presentation_type": "link" + } + } + ] +} diff --git a/tests/invalid/presented-missing-output-id.json b/tests/invalid/presented-missing-output-id.json new file mode 100644 index 0000000..05c2db5 --- /dev/null +++ b/tests/invalid/presented-missing-output-id.json @@ -0,0 +1,20 @@ +{ + "_test_description": "content_presented event with an event id and presentation data but no output_id. output_id is required on presentation events so presentation can be correlated with the output artifact it delivers (section 6.7).", + "_expected_error": "/events/0 'output_id' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440201", + "started_at": "2026-08-05T09:10:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440202", + "type": "content_presented", + "timestamp": "2026-08-05T09:10:01Z", + "turn_id": "1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + } + ] +} diff --git a/tests/invalid/presented-missing-presentation-type.json b/tests/invalid/presented-missing-presentation-type.json new file mode 100644 index 0000000..15cebb0 --- /dev/null +++ b/tests/invalid/presented-missing-presentation-type.json @@ -0,0 +1,19 @@ +{ + "_test_description": "content_presented event carrying presentation_kind but no presentation_type. Both are required: the kind says what crossed the presentation boundary, the type says how it was made perceivable.", + "_expected_error": "/events/0/data 'presentation_type' is a required property", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440163", + "started_at": "2026-08-01T15:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440164", + "type": "content_presented", + "timestamp": "2026-08-01T15:00:01Z", + "output_id": "response:1", + "content_id": "publisher:article:4", + "data": { + "presentation_kind": "source_reference" + } + } + ] +} diff --git a/tests/invalid/privacy-violation-ad-rendered-at-minimal.json b/tests/invalid/privacy-violation-ad-rendered-at-minimal.json index 913540d..c4cf0c1 100644 --- a/tests/invalid/privacy-violation-ad-rendered-at-minimal.json +++ b/tests/invalid/privacy-violation-ad-rendered-at-minimal.json @@ -1,6 +1,7 @@ { "_test_description": "APPLICATION-LAYER VIOLATION: Turn at minimal privacy with ad_rendered present. Passes JSON Schema validation but violates section 5.5: platform metadata (including ad_rendered) is not available at minimal level.", - "schema_version": "0.1", + "_expected_error": "'ad_rendered' present on turn with privacy_level 'minimal'", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440012", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/privacy-violation-model-id-at-minimal.json b/tests/invalid/privacy-violation-model-id-at-minimal.json new file mode 100644 index 0000000..f585b04 --- /dev/null +++ b/tests/invalid/privacy-violation-model-id-at-minimal.json @@ -0,0 +1,19 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: Turn at minimal privacy with model_id present. Passes JSON Schema validation but violates the privacy field gating rule (section 5.5): model_id MUST NOT be present when privacy_level is minimal.", + "_expected_error": "'model_id' present on turn with privacy_level 'minimal'", + "schema_version": "1.0", + "session_id": "770e8400-e29b-41d4-a716-446655440225", + "started_at": "2026-08-05T11:00:00Z", + "events": [ + { + "type": "turn_completed", + "timestamp": "2026-08-05T11:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "minimal", + "model_id": "claude-4-sonnet", + "response_tokens": 150 + } + } + ] +} diff --git a/tests/invalid/privacy-violation-query-at-intent.json b/tests/invalid/privacy-violation-query-at-intent.json index c08d585..6428d54 100644 --- a/tests/invalid/privacy-violation-query-at-intent.json +++ b/tests/invalid/privacy-violation-query-at-intent.json @@ -1,6 +1,7 @@ { "_test_description": "APPLICATION-LAYER VIOLATION: Turn at intent privacy with query_text present. Passes JSON Schema validation but violates section 5.5: query_text MUST NOT be present at intent level.", - "schema_version": "0.1", + "_expected_error": "'query_text' present on turn with privacy_level 'intent'", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440013", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/privacy-violation-query-at-minimal-batch.json b/tests/invalid/privacy-violation-query-at-minimal-batch.json new file mode 100644 index 0000000..77b0b70 --- /dev/null +++ b/tests/invalid/privacy-violation-query-at-minimal-batch.json @@ -0,0 +1,21 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: event batch whose turn_completed at minimal privacy carries query_text (section 5.5).", + "_expected_error": "'query_text' present on turn with privacy_level 'minimal'", + "document_type": "event_batch", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440644", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "turn_completed", + "timestamp": "2026-08-20T10:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "minimal", + "query_text": "What is the best coffee grinder under 100 pounds?", + "response_tokens": 120 + } + } + ] +} diff --git a/tests/invalid/privacy-violation-query-at-minimal-standalone.json b/tests/invalid/privacy-violation-query-at-minimal-standalone.json new file mode 100644 index 0000000..4bacc20 --- /dev/null +++ b/tests/invalid/privacy-violation-query-at-minimal-standalone.json @@ -0,0 +1,19 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: standalone envelope carrying a turn_completed at minimal privacy with query_text. The privacy gate applies wherever turns are emitted (section 5.5).", + "_expected_error": "'query_text' present on turn with privacy_level 'minimal'", + "document_type": "event", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440643", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "event": { + "type": "turn_completed", + "timestamp": "2026-08-20T10:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "minimal", + "query_text": "What is the best coffee grinder under 100 pounds?", + "response_tokens": 120 + } + } +} diff --git a/tests/invalid/privacy-violation-query-at-minimal.json b/tests/invalid/privacy-violation-query-at-minimal.json index 950a5bb..c0efce5 100644 --- a/tests/invalid/privacy-violation-query-at-minimal.json +++ b/tests/invalid/privacy-violation-query-at-minimal.json @@ -1,6 +1,7 @@ { "_test_description": "APPLICATION-LAYER VIOLATION: Turn at minimal privacy with query_text present. Passes JSON Schema validation but violates the privacy field gating rule (section 5.5): query_text MUST NOT be present when privacy_level is minimal.", - "schema_version": "0.1", + "_expected_error": "'query_text' present on turn with privacy_level 'minimal'", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440011", "started_at": "2026-03-28T10:00:00Z", "events": [ diff --git a/tests/invalid/privacy-violation-query-intent-at-minimal.json b/tests/invalid/privacy-violation-query-intent-at-minimal.json new file mode 100644 index 0000000..e29b62b --- /dev/null +++ b/tests/invalid/privacy-violation-query-intent-at-minimal.json @@ -0,0 +1,19 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: Turn at minimal privacy with query_intent present. Passes JSON Schema validation but violates the privacy field gating rule (section 5.5): query_intent MUST NOT be present when privacy_level is minimal.", + "_expected_error": "'query_intent' present on turn with privacy_level 'minimal'", + "schema_version": "1.0", + "session_id": "770e8400-e29b-41d4-a716-446655440221", + "started_at": "2026-08-05T11:00:00Z", + "events": [ + { + "type": "turn_completed", + "timestamp": "2026-08-05T11:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "minimal", + "query_intent": "purchase_intent", + "response_tokens": 150 + } + } + ] +} diff --git a/tests/invalid/privacy-violation-response-mode-at-minimal.json b/tests/invalid/privacy-violation-response-mode-at-minimal.json new file mode 100644 index 0000000..5cf7c78 --- /dev/null +++ b/tests/invalid/privacy-violation-response-mode-at-minimal.json @@ -0,0 +1,19 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: Turn at minimal privacy with response_mode present. Passes JSON Schema validation but violates the privacy field gating rule (section 5.5): response_mode MUST NOT be present when privacy_level is minimal.", + "_expected_error": "'response_mode' present on turn with privacy_level 'minimal'", + "schema_version": "1.0", + "session_id": "770e8400-e29b-41d4-a716-446655440224", + "started_at": "2026-08-05T11:00:00Z", + "events": [ + { + "type": "turn_completed", + "timestamp": "2026-08-05T11:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "minimal", + "response_mode": "standard", + "response_tokens": 150 + } + } + ] +} diff --git a/tests/invalid/privacy-violation-response-text-at-intent.json b/tests/invalid/privacy-violation-response-text-at-intent.json new file mode 100644 index 0000000..a02204c --- /dev/null +++ b/tests/invalid/privacy-violation-response-text-at-intent.json @@ -0,0 +1,21 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: turn at intent privacy with response_text present (section 5.5).", + "_expected_error": "'response_text' present on turn with privacy_level 'intent'", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440642", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "turn_completed", + "timestamp": "2026-08-20T10:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "response_text": "The Baratza Encore ESP is the best grinder under 100 pounds.", + "query_intent": "comparison", + "response_tokens": 120 + } + } + ] +} diff --git a/tests/invalid/privacy-violation-response-text-at-minimal.json b/tests/invalid/privacy-violation-response-text-at-minimal.json new file mode 100644 index 0000000..7390bbe --- /dev/null +++ b/tests/invalid/privacy-violation-response-text-at-minimal.json @@ -0,0 +1,19 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: Turn at minimal privacy with response_text present. Passes JSON Schema validation but violates the privacy field gating rule (section 5.5): response_text MUST NOT be present when privacy_level is minimal.", + "_expected_error": "'response_text' present on turn with privacy_level 'minimal'", + "schema_version": "1.0", + "session_id": "770e8400-e29b-41d4-a716-446655440220", + "started_at": "2026-08-05T11:00:00Z", + "events": [ + { + "type": "turn_completed", + "timestamp": "2026-08-05T11:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "minimal", + "response_text": "The Baratza Encore ESP is the best grinder under 100 pounds for espresso and filter alike.", + "response_tokens": 150 + } + } + ] +} diff --git a/tests/invalid/privacy-violation-response-type-at-minimal.json b/tests/invalid/privacy-violation-response-type-at-minimal.json new file mode 100644 index 0000000..a7101a2 --- /dev/null +++ b/tests/invalid/privacy-violation-response-type-at-minimal.json @@ -0,0 +1,19 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: Turn at minimal privacy with response_type present. Passes JSON Schema validation but violates the privacy field gating rule (section 5.5): response_type MUST NOT be present when privacy_level is minimal.", + "_expected_error": "'response_type' present on turn with privacy_level 'minimal'", + "schema_version": "1.0", + "session_id": "770e8400-e29b-41d4-a716-446655440223", + "started_at": "2026-08-05T11:00:00Z", + "events": [ + { + "type": "turn_completed", + "timestamp": "2026-08-05T11:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "minimal", + "response_type": "recommendation", + "response_tokens": 150 + } + } + ] +} diff --git a/tests/invalid/privacy-violation-topics-at-minimal.json b/tests/invalid/privacy-violation-topics-at-minimal.json new file mode 100644 index 0000000..a78a26a --- /dev/null +++ b/tests/invalid/privacy-violation-topics-at-minimal.json @@ -0,0 +1,22 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: Turn at minimal privacy with topics present. Passes JSON Schema validation but violates the privacy field gating rule (section 5.5): topics MUST NOT be present when privacy_level is minimal.", + "_expected_error": "'topics' present on turn with privacy_level 'minimal'", + "schema_version": "1.0", + "session_id": "770e8400-e29b-41d4-a716-446655440222", + "started_at": "2026-08-05T11:00:00Z", + "events": [ + { + "type": "turn_completed", + "timestamp": "2026-08-05T11:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "minimal", + "topics": [ + "coffee grinders", + "burr grinders" + ], + "response_tokens": 150 + } + } + ] +} diff --git a/tests/invalid/retrieved-bad-country.json b/tests/invalid/retrieved-bad-country.json new file mode 100644 index 0000000..df9ed41 --- /dev/null +++ b/tests/invalid/retrieved-bad-country.json @@ -0,0 +1,19 @@ +{ + "_test_description": "Edge retrieval with country 'usa'; the field is an ISO 3166-1 alpha-2 code (section 6.2).", + "_expected_error": "/events/0/data/country 'usa' does not match '^[A-Z]{2}$'", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440624", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "edge", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "country": "usa" + } + } + ] +} diff --git a/tests/invalid/retrieved-missing-identifier.json b/tests/invalid/retrieved-missing-identifier.json new file mode 100644 index 0000000..9d9b88a --- /dev/null +++ b/tests/invalid/retrieved-missing-identifier.json @@ -0,0 +1,19 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: content_retrieved event carrying neither content_url nor content_id. Passes JSON Schema (both are individually optional) but violates section 5.7.5: at least one MUST be present on every content event.", + "_expected_error": "'content_retrieved' carries neither content_url nor content_id", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440232", + "started_at": "2026-08-05T12:10:00Z", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-05T12:10:01Z", + "source_role": "agent", + "content_telemetry_id": "880e8400-e29b-41d4-a716-446655440233", + "data": { + "response_status": 200, + "media_type": "text" + } + } + ] +} diff --git a/tests/invalid/retrieved-missing-source-role.json b/tests/invalid/retrieved-missing-source-role.json new file mode 100644 index 0000000..055a19c --- /dev/null +++ b/tests/invalid/retrieved-missing-source-role.json @@ -0,0 +1,15 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: content_retrieved without source_role. Every retrieval event MUST say who observed it (sections 5.2.2, 5.7.1).", + "_expected_error": "content_retrieved event carries no source_role", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440630", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + } + ] +} diff --git a/tests/invalid/retrieved-response-status-out-of-range.json b/tests/invalid/retrieved-response-status-out-of-range.json new file mode 100644 index 0000000..cebdceb --- /dev/null +++ b/tests/invalid/retrieved-response-status-out-of-range.json @@ -0,0 +1,19 @@ +{ + "_test_description": "Edge retrieval with response_status 999, outside the HTTP status range 100-599 (section 6.2).", + "_expected_error": "/events/0/data/response_status 999 is greater than the maximum of 599", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440623", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "edge", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "response_status": 999 + } + } + ] +} diff --git a/tests/invalid/shared-ctx-token-two-presentations.json b/tests/invalid/shared-ctx-token-two-presentations.json new file mode 100644 index 0000000..fec7720 --- /dev/null +++ b/tests/invalid/shared-ctx-token-two-presentations.json @@ -0,0 +1,56 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: one event-level ctx_token appears on engagements bound to two different presentations. A token is minted for exactly one presentation occurrence (section 7.4.1).", + "_expected_error": "appears on engagements bound to two presentations", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440640", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440604", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:05Z", + "turn_id": "1", + "output_id": "response:2", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:06Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440602", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + }, + "ctx_token": "ct_5d1c9e4b2a7f8036" + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:07Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440604", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "engagement_type": "link_click" + }, + "ctx_token": "ct_5d1c9e4b2a7f8036" + } + ] +} diff --git a/tests/invalid/standalone-ctx-token-non-engagement.json b/tests/invalid/standalone-ctx-token-non-engagement.json new file mode 100644 index 0000000..8c9f427 --- /dev/null +++ b/tests/invalid/standalone-ctx-token-non-engagement.json @@ -0,0 +1,17 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: standalone envelope whose only session reference is a ctx_token but whose event is a content_grounded. An envelope ctx_token accompanies content_engaged events only (section 7.1).", + "_expected_error": "Envelope ctx_token accompanies a 'content_grounded' event", + "document_type": "event", + "schema_version": "1.0", + "ctx_token": "ct_9f3a1c7e2b8d4a06", + "event": { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 4200 + } + } +} diff --git a/tests/invalid/standalone-malformed-session-id.json b/tests/invalid/standalone-malformed-session-id.json new file mode 100644 index 0000000..03989f4 --- /dev/null +++ b/tests/invalid/standalone-malformed-session-id.json @@ -0,0 +1,13 @@ +{ + "_test_description": "Standalone envelope whose session_id is not a UUID (telemetry-event.json format: uuid).", + "_expected_error": "/session_id 'session-42' is not a 'uuid'", + "document_type": "event", + "schema_version": "1.0", + "session_id": "session-42", + "event": { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "agent", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + } +} diff --git a/tests/invalid/standalone-missing-document-type.json b/tests/invalid/standalone-missing-document-type.json index b0a0186..63508ae 100644 --- a/tests/invalid/standalone-missing-document-type.json +++ b/tests/invalid/standalone-missing-document-type.json @@ -1,6 +1,7 @@ { "_test_description": "Envelope-shaped document with an event key but no document_type. Per section 7.1 a document without document_type is treated as a session, and it fails the session schema.", - "schema_version": "0.1", + "_expected_error": "/ 'session_id' is a required property", + "schema_version": "1.0", "event": { "type": "content_retrieved", "timestamp": "2026-03-28T10:00:00Z", diff --git a/tests/invalid/standalone-missing-event.json b/tests/invalid/standalone-missing-event.json index c75bdfc..22ad0cc 100644 --- a/tests/invalid/standalone-missing-event.json +++ b/tests/invalid/standalone-missing-event.json @@ -1,6 +1,7 @@ { "_test_description": "Standalone event envelope (document_type 'event') with no event field. Fails the envelope schema, which requires event.", + "_expected_error": "/ 'event' is a required property", "document_type": "event", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "770e8400-e29b-41d4-a716-446655440010" } diff --git a/tests/invalid/standalone-missing-schema-version.json b/tests/invalid/standalone-missing-schema-version.json new file mode 100644 index 0000000..61b474b --- /dev/null +++ b/tests/invalid/standalone-missing-schema-version.json @@ -0,0 +1,12 @@ +{ + "_test_description": "Standalone event envelope missing schema_version, which the envelope schema requires.", + "_expected_error": "/ 'schema_version' is a required property", + "document_type": "event", + "session_id": "660e8400-e29b-41d4-a716-446655440626", + "event": { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "agent", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis" + } +} diff --git a/tests/invalid/standalone-missing-session-and-ctx-token.json b/tests/invalid/standalone-missing-session-and-ctx-token.json index 843c557..a4016e7 100644 --- a/tests/invalid/standalone-missing-session-and-ctx-token.json +++ b/tests/invalid/standalone-missing-session-and-ctx-token.json @@ -1,7 +1,8 @@ { "_test_description": "Standalone event envelope carrying neither session_id nor ctx_token. Passes JSON Schema (both are individually optional on the envelope) but violates section 5.7.5: an event MUST carry either session_id or ctx_token at Grounding conformance and above (section 7.1).", + "_expected_error": "neither session_id nor ctx_token", "document_type": "event", - "schema_version": "0.1", + "schema_version": "1.0", "event": { "type": "content_grounded", "timestamp": "2026-03-28T08:20:00Z", diff --git a/tests/invalid/turn-on-content-event.json b/tests/invalid/turn-on-content-event.json new file mode 100644 index 0000000..f6a4301 --- /dev/null +++ b/tests/invalid/turn-on-content-event.json @@ -0,0 +1,23 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: a turn object on a content_grounded event. Turn data is carried on turn_started and turn_completed only (sections 5.2, 5.4).", + "_expected_error": "Field 'turn' present on 'content_grounded' event", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440634", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 4200 + }, + "turn": { + "privacy_level": "minimal" + } + } + ] +} diff --git a/tests/invalid/withdrawn-ip-hash-batch.json b/tests/invalid/withdrawn-ip-hash-batch.json new file mode 100644 index 0000000..a79b8d5 --- /dev/null +++ b/tests/invalid/withdrawn-ip-hash-batch.json @@ -0,0 +1,18 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: event batch whose retrieval carries the withdrawn ip_hash (section 9.1), in the batch shape.", + "_expected_error": "'ip_hash' in data", + "document_type": "event_batch", + "schema_version": "1.0", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "edge", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "purpose": "inference", + "ip_hash": "sha256:e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6" + } + } + ] +} diff --git a/tests/invalid/withdrawn-ip-hash-session.json b/tests/invalid/withdrawn-ip-hash-session.json new file mode 100644 index 0000000..fc09f04 --- /dev/null +++ b/tests/invalid/withdrawn-ip-hash-session.json @@ -0,0 +1,20 @@ +{ + "_test_description": "APPLICATION-LAYER VIOLATION: session document whose retrieval carries the withdrawn ip_hash (section 9.1), in the session shape.", + "_expected_error": "'ip_hash' in data", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440647", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "edge", + "content_url": "https://www.example-news.com/economy/rate-decision-analysis", + "data": { + "purpose": "inference", + "ip_hash": "sha256:e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6" + } + } + ] +} diff --git a/tests/invalid/withdrawn-ip-hash.json b/tests/invalid/withdrawn-ip-hash.json new file mode 100644 index 0000000..9ff7f2a --- /dev/null +++ b/tests/invalid/withdrawn-ip-hash.json @@ -0,0 +1,22 @@ +{ + "_test_description": "Standalone event envelope from a CDN carrying ip_hash in data. The field was withdrawn in v1 (section 9.1): hashing does not anonymise a value drawn from a space small enough to enumerate. Passes JSON Schema because event data accepts additional properties, so the rule is enforced at the application layer.", + "_expected_error": "'ip_hash' in data", + "document_type": "event", + "schema_version": "1.0", + "event": { + "type": "content_retrieved", + "timestamp": "2026-03-28T08:15:00Z", + "source_role": "edge", + "content_telemetry_id": "990e8400-e29b-41d4-a716-446655440051", + "content_url": "https://www.telegraph.co.uk/business/2026/03/28/ftse-100-markets-live", + "data": { + "user_agent": "PerplexityBot/1.0", + "purpose": "inference", + "response_status": 200, + "asn": 396982, + "asn_org": "Perplexity AI", + "country": "US", + "ip_hash": "sha256:e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6" + } + } +} diff --git a/tests/mutation_smoke.py b/tests/mutation_smoke.py new file mode 100644 index 0000000..ec6ba19 --- /dev/null +++ b/tests/mutation_smoke.py @@ -0,0 +1,321 @@ +#!/usr/bin/env python3 +"""Mutation smoke test for the conformance suite. + +Replays suite-weakening mutations - each of which the suite once missed - +against a scratch copy of the repository and confirms validate.py fails under +every one of them. The working tree is never modified. + +A mutation "survives" when the mutated suite still exits 0; any survivor +is a real detection gap and fails this script. + +Usage: + python tests/mutation_smoke.py +""" + +import json +import shutil +import subprocess +import sys +import tempfile +from pathlib import Path + +REPO = Path(__file__).resolve().parent.parent + + +def mutate_drop_format_checker(root): + """Build every validator without a format checker (formats become no-ops).""" + p = root / "tests" / "validate.py" + src = p.read_text() + mutated = src.replace(", format_checker=FORMAT_CHECKER", "") + assert mutated != src + p.write_text(mutated) + return "malformed-parent-session-id.json" + + +def mutate_gut_withdrawn_ip_hash(root): + """Gut the withdrawn-ip-hash fixture: drop ip_hash and break the event + some other way, so it fails schema for a reason unrelated to its rule.""" + p = root / "tests" / "invalid" / "withdrawn-ip-hash.json" + doc = json.loads(p.read_text()) + del doc["event"]["data"]["ip_hash"] + del doc["event"]["timestamp"] + p.write_text(json.dumps(doc, indent=2) + "\n") + return "withdrawn-ip-hash.json" + + +def mutate_shrink_content_event_types(root): + """Remove content_presented from CONTENT_EVENT_TYPES.""" + p = root / "tests" / "validate.py" + src = p.read_text() + mutated = src.replace('"content_cited", "content_presented", "content_engaged",', + '"content_cited", "content_engaged",') + assert mutated != src + p.write_text(mutated) + return "presented-missing-identifier.json" + + +def mutate_shrink_privacy_forbidden_fields(root): + """Remove topics from the fields forbidden at minimal privacy.""" + p = root / "tests" / "validate.py" + src = p.read_text() + mutated = src.replace('"query_text", "response_text", "query_intent", "topics",', + '"query_text", "response_text", "query_intent",') + assert mutated != src + p.write_text(mutated) + return "privacy-violation-topics-at-minimal.json" + + +def mutate_zero_uuid_engagement(root): + """Point a valid session's engagement at an all-zeros presentation_id.""" + p = root / "tests" / "valid" / "session-citation-tier.json" + doc = json.loads(p.read_text()) + mutated = False + for event in doc["events"]: + if event.get("type") == "content_engaged": + event["presentation_id"] = "00000000-0000-0000-0000-000000000000" + mutated = True + assert mutated + p.write_text(json.dumps(doc, indent=2) + "\n") + return "session-citation-tier.json" + + +MUTATIONS = [ + ("drop format_checker", mutate_drop_format_checker), + ("gut withdrawn-ip-hash fixture", mutate_gut_withdrawn_ip_hash), + ("remove content_presented from CONTENT_EVENT_TYPES", mutate_shrink_content_event_types), + ("shrink PRIVACY_FORBIDDEN_FIELDS", mutate_shrink_privacy_forbidden_fields), + ("engagement at all-zeros presentation_id", mutate_zero_uuid_engagement), +] + + +# --------------------------------------------------------------------------- +# Review mutations (22 August pre-freeze review). Each weakens one rule the +# suite previously pinned in a single document shape, or not at all, and names +# the fixture that now fails under it. +# --------------------------------------------------------------------------- + +def _edit_text(root, rel, old, new): + path = root / rel + src = path.read_text() + assert old in src, (rel, old) + path.write_text(src.replace(old, new)) + + +def _edit_json(root, rel, fn): + path = root / rel + doc = json.loads(path.read_text()) + fn(doc) + path.write_text(json.dumps(doc, indent=2) + "\n") + + +def _event_branch(schema, etype): + for branch in schema["$defs"]["TelemetryEvent"]["allOf"]: + if branch["if"]["properties"]["type"]["const"] == etype: + return branch["then"] + raise KeyError(etype) + + +def _text(rel, old, new, fixture): + def mutate(root): + _edit_text(root, rel, old, new) + return fixture + return mutate + + +def _json(rel, fn, fixture): + def mutate(root): + _edit_json(root, rel, fn) + return fixture + return mutate + + +_V = "tests/validate.py" +_S = "telemetry-session.json" +_E = "telemetry-event.json" +_B = "telemetry-event-batch.json" +_M = "manifest.json" + +REVIEW_MUTATIONS = [ + # validate.py application-layer weakenings + ("exempt any envelope containing a retrieval from the session/ctx_token rule", + _text(_V, 'if types <= {"content_retrieved"}:', 'if "content_retrieved" in types:', + "batch-missing-session-mixed-retrieval.json")), + ("stop forbidding response_text at intent", + _text(_V, '"intent": {\n "query_text", "response_text",\n },', '"intent": {\n "query_text",\n },', + "privacy-violation-response-text-at-intent.json")), + ("drop the agent_cached -> cached:true half of the provenance rule", + _text(_V, 'if provenance == "agent_cached" and cached is not True:', 'if False:', + "grounding-provenance-cached-conflict-agent-cached.json")), + ("skip referential integrity when the session has no content_presented events", + _text(_V, 'if pid and pid not in presented_ids:', 'if presented_ids and pid and pid not in presented_ids:', + "engaged-presentation-id-no-presentations.json")), + ("privacy gating on session documents only", + _text(_V, ' violations = []\n\n for event in _iter_events(data):\n turn = event.get("turn")', + ' violations = []\n if is_standalone_event(data) or is_event_batch(data):\n return []\n for event in _iter_events(data):\n turn = event.get("turn")', + "privacy-violation-query-at-minimal-standalone.json")), + ("content-identifier rule on session documents only", + _text(_V, ' violations = []\n for event in _iter_events(data):\n if event.get("type") not in CONTENT_EVENT_TYPES:', + ' violations = []\n if is_standalone_event(data) or is_event_batch(data):\n return []\n for event in _iter_events(data):\n if event.get("type") not in CONTENT_EVENT_TYPES:', + "grounded-missing-identifier-standalone.json")), + ("ip_hash prohibition on standalone envelopes only", + _text(_V, ' violations = []\n for event in _iter_events(data):\n event_data = event.get("data")\n if not isinstance(event_data, dict):\n continue\n for field, reason', + ' violations = []\n if not is_standalone_event(data):\n return []\n for event in _iter_events(data):\n event_data = event.get("data")\n if not isinstance(event_data, dict):\n continue\n for field, reason', + "withdrawn-ip-hash-session.json")), + ("domains check accepts lookalike hosts (endswith without the dot)", + _text(_V, 'not bare.endswith("." + host)', 'not bare.endswith(host)', "manifest-lookalike-domain.json")), + ("drop the source_role requirement on retrievals", + _text(_V, 'if etype == "content_retrieved" and not event.get("source_role"):', 'if False:', + "retrieved-missing-source-role.json")), + ("drop the field-placement rules", + _text(_V, ' if event.get(field) is not None and etype not in allowed:', ' if False:', + "ctx-token-on-grounded.json")), + ("drop the duplicate event id rule", + _text(_V, ' if eid in seen_ids:', ' if False:', "duplicate-event-id.json")), + ("drop the same-content rule for engagements", + _text(_V, 'if pid in presented and not _same_content(e, presented[pid]):', 'if False:', + "engaged-presentation-content-mismatch.json")), + ("drop the one-token-one-presentation rule", + _text(_V, ' if bound != pid:', ' if False:', "shared-ctx-token-two-presentations.json")), + ("drop the envelope ctx_token rule", + _text(_V, " if not data.get(\"ctx_token\"):\n return []\n violations = []", " return []\n violations = []", + "standalone-ctx-token-non-engagement.json")), + ("drop the root-manifest rule for domains", + _text(_V, 'if "domains" in data and parsed.path != "/.well-known/content-telemetry.json":', 'if False:', + "manifest-domains-on-path-manifest.json")), + ("drop the agent/platform rule for ctx_resolution", + _text(_V, 'if telemetry.get("ctx_resolution") and not roles & {"agent", "platform"}:', 'if False:', + "manifest-ctx-resolution-on-content-owner.json")), + # telemetry-session.json + ("accept the v0.1 wire version", + _text(_S, '"const": "1.0",\n "description": "Content Telemetry schema version"', + '"enum": ["0.1", "1.0"],\n "description": "Content Telemetry schema version"', + "legacy-schema-version-0-1.json")), + ("remove the event-level ctx_token pattern", + _json(_S, lambda d: d["$defs"]["TelemetryEvent"]["properties"]["ctx_token"].pop("pattern"), + "event-level-ctx-token-bad-pattern.json")), + ("remove format:uuid from event ids", + _json(_S, lambda d: [d["$defs"]["TelemetryEvent"]["properties"][k].pop("format") for k in ("id", "citation_id", "presentation_id")], + "malformed-event-id.json")), + ("open the CitationType enum", + _json(_S, lambda d: d["$defs"]["CitationType"].pop("enum"), "citation-type-invalid.json")), + ("open the CitationPosition enum", + _json(_S, lambda d: d["$defs"]["CitationPosition"].pop("enum"), "citation-position-invalid.json")), + ("open the GroundingScope enum", + _json(_S, lambda d: d["$defs"]["GroundingScope"].pop("enum"), "grounding-scope-invalid.json")), + ("open the SourceProvenance enum", + _json(_S, lambda d: d["$defs"]["SourceProvenance"].pop("enum"), "grounding-provenance-invalid.json")), + ("drop required scope on content_grounded", + _json(_S, lambda d: _event_branch(d, "content_grounded")["properties"]["data"].pop("required"), + "grounding-scope-missing.json")), + ("drop required data on content_grounded", + _json(_S, lambda d: _event_branch(d, "content_grounded")["required"].remove("data"), "grounding-missing-data.json")), + ("drop required data on content_presented", + _json(_S, lambda d: _event_branch(d, "content_presented")["required"].remove("data"), "presented-missing-data.json")), + ("drop required data on content_cited", + _json(_S, lambda d: _event_branch(d, "content_cited")["required"].remove("data"), "cited-missing-data.json")), + ("remove the sha256 hash patterns", + _text(_S, '"pattern": "^sha256:[a-f0-9]{64}$"', '"type": "string"', "malformed-content-hash.json")), + ("remove minLength on identifier-scheme prefix", + _json(_M, lambda d: d["properties"]["identifier_schemes"]["items"]["properties"]["prefix"].pop("minLength"), "manifest-identifier-scheme-empty-prefix.json")), + ("drop the identifier_schemes owner-role placement check", + _text(_V, 'if data.get("identifier_schemes") and "content_owner" not in roles:', 'if False:', + "manifest-identifier-schemes-on-non-owner.json")), + ("remove minLength on output_id", + _json(_S, lambda d: d["$defs"]["TelemetryEvent"]["properties"]["output_id"].pop("minLength"), "empty-output-id.json")), + ("remove minimum on tokens_ingested", + _json(_S, lambda d: _event_branch(d, "content_grounded")["properties"]["data"]["properties"]["tokens_ingested"].pop("minimum"), + "grounded-negative-tokens-ingested.json")), + ("remove minimum on excerpt_chars", + _json(_S, lambda d: _event_branch(d, "content_cited")["properties"]["data"]["properties"]["excerpt_chars"].pop("minimum"), + "cited-negative-excerpt-chars.json")), + ("remove the response_status bounds", + _json(_S, lambda d: _event_branch(d, "content_retrieved")["properties"]["data"]["properties"]["response_status"].pop("maximum"), + "retrieved-response-status-out-of-range.json")), + ("remove the country pattern", + _json(_S, lambda d: _event_branch(d, "content_retrieved")["properties"]["data"]["properties"]["country"].pop("pattern"), + "retrieved-bad-country.json")), + ("drop required scheme on content_fingerprint", + _json(_S, lambda d: d["$defs"]["ContentFingerprint"]["required"].remove("scheme"), "grounding-fingerprint-missing-scheme.json")), + ("drop format:uri from the turn URL arrays", + _json(_S, lambda d: [d["$defs"]["ConversationTurn"]["properties"][k]["items"].pop("format") for k in ("content_urls_retrieved", "content_urls_cited")], + "malformed-content-urls-cited.json")), + ("drop format:date-time from started_at", + _json(_S, lambda d: d["properties"]["started_at"].pop("format"), "malformed-started-at.json")), + ("drop format:uuid from session_id", + _json(_S, lambda d: d["properties"]["session_id"].pop("format"), "malformed-session-id.json")), + # envelope schemas + ("batch: drop the presentation_id-unless-ctx_token conditional", + _json(_B, lambda d: d.pop("allOf"), "batch-engaged-missing-presentation-id.json")), + ("batch: remove the envelope ctx_token pattern", + _json(_B, lambda d: d["properties"]["ctx_token"].pop("pattern"), "batch-ctx-token-bad-pattern.json")), + ("event envelope: drop schema_version from required", + _json(_E, lambda d: d["required"].remove("schema_version"), "standalone-missing-schema-version.json")), + ("event envelope: drop format:uuid from session_id", + _json(_E, lambda d: d["properties"]["session_id"].pop("format"), "standalone-malformed-session-id.json")), + # manifest.json + ("manifest: drop required mode on coverage entries", + _json(_M, lambda d: d["properties"]["telemetry"]["properties"]["coverage"]["additionalProperties"].pop("required"), + "manifest-coverage-missing-mode.json")), + ("manifest: drop required endpoint", + _json(_M, lambda d: d["properties"]["telemetry"].pop("required"), "manifest-missing-endpoint.json")), + ("manifest: drop minItems on roles", + _json(_M, lambda d: d["properties"]["roles"].pop("minItems"), "manifest-empty-roles.json")), + ("manifest: drop uniqueItems on roles", + _json(_M, lambda d: d["properties"]["roles"].pop("uniqueItems"), "manifest-duplicate-roles.json")), + ("manifest: drop the https pattern on ctx_resolution", + _json(_M, lambda d: d["properties"]["telemetry"]["properties"]["ctx_resolution"].pop("pattern"), "manifest-ctx-resolution-http.json")), + ("manifest: drop the https pattern on endpoint", + _json(_M, lambda d: d["properties"]["telemetry"]["properties"]["endpoint"].pop("pattern"), "manifest-endpoint-http.json")), + ("manifest: drop the https pattern on id", + _json(_M, lambda d: d["properties"]["id"].pop("pattern"), "manifest-id-http.json")), + ("manifest: drop the well-known location pattern on id", + _json(_M, lambda d: d["properties"]["id"].pop("pattern"), "manifest-id-not-well-known.json")), + ("manifest: drop format:uri on id", + _json(_M, lambda d: d["properties"]["id"].pop("format"), "manifest-id-not-a-uri.json")), + ("manifest: drop const Ed25519 on keys[].type", + _json(_M, lambda d: d["properties"]["keys"]["items"]["properties"]["type"].pop("const"), "manifest-bad-key-type.json")), + ("manifest: drop required id on keys[]", + _json(_M, lambda d: d["properties"]["keys"]["items"]["required"].remove("id"), "manifest-key-missing-id.json")), +] + +MUTATIONS += REVIEW_MUTATIONS + + +def run_one(name, mutate): + with tempfile.TemporaryDirectory(prefix="ct-mutation-") as tmp: + root = Path(tmp) / "repo" + root.mkdir() + for schema in ("telemetry-session.json", "telemetry-event.json", + "telemetry-event-batch.json", "manifest.json"): + shutil.copy(REPO / schema, root / schema) + shutil.copytree(REPO / "tests", root / "tests") + expected_fixture = mutate(root) + proc = subprocess.run( + [sys.executable, str(root / "tests" / "validate.py")], + capture_output=True, text=True, + ) + caught = proc.returncode != 0 and f"FAIL {expected_fixture}" in proc.stdout + return caught, expected_fixture, proc + + +def main(): + survivors = 0 + for name, mutate in MUTATIONS: + caught, fixture, proc = run_one(name, mutate) + if caught: + print(f" CAUGHT {name} (failed via {fixture})") + else: + survivors += 1 + print(f" SURVIVED {name} (expected {fixture} to fail)") + print(f" exit={proc.returncode}") + for line in proc.stdout.splitlines()[-5:]: + print(f" {line}") + print() + print("=" * 60) + print(f"SUMMARY: {len(MUTATIONS) - survivors}/{len(MUTATIONS)} mutations caught") + print("=" * 60) + return 1 if survivors else 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/tests/valid/event-batch-agent.json b/tests/valid/event-batch-agent.json index de25dde..3351c5b 100644 --- a/tests/valid/event-batch-agent.json +++ b/tests/valid/event-batch-agent.json @@ -1,7 +1,7 @@ { "_test_description": "Event batch envelope from an agent at Grounding conformance, carrying the session-level fields (session_id, agent_id, started_at) that apply to every event in the batch. The agent buffers events within a session and flushes them together.", "document_type": "event_batch", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440007", "agent_id": "assistant.example.com", "started_at": "2026-03-28T08:19:55Z", @@ -17,13 +17,35 @@ "type": "content_grounded", "timestamp": "2026-03-28T08:20:03Z", "source_role": "agent", - "content_url": "https://arstechnica.com/science/2026/03/new-battery-chemistry-breakthrough/" + "content_url": "https://arstechnica.com/science/2026/03/new-battery-chemistry-breakthrough/", + "data": { + "scope": "session", + "chars_ingested": 5200 + } }, { + "id": "990e8400-e29b-41d4-a716-446655440062", "type": "content_cited", "timestamp": "2026-03-28T08:20:05Z", + "output_id": "response:1", "source_role": "agent", - "content_url": "https://arstechnica.com/science/2026/03/new-battery-chemistry-breakthrough/" + "content_url": "https://arstechnica.com/science/2026/03/new-battery-chemistry-breakthrough/", + "data": { + "citation_type": "reference" + } + }, + { + "id": "990e8400-e29b-41d4-a716-446655440063", + "type": "content_presented", + "timestamp": "2026-03-28T08:20:06Z", + "output_id": "response:1", + "citation_id": "990e8400-e29b-41d4-a716-446655440062", + "source_role": "agent", + "content_url": "https://arstechnica.com/science/2026/03/new-battery-chemistry-breakthrough/", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } } ] } diff --git a/tests/valid/event-batch-destination-engagements.json b/tests/valid/event-batch-destination-engagements.json new file mode 100644 index 0000000..1fdfc2b --- /dev/null +++ b/tests/valid/event-batch-destination-engagements.json @@ -0,0 +1,25 @@ +{ + "_test_description": "Destination-reported engagements delivered as a batch under one ctx_token: a direct-link surface minted the token per presentation, so two clicks on the same presentation share it and are distinguished by timestamp at resolution (section 7.4.1). No presentation_id and no session_id: the consumer restores both from the token.", + "document_type": "event_batch", + "schema_version": "1.0", + "ctx_token": "ct_77b41f0ac93e5d28", + "manifest_ref": "https://shop.example.com/.well-known/content-telemetry.json", + "events": [ + { + "type": "content_engaged", + "timestamp": "2026-08-20T11:15:42Z", + "content_url": "https://shop.example.com/products/anc-headphones-x9", + "data": { + "engagement_type": "link_click" + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T11:21:09Z", + "content_url": "https://shop.example.com/products/anc-headphones-x9", + "data": { + "engagement_type": "link_click" + } + } + ] +} diff --git a/tests/valid/event-batch-edge.json b/tests/valid/event-batch-edge.json index 8ae083d..8611ff4 100644 --- a/tests/valid/event-batch-edge.json +++ b/tests/valid/event-batch-edge.json @@ -1,7 +1,7 @@ { "_test_description": "Event batch envelope from a CDN at Retrieval conformance level. No session_id (content owner has no session context); the edge platform buffers detections across requests and flushes them as one batch. Validated against telemetry-event-batch.json envelope schema.", "document_type": "event_batch", - "schema_version": "0.1", + "schema_version": "1.0", "events": [ { "type": "content_retrieved", @@ -11,7 +11,7 @@ "content_url": "https://www.telegraph.co.uk/business/2026/03/28/ftse-100-markets-live", "data": { "user_agent": "PerplexityBot/1.0", - "bot_category": "inference", + "purpose": "inference", "verified": false, "cache_status": "miss", "response_status": 200 @@ -24,7 +24,7 @@ "content_url": "https://www.telegraph.co.uk/business/2026/03/28/bank-of-england-rates", "data": { "user_agent": "GPTBot/1.2", - "bot_category": "training", + "purpose": "training", "verified": true, "cache_status": "hit", "response_status": 200 diff --git a/tests/valid/event-batch-grounding-provenance.json b/tests/valid/event-batch-grounding-provenance.json new file mode 100644 index 0000000..9184ef2 --- /dev/null +++ b/tests/valid/event-batch-grounding-provenance.json @@ -0,0 +1,36 @@ +{ + "_test_description": "Batch envelope carrying grounding events from cached and third-party-sourced representations, exercising the shared provenance schema across batch delivery.", + "document_type": "event_batch", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440181", + "agent_id": "enterprise-rag-v1", + "started_at": "2026-08-12T10:10:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-12T10:10:01Z", + "content_id": "publisher:cached:7", + "data": { + "scope": "session", + "cached": true, + "provenance": "agent_cached", + "chars_ingested": 900 + } + }, + { + "type": "content_grounded", + "timestamp": "2026-08-12T10:10:02Z", + "content_id": "repository:item:9", + "data": { + "scope": "turn", + "cached": false, + "provenance": "third_party_sourced", + "chars_ingested": 500, + "content_fingerprint": { + "scheme": "https://example.org/fingerprint/v1", + "detected": false + } + } + } + ] +} diff --git a/tests/valid/event-standalone-agent.json b/tests/valid/event-standalone-agent.json index c5a58a0..095945c 100644 --- a/tests/valid/event-standalone-agent.json +++ b/tests/valid/event-standalone-agent.json @@ -1,7 +1,7 @@ { "_test_description": "Standalone event envelope from an agent, with session_id as a foreign key. The agent reports its own retrieval as a standalone event for streaming delivery.", "document_type": "event", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440006", "event": { "type": "content_retrieved", diff --git a/tests/valid/event-standalone-child-session.json b/tests/valid/event-standalone-child-session.json new file mode 100644 index 0000000..59c9f37 --- /dev/null +++ b/tests/valid/event-standalone-child-session.json @@ -0,0 +1,20 @@ +{ + "_test_description": "Standalone event envelope from a child agent session in a multi-agent topology: parent_session_id on the envelope identifies the orchestrating session. The grounded content entered the sub-agent's generation context.", + "document_type": "event", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440170", + "parent_session_id": "660e8400-e29b-41d4-a716-446655440171", + "agent_id": "research-subagent-v2", + "started_at": "2026-08-01T16:00:00Z", + "event": { + "type": "content_grounded", + "timestamp": "2026-08-01T16:00:02Z", + "content_url": "https://www.example-journal.org/analysis/press-freedom-ruling", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 4100, + "tokens_ingested": 1000 + } + } +} diff --git a/tests/valid/event-standalone-edge.json b/tests/valid/event-standalone-edge.json index 4acec96..5a7c415 100644 --- a/tests/valid/event-standalone-edge.json +++ b/tests/valid/event-standalone-edge.json @@ -1,7 +1,7 @@ { "_test_description": "Standalone event envelope from a CDN at Retrieval conformance level. No session_id (content owner has no session context). source_role is edge with full edge enrichment data. Validated against telemetry-event.json envelope schema.", "document_type": "event", - "schema_version": "0.1", + "schema_version": "1.0", "event": { "type": "content_retrieved", "timestamp": "2026-03-28T08:15:00Z", @@ -10,7 +10,7 @@ "content_url": "https://www.telegraph.co.uk/business/2026/03/28/ftse-100-markets-live", "data": { "user_agent": "PerplexityBot/1.0", - "bot_category": "inference", + "purpose": "inference", "verified": false, "cache_status": "miss", "response_status": 200, @@ -18,8 +18,7 @@ "ja4": "t13d1516h2_5b57614c22b0_7cb938dcc8ab", "asn": 396982, "asn_org": "Perplexity AI", - "country": "US", - "ip_hash": "sha256:e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6" + "country": "US" } } } diff --git a/tests/valid/event-standalone-engaged-ctx-token.json b/tests/valid/event-standalone-engaged-ctx-token.json index 659609c..9a0d9cc 100644 --- a/tests/valid/event-standalone-engaged-ctx-token.json +++ b/tests/valid/event-standalone-engaged-ctx-token.json @@ -1,7 +1,7 @@ { - "_test_description": "Standalone content_engaged (link_click) reported from a landing page after a click-out. Carries ctx_token in place of session_id (section 7.1): the destination corroborates the click without receiving the session UUID. The telemetry consumer resolves ctx_token to the originating session's click manifest.", + "_test_description": "Standalone content_engaged (link_click) reported from a landing page after a click-out. Carries ctx_token in place of session_id and no presentation_id: the destination cannot know the presentation UUID, and the consumer restores the token's presentation binding at resolution (section 7.4).", "document_type": "event", - "schema_version": "0.1", + "schema_version": "1.0", "ctx_token": "ct_9f3a1c7e2b8d4a06", "event": { "type": "content_engaged", diff --git a/tests/valid/event-standalone-engaged-redirect.json b/tests/valid/event-standalone-engaged-redirect.json new file mode 100644 index 0000000..6dd4815 --- /dev/null +++ b/tests/valid/event-standalone-engaged-redirect.json @@ -0,0 +1,14 @@ +{ + "_test_description": "Destination-reported click after a redirect chain: the presented short link redirected to the canonical URL, and the ctx_token and ctx_iss query parameters were propagated through the same-domain hops (section 7.4). The destination reports the canonical URL it serves; correlation runs through the token, not the URL.", + "document_type": "event", + "schema_version": "1.0", + "ctx_token": "ct_77b41f0ac93e5d28", + "event": { + "type": "content_engaged", + "timestamp": "2026-08-10T11:15:42Z", + "content_url": "https://shop.example.com/products/anc-headphones-x9", + "data": { + "engagement_type": "link_click" + } + } +} diff --git a/tests/valid/event-standalone-grounded-provenance.json b/tests/valid/event-standalone-grounded-provenance.json new file mode 100644 index 0000000..0bfff6b --- /dev/null +++ b/tests/valid/event-standalone-grounded-provenance.json @@ -0,0 +1,25 @@ +{ + "_test_description": "Standalone grounding of a third-party-sourced representation that the agent had cached (cached true is unconstrained for third_party_sourced, section 6.4), identified by an ISCC content_id alongside the URL, with an iscc fingerprint detection.", + "document_type": "event", + "schema_version": "1.0", + "session_id": "550e8400-e29b-41d4-a716-446655440000", + "agent_id": "agent-example", + "started_at": "2026-06-20T10:00:00Z", + "event": { + "id": "6ba7b810-9dad-11d1-80b4-00c04fd430c8", + "type": "content_grounded", + "timestamp": "2026-06-20T10:00:01Z", + "content_url": "https://example.com/article", + "content_id": "ISCC:KACYPXW445FTYNJ3", + "data": { + "scope": "turn", + "cached": true, + "provenance": "third_party_sourced", + "tokens_ingested": 512, + "content_fingerprint": { + "scheme": "iscc", + "detected": true + } + } + } +} diff --git a/tests/valid/event-standalone-grounding-fingerprint.json b/tests/valid/event-standalone-grounding-fingerprint.json new file mode 100644 index 0000000..45a5db8 --- /dev/null +++ b/tests/valid/event-standalone-grounding-fingerprint.json @@ -0,0 +1,26 @@ +{ + "_test_description": "Standalone grounding event from a publisher-authorised live fetch with an emitter-reported generic fingerprint detection. The event schema is shared by session, standalone and batch envelopes.", + "document_type": "event", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440180", + "agent_id": "research-agent-v1", + "started_at": "2026-08-12T10:00:00Z", + "event": { + "type": "content_grounded", + "timestamp": "2026-08-12T10:00:01Z", + "content_url": "https://publisher.example/articles/42", + "content_id": "publisher:article:42", + "data": { + "scope": "turn", + "cached": false, + "provenance": "agent_fetched", + "chars_ingested": 2400, + "content_hash": "sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa", + "content_fingerprint": { + "scheme": "https://example.org/fingerprint/v1", + "detected": true, + "value": "fp_42" + } + } + } +} diff --git a/tests/valid/event-standalone-index.json b/tests/valid/event-standalone-index.json new file mode 100644 index 0000000..6441b16 --- /dev/null +++ b/tests/valid/event-standalone-index.json @@ -0,0 +1,18 @@ +{ + "_test_description": "Standalone retrieval reported by a marketplace index (source_role index) that served the content to the agent: identified by the marketplace catalogue id with no canonical URL, resolved by content_id prefix registration (sections 4.4, 7.3).", + "document_type": "event", + "schema_version": "1.0", + "manifest_ref": "https://marketplace.example.com/.well-known/content-telemetry.json", + "event": { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "index", + "content_telemetry_id": "990e8400-e29b-41d4-a716-446655440602", + "content_id": "mkt:gridnews:5520", + "license_ref": "agreement-2026-017:gridnews", + "data": { + "media_type": "text", + "content_depth": "full" + } + } +} diff --git a/tests/valid/event-standalone-origin.json b/tests/valid/event-standalone-origin.json new file mode 100644 index 0000000..6c3f0f6 --- /dev/null +++ b/tests/valid/event-standalone-origin.json @@ -0,0 +1,19 @@ +{ + "_test_description": "Standalone retrieval reported by the content owner's own web server (source_role origin) at Retrieval conformance: no session context, origin enrichment fields (6.3) and content_depth recording that the agent reached the abstract only.", + "document_type": "event", + "schema_version": "1.0", + "event": { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "origin", + "content_telemetry_id": "990e8400-e29b-41d4-a716-446655440601", + "content_url": "https://journals.example.com/article/10.1000/xyz123", + "content_id": "doi:10.1000/xyz123", + "data": { + "user_agent": "ClaudeBot/1.0", + "response_status": 200, + "media_type": "text", + "content_depth": "abstract" + } + } +} diff --git a/tests/valid/event-standalone-presented.json b/tests/valid/event-standalone-presented.json new file mode 100644 index 0000000..58ce6bf --- /dev/null +++ b/tests/valid/event-standalone-presented.json @@ -0,0 +1,20 @@ +{ + "_test_description": "Standalone presentation envelope using the shared modality-neutral presentation schema.", + "document_type": "event", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440080", + "agent_id": "voice-assistant-v1", + "started_at": "2026-07-18T12:30:00Z", + "event": { + "id": "770e8400-e29b-41d4-a716-446655440081", + "type": "content_presented", + "timestamp": "2026-07-18T12:30:05Z", + "output_id": "audio-response:1", + "content_id": "publisher:article:1", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "spoken_credit", + "media_type": "audio" + } + } +} diff --git a/tests/valid/event-standalone-terms-ref.json b/tests/valid/event-standalone-terms-ref.json new file mode 100644 index 0000000..4a3ac4e --- /dev/null +++ b/tests/valid/event-standalone-terms-ref.json @@ -0,0 +1,15 @@ +{ + "_test_description": "Standalone retrieval event reported under governing terms with no grant: terms_ref names the terms (5.2.4), license_ref is absent, and the envelope carries manifest_ref (7.1) to tie the emitter to a domain without a session document.", + "document_type": "event", + "schema_version": "1.0", + "agent_id": "assistant.example.com", + "manifest_ref": "https://assistant.example.com/.well-known/content-telemetry.json", + "event": { + "type": "content_retrieved", + "timestamp": "2026-08-13T15:00:04Z", + "source_role": "agent", + "content_telemetry_id": "990e8400-e29b-41d4-a716-446655440071", + "content_url": "https://example.com/2026/08/13/markets-live", + "terms_ref": "https://terms.example.com/search-only/v2" + } +} diff --git a/tests/valid/manifest-agent-ctx-resolution.json b/tests/valid/manifest-agent-ctx-resolution.json new file mode 100644 index 0000000..0dc8e87 --- /dev/null +++ b/tests/valid/manifest-agent-ctx-resolution.json @@ -0,0 +1,12 @@ +{ + "_test_description": "Agent manifest declaring a click-token resolution endpoint in telemetry.ctx_resolution. A destination receiving ctx_iss=searchco.com/agents/web-search resolves this manifest and presents the token at the declared endpoint (sections 7.4, 8.5).", + "schema_version": "1.0", + "id": "https://searchco.com/agents/web-search/.well-known/content-telemetry.json", + "roles": ["agent"], + "operator": { "name": "SearchCo" }, + "telemetry": { + "endpoint": "https://telemetry.example.com/v1/events", + "conformance_level": "citation", + "ctx_resolution": "https://telemetry.example.com/v1/ctx/resolve" + } +} diff --git a/tests/valid/manifest-agent-with-keys.json b/tests/valid/manifest-agent-with-keys.json index 7b90977..fbb629e 100644 --- a/tests/valid/manifest-agent-with-keys.json +++ b/tests/valid/manifest-agent-with-keys.json @@ -1,6 +1,6 @@ { "_test_description": "Agent manifest served under a path prefix, with an Ed25519 signing key (including expires) and a telemetry endpoint advertising grounding conformance.", - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://searchco.com/agents/web-search/.well-known/content-telemetry.json", "roles": ["agent"], "operator": { "name": "SearchCo" }, diff --git a/tests/valid/manifest-content-owner-full.json b/tests/valid/manifest-content-owner-full.json index b71b116..533fb38 100644 --- a/tests/valid/manifest-content-owner-full.json +++ b/tests/valid/manifest-content-owner-full.json @@ -1,6 +1,6 @@ { "_test_description": "Content-owner manifest with operator.domain, a telemetry endpoint (no conformance_level), and a domains array with literal and wildcard subdomains.", - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://example.com/.well-known/content-telemetry.json", "roles": ["content_owner"], "operator": { "name": "Example Media", "domain": "example.com" }, diff --git a/tests/valid/manifest-content-owner-minimal.json b/tests/valid/manifest-content-owner-minimal.json index bc9a7f1..3fcdae6 100644 --- a/tests/valid/manifest-content-owner-minimal.json +++ b/tests/valid/manifest-content-owner-minimal.json @@ -1,6 +1,6 @@ { "_test_description": "Minimal valid content-owner manifest: schema_version, id, roles, operator only. No keys, telemetry, or domains.", - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://example.com/.well-known/content-telemetry.json", "roles": ["content_owner"], "operator": { "name": "Example Media" } diff --git a/tests/valid/manifest-coverage-declaration.json b/tests/valid/manifest-coverage-declaration.json new file mode 100644 index 0000000..e6332c1 --- /dev/null +++ b/tests/valid/manifest-coverage-declaration.json @@ -0,0 +1,18 @@ +{ + "_test_description": "Agent manifest declaring per-event-type coverage in telemetry.coverage (8.5, 5.7.6): content_grounded reported complete, content_retrieved sampled under a rule stated in the referenced terms.", + "schema_version": "1.0", + "id": "https://assistant.example.com/.well-known/content-telemetry.json", + "roles": ["agent"], + "operator": { "name": "Assistant Example" }, + "telemetry": { + "endpoint": "https://telemetry.assistant.example.com/v1/events", + "conformance_level": "grounding", + "coverage": { + "content_grounded": { "mode": "complete" }, + "content_retrieved": { + "mode": "sampled", + "terms_ref": "https://terms.example.com/reporting/v1" + } + } + } +} diff --git a/tests/valid/manifest-identifier-schemes.json b/tests/valid/manifest-identifier-schemes.json new file mode 100644 index 0000000..b919b3d --- /dev/null +++ b/tests/valid/manifest-identifier-schemes.json @@ -0,0 +1,26 @@ +{ + "_test_description": "Content owner root manifest declaring two identifier-scheme prefixes, one with a resolution endpoint and one without. Supports content_id prefix routing as a co-primary resolution path (sections 7.3, 8.6).", + "schema_version": "1.0", + "id": "https://example.com/.well-known/content-telemetry.json", + "roles": [ + "content_owner" + ], + "operator": { + "name": "Example Media" + }, + "telemetry": { + "endpoint": "https://telemetry.example.com/v1/events" + }, + "domains": [ + "example.com" + ], + "identifier_schemes": [ + { + "prefix": "exm:", + "resolution": "https://id.example.com/resolve" + }, + { + "prefix": "mkt:example:" + } + ] +} diff --git a/tests/valid/manifest-multi-role.json b/tests/valid/manifest-multi-role.json index 034e65f..8d185c8 100644 --- a/tests/valid/manifest-multi-role.json +++ b/tests/valid/manifest-multi-role.json @@ -1,6 +1,6 @@ { "_test_description": "Single manifest declaring multiple roles (content_owner and agent) with a key, telemetry endpoint, and domains.", - "schema_version": "0.1", + "schema_version": "1.0", "id": "https://publisher.com/.well-known/content-telemetry.json", "roles": ["content_owner", "agent"], "operator": { "name": "Publisher Co" }, diff --git a/tests/valid/manifest-platform.json b/tests/valid/manifest-platform.json new file mode 100644 index 0000000..e658529 --- /dev/null +++ b/tests/valid/manifest-platform.json @@ -0,0 +1,15 @@ +{ + "_test_description": "Platform manifest: an intermediary operating a telemetry consumer declares its inbound endpoint and a click-token resolution endpoint (sections 7.4.3, 8.5).", + "schema_version": "1.0", + "id": "https://telemetry.example.com/.well-known/content-telemetry.json", + "roles": [ + "platform" + ], + "operator": { + "name": "Telemetry Example" + }, + "telemetry": { + "endpoint": "https://telemetry.example.com/v1/events", + "ctx_resolution": "https://telemetry.example.com/v1/ctx/resolve" + } +} diff --git a/tests/valid/session-access-context.json b/tests/valid/session-access-context.json new file mode 100644 index 0000000..b2c1422 --- /dev/null +++ b/tests/valid/session-access-context.json @@ -0,0 +1,64 @@ +{ + "_test_description": "Session-level data container (5.1.3): access_context identifies the institution whose entitlement the agent used, required by the governing terms, with turn data at intent level per 5.5. The retrieval event carries content_depth: full (6.1).", + "document_type": "session", + "schema_version": "1.0", + "session_id": "880e8400-e29b-41d4-a716-446655440080", + "agent_id": "scholar-assistant.example.com", + "content_scope": "consortium-agreement-4471", + "started_at": "2026-08-13T14:02:10Z", + "ended_at": "2026-08-13T14:03:44Z", + "data": { + "access_context": { + "identifiers": [ + { "scheme": "ror", "value": "https://ror.org/013meh722" }, + { "scheme": "saml_entity_id", "value": "https://idp.example.ac.uk/shibboleth" } + ] + } + }, + "events": [ + { + "type": "turn_started", + "timestamp": "2026-08-13T14:02:11Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "query_intent": "question", + "topics": ["materials science"] + } + }, + { + "type": "content_retrieved", + "timestamp": "2026-08-13T14:02:14Z", + "source_role": "agent", + "content_telemetry_id": "990e8400-e29b-41d4-a716-446655440081", + "content_url": "https://journals.example.com/article/10.1000/xyz123", + "content_id": "doi:10.1000/xyz123", + "license_ref": "grant-4471-2026", + "data": { + "media_type": "text", + "content_depth": "full" + } + }, + { + "type": "content_grounded", + "timestamp": "2026-08-13T14:02:16Z", + "source_role": "agent", + "turn_id": "1", + "content_id": "doi:10.1000/xyz123", + "data": { + "scope": "turn", + "chars_ingested": 18400 + } + }, + { + "type": "turn_completed", + "timestamp": "2026-08-13T14:03:40Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "query_intent": "question", + "topics": ["materials science"] + } + } + ] +} diff --git a/tests/valid/session-cached-grounding.json b/tests/valid/session-cached-grounding.json index 59e8b09..11423fe 100644 --- a/tests/valid/session-cached-grounding.json +++ b/tests/valid/session-cached-grounding.json @@ -1,6 +1,6 @@ { - "_test_description": "Grounding from agent-side cache with no retrieval event. license_ref preserved from the original retrieval, as recommended by section 6.6.", - "schema_version": "0.1", + "_test_description": "Session document with agent-cached grounding and emitter-reported generic fingerprint detection. license_ref is preserved from the original retrieval, as recommended by section 6.4.", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440004", "agent_id": "copilot-v3", "started_at": "2026-03-28T09:00:00Z", @@ -15,9 +15,16 @@ "data": { "scope": "session", "cached": true, + "provenance": "agent_cached", + "chars_ingested": 11200, "tokens_ingested": 2800, "content_last_modified": "2026-03-27T16:00:00Z", "content_hash": "sha256:c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4", + "content_fingerprint": { + "scheme": "org.example.fingerprint.v1", + "detected": true, + "value": "fp_cached_42" + }, "media_type": "text" } }, diff --git a/tests/valid/session-citation-contradiction.json b/tests/valid/session-citation-contradiction.json new file mode 100644 index 0000000..ede6f51 --- /dev/null +++ b/tests/valid/session-citation-contradiction.json @@ -0,0 +1,79 @@ +{ + "_test_description": "Negative attribution and unclassified values: the response explicitly disagrees with a retrieved source (citation_type contradiction, section 6.5), a second association the agent could not classify carries citation_type and position unclassified, the grounding carries a content_version (ETag), and the turn ran in deep_research mode.", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440651", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "turn_started", + "timestamp": "2026-08-20T10:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "query_intent": "fact_check", + "topics": [ + "battery recycling" + ] + } + }, + { + "type": "content_retrieved", + "timestamp": "2026-08-20T10:00:01Z", + "source_role": "agent", + "content_url": "https://www.example-news.com/energy/battery-recycling-claims", + "data": { + "media_type": "text" + } + }, + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.example-news.com/energy/battery-recycling-claims", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 5100, + "content_version": "W/\"a1b2c3\"", + "content_last_modified": "2026-08-19T07:30:00Z" + } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/energy/battery-recycling-claims", + "data": { + "citation_type": "contradiction", + "position": "primary" + }, + "output_element_id": "answer:claim:1" + }, + { + "id": "770e8400-e29b-41d4-a716-446655440604", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.example-news.com/energy/battery-recycling-claims", + "data": { + "citation_type": "unclassified", + "position": "unclassified" + }, + "output_element_id": "answer:aside:1" + }, + { + "type": "turn_completed", + "timestamp": "2026-08-20T10:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "response_type": "fact_check", + "response_mode": "deep_research", + "response_tokens": 900 + } + } + ] +} diff --git a/tests/valid/session-citation-tier.json b/tests/valid/session-citation-tier.json index 24ff208..c2604d4 100644 --- a/tests/valid/session-citation-tier.json +++ b/tests/valid/session-citation-tier.json @@ -1,6 +1,6 @@ { - "_test_description": "Citation conformance level with optional display and engagement lifecycle signals.", - "schema_version": "0.1", + "_test_description": "Citation conformance level with optional presentation and engagement lifecycle signals.", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440003", "agent_id": "shopping-assistant-v2", "content_scope": "electronics-reviews", @@ -34,6 +34,7 @@ "data": { "scope": "session", "cached": false, + "chars_ingested": 24800, "tokens_ingested": 6200, "content_last_modified": "2026-03-20T10:00:00Z", "content_hash": "sha256:b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3", @@ -41,9 +42,12 @@ } }, { + "id": "880e8400-e29b-41d4-a716-446655440011", "type": "content_cited", "timestamp": "2026-03-28T14:00:06Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:recommendation:1", "content_url": "https://www.runnersworld.com/gear/best-trail-running-shoes", "content_id": "rw:best-trail-shoes-2026", "data": { @@ -56,13 +60,18 @@ } }, { - "type": "content_displayed", + "id": "880e8400-e29b-41d4-a716-446655440012", + "type": "content_presented", "timestamp": "2026-03-28T14:00:06Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:recommendation:1", + "citation_id": "880e8400-e29b-41d4-a716-446655440011", "content_url": "https://www.runnersworld.com/gear/best-trail-running-shoes", "content_id": "rw:best-trail-shoes-2026", "data": { - "display_type": "card" + "presentation_kind": "source_reference", + "presentation_type": "card" } }, { @@ -88,6 +97,7 @@ "type": "content_engaged", "timestamp": "2026-03-28T14:02:30Z", "turn_id": "1", + "presentation_id": "880e8400-e29b-41d4-a716-446655440012", "content_url": "https://www.runnersworld.com/gear/best-trail-running-shoes", "content_id": "rw:best-trail-shoes-2026", "data": { diff --git a/tests/valid/session-custom-media-type.json b/tests/valid/session-custom-media-type.json index 1d802f9..97e965c 100644 --- a/tests/valid/session-custom-media-type.json +++ b/tests/valid/session-custom-media-type.json @@ -1,6 +1,6 @@ { "_test_description": "Custom media_type value not in the core set (text, image, video, audio). The spec says emitters MAY use custom string values for media outside the core set and telemetry consumers MUST tolerate unknown media_type values. Exercises the custom value on both content_retrieved and content_grounded data.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440010", "agent_id": "cad-assistant-v2", "started_at": "2026-03-28T19:00:00Z", diff --git a/tests/valid/session-custom-response-mode.json b/tests/valid/session-custom-response-mode.json index 9d7b3a6..0d5b72e 100644 --- a/tests/valid/session-custom-response-mode.json +++ b/tests/valid/session-custom-response-mode.json @@ -1,6 +1,6 @@ { "_test_description": "Custom response_mode value not in the recommended set. The spec says platforms MAY use custom string values and telemetry consumers MUST tolerate unknown values.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440009", "agent_id": "podcast-gen-v1", "started_at": "2026-03-28T18:00:00Z", @@ -28,6 +28,7 @@ "data": { "scope": "turn", "cached": false, + "chars_ingested": 18000, "tokens_ingested": 4500 } }, diff --git a/tests/valid/session-earlier-turn-click.json b/tests/valid/session-earlier-turn-click.json new file mode 100644 index 0000000..2227dca --- /dev/null +++ b/tests/valid/session-earlier-turn-click.json @@ -0,0 +1,53 @@ +{ + "_test_description": "A click in a later turn on a presentation from an earlier turn: the content_engaged event carries the turn of the click but binds by presentation_id to the turn-1 presentation. Resolution returns the clicked content's lineage across turns by content identity, not a turn or timestamp cut (section 7.4).", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440220", + "agent_id": "assistant-v2", + "started_at": "2026-08-10T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-10T10:00:01Z", + "turn_id": "1", + "content_id": "publisher:feature:512", + "data": { "scope": "session", "chars_ingested": 15200 } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440221", + "type": "content_cited", + "timestamp": "2026-08-10T10:00:05Z", + "turn_id": "1", + "output_id": "response:1", + "content_id": "publisher:feature:512", + "data": { "citation_type": "reference", "position": "primary" } + }, + { + "id": "880e8400-e29b-41d4-a716-446655440222", + "type": "content_presented", + "timestamp": "2026-08-10T10:00:05Z", + "turn_id": "1", + "output_id": "response:1", + "citation_id": "770e8400-e29b-41d4-a716-446655440221", + "content_id": "publisher:feature:512", + "content_url": "https://publisher.example.com/features/512", + "data": { "presentation_kind": "source_reference", "presentation_type": "link" } + }, + { + "type": "content_grounded", + "timestamp": "2026-08-10T10:03:00Z", + "turn_id": "3", + "content_id": "otherpub:brief:77", + "data": { "scope": "turn", "chars_ingested": 2100 } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-10T10:04:12Z", + "turn_id": "3", + "presentation_id": "880e8400-e29b-41d4-a716-446655440222", + "ctx_token": "ct_e5a0b6d92c17f480", + "content_id": "publisher:feature:512", + "content_url": "https://publisher.example.com/features/512", + "data": { "engagement_type": "link_click" } + } + ] +} diff --git a/tests/valid/session-engagement-types.json b/tests/valid/session-engagement-types.json new file mode 100644 index 0000000..24dcc2e --- /dev/null +++ b/tests/valid/session-engagement-types.json @@ -0,0 +1,113 @@ +{ + "_test_description": "The non-click engagement types: a snippet presentation is expanded, and the detail view it opens is copied from and shared. Each action is one engagement occurrence on one presentation (sections 4.3, 6.7); three actions on two presentations, all in-product and agent-reported.", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440650", + "agent_id": "assistant-v4", + "started_at": "2026-08-20T10:00:00Z", + "events": [ + { + "type": "turn_started", + "timestamp": "2026-08-20T10:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "query_intent": "how_to", + "topics": [ + "sourdough" + ] + } + }, + { + "type": "content_grounded", + "timestamp": "2026-08-20T10:00:02Z", + "content_url": "https://www.kingarthurbaking.com/recipes/sourdough-starter-maintenance", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 7200 + } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440601", + "type": "content_cited", + "timestamp": "2026-08-20T10:00:03Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.kingarthurbaking.com/recipes/sourdough-starter-maintenance", + "data": { + "citation_type": "paraphrase", + "position": "primary" + }, + "output_element_id": "answer:steps:1" + }, + { + "id": "770e8400-e29b-41d4-a716-446655440602", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "citation_id": "770e8400-e29b-41d4-a716-446655440601", + "content_url": "https://www.kingarthurbaking.com/recipes/sourdough-starter-maintenance", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "snippet" + }, + "output_element_id": "answer:steps:1" + }, + { + "id": "770e8400-e29b-41d4-a716-446655440603", + "type": "content_presented", + "timestamp": "2026-08-20T10:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://www.kingarthurbaking.com/recipes/sourdough-starter-maintenance", + "data": { + "presentation_kind": "content", + "presentation_type": "detail_view", + "media_type": "text" + }, + "output_element_id": "panel:source:1" + }, + { + "type": "turn_completed", + "timestamp": "2026-08-20T10:00:05Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "response_type": "how_to", + "response_mode": "standard", + "response_tokens": 260 + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:06Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440602", + "content_url": "https://www.kingarthurbaking.com/recipes/sourdough-starter-maintenance", + "data": { + "engagement_type": "expand" + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:07Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440603", + "content_url": "https://www.kingarthurbaking.com/recipes/sourdough-starter-maintenance", + "data": { + "engagement_type": "copy" + } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-20T10:00:08Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440603", + "content_url": "https://www.kingarthurbaking.com/recipes/sourdough-starter-maintenance", + "data": { + "engagement_type": "share" + } + } + ] +} diff --git a/tests/valid/session-evidence-reference.json b/tests/valid/session-evidence-reference.json new file mode 100644 index 0000000..ae086db --- /dev/null +++ b/tests/valid/session-evidence-reference.json @@ -0,0 +1,30 @@ +{ + "_test_description": "Grounding event carrying a data.evidence array (section 6.8): one detached reference (ref plus digest) and one entry in an unknown scheme, which consumers tolerate. The slot attaches profile-defined evidence to an event's claim without raising its status.", + "schema_version": "1.0", + "session_id": "aa0e8400-e29b-41d4-a716-446655440600", + "agent_id": "assistant-v1", + "started_at": "2026-08-30T10:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-30T10:00:01Z", + "content_url": "https://example.com/articles/energy-prices", + "data": { + "scope": "session", + "cached": false, + "chars_ingested": 5400, + "evidence": [ + { + "scheme": "org.example.retrieval-receipt", + "ref": "https://evidence.example.com/receipts/8f3a", + "digest": "sha256:c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4" + }, + { + "scheme": "com.vendor.unknown-scheme", + "inline_note": "consumers tolerate unknown schemes and fields" + } + ] + } + } + ] +} diff --git a/tests/valid/session-funnel-exception-cited-no-grounded.json b/tests/valid/session-funnel-exception-cited-no-grounded.json index fad0e48..e84a4f1 100644 --- a/tests/valid/session-funnel-exception-cited-no-grounded.json +++ b/tests/valid/session-funnel-exception-cited-no-grounded.json @@ -1,6 +1,6 @@ { "_test_description": "Funnel exception: content_cited with no preceding content_grounded event. This is a hallucinated citation - the agent references content it never retrieved or loaded into context. Valid per section 4.3. Telemetry consumers SHOULD treat this as a lower-confidence signal.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440011", "agent_id": "assistant-v4", "started_at": "2026-03-28T20:00:00Z", @@ -16,9 +16,12 @@ } }, { + "id": "880e8400-e29b-41d4-a716-446655440021", "type": "content_cited", "timestamp": "2026-03-28T20:00:05Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:claim:1", "content_url": "https://www.ons.gov.uk/peoplepopulationandcommunity/populationandmigration", "data": { "citation_type": "reference", @@ -27,12 +30,17 @@ } }, { - "type": "content_displayed", + "id": "880e8400-e29b-41d4-a716-446655440022", + "type": "content_presented", "timestamp": "2026-03-28T20:00:05Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:claim:1", + "citation_id": "880e8400-e29b-41d4-a716-446655440021", "content_url": "https://www.ons.gov.uk/peoplepopulationandcommunity/populationandmigration", "data": { - "display_type": "link" + "presentation_kind": "source_reference", + "presentation_type": "link" } }, { diff --git a/tests/valid/session-funnel-exception-displayed-no-cited.json b/tests/valid/session-funnel-exception-presented-no-cited.json similarity index 74% rename from tests/valid/session-funnel-exception-displayed-no-cited.json rename to tests/valid/session-funnel-exception-presented-no-cited.json index d5ed91e..d4b3f4a 100644 --- a/tests/valid/session-funnel-exception-displayed-no-cited.json +++ b/tests/valid/session-funnel-exception-presented-no-cited.json @@ -1,6 +1,6 @@ { - "_test_description": "Funnel exception: content_displayed without content_cited. This is the Sources sidebar pattern - the agent shows source links without explicitly citing the content in the response text. Valid per section 4.3.", - "schema_version": "0.1", + "_test_description": "Funnel exception: content_presented without content_cited. This is the Sources sidebar pattern - the agent makes source links perceivable without semantically associating them with an output element. Valid per section 4.3.", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440010", "agent_id": "search-assistant-v1", "started_at": "2026-03-28T19:00:00Z", @@ -28,16 +28,21 @@ "data": { "scope": "turn", "cached": false, + "chars_ingested": 7200, "tokens_ingested": 1800 } }, { - "type": "content_displayed", + "id": "880e8400-e29b-41d4-a716-446655440031", + "type": "content_presented", "timestamp": "2026-03-28T19:00:06Z", "turn_id": "1", + "output_id": "response:1", + "output_element_id": "sources:1", "content_url": "https://www.kingarthurbaking.com/recipes/sourdough-starter-maintenance", "data": { - "display_type": "link" + "presentation_kind": "source_reference", + "presentation_type": "link" } }, { diff --git a/tests/valid/session-funnel-exception-displayed-no-grounded.json b/tests/valid/session-funnel-exception-presented-no-grounded.json similarity index 64% rename from tests/valid/session-funnel-exception-displayed-no-grounded.json rename to tests/valid/session-funnel-exception-presented-no-grounded.json index 308350f..9a6de97 100644 --- a/tests/valid/session-funnel-exception-displayed-no-grounded.json +++ b/tests/valid/session-funnel-exception-presented-no-grounded.json @@ -1,6 +1,6 @@ { - "_test_description": "Funnel exception: content_displayed without content_grounded. An agentic browser renders a publisher's page to the user (display_type: embed) without the content entering a generation context, and the user directs the agent to open a second page (engagement_type: agent_navigate). Valid per section 4.3. Also exercises media_type on display data and the open display_type/engagement_type vocabularies.", - "schema_version": "0.1", + "_test_description": "Funnel exception: content_presented without content_grounded. An agentic browser renders a publisher's page on a recipient-facing surface (presentation_type: embed) without the content entering a generation context, and the recipient directs the agent to open a second page (engagement_type: agent_navigate). Valid per section 4.3. Also exercises media_type on presentation data and the open presentation_type/engagement_type vocabularies.", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440020", "agent_id": "browser-agent-v1", "started_at": "2026-03-29T11:00:00Z", @@ -22,22 +22,28 @@ "content_url": "https://www.fitnessmedia.example/guides/home-workouts" }, { - "type": "content_displayed", + "id": "880e8400-e29b-41d4-a716-446655440041", + "type": "content_presented", "timestamp": "2026-03-29T11:00:02Z", "turn_id": "1", + "output_id": "browser-view:1", "content_url": "https://www.fitnessmedia.example/guides/home-workouts", "data": { - "display_type": "embed", + "presentation_kind": "content", + "presentation_type": "embed", "media_type": "text" } }, { - "type": "content_displayed", + "id": "880e8400-e29b-41d4-a716-446655440042", + "type": "content_presented", "timestamp": "2026-03-29T11:00:04Z", "turn_id": "1", + "output_id": "browser-view:1", "content_url": "https://www.fitnessmedia.example/videos/beginner-routine", "data": { - "display_type": "embed", + "presentation_kind": "content", + "presentation_type": "embed", "media_type": "video" } }, @@ -56,6 +62,7 @@ "type": "content_engaged", "timestamp": "2026-03-29T11:01:00Z", "turn_id": "1", + "presentation_id": "880e8400-e29b-41d4-a716-446655440041", "content_url": "https://www.fitnessmedia.example/guides/home-workouts", "data": { "engagement_type": "agent_navigate" diff --git a/tests/valid/session-grounding-tier.json b/tests/valid/session-grounding-tier.json index d3002ca..e10f147 100644 --- a/tests/valid/session-grounding-tier.json +++ b/tests/valid/session-grounding-tier.json @@ -1,6 +1,6 @@ { "_test_description": "Grounding conformance level: agent emitting content_retrieved, content_grounded, and turn events with privacy_level. Includes agent_id and data.scope as required by the Grounding tier.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440002", "agent_id": "research-assistant-v1", "started_at": "2026-03-28T12:00:00Z", @@ -30,6 +30,7 @@ "data": { "scope": "session", "cached": false, + "chars_ingested": 20400, "tokens_ingested": 5100, "content_last_modified": "2026-03-15T08:00:00Z", "media_type": "text" diff --git a/tests/valid/session-minimal.json b/tests/valid/session-minimal.json index b9c1ee0..a54498b 100644 --- a/tests/valid/session-minimal.json +++ b/tests/valid/session-minimal.json @@ -1,12 +1,13 @@ { - "_test_description": "Bare minimum conforming session: schema_version, session_id, started_at, and one event with type and timestamp.", - "schema_version": "0.1", + "_test_description": "Bare minimum conforming session: schema_version, session_id, started_at, and one content_retrieved event with type, timestamp, source_role and a content_url.", + "schema_version": "1.0", "session_id": "550e8400-e29b-41d4-a716-446655440000", "started_at": "2026-03-28T10:00:00Z", "events": [ { "type": "content_retrieved", "timestamp": "2026-03-28T10:00:01Z", + "source_role": "agent", "content_url": "https://example.com/article/getting-started" } ] diff --git a/tests/valid/session-multi-agent-child.json b/tests/valid/session-multi-agent-child.json new file mode 100644 index 0000000..400b4fa --- /dev/null +++ b/tests/valid/session-multi-agent-child.json @@ -0,0 +1,32 @@ +{ + "_test_description": "A delegated child session links to its immediate parent. The source enters the child agent's generation context and is grounded there; no citation or presentation is emitted merely because the child returns an internal response to its orchestrator.", + "document_type": "session", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440021", + "parent_session_id": "660e8400-e29b-41d4-a716-446655440020", + "agent_id": "research-subagent-v1", + "started_at": "2026-07-29T09:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-07-29T09:00:02Z", + "turn_id": "child-turn-1", + "content_url": "https://example.org/research/source", + "data": { + "scope": "turn", + "cached": false, + "chars_ingested": 3600, + "tokens_ingested": 900 + } + }, + { + "type": "turn_completed", + "timestamp": "2026-07-29T09:00:05Z", + "turn_id": "child-turn-1", + "turn": { + "privacy_level": "minimal", + "response_tokens": 180 + } + } + ] +} diff --git a/tests/valid/session-multi-owner-catalogue.json b/tests/valid/session-multi-owner-catalogue.json new file mode 100644 index 0000000..bd15288 --- /dev/null +++ b/tests/valid/session-multi-owner-catalogue.json @@ -0,0 +1,142 @@ +{ + "_test_description": "Multi-owner catalogue under one agreement (Annex B.5): content_scope identifies the marketplace agreement, owner resolution is per event - one owner by content_url domain, the other by a registered content_id prefix with no canonical URL at all. Two owners, each event resolving to its own.", + "schema_version": "1.0", + "session_id": "990e8400-e29b-41d4-a716-446655440500", + "agent_id": "research-assistant-v5", + "content_scope": "marketplace-agreement-2026-017", + "manifest_ref": "https://assistant.example.com/.well-known/content-telemetry.json", + "started_at": "2026-08-20T09:00:00Z", + "ended_at": "2026-08-20T09:00:09Z", + "events": [ + { + "type": "turn_started", + "timestamp": "2026-08-20T09:00:00Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "query_intent": "comparison", + "topics": [ + "electric vehicles", + "charging" + ] + } + }, + { + "type": "content_retrieved", + "timestamp": "2026-08-20T09:00:01Z", + "source_role": "index", + "content_telemetry_id": "990e8400-e29b-41d4-a716-446655440501", + "content_url": "https://www.autoreview.example/ev/charging-networks-2026", + "content_id": "mkt:autoreview:88213", + "license_ref": "agreement-2026-017:autoreview", + "data": { + "media_type": "text", + "content_depth": "full" + } + }, + { + "type": "content_retrieved", + "timestamp": "2026-08-20T09:00:01Z", + "source_role": "index", + "content_telemetry_id": "990e8400-e29b-41d4-a716-446655440502", + "content_id": "mkt:gridnews:5520", + "license_ref": "agreement-2026-017:gridnews", + "data": { + "media_type": "text", + "content_depth": "full" + } + }, + { + "type": "content_grounded", + "timestamp": "2026-08-20T09:00:02Z", + "turn_id": "1", + "content_url": "https://www.autoreview.example/ev/charging-networks-2026", + "content_id": "mkt:autoreview:88213", + "data": { + "scope": "turn", + "cached": false, + "provenance": "third_party_sourced", + "chars_ingested": 11200 + } + }, + { + "type": "content_grounded", + "timestamp": "2026-08-20T09:00:02Z", + "turn_id": "1", + "content_id": "mkt:gridnews:5520", + "data": { + "scope": "turn", + "cached": false, + "provenance": "third_party_sourced", + "chars_ingested": 6400 + } + }, + { + "id": "990e8400-e29b-41d4-a716-446655440503", + "type": "content_cited", + "timestamp": "2026-08-20T09:00:06Z", + "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:networks:1", + "content_url": "https://www.autoreview.example/ev/charging-networks-2026", + "content_id": "mkt:autoreview:88213", + "data": { + "citation_type": "paraphrase", + "position": "primary" + } + }, + { + "id": "990e8400-e29b-41d4-a716-446655440504", + "type": "content_cited", + "timestamp": "2026-08-20T09:00:06Z", + "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:tariffs:1", + "content_id": "mkt:gridnews:5520", + "data": { + "citation_type": "reference", + "position": "supporting" + } + }, + { + "id": "990e8400-e29b-41d4-a716-446655440505", + "type": "content_presented", + "timestamp": "2026-08-20T09:00:07Z", + "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:networks:1", + "citation_id": "990e8400-e29b-41d4-a716-446655440503", + "content_url": "https://www.autoreview.example/ev/charging-networks-2026", + "content_id": "mkt:autoreview:88213", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "link" + } + }, + { + "id": "990e8400-e29b-41d4-a716-446655440506", + "type": "content_presented", + "timestamp": "2026-08-20T09:00:07Z", + "turn_id": "1", + "output_id": "response:1", + "output_element_id": "answer:tariffs:1", + "citation_id": "990e8400-e29b-41d4-a716-446655440504", + "content_id": "mkt:gridnews:5520", + "data": { + "presentation_kind": "source_reference", + "presentation_type": "card" + } + }, + { + "type": "turn_completed", + "timestamp": "2026-08-20T09:00:09Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "response_type": "comparison", + "response_mode": "standard", + "response_tokens": 410 + } + } + ] +} diff --git a/tests/valid/session-multi-turn.json b/tests/valid/session-multi-turn.json index 27439ee..f3da34b 100644 --- a/tests/valid/session-multi-turn.json +++ b/tests/valid/session-multi-turn.json @@ -1,11 +1,21 @@ { - "_test_description": "Session-scoped grounding, 3 turns, citations in turns 1 and 3, zero-click outcome (browse with new_query exit). Turn 2 has no citation - the grounded content was in context but not explicitly referenced.", - "schema_version": "0.1", + "_test_description": "Session-scoped grounding, 3 turns, citations in turns 1 and 3, zero-click outcome. Turn 2 has no citation - the grounded content was in context but not explicitly referenced.", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440006", "agent_id": "copilot-v3", "started_at": "2026-03-28T16:00:00Z", "ended_at": "2026-03-28T16:10:00Z", "events": [ + { + "type": "turn_started", + "timestamp": "2026-03-28T16:00:00Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "query_intent": "question", + "topics": ["fusion energy", "ITER"] + } + }, { "type": "content_retrieved", "timestamp": "2026-03-28T16:00:01Z", @@ -20,24 +30,17 @@ "data": { "scope": "session", "cached": false, + "chars_ingested": 9600, "tokens_ingested": 2400, "media_type": "text" } }, { - "type": "turn_started", - "timestamp": "2026-03-28T16:00:00Z", - "turn_id": "1", - "turn": { - "privacy_level": "intent", - "query_intent": "question", - "topics": ["fusion energy", "ITER"] - } - }, - { + "id": "880e8400-e29b-41d4-a716-446655440051", "type": "content_cited", "timestamp": "2026-03-28T16:00:08Z", "turn_id": "1", + "output_id": "response:1", "content_url": "https://www.bbc.co.uk/news/science-environment-68234567", "content_id": "bbc:68234567", "data": { @@ -47,12 +50,16 @@ } }, { - "type": "content_displayed", + "id": "880e8400-e29b-41d4-a716-446655440052", + "type": "content_presented", "timestamp": "2026-03-28T16:00:08Z", "turn_id": "1", + "output_id": "response:1", + "citation_id": "880e8400-e29b-41d4-a716-446655440051", "content_url": "https://www.bbc.co.uk/news/science-environment-68234567", "data": { - "display_type": "link" + "presentation_kind": "source_reference", + "presentation_type": "link" } }, { @@ -100,9 +107,11 @@ } }, { + "id": "880e8400-e29b-41d4-a716-446655440053", "type": "content_cited", "timestamp": "2026-03-28T16:05:12Z", "turn_id": "3", + "output_id": "response:3", "content_url": "https://www.bbc.co.uk/news/science-environment-68234567", "content_id": "bbc:68234567", "data": { diff --git a/tests/valid/session-presentation-multimodal.json b/tests/valid/session-presentation-multimodal.json new file mode 100644 index 0000000..61da85e --- /dev/null +++ b/tests/valid/session-presentation-multimodal.json @@ -0,0 +1,96 @@ +{ + "_test_description": "Modality-neutral citation and presentation: an image excerpt is presented as content, an audio output speaks a source credit, a video is embedded without citation, a constructed citation is suppressed, and a repeated link presentation receives the engagement.", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440070", + "agent_id": "multimodal-assistant-v1", + "started_at": "2026-07-18T12:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-07-18T12:00:01Z", + "content_id": "museum:image:42", + "data": { "scope": "turn", "cached": true, "media_type": "image" } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440071", + "type": "content_cited", + "timestamp": "2026-07-18T12:00:04Z", + "turn_id": "1", + "output_id": "output:mixed:1", + "output_element_id": "image:region:credit", + "content_id": "museum:image:42", + "data": { "citation_type": "reference", "media_type": "image" } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440072", + "type": "content_presented", + "timestamp": "2026-07-18T12:00:05Z", + "turn_id": "1", + "output_id": "output:mixed:1", + "output_element_id": "image:region:1", + "content_id": "museum:image:42", + "data": { "presentation_kind": "content", "presentation_type": "embed", "media_type": "image" } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440073", + "type": "content_presented", + "timestamp": "2026-07-18T12:00:06Z", + "turn_id": "1", + "output_id": "output:mixed:1", + "output_element_id": "audio:credit:1", + "citation_id": "770e8400-e29b-41d4-a716-446655440071", + "content_id": "museum:image:42", + "data": { "presentation_kind": "source_reference", "presentation_type": "spoken_credit", "media_type": "audio" } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440074", + "type": "content_presented", + "timestamp": "2026-07-18T12:00:07Z", + "turn_id": "1", + "output_id": "output:mixed:1", + "output_element_id": "video:1", + "content_id": "archive:video:7", + "data": { "presentation_kind": "content", "presentation_type": "embed", "media_type": "video" } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440075", + "type": "content_cited", + "timestamp": "2026-07-18T12:00:08Z", + "turn_id": "1", + "output_id": "output:suppressed:1", + "output_element_id": "caption:1", + "content_id": "audio:interview:9", + "data": { "citation_type": "reference", "media_type": "audio" } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440076", + "type": "content_presented", + "timestamp": "2026-07-18T12:00:09Z", + "turn_id": "1", + "output_id": "output:mixed:1", + "output_element_id": "sources:link:1", + "citation_id": "770e8400-e29b-41d4-a716-446655440071", + "content_id": "museum:image:42", + "data": { "presentation_kind": "source_reference", "presentation_type": "link", "media_type": "text" } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440077", + "type": "content_presented", + "timestamp": "2026-07-18T12:00:10Z", + "turn_id": "1", + "output_id": "output:mixed:1", + "output_element_id": "sidebar:link:1", + "citation_id": "770e8400-e29b-41d4-a716-446655440071", + "content_id": "museum:image:42", + "data": { "presentation_kind": "source_reference", "presentation_type": "link", "media_type": "text" } + }, + { + "type": "content_engaged", + "timestamp": "2026-07-18T12:00:12Z", + "turn_id": "1", + "presentation_id": "770e8400-e29b-41d4-a716-446655440077", + "content_id": "museum:image:42", + "data": { "engagement_type": "link_click" } + } + ] +} diff --git a/tests/valid/session-repeated-link-engagement.json b/tests/valid/session-repeated-link-engagement.json new file mode 100644 index 0000000..ef7fea8 --- /dev/null +++ b/tests/valid/session-repeated-link-engagement.json @@ -0,0 +1,54 @@ +{ + "_test_description": "The same URL presented twice in one session: two content_presented events with distinct ids and distinct minted click tokens, and a content_engaged bound by presentation_id to the second occurrence. Matching on URL alone could not tell the two presentations apart (sections 6.7, 7.4).", + "schema_version": "1.0", + "session_id": "660e8400-e29b-41d4-a716-446655440210", + "agent_id": "assistant-v2", + "started_at": "2026-08-10T09:00:00Z", + "events": [ + { + "type": "content_grounded", + "timestamp": "2026-08-10T09:00:01Z", + "turn_id": "1", + "content_url": "https://news.example.com/markets/rate-decision", + "content_id": "newsex:article:8841", + "data": { "scope": "session", "chars_ingested": 9400 } + }, + { + "id": "770e8400-e29b-41d4-a716-446655440211", + "type": "content_cited", + "timestamp": "2026-08-10T09:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "content_url": "https://news.example.com/markets/rate-decision", + "data": { "citation_type": "paraphrase", "position": "primary" } + }, + { + "id": "880e8400-e29b-41d4-a716-446655440212", + "type": "content_presented", + "timestamp": "2026-08-10T09:00:04Z", + "turn_id": "1", + "output_id": "response:1", + "citation_id": "770e8400-e29b-41d4-a716-446655440211", + "content_url": "https://news.example.com/markets/rate-decision", + "data": { "presentation_kind": "source_reference", "presentation_type": "link" } + }, + { + "id": "880e8400-e29b-41d4-a716-446655440213", + "type": "content_presented", + "timestamp": "2026-08-10T09:02:10Z", + "turn_id": "2", + "output_id": "response:2", + "content_url": "https://news.example.com/markets/rate-decision", + "data": { "presentation_kind": "source_reference", "presentation_type": "link" } + }, + { + "type": "content_engaged", + "timestamp": "2026-08-10T09:02:31Z", + "turn_id": "2", + "presentation_id": "880e8400-e29b-41d4-a716-446655440213", + "ctx_token": "ct_r2d1a9c44be07f31", + "content_url": "https://news.example.com/markets/rate-decision", + "data": { "engagement_type": "link_click" } + } + ] +} diff --git a/tests/valid/session-retrieval-tier.json b/tests/valid/session-retrieval-tier.json index 494f0bb..3047baf 100644 --- a/tests/valid/session-retrieval-tier.json +++ b/tests/valid/session-retrieval-tier.json @@ -1,6 +1,6 @@ { "_test_description": "Retrieval conformance level: content owner CDN emitting content_retrieved with source_role and content_url. No session context beyond the minimum.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440001", "started_at": "2026-03-28T11:00:00Z", "events": [ @@ -12,7 +12,7 @@ "content_url": "https://www.theguardian.com/technology/2026/mar/28/ai-agents-roundup", "data": { "user_agent": "ClaudeBot/1.0", - "bot_category": "inference", + "purpose": "inference", "verified": true, "cache_status": "miss", "response_status": 200, @@ -20,8 +20,7 @@ "ja4": "t13d1517h2_8daaf6152771_02e4c6ae3e16", "asn": 14618, "asn_org": "Anthropic", - "country": "US", - "ip_hash": "sha256:a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4e5f6a1b2" + "country": "US" } } ] diff --git a/tests/valid/turn-privacy-full.json b/tests/valid/turn-privacy-full.json index 0cef475..b68a70b 100644 --- a/tests/valid/turn-privacy-full.json +++ b/tests/valid/turn-privacy-full.json @@ -1,6 +1,6 @@ { "_test_description": "turn_completed at full privacy level: query_text, response_text, intent, topics, model_id, ad_rendered all present. Maximum data sharing.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440007", "agent_id": "assistant-v4", "started_at": "2026-03-28T17:00:00Z", diff --git a/tests/valid/turn-privacy-intent.json b/tests/valid/turn-privacy-intent.json index 14a7dac..406a6e1 100644 --- a/tests/valid/turn-privacy-intent.json +++ b/tests/valid/turn-privacy-intent.json @@ -1,6 +1,6 @@ { "_test_description": "turn_completed at intent privacy level: query_intent, topics, response_type, response_mode, model_id, ad_rendered present. No query_text or response_text.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440012", "agent_id": "assistant-v4", "started_at": "2026-03-28T18:30:00Z", diff --git a/tests/valid/turn-privacy-minimal.json b/tests/valid/turn-privacy-minimal.json index 00e6b11..bc5d5db 100644 --- a/tests/valid/turn-privacy-minimal.json +++ b/tests/valid/turn-privacy-minimal.json @@ -1,6 +1,6 @@ { "_test_description": "turn_completed at minimal privacy level: only response_tokens and content_urls present. No intent, topics, query_text, response_text, model_id, ad_rendered, or response_mode.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440008", "agent_id": "assistant-v4", "started_at": "2026-03-28T17:30:00Z", diff --git a/tests/valid/turn-privacy-summary.json b/tests/valid/turn-privacy-summary.json index 38a2e99..44ff95d 100644 --- a/tests/valid/turn-privacy-summary.json +++ b/tests/valid/turn-privacy-summary.json @@ -1,6 +1,6 @@ { "_test_description": "turn_completed at summary privacy level: query_text and response_text are summarised (not verbatim). Includes query_intent, topics, query_tokens, response_tokens, model_id, ad_rendered, response_mode, response_type.", - "schema_version": "0.1", + "schema_version": "1.0", "session_id": "660e8400-e29b-41d4-a716-446655440009", "agent_id": "assistant-v4", "started_at": "2026-03-28T18:00:00Z", diff --git a/tests/validate.py b/tests/validate.py index ad51e01..ae3448a 100644 --- a/tests/validate.py +++ b/tests/validate.py @@ -1,6 +1,6 @@ #!/usr/bin/env python3 """ -Conformance test runner for Content Telemetry Specification v0.1. +Conformance test runner for Content Telemetry Specification v1. Validates JSON test fixtures against telemetry-session.json, telemetry-event.json, telemetry-event-batch.json, manifest.json, and application-layer conformance rules @@ -11,7 +11,7 @@ document. Usage: - pip install jsonschema + pip install "jsonschema[format-nongpl]" python validate.py """ @@ -25,7 +25,22 @@ from jsonschema import Draft202012Validator, ValidationError from referencing import Registry, Resource except ImportError: - print("ERROR: jsonschema package required. Install with: pip install jsonschema") + print('ERROR: jsonschema package required. Install with: pip install "jsonschema[format-nongpl]"') + sys.exit(1) + +# Format assertions (uuid, date-time, uri) are annotation-only unless a +# FormatChecker is attached to the validator. Every validator constructed in +# this suite MUST pass format_checker=FORMAT_CHECKER; guard at startup so a +# missing optional dependency (rfc3339-validator) hard-errors instead of +# silently downgrading every format assertion to a no-op. +FORMAT_CHECKER = Draft202012Validator.FORMAT_CHECKER +_missing_formats = {"uuid", "date-time"} - set(FORMAT_CHECKER.checkers) +if _missing_formats: + print( + "ERROR: format checker cannot enforce " + + ", ".join(sorted(_missing_formats)) + + '. Install with: pip install "jsonschema[format-nongpl]"' + ) sys.exit(1) @@ -38,7 +53,7 @@ # 1. Privacy level field gating (section 5.5): # - At "minimal" level: query_text, response_text, query_intent, topics, # model_id, ad_rendered, response_mode, and response_type MUST NOT be -# present. Only response_tokens and content_urls are allowed. +# present. Only query_tokens, response_tokens and content_urls are allowed. # - At "intent" level: query_text and response_text MUST NOT be present. # # 2. content_url or content_id requirement (section 5.7.5): @@ -55,6 +70,35 @@ # manifest's own host or a subdomain of it (section 8.6). JSON Schema # cannot compare values across array items or against the manifest's id. # +# 5. Grounding provenance and fingerprint migration (sections 5.7.5, 6.4, 12.1): +# agent_fetched requires cached false; agent_cached requires cached true. +# content_fingerprint MUST NOT carry preserved_in_output: v1 defines no +# output-side reuse reporting. +# +# 6. Referential integrity within a session document (sections 6.6-6.7): +# content_engaged.presentation_id references the exact content_presented +# event id, and citation_id on content_presented +# references a content_cited event id. JSON Schema cannot compare values +# across events. Session documents only: standalone envelopes and batch +# members may reference events delivered elsewhere. +# +# 7. Field placement and source_role (sections 5.2, 5.2.2, 5.7.1, 5.7.5): +# source_role is required on content_retrieved; presentation_id and the +# event-level ctx_token appear only on content_engaged, citation_id only on +# content_presented, turn only on turn events. +# +# 8. Session-document integrity beyond rule 6 (sections 6.6, 6.7, 7.4.1): +# event ids are distinct; the engagement and the presentation it references, +# and the presentation and the citation it references, identify +# the same content; one event-level ctx_token binds to one presentation. +# +# 9. Envelope ctx_token (section 7.1): an envelope carrying ctx_token carries +# content_engaged events only. +# +# 10. Manifest placement rules (sections 8.5, 8.6): domains only on a manifest +# served at the domain root; ctx_resolution only on agent/platform manifests; +# identifier_schemes only on content_owner manifests. +# # Not checked here: agent_id at Grounding/Citation conformance (section # 5.7) depends on the emitter's declared conformance level, which fixtures do # not carry, so it is out of scope for the fixture suite. @@ -75,10 +119,56 @@ "Turn at intent privacy includes query_text. " "Violates section 5.5: query_text MUST NOT be present at intent level." ), + "privacy-violation-response-text-at-minimal.json": ( + "Turn at minimal privacy includes response_text. " + "Violates section 5.5: response text MUST NOT be present at minimal level." + ), + "privacy-violation-query-intent-at-minimal.json": ( + "Turn at minimal privacy includes query_intent. " + "Violates section 5.5: intent classification MUST NOT be present at minimal level." + ), + "privacy-violation-topics-at-minimal.json": ( + "Turn at minimal privacy includes topics. " + "Violates section 5.5: topics MUST NOT be present at minimal level." + ), + "privacy-violation-response-type-at-minimal.json": ( + "Turn at minimal privacy includes response_type. " + "Violates section 5.5: response classification MUST NOT be present at minimal level." + ), + "privacy-violation-response-mode-at-minimal.json": ( + "Turn at minimal privacy includes response_mode. " + "Violates section 5.5: platform metadata MUST NOT be present at minimal level." + ), + "privacy-violation-model-id-at-minimal.json": ( + "Turn at minimal privacy includes model_id. " + "Violates section 5.5: platform metadata MUST NOT be present at minimal level." + ), "content-event-missing-identifier.json": ( "content_grounded event has neither content_url nor content_id. " "Violates section 5.7.5: every content event MUST carry at least one." ), + "presented-missing-identifier.json": ( + "content_presented event has neither content_url nor content_id. " + "Violates section 5.7.5: every content event MUST carry at least one." + ), + "retrieved-missing-identifier.json": ( + "content_retrieved event has neither content_url nor content_id. " + "Violates section 5.7.5: every content event MUST carry at least one." + ), + "engaged-missing-identifier.json": ( + "content_engaged event has neither content_url nor content_id. " + "Violates section 5.7.5: every content event MUST carry at least one." + ), + "engaged-presentation-id-unmatched.json": ( + "content_engaged.presentation_id matches no content_presented event id " + "in the session. Violates section 6.7: every engagement references the " + "exact content_presented.id on which the action occurred." + ), + "presented-citation-id-unmatched.json": ( + "content_presented.citation_id matches no content_cited event id in " + "the session. Violates section 6.7: citation_id references the " + "presented content_cited event's id." + ), "standalone-missing-session-and-ctx-token.json": ( "Standalone event envelope has neither session_id nor ctx_token. " "Violates section 5.7.5: an event MUST carry one at Grounding+ (section 7.1)." @@ -96,14 +186,135 @@ "Violates section 8.6: every entry MUST be the manifest's own host or a " "subdomain of it. Consumers reject the manifest as malformed (section 8.7)." ), + "withdrawn-ip-hash.json": ( + "Retrieval event carries ip_hash in data. " + "Violates section 9.1: the field was withdrawn in v1 and emitters " + "MUST NOT populate it." + ), + "grounding-fingerprint-preserved-in-output.json": ( + "Grounding fingerprint carries preserved_in_output. " + "Violates section 6.4: v1 defines no output-side reuse reporting and the field is withdrawn (section 12.1)." + ), + "grounding-provenance-cached-conflict.json": ( + "Grounding event declares agent_fetched with cached true. " + "Violates section 6.4: agent_fetched requires cached false." + ), + "grounding-provenance-cached-conflict-agent-cached.json": ( + "Grounding event declares agent_cached with cached false. " + "Violates section 6.4: agent_cached requires cached true." + ), + "privacy-violation-response-text-at-intent.json": ( + "Turn at intent privacy includes response_text. " + "Violates section 5.5: response_text MUST NOT be present at intent level." + ), + "privacy-violation-query-at-minimal-standalone.json": ( + "Standalone envelope turn at minimal privacy includes query_text. " + "Violates section 5.5: the gate applies wherever turns are emitted." + ), + "privacy-violation-query-at-minimal-batch.json": ( + "Batch member turn at minimal privacy includes query_text. " + "Violates section 5.5: the gate applies wherever turns are emitted." + ), + "grounded-missing-identifier-standalone.json": ( + "Standalone content_grounded event has neither content_url nor content_id. " + "Violates section 5.7.5 in the standalone envelope shape." + ), + "grounded-missing-identifier-batch.json": ( + "Batch member content_grounded event has neither content_url nor content_id. " + "Violates section 5.7.5 in the event batch shape." + ), + "withdrawn-ip-hash-session.json": ( + "Session document retrieval event carries ip_hash in data. " + "Violates section 9.1 in the session document shape." + ), + "withdrawn-ip-hash-batch.json": ( + "Batch member retrieval event carries ip_hash in data. " + "Violates section 9.1 in the event batch shape." + ), + "retrieved-missing-source-role.json": ( + "content_retrieved event carries no source_role. " + "Violates sections 5.2.2 and 5.7.1: source_role MUST be set on every retrieval." + ), + "ctx-token-on-grounded.json": ( + "Event-level ctx_token on a content_grounded event. " + "Violates section 5.2: ctx_token is valid only on content_engaged." + ), + "citation-id-on-grounded.json": ( + "citation_id on a content_grounded event. " + "Violates section 5.2: citation_id is valid only on content_presented." + ), + "presentation-id-on-cited.json": ( + "presentation_id on a content_cited event. " + "Violates section 5.2: presentation_id is valid only on content_engaged." + ), + "turn-on-content-event.json": ( + "turn object on a content_grounded event. " + "Violates section 5.2: turn data is carried on turn_started and turn_completed only." + ), + "duplicate-event-id.json": ( + "Two content_presented events share one id. " + "Violates section 6.6: repeated presentations receive distinct event IDs." + ), + "engaged-presentation-content-mismatch.json": ( + "content_engaged references a content_presented event of different content. " + "Violates section 6.7: the engagement identifies the same content as the presentation it acted on." + ), + "presented-citation-content-mismatch.json": ( + "content_presented.citation_id references a content_cited event of different content. " + "Violates section 6.6: the presentation and the citation it realises identify the same content." + ), + "engaged-presentation-id-no-presentations.json": ( + "content_engaged carries a presentation_id in a session with no content_presented events. " + "Violates section 6.7: every engagement references an exact presentation occurrence." + ), + "shared-ctx-token-two-presentations.json": ( + "One event-level ctx_token appears on engagements bound to two different presentations. " + "Violates section 7.4.1: a token is bound to exactly one presentation occurrence." + ), + "standalone-ctx-token-non-engagement.json": ( + "Standalone envelope carries ctx_token with a content_grounded event. " + "Violates section 7.1: an envelope ctx_token accompanies content_engaged events only." + ), + "batch-ctx-token-non-engagement.json": ( + "Event batch under ctx_token carries a content_grounded event. " + "Violates section 7.1: an envelope ctx_token accompanies content_engaged events only." + ), + "batch-missing-session-mixed-retrieval.json": ( + "Event batch mixing a retrieval with a grounding event carries neither session_id nor ctx_token. " + "Violates section 7.1: the retrieval-only exemption does not extend to a batch that carries other events." + ), + "manifest-domains-on-path-manifest.json": ( + "Manifest served under a path prefix carries domains. " + "Violates section 8.6: domains MAY appear only on manifests served from the domain root." + ), + "manifest-ctx-resolution-on-content-owner.json": ( + "content_owner manifest declares telemetry.ctx_resolution. " + "Violates section 8.5: ctx_resolution is valid on agent and platform manifests." + ), + "manifest-identifier-schemes-on-non-owner.json": ( + "Agent manifest carries identifier_schemes. " + "Violates section 8.6: identifier_schemes only on content_owner manifests." + ), + "manifest-lookalike-domain.json": ( + "Manifest at example.com claims evilexample.com in domains. " + "Violates section 8.6: a lookalike host is not a subdomain of the manifest host." + ), +} + +# V0.1 fields prohibited by the v1 migration rule (section 9.1). This is a +# compatibility check for the v1 transition, not a general registry of every +# field the specification may ever withdraw. The schemas cannot catch it: +# event `data` accepts additional properties by design. +V1_PROHIBITED_V0_1_EVENT_DATA_FIELDS = { + "ip_hash": "prohibited by the v1 migration rule; hashing does not anonymise an IP address", } # Event types that carry content and therefore require an identifier # (content_url or content_id) under section 5.7.5. turn_started and # turn_completed are turn events, not content events, and are exempt. CONTENT_EVENT_TYPES = { - "content_retrieved", "content_grounded", "content_cited", - "content_displayed", "content_engaged", + "content_retrieved", "content_grounded", + "content_cited", "content_presented", "content_engaged", } # Fields that MUST NOT appear at each privacy level (section 5.5). @@ -162,11 +373,11 @@ def load_schema(schema_path): manifest_schema_id = manifest_schema.get("$id", "") manifest_resource = Resource.from_contents(manifest_schema) registry = registry.with_resource(manifest_schema_id, manifest_resource) - manifest_validator = Draft202012Validator(manifest_schema, registry=registry) + manifest_validator = Draft202012Validator(manifest_schema, registry=registry, format_checker=FORMAT_CHECKER) else: manifest_validator = None - validator = Draft202012Validator(schema, registry=registry) + validator = Draft202012Validator(schema, registry=registry, format_checker=FORMAT_CHECKER) return schema, event_schema, batch_schema, validator, manifest_validator, registry @@ -204,12 +415,12 @@ def validate_standalone_event(data, session_schema, event_schema, registry): just the event body against the TelemetryEvent definition. """ if event_schema is not None: - validator = Draft202012Validator(event_schema, registry=registry) + validator = Draft202012Validator(event_schema, registry=registry, format_checker=FORMAT_CHECKER) errors = list(validator.iter_errors(data)) else: schema_id = session_schema.get("$id", "") wrapper = {"$ref": f"{schema_id}#/$defs/TelemetryEvent"} - validator = Draft202012Validator(wrapper, registry=registry) + validator = Draft202012Validator(wrapper, registry=registry, format_checker=FORMAT_CHECKER) errors = list(validator.iter_errors(data["event"])) return errors @@ -222,11 +433,11 @@ def validate_event_batch(data, session_schema, batch_schema, registry): event body against the TelemetryEvent definition. """ if batch_schema is not None: - validator = Draft202012Validator(batch_schema, registry=registry) + validator = Draft202012Validator(batch_schema, registry=registry, format_checker=FORMAT_CHECKER) return list(validator.iter_errors(data)) schema_id = session_schema.get("$id", "") wrapper = {"$ref": f"{schema_id}#/$defs/TelemetryEvent"} - validator = Draft202012Validator(wrapper, registry=registry) + validator = Draft202012Validator(wrapper, registry=registry, format_checker=FORMAT_CHECKER) errors = [] for event in data.get("events", []): errors.extend(validator.iter_errors(event)) @@ -237,12 +448,15 @@ def check_privacy_conformance(data): """ Check application-layer privacy conformance rules. + The privacy field gating of section 5.5 is a property of privacy_level + itself: it applies wherever conversation turns are emitted - session + documents, event batches, and standalone event envelopes alike. + Returns a list of violation descriptions, empty if conforming. """ violations = [] - events = data.get("events", []) - for event in events: + for event in _iter_events(data): turn = event.get("turn") if turn is None: continue @@ -262,8 +476,9 @@ def check_privacy_conformance(data): def _iter_events(data): - """Yield the content/turn events in a document, whether it is a session - (events list) or a standalone envelope (single event under 'event').""" + """Yield the content/turn events in a document, whatever its shape: a + session or event batch (events list) or a standalone envelope (single + event under 'event').""" if is_standalone_event(data): event = data.get("event") if isinstance(event, dict): @@ -312,21 +527,215 @@ def check_session_or_ctx_token(data): return [] +def check_referential_integrity(data): + """ + Check the intra-document event references of a session document: + + - Every content_engaged.presentation_id MUST reference the exact + content_presented.id on which the action occurred (section 6.7). + - Every citation_id on a content_presented event + references that content_cited event's id (section 6.6). + + Applies only to session documents, where the referenced events live in + the same document. Standalone envelopes and batch members legitimately + reference events delivered elsewhere (e.g. a click-out engagement carrying + a ctx_token), so they are exempt here; the corroborating click-out flow + is out of scope for this suite. Returns a list of violations. + """ + if is_standalone_event(data) or is_event_batch(data): + return [] + events = data.get("events", []) + presented_ids = { + e.get("id") for e in events + if e.get("type") == "content_presented" and e.get("id") + } + cited_ids = { + e.get("id") for e in events + if e.get("type") == "content_cited" and e.get("id") + } + violations = [] + for event in events: + etype = event.get("type") + if etype == "content_engaged": + pid = event.get("presentation_id") + if pid and pid not in presented_ids: + violations.append( + f"content_engaged presentation_id '{pid}' does not match " + "any content_presented event id in the session" + ) + if etype == "content_presented": + cid = event.get("citation_id") + if cid and cid not in cited_ids: + violations.append( + f"{etype} citation_id '{cid}' does not match any " + "content_cited event id in the session" + ) + return violations + + +def check_v1_migration_prohibitions(data): + """ + Check v1's explicit prohibitions on fields carried forward from v0.1. + Returns a list of violation descriptions. + """ + violations = [] + for event in _iter_events(data): + event_data = event.get("data") + if not isinstance(event_data, dict): + continue + for field, reason in V1_PROHIBITED_V0_1_EVENT_DATA_FIELDS.items(): + if field in event_data: + violations.append( + f"Event '{event.get('type')}' carries '{field}' in data ({reason})" + ) + return violations + + +def check_grounding_provenance(data): + """Check the provenance/cached pairings required by section 6.4.""" + violations = [] + for event in _iter_events(data): + if event.get("type") != "content_grounded": + continue + event_data = event.get("data") + if not isinstance(event_data, dict): + continue + provenance = event_data.get("provenance") + cached = event_data.get("cached") + if provenance == "agent_fetched" and cached is not False: + violations.append("content_grounded with agent_fetched does not carry cached false") + if provenance == "agent_cached" and cached is not True: + violations.append("content_grounded with agent_cached does not carry cached true") + fingerprint = event_data.get("content_fingerprint") + if isinstance(fingerprint, dict) and "preserved_in_output" in fingerprint: + violations.append( + "content_fingerprint carries preserved_in_output; withdrawn in v1" + ) + return violations + + +# Fields that belong to one event type (section 5.2). A key present with a +# non-null value on any other type is a placement violation. +FIELD_PLACEMENT = { + "presentation_id": {"content_engaged"}, + "ctx_token": {"content_engaged"}, + "citation_id": {"content_presented"}, + "turn": {"turn_started", "turn_completed"}, +} + + +def check_field_placement(data): + """Check source_role on retrievals (sections 5.2.2, 5.7.1) and the + event-type scoping of presentation_id, ctx_token, citation_id and turn + (section 5.2). Applies to every document shape.""" + violations = [] + for event in _iter_events(data): + etype = event.get("type") + if etype == "content_retrieved" and not event.get("source_role"): + violations.append("content_retrieved event carries no source_role") + for field, allowed in FIELD_PLACEMENT.items(): + if event.get(field) is not None and etype not in allowed: + violations.append( + f"Field '{field}' present on '{etype}' event; valid only on " + + ", ".join(sorted(allowed)) + ) + return violations + + +def _same_content(a, b): + """Two events identify the same content unless a shared identifier field + (content_id or content_url) carries different non-null values.""" + for field in ("content_id", "content_url"): + x, y = a.get(field), b.get(field) + if x is not None and y is not None and x != y: + return False + return True + + +def check_session_integrity(data): + """Session-document rules beyond check_referential_integrity (sections + 6.6, 6.7, 7.4.1): distinct event ids; engagement/presentation and + presentation/citation pairs identify the same content; one + event-level ctx_token binds to one presentation. Session documents only.""" + if is_standalone_event(data) or is_event_batch(data): + return [] + events = data.get("events", []) + violations = [] + seen_ids = set() + for e in events: + eid = e.get("id") + if eid: + if eid in seen_ids: + violations.append(f"Duplicate event id '{eid}' in session") + seen_ids.add(eid) + presented = {e.get("id"): e for e in events if e.get("type") == "content_presented" and e.get("id")} + cited = {e.get("id"): e for e in events if e.get("type") == "content_cited" and e.get("id")} + token_binding = {} + for e in events: + etype = e.get("type") + if etype == "content_engaged": + pid = e.get("presentation_id") + if pid in presented and not _same_content(e, presented[pid]): + violations.append( + f"content_engaged references presentation '{pid}' but identifies different content" + ) + token = e.get("ctx_token") + if token and pid: + bound = token_binding.setdefault(token, pid) + if bound != pid: + violations.append( + f"ctx_token '{token}' appears on engagements bound to two presentations" + ) + if etype == "content_presented": + cid = e.get("citation_id") + if cid in cited and not _same_content(e, cited[cid]): + violations.append( + f"{etype} references citation '{cid}' but identifies different content" + ) + return violations + + +def check_envelope_ctx_token(data): + """An envelope ctx_token accompanies content_engaged events only (section + 7.1); any other event type on such an envelope needs session_id instead.""" + if not (is_standalone_event(data) or is_event_batch(data)): + return [] + if not data.get("ctx_token"): + return [] + violations = [] + for event in _iter_events(data): + etype = event.get("type") + if etype != "content_engaged": + violations.append( + f"Envelope ctx_token accompanies a '{etype}' event; ctx_token is carried only for content_engaged" + ) + return violations + + def check_application_layer(data): """Run every application-layer conformance rule and return all violations.""" return ( check_privacy_conformance(data) + check_content_identifier(data) + check_session_or_ctx_token(data) + + check_referential_integrity(data) + + check_v1_migration_prohibitions(data) + + check_grounding_provenance(data) + + check_field_placement(data) + + check_session_integrity(data) + + check_envelope_ctx_token(data) ) def check_manifest_application_layer(data): """ - Check the manifest rejection rules of section 8.7 that JSON Schema cannot - express: duplicate keys[].id values, and domains entries that are not the - manifest's own host or a subdomain of it (section 8.6). - Returns a list of violation descriptions. + Check the manifest rejection rules that JSON Schema cannot express: + duplicate keys[].id values (section 8.7), domains entries that are not the + manifest's own host or a subdomain of it, domains on a path-prefixed + manifest (section 8.6), ctx_resolution on a manifest without the agent + or platform role (section 8.5), and identifier_schemes on a manifest + without the content_owner role (section 8.6). Returns a list of violation + descriptions. """ violations = [] @@ -337,7 +746,8 @@ def check_manifest_application_layer(data): violations.append(f"Duplicate keys[].id '{kid}'") seen.add(kid) - host = urlparse(data.get("id", "")).hostname + parsed = urlparse(data.get("id", "")) + host = parsed.hostname if host: for entry in data.get("domains", []): bare = entry[2:] if entry.startswith("*.") else entry @@ -347,9 +757,56 @@ def check_manifest_application_layer(data): f"'{host}' or a subdomain of it" ) + # domains only on a root manifest (section 8.6) + if "domains" in data and parsed.path != "/.well-known/content-telemetry.json": + violations.append( + f"domains present on a manifest served under a path prefix " + f"('{parsed.path}'); only root manifests carry domains" + ) + + # ctx_resolution only on agent/platform manifests (section 8.5) + telemetry = data.get("telemetry") or {} + roles = set(data.get("roles") or []) + if telemetry.get("ctx_resolution") and not roles & {"agent", "platform"}: + violations.append( + f"telemetry.ctx_resolution on a manifest with roles {sorted(roles)}; " + "valid on agent and platform manifests" + ) + + # identifier_schemes only on content_owner manifests (section 8.6) + if data.get("identifier_schemes") and "content_owner" not in roles: + violations.append( + f"identifier_schemes on a manifest with roles {sorted(roles)}; " + "identifier_schemes only on content_owner manifests" + ) + return violations +def schema_error_haystack(errors): + """ + Render the first schema error (deterministically chosen) as searchable + text: its JSON pointer plus its message, recursing into sub-errors of + combinators like anyOf. Invalid fixtures pin their intended violation by + requiring an _expected_error substring to appear in this text. + """ + first = sorted( + errors, + key=lambda e: ([str(p) for p in e.absolute_path], e.message), + )[0] + parts = [] + + def walk(error): + pointer = "/" + "/".join(str(p) for p in error.absolute_path) + parts.append(pointer) + parts.append(error.message) + for sub in error.context or []: + walk(sub) + + walk(first) + return " ".join(parts) + + def run_tests(): """Run all conformance tests and return (passed, failed, results).""" tests_dir = Path(__file__).parent @@ -426,9 +883,21 @@ def run_tests(): data = load_test_file(path) name = path.name desc = data.get("_test_description", "") + expected_error = data.get("_expected_error") is_app_layer = name in APPLICATION_LAYER_VIOLATIONS + # Every invalid fixture must pin its intended violation: a substring + # that must appear in the actual error. Without it, a fixture that + # fails for the wrong reason (e.g. after an unrelated edit) would + # still count as a pass. + if not isinstance(expected_error, str) or not expected_error: + print(f" FAIL {name}") + print(" Fixture missing required _expected_error field") + failed += 1 + results.append((name, False, "missing _expected_error")) + continue + if is_manifest_fixture(path): if manifest_validator is None: print(f" FAIL {name}") @@ -445,12 +914,19 @@ def run_tests(): schema_errors = list(session_validator.iter_errors(data)) if schema_errors: - # Failed JSON Schema - good - msg = schema_errors[0].message - print(f" PASS {name}") - print(f" Schema error: {msg}") - passed += 1 - results.append((name, True, None)) + # Failed JSON Schema - but only for the pinned reason + haystack = schema_error_haystack(schema_errors) + if expected_error in haystack: + print(f" PASS {name}") + print(f" Schema error: {schema_errors[0].message}") + passed += 1 + results.append((name, True, None)) + else: + print(f" FAIL {name}") + print(f" Schema error does not match _expected_error {expected_error!r}") + print(f" Actual: {haystack[:200]}") + failed += 1 + results.append((name, False, "wrong schema error")) elif is_app_layer: # Passes JSON Schema but should fail conformance @@ -459,16 +935,23 @@ def run_tests(): if is_manifest_fixture(path) else check_application_layer(data) ) - if conformance_violations: - print(f" PASS {name} [application-layer]") - print(f" {APPLICATION_LAYER_VIOLATIONS[name]}") - passed += 1 - results.append((name, True, None)) - else: + haystack = "; ".join(conformance_violations) + if not conformance_violations: print(f" FAIL {name}") print(f" Expected application-layer violation but none found") failed += 1 results.append((name, False, "Expected conformance violation")) + elif expected_error not in haystack: + print(f" FAIL {name}") + print(f" Violation does not match _expected_error {expected_error!r}") + print(f" Actual: {haystack[:200]}") + failed += 1 + results.append((name, False, "wrong conformance violation")) + else: + print(f" PASS {name} [application-layer]") + print(f" {APPLICATION_LAYER_VIOLATIONS[name]}") + passed += 1 + results.append((name, True, None)) else: # Should have failed schema but didn't @@ -477,6 +960,16 @@ def run_tests(): failed += 1 results.append((name, False, "Expected schema error")) + # --- Reconcile APPLICATION_LAYER_VIOLATIONS against invalid/ --- + # A dict key with no matching fixture file means an expectation silently + # dropped out of the suite (e.g. a renamed fixture). Fail the run. + invalid_names = {p.name for p in invalid_dir.glob("*.json")} + for name in sorted(set(APPLICATION_LAYER_VIOLATIONS) - invalid_names): + print(f" FAIL {name}") + print(" APPLICATION_LAYER_VIOLATIONS entry has no matching file in invalid/") + failed += 1 + results.append((name, False, "orphaned APPLICATION_LAYER_VIOLATIONS entry")) + # --- Summary --- total = passed + failed print()