diff --git a/SPECIFICATION.md b/SPECIFICATION.md index fbf71a2..8d6ced3 100644 --- a/SPECIFICATION.md +++ b/SPECIFICATION.md @@ -111,6 +111,9 @@ For the purposes of this specification, the following terms apply. | **privacy level** | data sharing tier controlling which conversation fields are populated: `full`, `summary`, `intent`, or `minimal` (section 5.5) | | **conformance level** | emitter capability tier: Retrieval, Grounding, or Citation (section 5.7) | | **content scope** | opaque identifier grouping sessions by their content access context (section 5.1.1) | +| **governing terms** | licence, contract or other terms selecting which events a relationship requires and the coverage, cadence, delivery, privacy and reports owed (sections 5.2.4, 5.7.6) | +| **qualifying occurrence** | occurrence satisfying an event type's core definition and occurrence boundary within the relationship scope reported under (section 5.7.6) | +| **coverage** | declared relationship between qualifying occurrences and emitted events: `complete`, `sampled`, `aggregated` or `selected` (section 5.7.6) | ## 4. Concepts @@ -195,26 +198,38 @@ Content moves through six stages during an agent interaction: 1. **Retrieved** - Content fetched over HTTP from an origin server, CDN, marketplace, or index. This is an infrastructure event observable by the content owner's infrastructure (origin server, edge network) and the agent. A retrieval may be cached by the agent for use across multiple sessions. + One retrieval occurrence is one completed fetch of a content representation as observed by the reporting party: a redirect chain resolving to one representation is one occurrence, and a revalidation returning no new representation (an HTTP 304) is not a new occurrence. Serving content from the agent's own cache is not a new retrieval; the reuse surfaces as grounding (stage 2), not as a repeated `content_retrieved` event. + 2. **Grounded** - Content used in the agent's generation context for this session or turn. The boundary is "this content entered the generation model's context" - the point where content can directly influence the model's output. Content used only for retrieval selection (embedding similarity search, re-ranking scores, routing decisions) without entering the generation context is not grounded. Grounding is architecture-neutral: same event whether the agent uses RAG, chain-of-thought reasoning, embeddings, or multi-step delegation (see section 6.4 for architecture-specific guidance). Grounding is decoupled from retrieval: content may be grounded from a live fetch, from agent-side cache, or from a pre-loaded index. Only the agent can report grounding events. + One grounding occurrence is one distinct content item entering a generation context at the declared `data.scope`: at `session` scope, a content item grounds once per session; at `turn` scope, once per turn it enters. A distinct content item is a distinct `content_id`, or its canonical `content_url` where no stable identifier exists (section 4.5). Continued presence within the declared scope is not a further occurrence; re-entry in a later turn is, when the scope is `turn`, and a change of `content_version` is a new occurrence at either scope. An emitter that ingests a content item in chunks MAY emit one grounding event per chunk, preserving the chunk-level hashes of section 6.4; events sharing content identity within one scope describe one occurrence, and consumers count occurrences by deduplicating on content identity and scope, not by counting events. + 3. **Reproduced** - The output artifact contains identified source content: a quotation, an excerpt, or a full copy, verbatim or near-verbatim. Reproduction is an output-construction claim by the system that built the output. It is independent of credit and of delivery: reproduced content may or may not also be cited, and the artifact may or may not later be presented. An uncredited excerpt in a response delivered through an API produces a `content_reproduced` event and nothing else - without this event, that use would be unreportable. + One reproduction occurrence is one identified source content item appearing in one output element (or in the output artifact, where no element identity exists), so a credited quotation and its companion citation share an `output_element_id` (section 6.6). Three quoted passages from one source in three elements are three occurrences; repeated appearance within one element is not a further occurrence. + 4. **Cited** - An output artifact explicitly associates identified source content with a response, claim, passage, quotation, or other output element. Citation is an output-construction relationship, not evidence that the output was delivered. A subset of grounded content is commonly cited, but a citation can also be emitted without a matching grounding event when an agent produces an uncorroborated or hallucinated source association. A citation MUST carry a resolvable reference to the source it associates: a `content_url` or a `content_id`. A source association with no resolvable reference is not a citation and MUST NOT be emitted as `content_cited`. Unlike other content events, where the identifier requirement is an application-layer rule (section 5.7.5), for `content_cited` and `content_reproduced` it is enforced by the JSON Schema. Reproduction and citation are sibling claims about the same artifact: reproduction records the material, citation records the credit. Neither implies the other. A credited quotation produces both events; an uncredited excerpt produces only a reproduction; a reference citation with no quoted material produces only a citation. + One citation occurrence is one distinct association between a source and an output element (or the output artifact, where no element identity exists). Associating the same source with three separate output elements produces three citation events; repeating the same association is not a further occurrence. + 5. **Presented** - Content or a source reference was rendered, played, spoken, embedded, or otherwise made perceivable on a recipient-facing surface. Presentation does not assert that a person noticed or attended to it. `presentation_kind` distinguishes source content (including a reproduced excerpt or media) from a source reference (such as a link, credit, or card). Not all citations are presented: an output can be stored, suppressed, or passed to another system before delivery. Grounding and presentation record different boundary crossings: grounding records entry into a generation context, while presentation records a recipient-facing delivery occurrence. As agent experiences evolve beyond the chat window the two diverge - content can shape an answer whose source is never presented, and an agent can present content that never entered a generation context (see *Departures from the funnel model* below). + One presentation occurrence is one rendering of content or a source reference on a recipient-facing surface; the event's `id` names that occurrence. Presenting the same artifact again - on a new surface, or in a new delivery - is a new occurrence. + 6. **Engaged** - The recipient or agent performed an observable action on a presentation: clicked a link, expanded a preview, copied text, shared the response, or directed the agent to act on the content. It does not imply attention beyond the reported action. `presentation_id` links the action to the exact presentation occurrence; a click-out can also carry a `ctx_token` that a destination resolves to the click context (section 7.4). + One engagement occurrence is one observed action on one presentation occurrence. + ``` Retrieved (HTTP layer, cacheable) → Grounded (influence layer, per-session or per-turn) @@ -232,6 +247,8 @@ Each stage after retrieval is typically a progressively narrower subset, except - **Citation-to-presentation** measures source associations constructed in output but not made perceivable - **Presentation-to-engagement** measures observable actions on exact presentation occurrences +These ratios are computed over reported events and are comparable across emitters only at known coverage (section 5.7.6). + #### Departures from the funnel model Five cases break the strict subset model: @@ -323,6 +340,7 @@ Additional content metadata - version, last-modified timestamp, content hash, me | `ended_at` | datetime | No | Session end (UTC) | | `conformance_level` | string | No | Informational conformance level advertised by the emitter (see section 5.7). Values: `retrieval`, `grounding`, `citation` | | `document_type` | string | No | `"session"` for session documents (see section 7.1 for the standalone event and event batch formats) | +| `data` | object | No | Session-level extension container, including access context (see 5.1.3) | | `events` | Event[] | No | Ordered list of events | `parent_session_id` links a delegated session to its immediate parent without @@ -357,6 +375,36 @@ The `manifest_ref` field optionally references a manifest (section 8), identifyi Format: the URL of a manifest served at `/.well-known/content-telemetry.json` under a path the participant controls. +#### 5.1.3 Session data and access context + +Sessions carry an optional `data` object mirroring the event-level `data` field (section 11.1): an extension container for session-scoped metadata. Extensions SHOULD namespace custom fields or use containers documented in this specification, and consumers MUST tolerate unknown fields within it. The session root itself is not an extension point: custom top-level siblings of `events` are not defined by this specification, and consumers are not required to preserve or interpret them. + +One container is defined in core. `access_context` records the context from which the session's access rights derive - the institution, not the individual: + +```json +{ + "schema_version": "0.1", + "session_id": "770e8400-e29b-41d4-a716-446655440000", + "content_scope": "consortium-agreement-4471", + "started_at": "2026-08-13T14:02:10Z", + "data": { + "access_context": { + "identifiers": [ + { "scheme": "ror", "value": "https://ror.org/013meh722" }, + { "scheme": "saml_entity_id", "value": "https://idp.example.ac.uk/shibboleth" } + ] + } + }, + "events": [] +} +``` + +`identifiers` is an array of typed identifiers, each a `scheme` and a `value`. `ror`, `saml_entity_id` and `isni` are the core scheme values; emitters MAY use other schemes and telemetry consumers MUST tolerate unknown ones, as with `media_type` (section 6.1). Access rights can derive through consortia, federated identity and proxies at once, so a session may carry both a SAML entity ID and the ROR ID it maps to. + +`access_context` identifies an institution, never an individual, and like every session field it is a claim by the emitter. The field serves the third-party agent that holds the entitlement and asserts the affiliation to the content owner; where the owner authenticated the session itself (`source_role` of `origin` or `edge`), it already knows the institution. Corroborating an asserted affiliation is verification-layer work, outside core. + +Emitters MUST NOT populate `access_context` unless the governing terms of the relationship require it, and SHOULD pair it with `intent` or `minimal` conversation-turn data (section 5.5): an identified institution combined with query text can come close to identifying an individual at a small subscriber. + ### 5.2 Event | Field | Type | Required | Description | @@ -375,6 +423,7 @@ Format: the URL of a manifest served at `/.well-known/content-telemetry.json` un | `content_url` | string | No | Content URL as fetched or canonical URL | | `content_id` | string | No | Content owner's stable content identifier (see 4.5) | | `license_ref` | string | No | Reference to a licence or grant the emitter associates with this event (see 5.2.3) | +| `terms_ref` | string | No | Reference to the governing terms the emitter associates with this event (see 5.2.4) | | `turn` | ConversationTurn | No | Conversation data (for turn events) | | `data` | object | No | Type-specific metadata (see section 6) | @@ -398,6 +447,16 @@ Core does not resolve, validate or interpret the reference. `license_ref` is par `license_ref` also does not identify the party whose entitlement was used. Where a publisher issues one grant per subscriber the value may work as a proxy for that subscriber, but only within the issuing publisher's namespace: nothing here requires the value to be typed, stable across sessions, or comparable between emitters. +#### 5.2.4 Terms reference + +The `terms_ref` field associates a telemetry event with the governing terms under which it is reported: a licence agreement, a standard-form contract, a tariff, a profile's terms, or any other terms document. Like `license_ref`, the value MAY be a public URL or an opaque identifier that both parties can resolve. Nothing requires the terms to be published: per-relationship terms are often confidential, and an opaque identifier resolved privately is a conforming reference. + +`license_ref` records which grant the emitter says applied (5.2.3). `terms_ref` records which terms govern the event's commercial consequences and the emitter's reporting obligations. Either may appear without the other: an access outside any grant carries no `license_ref`, and can still carry the `terms_ref` of the terms that attach consequences to that access. + +Core does not resolve, validate or interpret the reference, and `terms_ref` does not redefine core event semantics. Governing terms select which events a relationship requires and at what coverage (section 5.7.6), together with the cadence, delivery, privacy and reports owed (SCOPE.md); the meaning and occurrence boundary of each event remain those defined in sections 4.3 and 6, whatever `terms_ref` points to. + +A processor that stores, forwards or transforms a document MUST preserve `terms_ref` unchanged and MUST NOT remove or rewrite it. A `terms_ref` value MUST always refer to the same terms: when terms change, the emitter references them with a new value, so events emitted under earlier terms remain resolvable to them. A consumer MUST NOT read the absence of `terms_ref` as a statement that no terms governed the event (the same rule as event absence under coverage, section 5.7.6). + ### 5.3 Event types #### Content events @@ -466,6 +525,8 @@ These are the recommended values. Platforms with additional product surfaces (co An emitter that populates a conversation turn MUST NOT include a field above that turn's declared `privacy_level` - for example, `query_text` MUST NOT be present when `privacy_level` is `intent` or `minimal`. This restriction is a property of `privacy_level` itself: it applies wherever conversation turns are emitted, independent of the emitter's conformance level. +The session-level `access_context` container (section 5.1.3) does not appear in this table but is subject to the privacy model. It is populated only where governing terms require it, and pairing it with `full` or `summary` turn data is discouraged (section 5.1.3): `privacy_level` controls how much of the query and response is visible, `access_context` identifies whose access rights the session used, and populating both makes re-identification easier. + **Token counts** includes `query_tokens` and `response_tokens`. These are available at all levels because they are needed for token-based counting models and do not reveal user intent or platform strategy. They carry the same portability limit as `tokens_ingested` (section 6.4): both are measured in the emitter's own tokeniser and are not comparable between agents. Version 1 does not define corresponding turn-level character counts; `chars_ingested` measures source content placed in a generation context, not query or response length. **Response classification** includes `response_type` (e.g., `"recommendation"`, `"explanation"`). Available at `intent` level and above, as it can reveal the nature of the user's query. @@ -492,7 +553,7 @@ These are the core values. Extensions MAY define additional intent category valu Emitters that advertise a standard capability tier use one of three conformance levels. The authoritative declaration lives in the emitter's manifest (section 8). Emitters MAY also include an optional `conformance_level` field on individual session documents; when present it is informational and consumers MUST NOT treat it as a substitute for verifying the manifest's declaration. -Each level is named for the event it adds: a level proves the emitter produces that event and everything below it. These levels describe what an emitter reports, not what a consumer computes from it; attribution - the apportioning of credit across content - is performed by a telemetry consumer at whatever funnel level the parties agree (section 10), and can be computed from grounding alone, without citation. An emitter does not need to reach the Citation level for its telemetry to support attribution. +Each level is named for the event it adds: a level proves the emitter produces that event and everything below it. A level does not assert that every qualifying occurrence was reported (section 5.7.6). These levels describe what an emitter reports, not what a consumer computes from it; attribution - the apportioning of credit across content - is performed by a telemetry consumer at whatever funnel level the parties agree (section 10), and can be computed from grounding alone, without citation. An emitter does not need to reach the Citation level for its telemetry to support attribution. | Level | Events | What it proves | Typical emitter | |-------|--------|----------------|-----------------| @@ -566,6 +627,25 @@ The JSON Schema (`telemetry-session.json`) validates structure and types but can The `tests/` directory provides an informative reference suite for these rules. A consumer that receives a privacy-violating turn (e.g., `query_text` present at `minimal` level) SHOULD strip the offending fields rather than reject the document carrying them. +#### 5.7.6 Occurrence, qualifying events and coverage + +Each core event type has the meaning and occurrence boundary defined in sections 4.3 and 6. Profiles, deployment configurations and governing terms MUST NOT redefine them. A relationship that needs a different assertion defines a namespaced extension event (sections 5.3 and 11.1); it does not reuse a core type with altered semantics. + +An occurrence is **qualifying** for an emitter when it satisfies the core definition and occurrence boundary of its event type and falls within the relationship scope the emitter reports under - the content, domains or relationships selected by the applicable governing terms or deployment configuration. + +A conformance level (sections 5.7.1 to 5.7.3) does not assert that every qualifying occurrence was reported. Reporting coverage is a separate, explicit declaration, stated as one of four modes: + +- **complete** - every qualifying occurrence is emitted +- **sampled** - qualifying occurrences are emitted under a stated sampling rule +- **aggregated** - qualifying occurrences are reported only through a stated aggregation rule +- **selected** - only qualifying occurrences satisfying a further stated condition are emitted + +A coverage declaration states its mode together with the relationship scope it applies over; both MUST be disclosed to the receiving party. The rule or condition for `sampled`, `aggregated` and `selected` MUST be objectively decidable from information available at emission time and MUST NOT depend on the emitter's discretion at the moment of emission. An emitter reporting under governing terms that state a coverage mode MUST report at that mode, and an emitter MUST NOT declare or describe its reporting as `complete` for an event type unless every qualifying occurrence is emitted. A consumer MUST NOT treat the absence of an event as evidence that no occurrence happened except where complete coverage applies. + +An emitter MAY declare its coverage modes machine-readably in its manifest (`telemetry.coverage`, section 8.5); a manifest declaration is subject to the same rules, and where governing terms and a manifest declaration conflict, the governing terms take precedence for the relationships they cover. + +Whether an emitter's reporting in fact met its declared coverage is the completeness question of SCOPE.md's conformance list: it is answered by verification and audit mechanisms outside core, not by the declaration itself. + ## 6. Data profiles The `data` field on events carries type-specific metadata. These profiles document the recommended fields by event type and source role, in lifecycle order. None are required except where a section states otherwise (`reproduction_type` in 6.6; `presentation_kind` and `presentation_type` in 6.7), but emitting them enables richer attribution. @@ -577,11 +657,16 @@ When the reporter is the agent (`source_role: agent`), the following fields are | Field | Type | Description | |-------|------|-------------| | `media_type` | string | Content medium: `text`, `image`, `video`, `audio` (see below) | +| `content_depth` | string | Depth of the content record reached: `metadata`, `abstract`, `full` (see below) | `media_type` on retrieval events allows content owners to see what types of content are being fetched, independent of whether those retrievals result in grounding or citation. Defaults to `text` when absent. `text`, `image`, `video`, and `audio` are the core values. Emitters MAY use custom string values for media outside the core set (for example `3d` or `dataset`). Telemetry consumers MUST tolerate unknown `media_type` values. This rule applies to `media_type` on every event type that carries it (sections 6.4, 6.5, 6.6, 6.7). +`content_depth` records how much of the content record the retrieval reached: `metadata` for a bibliographic or descriptive record only, `abstract` for an abstract or summary record, `full` for the full content record. These are the core values; emitters MAY use custom values and telemetry consumers MUST tolerate unknown ones. Where entitlement gates depth, a retrieval that reached only an abstract and a retrieval of full text are otherwise indistinguishable at the retrieval layer. Depth records what was reachable at retrieval, independent of what portion later entered a generation context. + +Although listed in the agent profile, `content_depth` applies to `content_retrieved` events from any `source_role`. The origin that served the response knows the depth authoritatively, and origin and edge reporters SHOULD include it alongside their fields in sections 6.2 and 6.3 where entitlement gates depth. + ### 6.2 Edge enrichment (`content_retrieved` + `source_role: edge`) CDN and edge network integrations SHOULD include these fields: @@ -862,7 +947,7 @@ A standalone event carries `document_type`, `schema_version`, and optionally `se } ``` -An event batch carries the same envelope fields with `"document_type": "event_batch"` and an `events` array. Envelope-level fields (`session_id`, `parent_session_id`, `ctx_token`, `agent_id`, `started_at`) apply to every event in the batch; events belonging to different sessions MUST be delivered in separate batches or as session documents. +An event batch carries the same envelope fields with `"document_type": "event_batch"` and an `events` array. Envelope-level fields (`session_id`, `parent_session_id`, `ctx_token`, `agent_id`, `started_at`, `manifest_ref`) apply to every event in the batch; events belonging to different sessions MUST be delivered in separate batches or as session documents. ```json { @@ -900,6 +985,8 @@ The primary schema (`telemetry-session.json`) validates session documents. A sta An agent emitter that uses standalone events or event batches for streaming delivery and wants to achieve Grounding or Citation conformance MUST include the optional `agent_id` and `started_at` fields on the envelope. Each envelope MUST also carry `session_id`, except for click-out engagement events where `ctx_token` is used instead. Consumers reconstruct the session from the stream of envelopes sharing the same `session_id`, or resolve the owning session from `ctx_token`. +The optional `manifest_ref` field is available on standalone event and event batch envelopes, mirroring the session-level field (section 5.1.2). It identifies the emitter's manifest where no session document carries one. An event delivered standalone in support of settlement, audit or other obligations under governing terms (section 5.2.4) SHOULD carry `manifest_ref`, since it is the only envelope field that names the manifest - and so the domain - under which the emitter claims to report; verifying that claim uses the manifest mechanisms of section 8. + Origin-side emitters (source role `origin` or `edge`) are not expected to achieve Grounding conformance and do not need these fields. Telemetry consumers MUST accept all three delivery formats, reconstructing sessions from standalone events and event batches where needed. @@ -1093,6 +1180,9 @@ Public keys used to sign telemetry events emitted by this participant. Per-event | `endpoint` | string | Yes | HTTPS URL. For agents and platforms, the outbound submission endpoint. For content owners, the inbound destination for events about the content owner's content. | | `conformance_level` | string | No | Conformance level advertised by this participant's own emitter(s). One of `retrieval`, `grounding`, `citation` (see 5.7). | | `ctx_resolution` | string | No | HTTPS URL of the click-token resolution endpoint operated by or for this participant (see 7.4). Valid on `agent` and `platform` manifests. | +| `coverage` | object | No | Per-event-type coverage declaration: a map from event type to `{ "mode": …, "terms_ref": … }`, where `mode` is one of `complete`, `sampled`, `aggregated`, `selected` (see 5.7.6) and `terms_ref` optionally names the terms stating the rule or condition. | + +`coverage` makes the emitter's declared coverage machine-visible. It is a claim like the rest of the manifest, subject to the rules of section 5.7.6: a `complete` entry asserts that every qualifying occurrence of that event type within the declared relationship scope is emitted, and the other modes are meaningful only with their rule or condition reachable through `terms_ref` or otherwise disclosed to the receiving party. `conformance_level` is informational. It advertises the level of telemetry the manifest's participant emits. It does **not** constrain what an inbound `endpoint` accepts - an endpoint accepts whatever events it is configured to accept, regardless of any level declared here - and it places **no requirement** on other emitters. On a `content_owner` manifest it describes only the events the owner's own infrastructure emits (typically a CDN edge worker at `retrieval`); it says nothing about what agents or platforms report about the owner's content, which those parties advertise in their own manifests. A `content_owner` manifest SHOULD omit `conformance_level` unless the owner operates its own emitter. There is no field for a content owner to *request* a minimum level from agents; consumers tolerate events from any level (see 5.7), and the protocol does not give a manifest a way to demand more. @@ -1292,6 +1382,8 @@ Implementations MAY extend core event types with custom fields in the `data` obj New event types (e.g., a commerce extension's `checkout_completed`) require a schema extension. The core schema validates only the event types listed in section 5.3. +Sessions carry a parallel extension container: the session-level `data` object (section 5.1.3). Session-scoped extension metadata belongs there, not in custom top-level fields on the session document. + ### 11.2 Custom intent categories `query_intent` accepts custom string values beyond the core set. Extensions SHOULD namespace their values to avoid collisions (e.g., `price_check` for a commerce extension). For ad-hoc categories that don't warrant a formal extension, use `other` with details in `topics`. @@ -1377,6 +1469,8 @@ gain the `ct_` pattern; presentation binding moves from the URL-carried `presentation_id` a destination could never legitimately know to issuer state restored at resolution. +V1 tightens occurrence boundaries (section 4.3). Each core event now has a stated occurrence and cardinality: retrieval per completed fetch (a cache serve is not a retrieval), grounding per distinct content item per declared scope (chunk-level events deduplicate to one occurrence by content identity), reproduction per source item per output element, presentation per rendering occurrence, engagement per observed action. Preview emitters that emitted per chunk, per passage, or re-emitted `content_retrieved` on cache serves remain schema-valid but SHOULD re-map to the stated boundaries; consumers comparing preview and v1 volumes should expect counts to shift where emitters previously chose finer or coarser units. Coverage becomes an explicit declaration (section 5.7.6) rather than an implication of conformance level. Session-scoped extension metadata belongs in the session-level `data` container (section 5.1.3); custom top-level siblings of `events`, accepted by the preview schema, are undefined. + Preview versions (0.x) use two-component version numbers. From 1.0.0 onward, versions follow [semantic versioning](https://semver.org/): - **Major** (1.0.0 → 2.0.0) - breaking changes to required fields diff --git a/manifest.json b/manifest.json index ddaa8d2..fd69a7b 100644 --- a/manifest.json +++ b/manifest.json @@ -88,6 +88,25 @@ "format": "uri", "pattern": "^https://", "description": "HTTPS URL of the click-token resolution endpoint operated by or for this participant. Valid on agent and platform manifests. Destinations resolve ctx_iss to this manifest and present the token here. (Section 7.4)" + }, + "coverage": { + "type": "object", + "description": "Per-event-type coverage declaration, a map from event type to a mode object. A claim subject to the rules of section 5.7.6: complete asserts every qualifying occurrence within the declared relationship scope is emitted. (Sections 8.5, 5.7.6)", + "additionalProperties": { + "type": "object", + "required": ["mode"], + "properties": { + "mode": { + "type": "string", + "enum": ["complete", "sampled", "aggregated", "selected"], + "description": "Coverage mode (section 5.7.6)" + }, + "terms_ref": { + "type": ["string", "null"], + "description": "Reference to the terms stating the sampling, aggregation or selection rule and the relationship scope (section 5.2.4)" + } + } + } } }, "description": "Telemetry endpoint declaration. (Section 8.5)" diff --git a/telemetry-event-batch.json b/telemetry-event-batch.json index 2c2d987..ceef522 100644 --- a/telemetry-event-batch.json +++ b/telemetry-event-batch.json @@ -40,6 +40,10 @@ "format": "date-time", "description": "Session start timestamp (UTC). Mirrors the session-level started_at field. REQUIRED for emitters at Grounding conformance or above when using event batch delivery." }, + "manifest_ref": { + "type": ["string", "null"], + "description": "Manifest reference - the URL of a manifest at /.well-known/content-telemetry.json identifying the emitter. Applies to every event in the batch and mirrors the session-level manifest_ref field. See sections 7.1 and 8." + }, "events": { "type": "array", "minItems": 1, diff --git a/telemetry-event.json b/telemetry-event.json index 6ec0b53..144ff75 100644 --- a/telemetry-event.json +++ b/telemetry-event.json @@ -40,6 +40,10 @@ "format": "date-time", "description": "Session start timestamp (UTC). Mirrors the session-level started_at field. REQUIRED for emitters at Grounding conformance or above when using standalone event delivery." }, + "manifest_ref": { + "type": ["string", "null"], + "description": "Manifest reference - the URL of a manifest at /.well-known/content-telemetry.json identifying the emitter. Mirrors the session-level manifest_ref field. RECOMMENDED on standalone events supporting settlement or audit obligations under governing terms. See sections 7.1 and 8." + }, "event": { "$ref": "telemetry-session.json#/$defs/TelemetryEvent", "description": "The telemetry event" diff --git a/telemetry-session.json b/telemetry-session.json index 2f833a6..fb90e62 100644 --- a/telemetry-session.json +++ b/telemetry-session.json @@ -54,6 +54,38 @@ "format": "date-time", "description": "Session end timestamp (UTC)" }, + "data": { + "type": ["object", "null"], + "additionalProperties": true, + "description": "Session-level extension container mirroring the event-level data field (section 5.1.3). Custom fields SHOULD be namespaced; consumers MUST tolerate unknown fields.", + "properties": { + "access_context": { + "type": "object", + "additionalProperties": true, + "description": "Context from which the session's access rights derive - an institution, never an individual. Populated only where the governing terms of the relationship require it. See section 5.1.3.", + "properties": { + "identifiers": { + "type": "array", + "description": "Typed institutional identifiers. Access rights can derive through consortia, federated identity and proxies at once.", + "items": { + "type": "object", + "required": ["scheme", "value"], + "properties": { + "scheme": { + "type": "string", + "description": "Identifier scheme. Core values: ror, saml_entity_id, isni. Consumers MUST tolerate unknown schemes." + }, + "value": { + "type": "string", + "description": "Identifier value in the scheme's own format" + } + } + } + } + } + } + } + }, "events": { "type": "array", "items": { @@ -141,6 +173,10 @@ "type": ["string", "null"], "description": "Reference to a licence or grant the emitter associates with this event (JWT jti, CoMP package ID, or opaque identifier). Core does not resolve or interpret it. See section 5.2.3." }, + "terms_ref": { + "type": ["string", "null"], + "description": "Reference to the governing terms the emitter associates with this event (public URL or opaque identifier both parties can resolve). Distinct from license_ref, which records the grant that applied. Core does not resolve or interpret it, and it does not redefine core event semantics. See section 5.2.4." + }, "turn": { "oneOf": [{ "$ref": "#/$defs/ConversationTurn" }, { "type": "null" }], "description": "Conversation turn data (for turn_started/turn_completed)" diff --git a/tests/invalid/access-context-identifier-missing-value.json b/tests/invalid/access-context-identifier-missing-value.json new file mode 100644 index 0000000..6e00daa --- /dev/null +++ b/tests/invalid/access-context-identifier-missing-value.json @@ -0,0 +1,15 @@ +{ + "_test_description": "Session access_context identifier missing its value. The schema requires both scheme and value on every identifier (5.1.3).", + "document_type": "session", + "schema_version": "0.1", + "session_id": "880e8400-e29b-41d4-a716-446655440090", + "started_at": "2026-08-13T14:02:10Z", + "data": { + "access_context": { + "identifiers": [ + { "scheme": "ror" } + ] + } + }, + "events": [] +} diff --git a/tests/invalid/access-context-identifiers-not-array.json b/tests/invalid/access-context-identifiers-not-array.json new file mode 100644 index 0000000..2588b9f --- /dev/null +++ b/tests/invalid/access-context-identifiers-not-array.json @@ -0,0 +1,13 @@ +{ + "_test_description": "Session access_context.identifiers as a bare string. The schema requires an array of {scheme, value} objects (5.1.3).", + "document_type": "session", + "schema_version": "0.1", + "session_id": "880e8400-e29b-41d4-a716-446655440091", + "started_at": "2026-08-13T14:02:10Z", + "data": { + "access_context": { + "identifiers": "https://ror.org/013meh722" + } + }, + "events": [] +} diff --git a/tests/invalid/manifest-coverage-bad-mode.json b/tests/invalid/manifest-coverage-bad-mode.json new file mode 100644 index 0000000..1e076e6 --- /dev/null +++ b/tests/invalid/manifest-coverage-bad-mode.json @@ -0,0 +1,13 @@ +{ + "_test_description": "Manifest coverage entry with a mode outside the enum. The schema requires one of complete, sampled, aggregated, selected (8.5, 5.7.6).", + "schema_version": "0.1", + "id": "https://assistant.example.com/.well-known/content-telemetry.json", + "roles": ["agent"], + "operator": { "name": "Assistant Example" }, + "telemetry": { + "endpoint": "https://telemetry.assistant.example.com/v1/events", + "coverage": { + "content_grounded": { "mode": "partial" } + } + } +} diff --git a/tests/valid/event-standalone-terms-ref.json b/tests/valid/event-standalone-terms-ref.json new file mode 100644 index 0000000..7293ad0 --- /dev/null +++ b/tests/valid/event-standalone-terms-ref.json @@ -0,0 +1,15 @@ +{ + "_test_description": "Standalone retrieval event reported under governing terms with no grant: terms_ref names the terms (5.2.4), license_ref is absent, and the envelope carries manifest_ref (7.1) to tie the emitter to a domain without a session document.", + "document_type": "event", + "schema_version": "0.1", + "agent_id": "assistant.example.com", + "manifest_ref": "https://assistant.example.com/.well-known/content-telemetry.json", + "event": { + "type": "content_retrieved", + "timestamp": "2026-08-13T15:00:04Z", + "source_role": "agent", + "content_telemetry_id": "990e8400-e29b-41d4-a716-446655440071", + "content_url": "https://example.com/2026/08/13/markets-live", + "terms_ref": "https://terms.example.com/search-only/v2" + } +} diff --git a/tests/valid/manifest-coverage-declaration.json b/tests/valid/manifest-coverage-declaration.json new file mode 100644 index 0000000..642d0eb --- /dev/null +++ b/tests/valid/manifest-coverage-declaration.json @@ -0,0 +1,18 @@ +{ + "_test_description": "Agent manifest declaring per-event-type coverage in telemetry.coverage (8.5, 5.7.6): content_grounded reported complete, content_retrieved sampled under a rule stated in the referenced terms.", + "schema_version": "0.1", + "id": "https://assistant.example.com/.well-known/content-telemetry.json", + "roles": ["agent"], + "operator": { "name": "Assistant Example" }, + "telemetry": { + "endpoint": "https://telemetry.assistant.example.com/v1/events", + "conformance_level": "grounding", + "coverage": { + "content_grounded": { "mode": "complete" }, + "content_retrieved": { + "mode": "sampled", + "terms_ref": "https://terms.example.com/reporting/v1" + } + } + } +} diff --git a/tests/valid/session-access-context.json b/tests/valid/session-access-context.json new file mode 100644 index 0000000..ee6d3c4 --- /dev/null +++ b/tests/valid/session-access-context.json @@ -0,0 +1,64 @@ +{ + "_test_description": "Session-level data container (5.1.3): access_context identifies the institution whose entitlement the agent used, required by the governing terms, with turn data at intent level per 5.5. The retrieval event carries content_depth: full (6.1).", + "document_type": "session", + "schema_version": "0.1", + "session_id": "880e8400-e29b-41d4-a716-446655440080", + "agent_id": "scholar-assistant.example.com", + "content_scope": "consortium-agreement-4471", + "started_at": "2026-08-13T14:02:10Z", + "ended_at": "2026-08-13T14:03:44Z", + "data": { + "access_context": { + "identifiers": [ + { "scheme": "ror", "value": "https://ror.org/013meh722" }, + { "scheme": "saml_entity_id", "value": "https://idp.example.ac.uk/shibboleth" } + ] + } + }, + "events": [ + { + "type": "turn_started", + "timestamp": "2026-08-13T14:02:11Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "query_intent": "question", + "topics": ["materials science"] + } + }, + { + "type": "content_retrieved", + "timestamp": "2026-08-13T14:02:14Z", + "source_role": "agent", + "content_telemetry_id": "990e8400-e29b-41d4-a716-446655440081", + "content_url": "https://journals.example.com/article/10.1000/xyz123", + "content_id": "doi:10.1000/xyz123", + "license_ref": "grant-4471-2026", + "data": { + "media_type": "text", + "content_depth": "full" + } + }, + { + "type": "content_grounded", + "timestamp": "2026-08-13T14:02:16Z", + "source_role": "agent", + "turn_id": "1", + "content_id": "doi:10.1000/xyz123", + "data": { + "scope": "turn", + "chars_ingested": 18400 + } + }, + { + "type": "turn_completed", + "timestamp": "2026-08-13T14:03:40Z", + "turn_id": "1", + "turn": { + "privacy_level": "intent", + "query_intent": "question", + "topics": ["materials science"] + } + } + ] +}