Skip to content

[Story]: Confluence connector v0 (pages, page tree) #269

Description

@adoLime

User Story

As an admin or project manager, I want to connect a Confluence Cloud space and ingest its pages (incl. page-tree hierarchy and structured content) into the canonical artifact store, so that users can ask the chat about real architecture docs, ADRs, runbooks and meeting notes — the enterprise onboarding knowledge that lives in the wiki.

Context & Motivation

Acceptance Criteria

  • Given an admin or project manager provides a Confluence Cloud base URL (https://{domain}.atlassian.net) + space key + read-only API token, when the connector runs, then it ingests all pages of that space (title + storage-format body) as canonical artifacts of type page, preserving the page tree as relationships (parent_of / child_of) (Data Sources MP Define dev team split (Backend, Retrieval/AI, Frontend) #3).
  • Given a page's storage-format XHTML body, when mapped, then the canonical artifact carries body_text (clean text), sections (headings/levels), tables (as markdown), and code_blocks (extracted from code macros) — no raw XHTML leaks into body_text (Data Sources HP Plan Scrum events #7).
  • Given a configured allowlist/denylist of page paths or space keys, when ingestion runs, then only in-scope pages are ingested (Data Sources HP Rules #5).
  • Given each ingested page, when stored, then it carries provenance (source_system=confluence, source_url, source_artifact_id=<page id>, source_version=<page version number>, content_hash, ingestion_run_id, anchors [Story]: Ingestion run history (PM view) #100); on a second run with no upstream changes no duplicate artifacts are created (hash- + page-id-based idempotency); on an edited page (bumped version.number) only that page is updated (incremental sync — Data Sources HP Individual Reports should state issue IDs (in the GH project) + dev, test or docu work #9).
  • Given Confluence Cloud rate limits (HTTP 429) or transient errors, when hit, then the run retries with Retry-After backoff and reports partial success (no full failure) (Data Sources MP Define "Definition of Done" #1).
  • Given the Confluence Cloud connect/discover/update endpoints, when called, then they are protected with @PreAuthorize("hasRole('PM') or hasRole('ADMIN')") — consistent with the GitHub ([Story]: GitHub repository connector v0 (commits, files, issues, PR metadata) #159) and Jira ([Story]: Jira CSV issue connector adapter #160) connector role model; plain USER role is rejected with 403. The read-only API token is stored via the secret mechanism (anchors [Story]: Keep secrets out of the repo #94/[Story]: GitHub Actions secrets configured for CI #126), never in repo (Data Sources HP Define dev team split (Backend, Retrieval/AI, Frontend) #3).
  • CONFLUENCE added to the SourceSystem Kotlin enum; ConfluenceConnector implements IConnector (auto-picked up by connector overview, batch-patch, project-scoped source lookup); ArtifactCommand (and JPA entity + DB schema) extended with sections, tables, code_blocks, relationships fields with existing Github/Jira/Upload mappers updated to emit empty values (no behaviour change). OpenAPI documented; one real Confluence Cloud space ingested end-to-end on dev and queryable via chat.

Sub-Tasks

  • Confluence Cloud REST API client with read-only token auth (Basic email:api_token), base URL https://{domain}.atlassian.net/wiki/rest/api/
  • ConfluenceConnectorController with connect/discover/update endpoints, all annotated @PreAuthorize("hasRole('PM') or hasRole('ADMIN')") (mirrors GithubConnectorController / JiraController)
  • ConfluenceSpaceConnection + config (analogous to GithubRepositoryConnection); implement IConnector (ConfluenceConnector.kt, id "confluence"); extend SourceSystem Kotlin enum with CONFLUENCE and frontend SOURCE_SYSTEMS array.
  • Extend ArtifactCommand + JPA entity + DB migration with sections, tables, code_blocks, relationships fields; update GithubArtifactMapper, JiraArtifactMapper, UploadArtifactMapper to emit empty lists (backward-compatible).
  • Storage-format XHTML parser → canonical body_text + sections + tables + code_blocks (extract code macros, strip other ac:structured-macro); Confluence page → canonical artifact + chunks mapper (via SS-035 / [Story]: Canonical artifact schema and metadata normalization #161).
  • Page-tree fetch (/space/{key}/content/page?depth=all) + parent/child relationships; scope/allowlist config (space key, page path globs); incremental sync tracking last-synced version.number / version.when per page.
  • Rate-limit (429 + Retry-After) + retry/backoff handling; provenance fields populated ([Story]: Ingestion run history (PM view) #100); page restrictions captured raw in access.source_acl (no enforcement); secret wiring ([Story]: Keep secrets out of the repo #94/[Story]: GitHub Actions secrets configured for CI #126).
  • OpenAPI annotation + smoke test against a real Confluence Cloud space.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

enhancementNew feature or requeststoryteam:aiAI / Retrieval team (sprintstart-ai, Python)team:backendBackend team (sprintstart-backend, Kotlin/Spring Boot)

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions