diff --git a/docs/en/rfcs/1803-artifact-search-projections.md b/docs/en/rfcs/1803-artifact-search-projections.md new file mode 100644 index 000000000..500e0430f --- /dev/null +++ b/docs/en/rfcs/1803-artifact-search-projections.md @@ -0,0 +1,221 @@ +--- +title: Artifact Search Projections and Join-Free Retrieval +--- + +- Proposal Name: `artifact_search_projections` +- Start Date: 2026-09-30 +- RFC PR: [oceanbase/powercontext#1803](https://github.com/oceanbase/powercontext/pull/1803) +- Migration dependency: [Unified Versioned Database Migrations, #1771](https://github.com/oceanbase/powercontext/pull/1771) +- Amends RFCs: [0014](0014_memory_layer_design.md), [0051](0051_experience_skill_artifact_families.md), + [0080](0080_memory_search_reranking.md), [1417](1417_topic_memory.md) +- Related RFCs: [1396](1396_handoff_access_control.md), [1467](1467_artifact_tags.md), + [1549](1549_artifact_family_unification.md), [1652](1652_memory_quality_and_lifecycle.md) + +# Summary + +This RFC addresses index problems in OceanBase full-text and vector search by proposing a common query boundary for +searchable Artifact Families: **full-text and vector retrieval statements must not use business-table joins**. Matching, +eligibility filtering, candidate ordering, and truncation cannot depend on joins to authoritative history, heads, tags, +or other business tables. + +OceanBase must follow this query boundary. Whether it is also mandatory for other backends remains an open question. + +Each Family should use a search wide table or separate current search projections. Tables may be separated by full-text +and vector capabilities or by search granularity, such as topics and chunks. All data need not occupy one physical table. +After candidates are selected, content and display fields may be fetched in batches using exact references. + +Authoritative revisions and heads retain their identity, history, and current-version responsibilities. Search projections +are rebuildable derived data maintained synchronously with authoritative state. The standard covers Memory, Topic Memory, +Experience, and Skill; each Family's adaptation proceeds independently. + +# Motivation + +## Joins affect search indexes and retrieval results + +Families already maintain partial search projections, but search still involves cross-table associations: + +| Family | Current search structure | +| --- | --- | +| Memory | Separate full-text and vector projections; retrieval joins entry versions for content and uses correlated tag filters | +| Topic Memory | Separate Topic/chunk full-text and vector projections; vector candidates join display content after truncation | +| Experience | Full-text matching on common heads, followed by a join to authoritative revision content | +| Skill | The same full-text retrieval path as Experience | + +In OceanBase, a distance-ordered vector scan participating in a merge join causes nearest-neighbor loss. Joining full-text +search with other business tables also degrades the full-text index path. Topic Memory already limits ANN candidates +before joining content to work around the vector merge-join issue, showing that candidate selection and content reads +can be organized separately. + +Families need a clear retrieval boundary that keeps business associations for content, current state, and tags out of +full-text and vector retrieval plans. + +## Search data has different maintenance costs + +Full text, vectors, and multivalued tags use different indexes. Topics and chunks also have different search granularities. +Requiring one physical table couples content duplication, vector configuration, tag updates, and index lifecycles. +Colocating fields alone does not ensure that multivalued tags and vector filtering combine efficiently. + +Wide tables suit direct return of complete results. Separate projections suit independently maintained capabilities and +granularities. The common contract should allow Families to choose an appropriate layout. + +## Eligibility and content completion need different boundaries + +Content can be fetched in batches after candidates are selected. Scope, tags, lifecycle, and access eligibility affect +which objects qualify as candidates. Applying them only after taking a global top-k lets ineligible objects consume +candidate slots and can exclude relevant eligible objects. + +Allowing separate tables therefore requires guarantees for eligibility, exact references, and read consistency. + +# Guide-level explanation + +## Use a wide table or separate projections + +A Family may store search fields and response content in its own wide table and return complete results directly. +Alternatively, it may maintain separate current full-text and vector projections, retrieve exact references, match +positions, and scores, then fetch content in batches. + +Topic Memory can retain separate topic and chunk search granularities. Chunk hits still reference the exact Topic +revision, preserving candidate budgets, result folding, and scoring semantics. Full-text-only Experience and Skill do +not need vector projections. + +The corresponding Family owns these layouts. Families do not have to share one search table. + +## Return complete, consistent results + +Search interfaces continue returning their promised content, metadata, and exact version references. With separate +projections, the service performs batch content reads; callers need not request details per hit. Content reads use the +version selected during retrieval and cannot replace it by resolving latest again. + +After publication, deactivation, or tag changes, search selects candidates from consistent current state. Historical +references remain readable under access control. Search projection layout does not affect exact historical reads. + +## Preserve tag filtering semantics + +Tag management may retain independent authoritative assignments. Search uses synchronously maintained tag projections +or complete, bounded filter inputs established before retrieval to apply existing all/any matching semantics during +candidate selection. + +It cannot arbitrarily truncate the tag-matching object set before ranking or join tags after retrieval to filter results. +Deployments without embeddings retain existing full-text search. Backend capability boundaries explicitly define the +supported combinations of filtering and retrieval modes. + +# Reference-level explanation + +## Core contract + +The following query organization and projection maintenance requirements apply to backends covered by this standard. +Whether non-OceanBase backends are included remains an unresolved question. Every backend must preserve existing +authorization, filtering, current-version, and returned-content semantics. + +1. **No business joins in retrieval statements.** Full-text and vector candidate selection must not join authoritative + history, heads, tags, or other business tables. Correlated subqueries cannot bring those associations back into retrieval. +2. **Multiple current projections are allowed.** Family-owned wide tables or separate projections are recommended. + The Family and backend determine physical table counts, whether full text and vectors share storage, and search granularity. +3. **Eligibility participates in candidate selection.** Scope, tags, lifecycle, and access control retain their semantics. + Filtering after truncation cannot replace retrieval of eligible candidates. +4. **Content completion runs separately.** Once candidates exist, separate batch reads may complete their content. + Business tables must not be joined in the same statement that performs full-text or vector retrieval. Content reads + do not redetermine eligibility, the current version, or relevance ordering. +5. **Current state is consistent.** Authoritative changes and affected projections update atomically. All retrieval + channels and content reads in one search share a consistent committed state. +6. **History and budgets are preserved.** Rebuilding changes no authoritative identities, historical content, exact + references, or Source processing positions. Candidates, filter-input transfer, content reads, and fusion respect + budgets, without per-hit queries or unbounded object lists connecting stages. + +## Responsibilities and scope + +`pc_artifacts` and `pc_artifact_heads` retain their existing roles. Projections maintained by Family writers establish +currentness; retrieval does not join heads to identify the latest revision. Candidates carry exact identities sufficient +to identify their original content and match positions. + +Content completion may read the corresponding current content projection or authoritative exact-version records. +It must not silently substitute another revision, omit missing content, or interpret inconsistencies as empty results. +Returned content must correspond to the candidates used for scoring. + +Independent changes to tags and other eligibility information also update every affected projection synchronously. +Metadata changes need not all create content revisions. Existing Scope-level authorization can run before search; +changing search storage does not require changing the authorization model. + +Backends covered by this standard may implement search with native or auxiliary indexes. Internal database index access +is not a business-table join prohibited by this RFC. Application queries still respect the boundary between retrieval +and content reads. A shared standard allows backend-specific DDL and index implementations. + +This RFC covers existing and future searchable Families and defines only the common search standard. Each Family's +adaptation is designed and delivered independently; simultaneous completion is not required. Domain contracts own content +generation, evidence semantics, and semantic merging. Public return content, filters, and scoring contracts remain unchanged. + +## Upgrade and compatibility + +Schema changes and search projection rebuilding follow +[RFC #1771: Unified Versioned Database Migrations](https://github.com/oceanbase/powercontext/pull/1771), +using its unified maintenance entry point for the offline upgrade. + +Search projections are rebuilt from authoritative data without creating content revisions or re-extracting processed +Sources. Existing exact references and access control are preserved. Historical content reads do not depend on historical +vectors. + +# Drawbacks + +- **Projection maintenance costs.** Wide tables may duplicate content. Separate projections still duplicate some + eligibility attributes and require synchronous maintenance of multiple current representations. +- **Read coordination.** Batch content completion adds a read and must use the same versions and consistent state as retrieval. +- **Index combination limits.** Backend capabilities determine how multivalued tags combine with vector retrieval. + Removing joins does not eliminate filtering or scanning costs. +- **Upgrade downtime.** Projection migration and index rebuilding require extra space and a maintenance window. + Authoritative history continues to grow independently. + +# Rationale and alternatives + +## Constrain retrieval while allowing projection choices + +Wide tables reduce content lookups and coordination between copies, suiting Families whose response fields align with +search units. Separate projections allow independent maintenance of full text, vectors, topics, and chunks, reducing +large-field duplication and separating index management. This proposal recommends both layouts under the same query +boundary and consistency responsibilities. + +## Require exactly one wide table per Family + +This simplifies direct result return but restricts independent evolution of search granularities, vector configurations, +and backend indexes. It also cannot replace multivalued-tag index design, so it is not a mandatory physical requirement. + +## Keep business associations in OceanBase retrieval + +This reduces some duplication but retains OceanBase nearest-neighbor loss and full-text index degradation. Expressing +associations as correlated subqueries or outer joins in the same retrieval statement still combines business tables +with the retrieval plan. Separate batch content reads establish an explicit boundary between the two stages. + +# Prior art + +[RFC 1417](1417_topic_memory.md) provides current Topic/chunk projections, separate retrieval channels, and atomic +publication as a foundation for independent projections. [RFC 0051](0051_experience_skill_artifact_families.md) defines +Experience/Skill content and admission. [RFC 1467](1467_artifact_tags.md) and +[RFC 1396](1396_handoff_access_control.md) define tag and authorization semantics that must be preserved. + +# Unresolved questions + +1. Determine the minimum supported OceanBase version. +2. Decide whether this standard is also mandatory for non-OceanBase backends. + +The second question has two options; the backend scope remains undecided: + +| Option | Benefits | Costs | +| --- | --- | --- | +| Require every backend to follow the standard | Each Family changes its shared data model once and reuses common projection maintenance and retrieval logic, reducing long-term maintenance branches. | A common retrieval design may sacrifice some retrieval performance on backends such as SQLite and limit backend-specific optimization. | +| Require OceanBase to follow the standard; let other backends choose | Each backend can choose queries and storage layouts suited to its index capabilities, retaining room for independent optimization. | Divergent retrieval paths add maintenance branches. If storage models also diverge, they require mappings to a common business model and long-term maintenance of multiple read, write, migration, and verification paths. | + +SQLite and seekDB currently primarily serve embedded, local use, with relatively small expected datasets. Given this +usage, a common retrieval design with fewer maintenance branches is preferred. Whether compliance is mandatory for +every backend remains undecided. + +A shared standard does not require identical DDL or indexes across backends. Backends that are not required to follow it +may still reuse the same models and retrieval implementation voluntarily. + +If backends adopt different storage models, how to isolate those differences through a common business model and storage +adapters remains a deferred question for separate discussion. This RFC does not design that architecture or make its +construction a prerequisite for the index changes. + +# Future possibilities + +A unified search could combine each Family's results, for example returning Memory, Experience, and Skill in one request +and ranking them together. Scoring and ranking across Families would need separate decisions. This capability is not a +delivery requirement of this proposal. diff --git a/docs/zh/rfcs/1803-artifact-search-projections.md b/docs/zh/rfcs/1803-artifact-search-projections.md new file mode 100644 index 000000000..f516adeba --- /dev/null +++ b/docs/zh/rfcs/1803-artifact-search-projections.md @@ -0,0 +1,185 @@ +--- +title: 制品检索投影与无 JOIN 召回 +--- + +- 提案名称:`artifact_search_projections` +- 起始日期:2026-09-30 +- RFC PR:[oceanbase/powercontext#1803](https://github.com/oceanbase/powercontext/pull/1803) +- 迁移依赖:[统一版本化数据库迁移,#1771](https://github.com/oceanbase/powercontext/pull/1771) +- 修订 RFC:[0014](0014_memory_layer_design.md)、[0051](0051_experience_skill_artifact_families.md)、 + [0080](0080_memory_search_reranking.md)、[1417](1417_topic_memory.md) +- 相关 RFC:[1396](1396_handoff_access_control.md)、[1467](1467_artifact_tags.md)、 + [1549](1549_artifact_family_unification.md)、[1652](1652_memory_quality_and_lifecycle.md) + +# 摘要 + +本 RFC 针对 OceanBase 全文与向量检索中的索引问题,为可检索的 Artifact Family 提出共同的查询边界: +**全文与向量召回语句禁止使用业务表 JOIN**。匹配、资格过滤、候选排序与截断不能依赖关联权威历史、head、标签等 +业务表完成。 + +OceanBase 必须遵守这一查询边界;是否对非 OceanBase 后端同样强制执行,列为待决问题。 + +推荐各 Family 使用检索宽表或独立的当前检索投影。Family 可以按检索通道或主题、片段等搜索粒度拆表,无需将全部 +数据合并到一张物理表。候选确定后,允许按精确引用批量读取正文及展示信息。 + +权威版本与 head 保留身份、历史和当前版本职责;检索投影是与权威状态同步维护、可重建的派生数据。本规范覆盖 +Memory、Topic Memory、Experience 和 Skill,各制品的改造独立推进。 + +# 动机 + +## JOIN 影响检索索引与召回结果 + +当前各制品已经维护部分检索投影,但搜索仍有跨表关联: + +| 制品 | 当前检索结构 | +| --- | --- | +| Memory | 全文与向量投影分离,召回关联 entry 版本取得正文,标签过滤依赖关联查询 | +| Topic Memory | Topic、Chunk 全文投影与向量投影分离,向量候选截断后关联展示内容 | +| Experience | 在公共 head 上全文匹配,再关联权威版本取得内容 | +| Skill | 与 Experience 共用上述全文检索路径 | + +OceanBase 中,距离排序的向量扫描参与 merge join 会导致近邻丢失;全文检索与其他业务表的 JOIN 组合也会导致 +全文索引路径退化。Topic Memory 已通过先截断 ANN 候选、再关联内容的方式绕过向量 merge join 问题,说明候选 +选择与内容读取可以分开组织。 + +各 Family 需要明确的召回边界,避免正文、当前状态或标签等业务关联进入全文与向量的检索计划。 + +## 不同检索数据有不同维护成本 + +全文、向量和多值标签的索引方式不同;Topic 与片段的搜索粒度也不同。强制放进一张物理表,会把内容冗余、向量配置、 +标签更新及索引生命周期绑定在一起。仅把字段放到同一张表,也不能保证多值标签与向量过滤可以高效组合。 + +宽表适合直接返回完整结果,独立投影适合按检索能力和数据粒度分别维护。公共契约应允许各 Family 选择合适的布局。 + +## 资格过滤与正文补齐需要不同边界 + +正文可以在候选确定后批量读取。Scope、标签、生命周期和访问资格则会影响哪些对象应当成为候选,不能在取完全局 +前若干条之后才过滤。否则不合格对象会占用候选名额,符合条件的相关对象可能无法进入结果。 + +因此,允许分表必须同时约束候选资格、精确引用和读取一致性。 + +# 使用说明 + +## 使用宽表或独立投影 + +Family 可以将检索字段与返回内容放在自己的宽表中,直接返回完整结果;也可以分别维护全文、向量等当前投影, +召回时返回精确引用、命中位置和分数,再批量读取内容。 + +Topic Memory 可以保留主题与片段的独立搜索粒度。片段命中仍引用所属 Topic 的确切版本,保持原有的候选预算、 +结果折叠及评分语义。Experience、Skill 只支持全文时,无需建立向量投影。 + +这些布局都由对应 Family 负责,不要求所有制品共用一张检索表。 + +## 返回完整且一致的结果 + +搜索接口继续返回契约承诺的正文、元数据及精确版本引用。采用独立投影时,由服务完成批量内容读取,调用方无需 +逐条请求详情。补齐内容使用候选确定的版本,不能再次读取 latest 后替换候选的内容。 + +制品发布、停用或修改标签后,搜索基于一致的当前状态选择候选;历史引用继续按访问控制规则读取。 +指定历史版本的查询不受检索投影布局影响。 + +## 保持标签过滤语义 + +标签管理可以继续使用独立的权威关联记录。搜索通过同步维护的标签投影,或召回前已经完整确定的有界过滤条件, +在选择候选时执行既有的全部满足、任意满足语义。 + +不能先随意截断符合标签的对象集合再计算相关性,也不能召回之后才关联标签筛选。没有 embedding 的部署仍保留 +已有的全文搜索能力;支持哪些过滤与检索模式组合,由后端能力边界明确规定。 + +# 参考级说明 + +## 核心契约 + +以下查询组织与投影维护要求适用于纳入本规范的后端。非 OceanBase 后端是否纳入,见“未解决问题”;所有后端 +均须保持既有的权限、过滤、当前版本和返回内容语义。 + +1. **召回语句不含业务 JOIN。** 全文与向量候选选择不得关联权威历史、head、标签或其他业务表;相关子查询也不能 + 将这些关联带回召回路径。 +2. **允许多张当前投影。** 推荐 Family 拥有的宽表或独立投影;物理表数量、全文与向量是否共表、搜索单元粒度由 + Family 和后端共同确定。 +3. **资格参与候选选择。** Scope、标签、生命周期及访问控制保持既有语义,不能以截断后的过滤替代合格候选的召回。 +4. **内容补齐独立执行。** 候选产生后,允许通过独立的批量读取补齐内容;不得在包含全文或向量召回的同一语句中 + 再关联业务表。补齐阶段不重新决定候选资格、当前版本或相关性排序。 +5. **当前状态一致。** 权威变更与受影响投影原子更新;同一次搜索的各召回通道和内容补齐共享一致的已提交状态。 +6. **历史与预算保持。** 重建不改变权威身份、历史内容、精确引用或 Source 处理位置。候选、过滤条件传递、内容 + 补齐和融合遵守预算,不使用逐命中查询或无界对象清单连接各阶段。 + +## 职责与适用范围 + +`pc_artifacts` 与 `pc_artifact_heads` 保留现有职责。当前性由 Family 写入时维护的投影保证,召回不再 JOIN head +判断最新版本。候选必须携带足以确定原文及命中位置的精确身份。 + +内容补齐可以读取对应的当前内容投影,也可以读取权威的精确版本记录。它不能静默采用其他版本、遗漏缺失内容, +或把不一致解释为空结果。返回给用户的内容与评分所依据的候选必须对应。 + +标签及其他资格信息的独立变化也要同步反映到所有受影响投影;不要求所有元数据变化都生成内容 revision。 +当前的 Scope 级授权可以在进入搜索前完成,无需因存储改造改变授权模型。 + +适用本规范的后端可以使用原生索引或辅助索引表达检索能力。数据库内部的索引访问不属于本 RFC 禁止的业务表 JOIN; +应用查询仍需遵守召回与内容读取的边界。统一规范允许各后端保留不同的 DDL 和索引实现。 + +本 RFC 覆盖现有及新增的可检索 Family,只定义公共检索规范。各制品的改造分别设计、独立推进,不要求同时完成。 +内容生成、证据含义和语义合并由各自领域契约定义,公开的返回内容、过滤和评分契约继续保持。 + +## 升级与兼容 + +数据库结构变更和检索投影重建遵守 [RFC #1771:统一版本化数据库迁移](https://github.com/oceanbase/powercontext/pull/1771), +通过其统一维护入口完成停服升级。 + +检索投影从权威数据重建,不生成新的内容 revision,不重新抽取已经处理的 Source。已有精确引用与访问控制保留, +历史正文读取不依赖历史向量。 + +# 缺点 + +- **增加投影维护成本。** 宽表可能重复保存正文;分离投影仍需重复保存部分过滤属性,并同步维护多份当前状态。 +- **增加读取协调。** 批量补齐多一次内容读取,需要保证它与召回使用相同的版本及一致状态。 +- **索引组合仍有边界。** 多值标签与向量检索的组合能力取决于后端,移除 JOIN 不会消除其过滤和扫描成本。 +- **升级需要停服。** 投影迁移与索引重建需要额外空间和维护窗口,也不会消除权威历史本身的增长。 + +# 设计理由与替代方案 + +## 约束召回,允许选择投影布局 + +宽表可减少结果补齐和副本协调,适合返回字段与检索单元较为一致的 Family。独立投影可以分别维护全文、向量及 +主题、片段,减少大字段复制并独立管理索引。本提案推荐两种方式,统一其查询边界和一致性责任。 + +## 强制每个 Family 只有一张宽表 + +可以简化直接返回结果的路径,但会限制不同搜索粒度、向量配置和后端索引的独立演进。它也不能替代多值标签的索引 +设计,因此不作为所有 Family 必须遵守的物理要求。 + +## 在 OceanBase 召回中保留业务关联 + +能减少部分数据冗余,但会继续带来 OceanBase 向量近邻丢失及全文索引路径退化。将关联写成相关子查询,或者在同一 +召回语句外层补内容,仍将业务表与检索计划组合在一起。独立批量内容读取可以明确隔开这两个阶段。 + +# 先例 + +[RFC 1417](1417_topic_memory.md) 的 Topic/Chunk 当前投影、分通道召回与原子发布,为独立投影提供基础。 +[RFC 0051](0051_experience_skill_artifact_families.md) 定义 Experience/Skill 内容与准入; +[RFC 1467](1467_artifact_tags.md) 和 [RFC 1396](1396_handoff_access_control.md) 定义需要保留的标签与授权语义。 + +# 未解决问题 + +1. 确定 OceanBase 最低支持版本。 +2. 非 OceanBase 是否也必须遵循本规范。 + +第二项有以下两种方案,适用范围尚待确定: + +| 方案 | 优点 | 代价 | +| --- | --- | --- | +| 所有后端强制遵守 | 同一种制品在共享数据模型上完成一次改造,复用投影维护和检索的公共逻辑,减少长期维护分支。 | 统一检索设计可能牺牲 SQLite 等后端的部分检索性能,限制针对单一后端的优化空间。 | +| 仅 OceanBase 强制遵守,其他后端自行选择 | 各后端可以按自身索引能力选择查询方式和存储布局,保留单独优化的空间。 | 检索路径分化会增加维护分支;存储模型进一步分化时,还需处理到统一业务模型的映射,长期维护多套读写、迁移和验证路径。 | + +SQLite 和 seekDB 当前主要用于嵌入式、本地场景,预期数据规模较小。基于这一使用定位,倾向于优先采用统一检索 +设计、减少维护分支;是否对所有后端强制执行仍待确定。 + +统一遵守规范不要求各后端使用完全相同的 DDL 和索引;不强制遵守的后端也可以主动复用同一套模型与检索实现。 + +如果不同后端采用不同存储模型,如何通过统一业务模型和存储适配隔离差异,作为遗留问题单独讨论。本提案不设计 +这套架构,也不将其建设作为索引改造的前置条件。 + +# 后续可能性 + +可以在各 Family 的检索结果之上提供统一搜索,例如一次搜索同时返回 Memory、Experience 和 Skill,再合并排序。 +这需要另外确定不同制品之间的评分和排序规则,不作为本提案的交付要求。