Skip to content

ci: Node 24 전환과 pre-Pages 문서 보존 감사 - #66

Merged
hellices merged 26 commits into
mainfrom
docs/pre-pages-preservation-audit
Sep 14, 2026
Merged

hellices merged 26 commits into
mainfrom
docs/pre-pages-preservation-audit

Conversation

@hellices

Copy link
Copy Markdown
Owner

변경 내용

GitHub Actions Node 24 전환

  • actions/checkout@v7
  • actions/setup-python@v7
  • actions/configure-pages@v6
  • actions/upload-pages-artifact@v5
  • actions/deploy-pages@v5

Documentation CI와 Pages는 역사 감사를 위해 fetch-depth: 0을 사용하고, Oryx workflow는 shallow checkout을 유지합니다.

pre-Pages 보존 감사

Pages 구축 직전 커밋 a4e6801과 최초 Pages 커밋 9ace9667을 고정 기준으로 재실행 가능한 감사를 추가했습니다.

  • 기준 tracked 파일: 359개
  • 기준 Markdown: 73개
  • 기준 공개 문서: 62개
  • 현재 canonical 문서: 66개(MCP PR #62의 신규 4편 포함)
  • 현재 선언된 redirect: 68개

감사는 다음을 검증합니다.

  • 62개 기준 문서의 one-to-one Git/redirect 계보
  • 제목, heading, prose, code fence, table, image, 실제 link destination과 multiplicity
  • 검토 예외의 정확한 누락 횟수 및 현재 replacement fingerprint/file SHA
  • evidence가 tracked public source인지 여부
  • canonical HTML, redirect, 검색, 서비스/글 찾기 reachability
  • 렌더링 이미지·download·script·stylesheet·srcset와 sample publish byte
  • hidden/inert/details/SVG/inline CSS 가시성
  • URL origin, Pages prefix, encoding, traversal, symlink 및 MkDocs Markdown alias
  • sample source가 Pages artifact에 노출되지 않는지 여부

Documentation CI와 Pages 배포가 strict build/search 이후 이 감사를 실행합니다.

감사에서 발견해 수정한 실제 문제

  1. Agent Memory 종합 문서에서 pre-Pages의 문서별 목적·대상 독자 표가 사라져 있었습니다. 역사적 reader map으로 owning sample README에 복원하고 canonical 진입 문서에서 연결했습니다.
  2. Azure MySQL Blue/Green 문서의 <details> 3개가 markdown="1" 없이 사용돼 제목·표가 literal text로 보이고 placeholder가 HTML 태그로 해석됐습니다. wrapper 렌더링 설정만 수정했습니다.

그 외 35개 문서의 변경 248건은 안전성 치환, canonical 경로, 현재 실행 명령, 이미지 alias, 검증 상태 등의 현재 replacement/evidence 415건에 결속해 재검증합니다. 27개 문서는 예외 없이 구조가 보존됐습니다.

새 upstream 반영

작업 중 병합된 MCP PR #62의 72개 변경 파일은 그대로 유지했습니다. 기준 보존 집합은 62편으로 고정하고, 현재 문서 수는 동적으로 처리해 신규 4편을 포함한 66편 전체의 표시·검색·자산을 감사합니다.

검증

  • documentation tests: 1,459 passed
  • JavaScript tests: 12 passed
  • Azure SRE Agent sample: 440 passed
  • metadata/source/link validation: 66 documents
  • public safety: 364 files
  • strict MkDocs build: passed
  • search coverage: 66 documents / 13 tags
  • preservation audit:
    • 359 baseline files / 73 Markdown
    • 62/62 baseline documents mapped and preserved
    • 66 current documents searchable and visible
    • declared redirects and local assets verified
  • final whole-branch review: Critical/Important findings 없음

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The lineage audit can overwrite Git-detected changes as reviewed; this critical validation issue must be fixed, and the Pages action reference should be updated to v5.

Get a fresh assessment by requesting another Copilot review.

Pull request overview

This PR upgrades documentation CI to Node 24-compatible Actions and adds a reproducible pre-Pages preservation audit.

Changes:

  • Updates workflow actions and checkout history settings.
  • Adds source, site, visibility, redirect, asset, and artifact auditing with tests.
  • Restores Agent Memory guidance and fixes MySQL details rendering.
File summaries
File Summary
tests/docs/test_workflows.py Tests workflow versions and audit placement.
tests/docs/test_pre_pages.py Tests baseline inventory and Git lineage validation.
tests/docs/test_pre_pages_site.py Tests rendered-site, link, and URL validation.
tests/docs/test_pages_artifact.py Tests published artifact contents and filtering.
tests/docs/fixtures/pre_pages_chromium_visibility.json Provides browser visibility fixtures.
scripts/docs/pre_pages.py Implements baseline inventory and lineage auditing.
scripts/docs/pre_pages_visibility.py Validates authored-content visibility.
scripts/docs/pre_pages_site.py Audits pages, redirects, links, and assets.
scripts/docs/pre_pages_css.py Parses supported visibility CSS.
scripts/docs/pre_pages_content.py Validates Markdown structure and evidence.
scripts/docs/audit_pre_pages.py Provides the audit CLI and reporting.
README.md Documents audit usage.
docs/services/microsoft-foundry/agent-memory/samples/research-artifacts/README.md Restores the historical reader map.
docs/services/microsoft-foundry/agent-memory/index.md Links the restored reader guidance.
docs/services/azure-database-for-mysql/blue-green-upgrade/index.md Fixes details-block Markdown rendering.
docs/contributing/index.md Documents audit requirements and baseline semantics.
CONTRIBUTING.md Updates contributor validation guidance.
AGENTS.md Updates repository workflow guidance.
.github/workflows/pages.yml Upgrades Pages actions and audits output before deployment.
.github/workflows/oryx-python-build-test.yml Updates checkout while retaining shallow history.
.github/workflows/docs-ci.yml Runs full-history documentation validation and auditing.
Review details

Suppressed comments (1)

.github/workflows/pages.yml:59

  • This changes the deployed artifact action to v5, but the contribution contract still links to actions/upload-pages-artifact/blob/v4/action.yml (docs/contributing/index.md:115). That leaves the documented hidden-file packaging behavior tied to a different action revision; update the reference and any version-specific wording to v5 so the audit guidance matches the workflow.
        uses: actions/upload-pages-artifact@v5
  • Files reviewed: 22/23 changed files
  • Comments generated: 1
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread scripts/docs/pre_pages.py

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Three unresolved findings remain, including one critical URL-scheme validation gap and two moderate audit-validation issues.

Get a fresh assessment by requesting another Copilot review.

Review details

Suppressed comments (1)

tests/docs/test_workflows.py:89

  • This counts raw substring occurrences in every run string, so a single real invocation plus a comment or echo mentioning python scripts/docs/audit_pre_pages.py is treated as a duplicate. Because this helper enforces the workflow contract, it can reject a workflow that executes the audit exactly once; count parsed command invocations rather than arbitrary text occurrences.
            run = step.get("run")
            if isinstance(run, str):
                occurrences += run.count(command)
  • Files reviewed: 22/23 changed files
  • Comments generated: 2
  • Review effort level: Lite

Comment thread scripts/docs/pre_pages_site.py Outdated
Comment thread scripts/docs/pre_pages.py Outdated

@my-reviewer-agent my-reviewer-agent Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

이슈 1건을 인라인 코멘트로 남겼습니다.

self.stack[-1].children.append(element)
self.semantic_events.append(_SemanticEvent("start", element))
if tag not in _VOID:
self.stack.append(element)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[warning] non-void 요소를 항상 현재 부모 아래에 push하므로, HTML에서 허용하는 종료 태그 생략이 visibility 상속을 왜곡합니다. 정적 재현 입력은 기본 AuthoredContent에 <body><p hidden><p>public</body>를 feed하고 close하는 경우입니다. HTML에서는 두 번째 <p>가 첫 번째 문단을 암시적으로 종료하므로 public은 보이는 별도 문단입니다. 그러나 HTMLParser는 암시적인 handle_endtag 호출을 생성하지 않고, 이 구현도 이를 보완하지 않으므로 두 번째 문단이 첫 번째 문단의 자식으로 저장됩니다. 이후 _resolve가 hidden의 display_hidden을 상속시키고, _collect는 public을 blocks에서 제외하면서 hidden을 public으로 설정합니다. 이는 이 클래스의 inherited HTML visibility 판정 계약을 위반합니다. 새 <p>를 처리하기 전에 열린 문단을 암시적으로 닫는 등 HTML tree-construction 규칙을 적용하고, 위 입력에서 blocks에 ('p', 'public')이 포함되고 hidden이 None인 회귀 테스트를 추가해 주세요. 실행 결과가 아니라 제공된 파서 콜백과 수집 경로에 근거한 정적 분석입니다.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in f8b8ddf. The visibility parser now applies bounded HTML implied-end rules for paragraphs/block starts, lists, definition lists, table rows/cells/sections, and options, with Chromium fixtures including the hidden-p/visible-p case.

@hellices

Copy link
Copy Markdown
Owner Author

The suppressed workflow-test finding is fixed in 8c18889. The partial shell parser was removed. Each docs/Pages workflow permits exactly one literal audit-script reference, and it must be the dedicated exact adjacent step; every other literal mention is rejected, while suffixes such as .backup and .pyc are not treated as the script.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Critical audit gaps remain in Git lineage and browser visibility validation, plus a moderate taxonomy redirect gap.

Get a fresh assessment by requesting another Copilot review.

Review details

Suppressed comments (1)

scripts/docs/pre_pages_site.py:514

  • The redirect audit only iterates catalog.redirects, but the site also generates taxonomy-owned redirects from filter_redirect_targets() for /tags/, /articles/, and every legacy_tag_redirects entry. Those pages are parsed for generic HTML/assets later, but their canonical, refresh, and visible fallback targets are never checked, so a broken legacy redirect can pass this audit. Include the generated taxonomy redirect set in the same exact-target checks (and keep them excluded from search).
    for pages_path, canonical in catalog.redirects.items():
        document = by_path.get(canonical)
        declarations = document.metadata.get("redirect_from", []) if document is not None else []
        if not isinstance(declarations, list) or declarations.count(str(pages_path)) != 1:
            errors.append(f"{canonical}: {pages_path} must occur exactly once in redirect_from")
        canonical_html = site.root / canonical.with_suffix(".html")
        redirect_path = site.root / pages_path.with_suffix(".html")
        if redirect_path in search_locations:
            errors.append(f"{pages_path}: redirect location is present in search")
        redirect_page = read_page(redirect_path, f"{pages_path}: redirect")
        if redirect_page is None:
  • Files reviewed: 25/26 changed files
  • Comments generated: 3
  • Review effort level: Lite

Comment thread scripts/docs/pre_pages.py Outdated
Comment thread scripts/docs/pre_pages_visibility.py
Comment thread scripts/docs/pre_pages_visibility.py

@my-reviewer-agent my-reviewer-agent Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

검토 범위에서 보고할 이슈가 없습니다.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

Moderate findings remain in the HTML, site, and visibility audit parsers and must be fixed before approval.

Review details

Suppressed comments (6)

scripts/docs/pre_pages_html.py:49

  • p is explicitly exempted from the table-context check, but this parser does not implement HTML's foster-parenting algorithm. For markup such as <table hidden><p>text</p></table>, Chromium foster-parents the paragraph outside the table while this tree keeps it under the hidden table, producing different visibility/parent semantics and allowing malformed table content to be audited against the wrong tree. Either model foster parenting here or reject this unsupported context instead of silently exempting it.
    if tag in _TABLE_CONTENT - {"p"}:
        table = _in_scope(stack, {"table"}, {"template", "svg", "math"})
        if table is None or any(name not in _TABLE_CONTENT for name in stack[table + 1:]):
            raise AuditFormatError(f"unsupported ambiguous table context for <{tag}>")

scripts/docs/pre_pages_site.py:70

  • _srcset treats the whole non-whitespace token as one URL unless it ends in a comma. Consequently a valid candidate list such as images/a.png,b.png is checked as the single path images/a.png,b.png; if that combined file exists, the audit can pass while either actual candidate is missing. Split normal comma-separated candidates without reintroducing the special handling for commas inside data: URLs, and add a fixture where only the combined filename exists but the two candidates do not.
        match = re.match(r"([^ \t\n\f\r]+)(.*)", remaining, re.DOTALL)
        assert match is not None
        token, remaining = match.groups()
        if token.endswith(","):
            if token.endswith(",,"):
                raise AuditFormatError("empty srcset candidate")
            token = token[:-1]
        else:
            descriptors, _, remaining = remaining.partition(",")

scripts/docs/pre_pages_site.py:36

  • The site HTML parser also treats descendants of an <iframe> as active page markup because _INERT omits iframe. Consequently fallback <img>, <link>, or <a> elements are added to page.targets and can produce rendered-asset/reachability errors for content the browser does not render in the parent document. Add iframe to this inert set as well.
_INERT = frozenset(("script", "style", "template", "noscript"))

scripts/docs/pre_pages_visibility.py:114

  • AuthoredContent relies on HTMLParser's default raw-text set, which only covers script and style; it does not enter raw/RCDATA mode for the other containers treated as raw elsewhere (plaintext, textarea, xmp, iframe, noembed, and noframes). For example, after <plaintext>...</plaintext> a later <a> is parsed as a real link here even though the browser renders the entire remainder as plaintext, so the audit can record or approve reachability/evidence that is not actually present. Configure the parser for the HTML raw/RCDATA containers (and add a trailing-content regression test) before relying on this visibility result.
    def handle_starttag(self, tag: str, attrs: list[tuple[str, str | None]]) -> None:
        parent = self.stack[-1]
        if not parent.svg or parent.tag == "foreignobject":
            index = implied_end_on_start([element.tag for element in self.stack], tag)
            if index is not None:
                del self.stack[index:]
        attributes = {}
        for name, value in attrs:
            attributes.setdefault(name, value)
        element = _Element(tag, attributes)
        parent = self.stack[-1]
        element.svg = tag == "svg" or parent.svg and parent.tag != "foreignobject"
        self.stack[-1].children.append(element)
        self.semantic_events.append(_SemanticEvent("start", element))
        if tag not in _VOID:
            self.stack.append(element)

scripts/docs/pre_pages_visibility.py:42

  • visibility is an inherited CSS property, so visibility: inherit and visibility: unset must retain a hidden ancestor's computed state. This branch currently treats both values as visible whenever they appear on a descendant, allowing content hidden by an ancestor to count as preserved/reachable; carry self.visibility_hidden through for these values (and add a regression case for both).
        visibility = style.get("visibility")
        return Visibility(
            self.display_hidden or "hidden" in attributes or style.get("display") == "none",
            visibility in {"hidden", "collapse"} if visibility in {"visible", "hidden", "collapse", "initial"} else self.visibility_hidden,

scripts/docs/pre_pages_visibility.py:18

  • _collect() descends into <iframe> fallback content because iframe is missing from _NONCONTENT. That makes fallback headings/text and links count as visible authored blocks/reachability even though they are not part of the parent page when the iframe is supported, so a hidden/deleted page can be incorrectly accepted (or a fallback link can be required). Treat iframe contents as non-content, consistent with the raw/hidden handling in pre_pages_content.py.
_NONCONTENT = frozenset(("head", "script", "style", "template", "noscript", "title"))
  • Files reviewed: 28/29 changed files
  • Comments generated: 0 new
  • Review effort level: Lite

@my-reviewer-agent my-reviewer-agent Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

이슈 2건을 인라인 코멘트로 남겼습니다.

Comment thread scripts/docs/pre_pages_visibility.py Outdated
if isinstance(child, _Element) and child.tag == "summary"
), None) if tag == "details" else None
closed_body = tag == "details" and "open" not in attributes and not (
self._summary_hit(summary) if summary else False

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[warning] 명시적인 summary가 없는 <details>도 브라우저가 기본 disclosure control을 제공하므로 사용자가 본문을 펼칠 수 있습니다. 그러나 이 분기는 summary is None이면 항상 False를 반환하여 closed_body=True로 만듭니다. 정적 추적상 AuthoredContent()에 <body><details><p>본문</p></details></body>를 공급하고 close()하면, _collect()가 p에 blocked=True를 전달하고 본문을 blocks에서 제외하며 self.hidden에 기록합니다. 같은 경로의 링크도 self.links에서 제외됩니다. 이는 펼칠 수 있는 명시적 summary가 있을 때 본문을 수집하는 현재 동작과 달리, 정상적인 기본 컨트롤로 접근 가능한 내용을 잘못 분류합니다. summary가 없는 경우에는 details 자체의 가시성·상호작용 상태를 바탕으로 기본 컨트롤의 접근 가능성을 처리하고, 기본 컨트롤로 펼칠 수 있는 본문과 링크가 수집되는 회귀 테스트를 추가해 주세요.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 592919e. Summaryless visible/interactable details now use the browser default disclosure control and their body/links count as accessible. Hidden, inert, display-none, aria-hidden, and already-open states have dedicated source/site/reachability regressions plus a browser-oracle fixture.

assert audit_step.get("name") == "Audit pre-Pages content preservation"
audit_run = audit_step.get("run")
assert isinstance(audit_run, str), workflow_name
assert audit_run.strip() == AUDIT_COMMAND, workflow_name

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[warning] 필수 audit 단계에 if: ${{ false }}를 추가하는 변이를 이 검사가 허용합니다. 정적 추적상 workflow_jobs()는 해당 단계의 dict를 그대로 보존하지만, 이 helper는 검색 단계 다음에 있는 audit의 name과 run만 확인합니다. count_run_blocks_with_audit_reference() 역시 조건과 무관하게 run 문자열을 세므로, 나머지 계약을 만족하는 workflow에서는 모든 assertion이 그대로 통과합니다. 반면 GitHub Actions는 이 조건을 평가하여 audit를 실행하지 않으므로, 테스트가 보장하려는 검색 검증 후 audit 실행 계약이 깨집니다. 이는 실행 재현 결과가 아니라 제공된 코드에 근거한 정적 분석입니다. 필수 audit 단계에 실행을 건너뛰는 조건을 허용하지 않도록 검사하고(조건 없는 단계가 계약이라면 assert "if" not in audit_step), 정상 fixture에 이 조건을 추가했을 때 거부되는 회귀 테스트를 추가하세요.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 592919e. Workflow contracts now reject if on the validate/build jobs or dedicated audit steps and reject continue-on-error: true, while preserving exact step, adjacency, and literal-reference checks.

@hellices

Copy link
Copy Markdown
Owner Author

Fourth Copilot review addressed in f80b66a:

  • direct table/section/row flow content requiring foster parenting now fails closed; valid cell/caption content remains accepted;
  • iframe fallback and plaintext/textarea/xmp/noembed/noframes content cannot provide parent-page authored/link/asset evidence;
  • visibility: inherit and unset retain the ancestor computed state;
  • iframe descendants are consistently non-content for authored and site reachability checks.

The suggested no-whitespace srcset split was not applied after Chromium verification. For srcset="https://srcset-test.invalid/a.png,https://srcset-test.invalid/b.png", Chromium requested the entire comma-containing URL and currentSrc was that same combined URL. A permanent browser-oracle fixture now protects this behavior; normal comma+ASCII-whitespace candidates and data URLs remain tested.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

The broad workflow, documentation, and preservation-audit changes require final human review.

Review details
  • Files reviewed: 29/30 changed files
  • Comments generated: 0 new
  • Review effort level: Lite

@my-reviewer-agent my-reviewer-agent Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

이슈 2건을 인라인 코멘트로 남겼습니다.

Comment thread scripts/docs/pre_pages_visibility.py Outdated
if isinstance(child, _Element) and child.tag == "summary"
), None) if tag == "details" else None
closed_body = tag == "details" and "open" not in attributes and not (
self._summary_hit(summary) if summary else False

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[warning] 명시적인 summary가 없는 닫힌 details도 브라우저가 기본 summary 컨트롤을 제공하므로 사용자가 본문을 펼칠 수 있습니다. 그러나 이 분기는 summary가 없으면 무조건 False를 사용합니다. 정적 추적상 <body><details><p>본문</p></details></body>를 수집하면 closed_body가 True가 되고, 모든 자식에 child_blocked가 전달되어 본문이 blocks에서 제외되고 self.hidden에 기록됩니다. 이는 펼칠 수 있는 details 본문을 수집하는 이 코드의 동작과 어긋납니다. summary가 없는 경우에는 details의 Visibility.interactive를 기준으로 기본 컨트롤의 접근 가능성을 판정하고, 명시적인 summary가 있는 경우에는 기존 _summary_hit 검사를 유지하세요. 기본 summary로 펼칠 수 있는 경우와 hidden 또는 inert로 차단된 경우를 구분하는 회귀 테스트도 추가하세요.

assert audit_step.get("name") == "Audit pre-Pages content preservation"
audit_run = audit_step.get("run")
assert isinstance(audit_run, str), workflow_name
assert audit_run.strip() == AUDIT_COMMAND, workflow_name

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[warning] 전용 audit step에 if: ${{ false }}를 추가하는 mutation을 이 검증이 허용합니다. 실행 결과가 아닌 정적 추적 근거입니다. yaml.safe_load와 workflow_jobs는 해당 step의 if를 보존하지만, 이 helper는 step의 위치, name, run 및 audit 참조 블록 수만 확인합니다. 따라서 기존에 통과하는 입력에서 audit step에 이 조건만 추가하면 모든 assertion의 결과가 그대로 유지됩니다. 반면 GitHub Actions는 그 step을 건너뛰므로 python scripts/docs/audit_pre_pages.py가 실행되지 않아, search validation 다음에 audit를 실행한다는 테스트 계약을 위반합니다. 전용 audit step의 실행을 차단하는 조건도 거부하도록 검증하고, 두 대상 workflow 각각에 if: ${{ false }}를 추가했을 때 AssertionError가 발생하는 회귀 테스트를 추가하세요.

@my-reviewer-agent my-reviewer-agent Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

검토 범위에서 보고할 이슈가 없습니다.

Copilot stopped reviewing on behalf of hellices due to an error September 14, 2026 05:55
hellices and others added 5 commits September 14, 2026 14:56
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Compare normalized source structures against the fixed historical baseline, with exact reviewed fingerprints rather than similarity approvals. Preserve the omitted Agent Memory reader/audience map as linked historical sample evidence.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
hellices and others added 21 commits September 14, 2026 14:56
Bound each reviewed baseline loss to an exact missing count and verified current structure or file evidence. Migrate all 235 approvals, preserve quoted code semantics, reject generic review reasons, and regress the Nginx and duplicate-link deletion loopholes.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Bind replacement-link approvals to explicit current destinations and counts, allowing only verified redundant self-link removals. Restrict all evidence to public HEAD-tracked nonignored sources and require concrete reasons naming exact evidence paths. Recheck all 235 approvals and retain the historical reader-map restoration.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Tokenize HTML tags, attributes, and raw containers before Markdown scans so hidden syntax cannot forge current evidence. Read genuine element attributes and rendered anchor text, and honor image escape parity across inline, reference, and nested links. Regress both real bypasses while preserving all 235 approvals unchanged.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Use the existing Python-Markdown renderer and an HTMLParser collector as the sole link/image evidence authority. Replace source code-span regexes with exact delimiter-run scanning and preserve literal HTML boundaries. Re-audit all 62 documents and retire only the renderer-confirmed false-positive Dockerfile link approval.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Enable attribute lists and the safe repository Markdown extensions through one immutable configuration used by production and test oracles. Fingerprint final href/src/alt overrides, document generated-TOC and snippet exclusions, and regress the real S1 override without changing existing approvals.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Audit the fixed pre-Pages corpus through canonical HTML, redirects, search, service/Explore links and rendered assets. Fail closed on malformed output or escaping paths, and preserve atomic JSON diagnostics.

Restore Markdown rendering in three MySQL details blocks without changing historical content or approval evidence.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Enforce deployment URL origins and prefixes before filesystem resolution, reject local Windows syntax, and strictly parse refresh quotes and ASCII srcset candidates.

Exclude inherited inert anchors from reachability evidence and reject generated HTML/search aliases of published Markdown downloads without weakening byte checks.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Reject repeated or encoded slash segments before URL normalization and unsupported authority spellings before external-origin classification.

Derive published Markdown page/search aliases from installed MkDocs File classification and destinations, including README mappings and exact extension case behavior.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Reject three-plus raw or percent-encoded leading slashes before URL parsing and preserve complete search locations through shared resolution.

Keep homepage search semantics and valid root/protocol-relative references, with WHATWG browser-origin and parser-order regression coverage.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Parse original Markdown before normalizing extracted values. Reject persistently hidden authored source and require visible authored structure in the canonical Material article. Validate fetched scripts, stylesheets and active preloads through the existing strict URL resolver.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Tokenize the supported inline CSS subset and reject unresolved visibility syntax. Finalize disclosure openability from complete summary descendants before collecting either authored content or reachability links, and retain visible authored SVG text without counting decorative icons.

Add Chromium-derived regressions and preserve independent frontend asset checks and strict URL resolution.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Keep the 62-document preservation baseline fixed while auditing every current canonical document, declared redirect, and local asset. Expose dynamic current-document counts in results, JSON, CLI output, and contributor guidance.

Project semantic link and image evidence through finalized shared visibility so excluded SVG foreignObject content cannot satisfy approvals, while retaining raw/code and nested-anchor rendering guarantees.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Require reviewed dispositions to match deleted baseline paths and align the contributor artifact reference with the workflow's v5 action.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Reject executable and file URL schemes before external classification. Require disposition replacements to be HEAD-tracked, unignored regular source files under explicit roots without symlink components.

Count actual shell audit invocations rather than comments or output strings. Keep Material search sharing functional with a safe trusted-template initializer instead of weakening URL validation.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Require every historical document to match its exact initial Git rename at the configured threshold, with strict NUL-record validation.

Share bounded implied-end handling across source and built HTML readers. Require an authored summary for closed disclosure evidence while retaining Chromium-verified late first-summary behavior and asset checks.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Fail closed on unsupported table flow and text, and keep iframe fallback and raw/RCDATA markup out of parent-page evidence and asset discovery. Preserve parsing after proper closes and plaintext-to-EOF behavior.

Record Chromium request/currentSrc evidence for comma-containing srcset URLs and retain the existing browser-correct tokenization and visibility inheritance semantics.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Use the browser's default legend for visible, interactable summaryless details while preserving inherited blockers and open-content behavior. Replace the superseded conservative policy with native Chromium oracle coverage.

Require unconditional, failure-gating audit steps and validate/build jobs without weakening exact placement or literal-reference checks.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note

Copilot was unable to run its full agentic suite in this review.

Pull request overview

Copilot reviewed 30 out of 31 changed files in this pull request and generated 6 comments.

Comment on lines +56 to +58
이관 전 문서별 상세 읽기 목표와 대상 독자는
[원문 독자 지도](https://github.com/hellices/devguidesample/blob/main/docs/services/microsoft-foundry/agent-memory/samples/research-artifacts/README.md#historical-reader-map-pre-pages)에
역사적 기록으로 보존한다.
Comment thread scripts/docs/hooks.py
Comment on lines +386 to +391
source, _, _ = env.loader.get_source(env, template)
safe_source = re.sub(
r'<a\b[^>]*data-md-component="search-share"[^>]*>',
lambda match: match[0].replace('href="javascript:void(0)"', 'href="#"'),
source,
)
Comment thread scripts/docs/pre_pages.py
Comment on lines +574 to +578
if canonical in used:
raise AuditFormatError(f"{canonical}: current public document used more than once")
if document.baseline_path in resolved:
raise AuditFormatError(f"duplicate baseline_path: {document.baseline_path}")
resolved[document.baseline_path] = by_path[canonical]
return any(digit != "0" for digit in width[1])
if re.fullmatch(r"-?(?:[0-9]+(?:\.[0-9]+)?|\.[0-9]+)(?:[eE][+-]?[0-9]+)?x", value):
density = float(value[:-1])
return math.isfinite(density) and density >= 0
Comment on lines +111 to +114
attributes = {}
for name, value in attrs:
attributes.setdefault(name, value)
element = _Element(tag, attributes)
target = tmp_path / relative
target.parent.mkdir(parents=True, exist_ok=True)
shutil.copyfile(root / relative, target)
source_map = next((Path(material.__file__).parent / "templates/assets/javascripts").glob("bundle*.js.map"))
@hellices
hellices force-pushed the docs/pre-pages-preservation-audit branch from 592919e to 1a2f09b Compare September 14, 2026 06:05
@hellices

Copy link
Copy Markdown
Owner Author

Rebased onto latest main (dffa720, sidebar toggle PR #68). The latest tree passed 1,881 documentation tests, 16 JavaScript tests, 440 SRE sample tests, strict build and all validators. Preservation audit remains 62/62 baseline documents and now checks all 66 current documents plus the new sidebar JavaScript asset.

@hellices

Copy link
Copy Markdown
Owner Author

최종 검증 및 리뷰

  • 최신 main (dffa720, sidebar toggle PR 데스크톱 좌측 내비게이션 햄버거 토글 추가 #68) 위로 rebase 완료
  • reviewer approval: 보고할 이슈 없음
  • GitHub Copilot의 모든 인라인 지적 반영; 마지막 agentic run은 20분 제한으로 종료됐고 새 댓글 없음
  • Documentation CI / Oryx: 성공, Node.js 20 annotation 없음
  • documentation tests: 1,881 passed
  • JavaScript tests: 16 passed
  • SRE sample tests: 440 passed
  • metadata/source/link: 66 documents
  • public safety: 365 files
  • strict build/search: 성공
  • preservation audit: 359 baseline files, 73 Markdown, 62/62 baseline documents preserved; 66 current documents visible/searchable

감사에서 실제로 확인된 누락은 Agent Memory historical reader map 1건이며 복원했습니다. Azure MySQL details 렌더링 오류 3개 wrapper도 수정했습니다.

@hellices
hellices merged commit 8ea4207 into main Sep 14, 2026
2 of 3 checks passed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note

Copilot was unable to run its full agentic suite in this review.

Pull request overview

Copilot reviewed 30 out of 31 changed files in this pull request and generated 1 comment.

Comment on lines +56 to +58
이관 전 문서별 상세 읽기 목표와 대상 독자는
[원문 독자 지도](https://github.com/hellices/devguidesample/blob/main/docs/services/microsoft-foundry/agent-memory/samples/research-artifacts/README.md#historical-reader-map-pre-pages)에
역사적 기록으로 보존한다.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants