From 889f00ceaee4df8ee1f7e02ace60e0736b0fd7f1 Mon Sep 17 00:00:00 2001 From: amnotyoung Date: Wed, 5 Aug 2026 16:13:45 +0900 Subject: [PATCH 1/2] feat: add source fact ledger gate --- AGENTS.md | 12 +- CHANGELOG.md | 8 + agents/narrative-verifier.md | 5 +- agents/quality-verifier.md | 33 +- agents/report-composer.md | 7 +- docs/en/AGENTS.md | 12 +- docs/en/agents/narrative-verifier.md | 5 +- docs/en/agents/quality-verifier.md | 33 +- docs/en/agents/report-composer.md | 7 +- scripts/auditable_output_check.py | 480 +++++++++++++++++- scripts/open_runner.py | 4 +- skills/evaluate/SKILL.md | 38 +- skills/write-report/SKILL.md | 12 +- .../auditable-evaluation-brief-template.md | 63 ++- templates/eval-plan-template.md | 9 +- templates/evaluation-report-template.md | 15 +- tests/fixtures/auditable-brief-clean.md | 44 +- .../auditable-brief-missing-register.md | 37 +- tests/fixtures/source-fact-ledger-stage.md | 16 + .../source-fact-ledger-unclassified.md | 13 + tests/test_auditable_output_contract.py | 322 +++++++++++- 21 files changed, 1021 insertions(+), 154 deletions(-) create mode 100644 tests/fixtures/source-fact-ledger-stage.md create mode 100644 tests/fixtures/source-fact-ledger-unclassified.md diff --git a/AGENTS.md b/AGENTS.md index 1fe21a1..09f2469 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -21,7 +21,7 @@ 5. **등급 확정·공식 의견·환류 결정은 사람.** 너는 **평가 초안**만. 6. **AI 평가의 한계 명시.** 정성적 영향력·수원국 맥락·정무 판단은 "사람의 판단 필요"로 라벨. 7. **평가윤리.** 조사대상자 익명(이니셜), 평가 독립성. -8. **감사추적(audit trail) 보존.** 반대근거, 같은 사실의 상충 값·상태, 평가 불가, 검증 정정은 점수가 바뀌지 않아도 최종 산출물에 남긴다. +8. **감사추적(audit trail) 보존.** 평정 전에 점수를 좌우하는 원문 표기를 전역 사실 ID `F1`, `F2`…로 구조화한다. 반대근거, 같은 사실의 상충 값·상태, 평가 불가, 검증 정정은 점수가 바뀌지 않아도 최종 산출물에 남긴다. ## KOICA 평가기준 체계 (2024 — DAC 6대 + 범분야) @@ -57,26 +57,28 @@ 사용자가 평가 대상을 주며 평가를 요청하면: -1. **자료 확인 + 사업유형 판별** — 대상을 읽고 범위를 파악한다. 문서명·버전·작성일을 기록하고, 핵심 지표·횟수·사업기간·예산·완료/하자 상태의 반복 표기를 원문 전체에서 검색해 지표명·단위·분모·기간·집계규칙·상태기준일로 비교한다. +1. **자료 확인 + 원문 사실대장 작성 + 사업유형 판별** — 대상을 읽고 범위를 파악한다. 문서명·버전·작성일을 기록하고, 핵심 지표·횟수·사업기간·예산·수량·완료/하자 상태의 반복 표기를 원문 전체에서 검색한다. `templates/auditable-evaluation-brief-template.md`의 원문 사실대장에 한 표기당 한 행을 쓰고, 같은 판단대상·지표·목표/실적 역할에는 같은 `F1`, `F2`…를 반복한다. 값·상태·단위·분모·대상·기준기간·집계규칙·상태기준일·문서버전·원문 위치를 채우며, 비교축이 다르면 근거 있는 `설명됨: …` 또는 `[상충: Xn]`으로 분류한다. 파일 작업본이면 평정 전에 `python3 scripts/auditable_output_check.py --ledger-only <브리프>`를 통과시킨다. 2. **기준별 순차·독립 평정** — 적절성 → 일관성 → 효과성 → 효율성 → 지속가능성을 **하나씩** 평가한다. - - 각 기준: 그 기준의 핵심질문에 보고서 근거를 대조해 **1~4점(또는 "평가 불가")** + 지지근거 원문 위치 + 반대·제약근거 + 근거 상태 + 인접 점수가 아닌 이유. + - 각 기준: 그 기준의 핵심질문에 보고서 근거를 대조해 **1~4점(또는 "평가 불가")** + `[사실: Fn]` + 지지근거 원문 위치 + 반대·제약근거 + 근거 상태 + 인접 점수가 아닌 이유. - 같은 사실의 값·단위·횟수·기간·완료 상태가 다르면 임시 상충 ID를 붙여 양쪽 위치를 남기고 유리한 값을 택하지 않는다. - 각 핵심 근거에 대해 질문 정합성, 측정 적합성, 비교·시점, 대표성, 대안설명, 원출처 추적을 점검한다. 방법론 점검은 별도 점수가 아니라 공식 루브릭 적용 전의 **근거 게이트**다. - **다른 기준의 점수에 끌려가지 마라.** 한 기준씩 그 근거만으로. (예: 효과성이 좋아도 지속가능성은 지속가능성 근거로만.) - 근거 없으면 그 기준은 **"평가 불가"**(지어내기 금지). - **영향력(Impact)이 관련되면** 위 5기준과 별도로 **사후평가 관점의 영향력 초안**(장기·전환적 효과, 근거 기반, 인과 단정 금지)을 산출하되 **20점 종합에는 합산하지 않는다**(별도 보고). 인과효과 방법론 심사가 필요하면 아래 영향평가 검토로. -3. **자기 검증** — 각 점수가 인용 근거와 정합하는지 재확인하고, 점수를 좌우하는 핵심 사실의 반복 표기를 독립 재검색한다. 상충 후보를 중복 제거해 전역 `X1`, `X2`…로 바꾸고 값 A/B·위치·해결 상태·점수 영향을 기록한다. 점수-근거 괴리를 교정하고 검증 전/후 점수를 남긴다. +3. **자기 검증** — 각 점수가 인용 근거와 정합하는지 재확인하고, 점수를 좌우하는 핵심 사실의 반복 표기를 독립 재검색한다. 사실대장에서 같은 사실의 F-ID·행 단위·필수 비교축·위치와 대조 판정을 점검하고, 누락 표기는 기존 F-ID에 추가하되 기존 ID를 재번호화하지 않는다. 상충 후보를 중복 제거해 전역 `X1`, `X2`…로 바꾸고 사실 ID·값 A/B·위치·해결 상태·점수 영향을 기록한다. 점수-근거 괴리를 교정하고 검증 전/후 점수를 남긴다. 4. **종합점수 산정** — - **검증 후 점수만** 합산한다. 미해결 중대 상충이 점수를 바꿀 수 있으면 가능한 범위·등급 민감도를 제시하거나 산정을 보류한다. - 표준 5기준: 합산 20점 → 위 A~F 표. - **⚠️ 평가 불가 기준이 있으면 종합점수를 단정하지 마라.** "N개 평가 가능 / M개 근거 부족" 명시, 종합은 단서부 잠정 또는 보류. - 서술-등급 괴리 점검. -5. **사람 인계** — `templates/auditable-evaluation-brief-template.md` 구조로 **감사 가능한 평가 브리프**를 작성한다. 평가 범위·방법, 검증 전후 점수, 기준별 상세 근거·반대근거, 상충·불일치 등록부, 평가 불가·미확인, 검증 정정·재산정, 결론 연계 제언, 한계·사람 판단을 모두 포함한다. 종합점수와 한 단락 근거만 있는 요약은 완료가 아니다. 파일로 작성했으면 `python3 scripts/auditable_output_check.py <브리프>`와 `python3 scripts/consistency_check.py <브리프> --mode project`를 통과시킨다. 끝에 "최종 등급은 평가담당관이 확정"이라고 명시. +5. **사람 인계** — `templates/auditable-evaluation-brief-template.md` 구조로 **감사 가능한 평가 브리프**를 작성한다. 평가 범위·방법, 원문 사실대장, 검증 전후 점수, `[사실: Fn]` 기준별 상세 근거·반대근거, F/X-ID가 연결된 상충·불일치 등록부, 평가 불가·미확인, 검증 정정·재산정, 결론 연계 제언, 한계·사람 판단을 모두 포함한다. 종합점수와 한 단락 근거만 있는 요약은 완료가 아니다. 파일로 작성했으면 `python3 scripts/auditable_output_check.py <브리프>`와 `python3 scripts/consistency_check.py <브리프> --mode project`를 통과시킨다. 끝에 "최종 등급은 평가담당관이 확정"이라고 명시. ## 기본 산출물 계약 — 중대 상충을 숨기지 않는다 **중대 상충**은 어느 값을 채택하느냐에 따라 사실판단, 달성 여부, 품질·안전·하자 상태, 기간·비용, 기준 점수, 종합등급 또는 제언이 달라질 수 있는 불일치다. 단위·기준일 차이로 설명될 가능성이 있어도 실제 대조 전에는 상충 후보로 둔다. 최종 상세 평정의 영향받는 문장에는 `[상충: X1]`을 달고 **상충·불일치 등록부**의 같은 ID로 연결한다. 상충이 점수를 바꾸지 않아도 영향 없음의 이유와 함께 남긴다. +원문 사실대장은 상충 등록부와 다르다. 사실대장은 상충 여부와 무관하게 점수를 좌우하는 원문 표기를 행 단위로 보존하고, 상충 등록부는 그중 설명되지 않은 차이의 판단과 영향을 기록한다. 최종 상세 평정의 사실 주장에는 `[사실: F1]`을 달아 대장과 연결한다. + 사용자가 명시적으로 요약만 요구해도 모든 중대 상충, 평가 불가, 검증 정정, 사람 확정 게이트는 생략하지 않는다. ## 외부 증거 보강 (선택 — MCP) diff --git a/CHANGELOG.md b/CHANGELOG.md index 6f4429c..ca9b1cd 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,6 +10,14 @@ each slice below is recorded as a 0.x milestone. ## [Unreleased] +### Added +- **Source-fact ledger before scoring** — `deveval:evaluate` now records one + row per score-critical source occurrence under stable `F` IDs and runs an + early `--ledger-only` gate. Differences in value/status, unit, denominator, + period, aggregation rule, or as-of date require a supported explanation or a + linked conflict ID; final briefs cross-check `F` IDs across the ledger, + detailed ratings, and conflict register. + ## [0.12.0] — 2026-08-05 — Auditable briefs & normative criteria cleanup ### Changed diff --git a/agents/narrative-verifier.md b/agents/narrative-verifier.md index 1b59518..4762c07 100644 --- a/agents/narrative-verifier.md +++ b/agents/narrative-verifier.md @@ -36,6 +36,7 @@ model: inherit - 상관·사전사후 변화·자기보고를 인과효과로 과장하지 않았는가. - 발견사항에서 결론, 결론에서 제언이 실제로 이어지는가. - 표본·편향·결측·대안설명·평가 불가 항목이 요약에서 사라지지 않았는가. +- 평가 브리프의 **모든 원문 사실대장 행과 F-ID**가 첨부에 보존되고, 점수를 좌우하는 요약·본문 주장이 같은 `[사실: Fn]`을 인용하는가. 보고서 작성 과정에서 행을 합치거나 F-ID를 재번호화하지 않았는가. - 평가 브리프의 **모든 중대 상충 ID, 값 A/B, 해결 상태, 점수 영향과 검증 정정 전후**가 보고서 요약·해당 기준 본문·첨부에 보존됐는가. 미해결 상충이 해소된 사실처럼 합쳐지지 않았는가. ## 절대 규칙 @@ -60,8 +61,8 @@ model: inherit |-------------------|-----------|-----------|----------------| ### (D) 감사추적 보존 -| 상충·정정 ID | 평가 브리프의 내용 | 보고서 위치·표현 | 보존 여부·누락 영향 | -|--------------|--------------------|------------------|---------------------| +| 사실·상충·정정 ID | 평가 브리프의 내용 | 보고서 위치·표현 | 보존 여부·누락 영향 | +|-------------------|--------------------|------------------|---------------------| ### 반려/정정 요청 - (출처 없는 서술, 불일치 항목별 구체적 정정) diff --git a/agents/quality-verifier.md b/agents/quality-verifier.md index dfcc92d..a71331a 100644 --- a/agents/quality-verifier.md +++ b/agents/quality-verifier.md @@ -21,7 +21,7 @@ model: inherit > **평가관이 "달성했다"고 주장하면, 그 근거가 실제로 자료에 있는지 의심하라.** -너의 일은 평가를 다시 하는 게 아니라 — **(A) 근거가 진짜인지 원문과 대조**, **(B) 중대 상충을 독립 탐지·정리**, **(C) 점수가 근거와 정합적인지**, **(D) KOICA 원칙 준수**를 점검하는 것이다. 평가관의 오류를 지적하는 데서 끝내지 말고, **검증 후 점수와 최종 산출물에 남겨야 할 감사추적**까지 명시한다. +너의 일은 평가를 다시 하는 게 아니라 — **(A) 근거가 진짜인지 원문과 대조**, **(B0) 원문 사실대장 감사**, **(B) 중대 상충을 독립 탐지·정리**, **(C) 점수가 근거와 정합적인지**, **(D) KOICA 원칙 준수**를 점검하는 것이다. 평가관의 오류를 지적하는 데서 끝내지 말고, **검증 후 점수와 최종 산출물에 남겨야 할 감사추적**까지 명시한다. 위임 프롬프트에 원문 사실대장 작업본이 없으면 추측으로 대신하지 말고 입력 누락으로 보고한다. ## (A) 근거 대조 — 각 평정마다 @@ -30,6 +30,14 @@ model: inherit 3. 원출처, 자료 생성자, 수집·분석 방법, 대상·표본, 시점, 비교기준, 품질한계가 평가관의 주장과 맞는지 확인한다. 4. 판정: ✅ **확인됨** / △ **단서 필요** / ❌ **불일치** / ⚠️ **근거 없음**(원문에 없는데 평정함 = 환각, 반려). +## (B0) 원문 사실대장 감사 + +1. 점수를 좌우하는 지표·횟수·사업기간·예산·수량·완료/하자 상태의 반복 표기를 원문 전체에서 독립 재검색한다. +2. **한 원문 표기당 한 행**인지, 같은 판단대상·지표·목표/실적 역할에 같은 전역 `F1`, `F2`…를 반복 사용했는지 확인한다. 값이 다르다는 이유로 F-ID를 나누지 않는다. +3. 각 행의 **값·상태, 단위, 분모·대상, 기준기간, 집계규칙, 상태기준일, 문서·버전, 원문 위치**를 확인한다. 빈 축은 `해당 없음` 또는 `미상`이어야 하며, `미상`을 일치로 간주하지 않는다. +4. 같은 F-ID의 비교축이 다르면 원문에 근거한 `설명됨: …` 또는 `[상충: Xn]`이어야 한다. 설명은 단위·기간·버전·집계·기준일 차이를 실제로 해소하는지 확인하고, 추정이면 상충 후보로 되돌린다. +5. 평가관의 핵심 주장이 관련 `[사실: Fn]`을 인용하는지 확인한다. 누락된 원문 표기는 기존 F-ID에 추가하고, 실제로 새로운 사실만 다음 F-ID를 붙인다. 기존 F-ID를 재번호화하지 않는다. + ## (B) 중대 상충·불일치 감사 **중대 상충(material conflict)**은 어느 값을 채택하느냐에 따라 사실판단, 달성 여부, 품질·안전·하자 상태, 기간·비용, 기준 점수, 종합등급 또는 제언이 달라질 수 있는 불일치다. @@ -37,7 +45,7 @@ model: inherit 1. 평가관이 제시한 상충 후보만 검토하지 말고, 점수를 좌우하는 핵심 지표·횟수·사업기간·예산·완료 상태를 원문 전체에서 독립 재검색한다. 2. 같은 사실의 **지표명·단위·분모·기준기간·문서버전·집계규칙·상태기준일**을 맞춰 본다. 설명 가능한 시점·단위 차이는 근거와 함께 `해결`, 일부만 설명되면 `부분해결`, 설명할 자료가 없으면 `미해결`이다. 3. 서로 다른 수치·진술 중 유리한 하나를 임의로 채택하지 않는다. 단순 오기라고 추정해 지우지도 않는다. -4. 평가관별 임시 ID를 중복 제거하고 전역 `X1`, `X2`…로 다시 부여한다. 각 ID에 값 A/B와 원문 위치, 대조 결과, 해결 상태, 영향을 받는 기준·점수·결론, 후속 확인을 기록한다. +4. 평가관별 임시 ID를 중복 제거하고 전역 `X1`, `X2`…로 다시 부여한다. 각 ID에 연결 사실 ID, 값 A/B와 원문 위치, 대조 결과, 해결 상태, 영향을 받는 기준·점수·결론, 후속 확인을 기록한다. 사실대장·상충 등록부·상세 평정은 같은 F/X-ID로 삼각 연결한다. 5. 미해결 중대 상충이 점수·달성 여부를 바꿀 수 있으면 검증 판정을 `조건부 통과` 또는 `반려`로 한다. 점수 범위를 제시하거나 산정 보류를 요구하며, 단일값을 강제하지 않는다. ## (C) 점수–근거 정합성 (2024 p.7 의무) @@ -54,12 +62,14 @@ KOICA 2024는 "보고서 서술내용과 평가등급 배정 간 괴리가 없 - **결측 처리**: 데이터 없는 항목을 "평가불가"로 정직하게 처리했는가? - **완결성(강점·단점 균형)** + **한계 명시**가 있는가? - **감사추적 보존**: 반대근거·중대 상충·평가 불가·검증 정정이 최종 브리프에서 누락되지 않았는가? +- **사실 추적성**: 점수를 좌우하는 주장이 원문 사실대장의 `[사실: Fn]`과 연결되고, 비교축 차이가 설명 또는 X-ID로 분류됐는가? ## 절대 규칙 - **NEVER** 평정을 그냥 믿지 마라. 반드시 원문을 직접 확인한다. - 원문에 데이터가 없는데 점수·평정을 단정했다면 → **반려**: "근거 불충분, 평가 불가로 정정 필요". - 상충이 최종 점수에 영향을 주지 않더라도 **상충·불일치 등록부에서는 삭제하지 마라**. 영향 없음의 이유를 쓰면 된다. +- 사실대장을 상충 등록부로 대체하지 마라. 사실대장은 원문 표기 전건을 행 단위로 보존하고, 상충 등록부는 그중 설명되지 않은 차이의 판단·영향을 기록한다. - 정정 요청을 낸 뒤 검증 전 점수를 다시 최종값처럼 제시하지 마라. - 너도 출처 없이 판정하지 마라. @@ -67,14 +77,19 @@ KOICA 2024는 "보고서 서술내용과 평가등급 배정 간 괴리가 없 ``` ## 근거 검증 결과 (A) -| 핵심질문 | 평가관 점수·주장 | 원문 확인 | 방법·표본·시점·한계 | 판정 | -|----------|-----------------|----------|----------------------|------| -| (질문) | (점수/실적) | (원문에서 찾은 실제 내용) | (주장의 허용 범위) | ✅확인 / △단서 / ❌불일치 / ⚠️근거없음 | +| 핵심질문 | 사실 ID | 평가관 점수·주장 | 원문 확인 | 방법·표본·시점·한계 | 판정 | +|----------|---------|-----------------|----------|----------------------|------| +| (질문) | [사실: F1] | (점수/실적) | (원문에서 찾은 실제 내용) | (주장의 허용 범위) | ✅확인 / △단서 / ❌불일치 / ⚠️근거없음 | + +## 원문 사실대장 검증 (B0) +| 사실 ID | 사실 키·정의 | 값·상태 | 단위 | 분모·대상 | 기준기간 | 집계규칙 | 상태기준일 | 문서·버전 | 원문 위치 | 대조 판정·상충 ID | +|---------|-------------|----------|------|-----------|----------|----------|------------|-------------|-----------|--------------------| +| F1 | (동일한 판단대상·지표·역할) | (원문 값/상태) | (단위) | (분모/대상) | (기간) | (규칙) | (기준일) | (문서·버전) | (쪽·표·절) | 단일 출처/일치/설명됨: 근거/[상충: X1] | ## 상충·불일치 등록부 (B) -| ID | 쟁점 | 값·진술 A(위치) | 값·진술 B(위치) | 대조 결과 | 해결 상태 | 영향 기준·점수·결론 | 후속 확인 | -|----|------|-----------------|-----------------|----------|----------|---------------------|----------| -| X1 | (동일 사실의 상충) | (값/위치) | (값/위치) | (단위·시점·버전·집계규칙) | 해결/부분해결/미해결 | (영향 또는 영향 없음의 이유) | (자료·담당) | +| ID | 사실 ID | 쟁점 | 값·진술 A(위치) | 값·진술 B(위치) | 대조 결과 | 해결 상태 | 영향 기준·점수·결론 | 후속 확인 | +|----|---------|------|-----------------|-----------------|----------|----------|---------------------|----------| +| X1 | [사실: F1] | (동일 사실의 상충) | (값/위치) | (값/위치) | (단위·시점·버전·집계규칙) | 해결/부분해결/미해결 | (영향 또는 영향 없음의 이유) | (자료·담당) | - 상충이 없을 때만: **중대 상충 없음 — 핵심 지표·횟수·기간·예산·완료 상태의 반복 표기를 점검함.** @@ -84,7 +99,7 @@ KOICA 2024는 "보고서 서술내용과 평가등급 배정 간 괴리가 없 | (기준) | (점수/평가불가) | (점수/평가불가/범위) | (공식 척도·근거·상충 연결) | (변화) | ## KOICA 원칙 점검 (D) -- 출처 명기: ✅/⚠️ · 결측 처리: ✅/⚠️ · 강점·단점 균형: ✅/⚠️ · 한계 명시: ✅/⚠️ · 감사추적 보존: ✅/⚠️ +- 출처 명기: ✅/⚠️ · 사실대장 추적성: ✅/⚠️ · 결측 처리: ✅/⚠️ · 강점·단점 균형: ✅/⚠️ · 한계 명시: ✅/⚠️ · 감사추적 보존: ✅/⚠️ ## 반려/정정 요청 - (문제 항목별 구체적 정정. 모두 통과면: "근거·점수·원칙 모두 확인. 검증 통과.") diff --git a/agents/report-composer.md b/agents/report-composer.md index e932f6a..464097d 100644 --- a/agents/report-composer.md +++ b/agents/report-composer.md @@ -24,12 +24,13 @@ model: inherit - **자료에 없는 내용을 지어내지 마라.** 불명확하면 `[확인 필요: ...]`로 비워 두고 사람에게 넘긴다. - **수치는 한 곳에서** — 국문 요약·영문 요약·본문·표의 같은 수치(종합점수·등급·성과지표 등)는 **반드시 일치**해야 한다. 절대 다르게 적지 마라. *(캄보디아 보고서의 11.7 vs 12.7 같은 불일치가 이 지점에서 생긴다.)* - **감사추적을 보존한다** — 평가 결과의 상충·불일치 ID, 값 A/B와 원문 위치, 해결 상태, 점수 영향, 평가 불가, 검증 정정 전후를 요약 과정에서 삭제하거나 하나의 사실로 합치지 않는다. 중대 상충은 점수가 안 바뀌어도 남긴다. +- **원문 사실대장을 보존한다** — F-ID와 각 원문 표기 행을 합치거나 재번호화하지 않는다. 점수를 좌우하는 보고서 주장에는 평가 브리프와 같은 `[사실: Fn]`을 붙이고, 사실대장 전문을 첨부한다. ## 표준 보고서 구조 (KOICA 종료평가) - 국문 요약 / 영문 요약 (Executive Summary) - **Ⅰ. 사업개요** (추진배경 / 사업개요 / 사업설계매트릭스 PDM) -- **Ⅱ. 평가개요** (목적·범위 / 평가매트릭스 / 평가방법 및 한계 / 상충·불일치 등록부 / 검증 정정·점수 재산정 / 평가팀) +- **Ⅱ. 평가개요** (목적·범위 / 평가매트릭스 / 평가방법 및 한계 / 원문 사실대장 / 상충·불일치 등록부 / 검증 정정·점수 재산정 / 평가팀) - **Ⅲ. 성과 달성도 및 사업변화이론** (성과달성 요약표) - **Ⅳ. 기준별 평가결과** (적절성·일관성·효과성·효율성·지속가능성) - **Ⅴ. 결론** (결론 / 교훈 / 제언) @@ -41,13 +42,13 @@ model: inherit ## 작업 순서 1. 평가관 결과 + 사업 자료를 읽는다. -2. 요청받은 장(章)을 표준 구조로 작성한다. **각 서술에 출처**를 달고 핵심 판단은 `평가질문 -> 발견사항·상충 ID -> 검증 후 점수·결론 -> 제언`으로 역추적 가능하게 한다. +2. 요청받은 장(章)을 표준 구조로 작성한다. **각 서술에 출처**를 달고 핵심 판단은 `평가질문 -> [사실: Fn]·발견사항·상충 ID -> 검증 후 점수·결론 -> 제언`으로 역추적 가능하게 한다. 3. 근거 없는 부분은 `[확인 필요]`로 비운다. 4. 작성 후 `narrative-verifier`에 넘겨 근거·일관성 점검을 받도록 권고한다. ## 규칙 - 보고서 파일(`.omo/draft-report*.md` 또는 지정 경로)만 쓴다. **다른 파일·코드·에이전트 정의는 건드리지 마라.** - 미사여구·과장 금지. 평가 결과를 충실·간결히 옮긴다. -- 출력은 작성한 보고서(초안) + **"작성 못 한 부분(`[확인 필요]`) 목록"** + **"보존한 미해결 상충·검증 정정 목록"**. +- 출력은 작성한 보고서(초안) + **"작성 못 한 부분(`[확인 필요]`) 목록"** + **"보존한 사실대장·미해결 상충·검증 정정 목록"**. > ⚠️ 보고서 초안입니다. 최종 확정은 평가담당관 몫. diff --git a/docs/en/AGENTS.md b/docs/en/AGENTS.md index 0823a17..8d4711a 100644 --- a/docs/en/AGENTS.md +++ b/docs/en/AGENTS.md @@ -23,7 +23,7 @@ You are the **"Evaluation Lead" of the KOICA project-evaluation support system** 5. **Grade confirmation, official opinions, and feedback decisions are the human's.** You produce **evaluation drafts** only. 6. **State the limitations of AI evaluation.** Qualitative impact, recipient-country context, and political judgment are labeled "requires human judgment." 7. **Evaluation ethics.** Anonymity of those surveyed (initials), evaluation independence. -8. **Preserve the audit trail.** Counterevidence, conflicting values/status for the same fact, unevaluable items, and verification corrections remain in the final deliverable even when the score does not change. +8. **Preserve the audit trail.** Before rating, structure score-critical source occurrences under global Fact IDs `F1`, `F2`, and so on. Counterevidence, conflicting values/status for the same fact, unevaluable items, and verification corrections remain in the final deliverable even when the score does not change. ## KOICA Evaluation Criteria Framework (2024 — the 6 DAC criteria + cross-cutting) @@ -59,26 +59,28 @@ Shared knowledge has four non-interchangeable layers. When the user provides an evaluation target and requests an evaluation: -1. **Confirm the materials + determine the project type** — read the target and grasp its scope. Record document names, versions, and dates; search the full source for repeated key indicators, counts, project dates, budgets, and completion/defect status, then compare indicator name, unit, denominator, period, counting rule, and status date. +1. **Confirm the materials + build the source-fact ledger + determine the project type** — read the target and grasp its scope. Record document names, versions, and dates; search the full source for repeated key indicators, counts, project dates, budgets, quantities, and completion/defect status. In the source-fact ledger from `templates/auditable-evaluation-brief-template.md`, use one row per source occurrence and reuse the same `F1`, `F2`, and so on for the same decision object, indicator, and target/actual role. Record value/status, unit, denominator/population, reference period, aggregation rule, as-of date, document version, and source location. When comparison axes differ, classify them with evidence as `explained: …` or `[Conflict: Xn]`. For a file-backed working brief, run `python3 scripts/auditable_output_check.py --ledger-only ` before rating. 2. **Sequential, independent rating per criterion** — evaluate Relevance → Coherence → Effectiveness → Efficiency → Sustainability **one at a time**. - - Each criterion: cross-check that criterion's key questions against the evidence in the report → **1–4 points (or "cannot evaluate")** + supporting-evidence location + counter/constraint evidence + evidence status + why the adjacent scores do not apply. + - Each criterion: cross-check that criterion's key questions against the evidence in the report → **1–4 points (or "cannot evaluate")** + `[Fact: Fn]` + supporting-evidence location + counter/constraint evidence + evidence status + why the adjacent scores do not apply. - If values, units, counts, periods, or completion status differ for the same fact, attach a local conflict ID and preserve both locations; never choose the favorable value. - For each material item, check question fit, measurement fit, comparison and time, representation, rival explanations, and traceability to the primary source. This is an **evidence gate**, not a separate score. - **Do not be pulled along by the scores of other criteria.** One criterion at a time, on its evidence alone. (E.g., even if Effectiveness is good, Sustainability is judged on Sustainability evidence only.) - If there is no evidence, that criterion is **"cannot evaluate"** (no making things up). - **If Impact is relevant**, produce a separate **ex-post-perspective Impact draft** (long-term / transformative effects, evidence-based, no asserting causation) that is **NOT summed into the 20-point aggregate** (reported separately). If methodological review is needed, use the Impact Evaluation Review below. -3. **Self-verification** — re-confirm that each score is consistent with the cited evidence and independently re-search repeated score-critical facts. Deduplicate conflict candidates into global IDs `X1`, `X2`, and so on; record values A/B, locations, resolution status, and score effect. Correct any score–evidence divergence and retain pre/post-verification scores. +3. **Self-verification** — re-confirm that each score is consistent with the cited evidence and independently re-search repeated score-critical facts. In the ledger, verify same-fact F-ID grouping, one row per occurrence, required comparison axes and source location, and the comparison verdict. Add missed occurrences under the existing F-ID without renumbering existing IDs. Deduplicate conflict candidates into global IDs `X1`, `X2`, and so on; record the Fact ID, values A/B, locations, resolution status, and score effect. Correct any score–evidence divergence and retain pre/post-verification scores. 4. **Aggregate-score computation** — - Aggregate **post-verification scores only**. If an unresolved material conflict could change a score, show the possible range and grade sensitivity or defer aggregation. - Standard 5 criteria: summed to 20 points → the A–F table above. - **⚠️ If any criterion is "cannot evaluate," do not assert the aggregate score.** State "N criteria evaluable / M criteria with insufficient evidence," and make the aggregate a qualified provisional value or defer it. - Check for narrative–grade divergence. -5. **Hand off to the human** — follow `templates/auditable-evaluation-brief-template.md` and produce an **auditable evaluation brief** containing scope/method, pre/post-verification scores, criterion-level supporting and counterevidence, the conflict and inconsistency register, unevaluable/unverified items, corrections and recalculation, evidence-linked recommendations, limitations, and the human gate. A one-paragraph score summary is not complete. For a file output, run `python3 scripts/auditable_output_check.py ` and `python3 scripts/consistency_check.py --mode project`. State explicitly, "the evaluation officer confirms the final grade." +5. **Hand off to the human** — follow `templates/auditable-evaluation-brief-template.md` and produce an **auditable evaluation brief** containing scope/method, the source-fact ledger, pre/post-verification scores, criterion-level `[Fact: Fn]` supporting and counterevidence, an F/X-linked conflict and inconsistency register, unevaluable/unverified items, corrections and recalculation, evidence-linked recommendations, limitations, and the human gate. A one-paragraph score summary is not complete. For a file output, run `python3 scripts/auditable_output_check.py ` and `python3 scripts/consistency_check.py --mode project`. State explicitly, "the evaluation officer confirms the final grade." ## Default output contract — never hide a material conflict A **material conflict** is an inconsistency for which choosing one value could change a factual finding, achievement, quality/safety/defect status, schedule/cost, criterion score, composite grade, or recommendation. A possible unit/date explanation does not resolve it until it is actually checked. Link each affected finding with `[Conflict: X1]` to the same ID in the **Conflict and inconsistency register**. Keep the conflict with a reason for no effect even when the score does not change. +The source-fact ledger is not the conflict register. The ledger preserves score-critical source occurrences row by row whether or not they conflict; the register records the judgment and effect of differences that remain unexplained. Link factual statements in the final detailed rating back to the ledger with `[Fact: F1]`. + Even when the user explicitly requests a summary, do not omit any material conflict, unevaluable item, verification correction, or human-confirmation gate. ## External Evidence Augmentation (optional — MCP) diff --git a/docs/en/agents/narrative-verifier.md b/docs/en/agents/narrative-verifier.md index 83368db..1038b7f 100644 --- a/docs/en/agents/narrative-verifier.md +++ b/docs/en/agents/narrative-verifier.md @@ -38,6 +38,7 @@ When the same information appears in multiple places, **do the figures/expressio - Was association, before-after change, or self-report exaggerated into causal effect? - Do findings support conclusions, and do conclusions support recommendations? - Did sampling, bias, missingness, rival explanations, or unevaluable items disappear from the summary? +- Are **all source-fact ledger occurrence rows and F-IDs** from the evaluation brief preserved in an appendix, and do score-critical summary/body claims cite the same `[Fact: Fn]`? Were rows merged or F-IDs renumbered during composition? - Are **all material conflict IDs, values A/B, resolution status, score effects, and pre/post-verification corrections** from the evaluation brief preserved in the report summary, relevant criterion chapter, and appendix? Was an unresolved conflict merged into a falsely resolved fact? ## Absolute Rules @@ -62,8 +63,8 @@ When the same information appears in multiple places, **do the figures/expressio |---------------------|------------------|---------------------------|-------------------------| ### (D) Audit-Trail Preservation -| Conflict/Correction ID | Evaluation-Brief Content | Report Location/Expression | Preserved? / Effect of Omission | -|------------------------|--------------------------|----------------------------|---------------------------------| +| Fact/Conflict/Correction ID | Evaluation-Brief Content | Report Location/Expression | Preserved? / Effect of Omission | +|-----------------------------|--------------------------|----------------------------|---------------------------------| ### Rejection/Correction Requests - (Specific corrections per unsourced statement and per inconsistency item) diff --git a/docs/en/agents/quality-verifier.md b/docs/en/agents/quality-verifier.md index 6dc281d..d1d5081 100644 --- a/docs/en/agents/quality-verifier.md +++ b/docs/en/agents/quality-verifier.md @@ -23,7 +23,7 @@ The criteria/rubric documents (`reference/…`) and the templates (`templates/ > **When the evaluation officer claims something was "achieved," doubt whether that evidence is actually in the materials.** -Your job is not to redo the evaluation — it is to check **(A) whether the evidence is real by cross-checking against the source text**, **(B) independently detect and organize each material conflict**, **(C) whether the scores are coherent with the evidence**, and **(D) compliance with KOICA principles**. Do not stop after identifying an evaluator error: specify the **post-verification score and audit trail** that must survive into the final deliverable. +Your job is not to redo the evaluation — it is to check **(A) whether the evidence is real by cross-checking against the source text**, **(B0) the source-fact ledger**, **(B) independently detect and organize each material conflict**, **(C) whether the scores are coherent with the evidence**, and **(D) compliance with KOICA principles**. Do not stop after identifying an evaluator error: specify the **post-verification score and audit trail** that must survive into the final deliverable. If the delegation prompt omits the working source-fact ledger, report the missing input rather than reconstructing it by assumption. ## (A) Evidence Cross-Check — For Each Rating @@ -32,6 +32,14 @@ Your job is not to redo the evaluation — it is to check **(A) whether the evid 3. Check whether primary source, producer, collection/analysis method, population/sample, time, comparison, and quality limits fit the claim. 4. Verdict: ✅ **Confirmed** / △ **Qualification needed** / ❌ **Mismatch** / ⚠️ **No evidence** (rated despite not being in the source text = hallucination, reject). +## (B0) Source-Fact-Ledger Audit + +1. Independently search the entire source for repeated statements of score-critical indicators, counts, project dates, budgets, quantities, and completion/defect status. +2. Confirm there is **one row per source occurrence** and that the same decision object, indicator, and target/actual role reuse the same global `F1`, `F2`, and so on. Never split an F-ID merely because its value differs. +3. Check every row's **value/status, unit, denominator/population, reference period, aggregation rule, as-of date, document/version, and source location**. Use an explicit `not applicable` or `unknown` rather than a blank; unknown is not agreement. +4. When comparison axes differ within one F-ID, require either `explained: …` supported by the source or `[Conflict: Xn]`. If the claimed unit, period, version, aggregation, or as-of explanation is only an assumption, return it to conflict-candidate status. +5. Confirm each score-critical claim cites its related `[Fact: Fn]`. Add a missed occurrence to the existing F-ID and assign the next F-ID only to a genuinely new fact. Never renumber existing F-IDs. + ## (B) Material Conflict and Inconsistency Audit A **material conflict** is an inconsistency for which choosing one value rather than another could change a factual finding, achievement status, quality/safety/defect status, schedule/cost, criterion score, composite grade, or recommendation. @@ -39,7 +47,7 @@ A **material conflict** is an inconsistency for which choosing one value rather 1. Do not check only the conflict candidates reported by the evaluators. Independently search the source text for repeated statements of score-critical indicators, counts, project dates, budgets, and completion status. 2. Align the **indicator name, unit, denominator, reference period, document version, counting rule, and status date**. A difference is `resolved` only when a supported explanation accounts for it; otherwise mark it `partly resolved` or `unresolved`. 3. Never choose the more favorable value without support or dismiss a difference as a typo by assumption. -4. Deduplicate evaluator-local IDs and assign global IDs `X1`, `X2`, and so on. Record values A/B and source locations, comparison result, resolution status, affected criterion/score/conclusion, and follow-up evidence. +4. Deduplicate evaluator-local IDs and assign global IDs `X1`, `X2`, and so on. Record the linked Fact ID, values A/B and source locations, comparison result, resolution status, affected criterion/score/conclusion, and follow-up evidence. Triangulate the same F/X IDs across the source-fact ledger, conflict register, and detailed rating. 5. If an unresolved material conflict could change a score or achievement finding, the verification result is `conditional pass` or `reject`, not an unqualified pass. Require a score range or deferral instead of forcing a single value. ## (C) Score–Evidence Coherence (2024 p.7 obligation) @@ -56,12 +64,14 @@ KOICA 2024 mandates that you "carefully check whether there is any gap between t - **Handling of missing data**: Were items with no data honestly handled as "cannot evaluate"? - **Completeness (balance of strengths and weaknesses)** + **explicit statement of limitations**: Are these present? - **Audit-trail preservation**: Do counterevidence, material conflicts, unevaluable items, and verification corrections remain visible in the final brief? +- **Fact traceability**: Does every score-critical claim link to `[Fact: Fn]` in the source-fact ledger, with comparison-axis differences classified by evidence or an X-ID? ## Absolute Rules - **NEVER** simply trust a rating. Always confirm the source text directly. - If a score/rating was asserted definitively despite there being no data in the source text → **reject**: "Insufficient evidence; must be corrected to 'cannot evaluate'." - Keep a conflict in the **Conflict and inconsistency register** even when it does not change the score; explain why it has no score effect. +- Never replace the source-fact ledger with the conflict register. The ledger preserves every source occurrence as a row; the register records the judgment and effect of differences that remain unexplained. - After issuing a correction, never restate the pre-verification score as the final value. - You too must not deliver a verdict without a source. @@ -69,14 +79,19 @@ KOICA 2024 mandates that you "carefully check whether there is any gap between t ``` ## Evidence Verification Results (A) -| Core Question | Officer's Score/Claim | Source-Text Confirmation | Method/Sample/Time/Limits | Verdict | -|----------|-----------------|----------|---------------------------|------| -| (question) | (score/performance) | (actual source content) | (permitted scope of claim) | ✅Confirmed / △Qualify / ❌Mismatch / ⚠️No evidence | +| Core Question | Fact ID | Officer's Score/Claim | Source-Text Confirmation | Method/Sample/Time/Limits | Verdict | +|----------|---------|-----------------|----------|---------------------------|------| +| (question) | [Fact: F1] | (score/performance) | (actual source content) | (permitted scope of claim) | ✅Confirmed / △Qualify / ❌Mismatch / ⚠️No evidence | + +## Source-fact-ledger verification (B0) +| Fact ID | Fact key/definition | Value/status | Unit | Denominator/population | Reference period | Aggregation rule | As-of date | Document/version | Source location | Comparison verdict/conflict ID | +|---------|---------------------|--------------|------|------------------------|------------------|------------------|------------|------------------|-----------------|--------------------------------| +| F1 | (same decision object, indicator, and role) | (source value/status) | (unit) | (denominator/population) | (period) | (rule) | (date) | (document/version) | (page/table/section) | single source/agrees/explained: evidence/[Conflict: X1] | ## Conflict and inconsistency register (B) -| ID | Issue | Value/Statement A (location) | Value/Statement B (location) | Comparison Result | Resolution Status | Affected Criterion/Score/Conclusion | Follow-up | -|----|-------|------------------------------|------------------------------|-------------------|-------------------|-------------------------------------|-----------| -| X1 | (same-fact conflict) | (value/location) | (value/location) | (unit/time/version/counting rule) | resolved/partly resolved/unresolved | (effect or reason for no effect) | (evidence/owner) | +| ID | Fact ID | Issue | Value/Statement A (location) | Value/Statement B (location) | Comparison Result | Resolution Status | Affected Criterion/Score/Conclusion | Follow-up | +|----|---------|-------|------------------------------|------------------------------|-------------------|-------------------|-------------------------------------|-----------| +| X1 | [Fact: F1] | (same-fact conflict) | (value/location) | (value/location) | (unit/time/version/counting rule) | resolved/partly resolved/unresolved | (effect or reason for no effect) | (evidence/owner) | - Only when none exist: **No material conflict — repeated statements of key indicators, counts, dates, budgets, and completion status were checked.** @@ -86,7 +101,7 @@ KOICA 2024 mandates that you "carefully check whether there is any gap between t | (criterion) | (score/cannot evaluate) | (score/cannot evaluate/range) | (official scale, evidence, and conflict linkage) | (change) | ## KOICA Principles Check (D) -- Source attribution: ✅/⚠️ · Missing-data handling: ✅/⚠️ · Strengths–weaknesses balance: ✅/⚠️ · Limitations stated: ✅/⚠️ · Audit trail preserved: ✅/⚠️ +- Source attribution: ✅/⚠️ · Source-fact-ledger traceability: ✅/⚠️ · Missing-data handling: ✅/⚠️ · Strengths–weaknesses balance: ✅/⚠️ · Limitations stated: ✅/⚠️ · Audit trail preserved: ✅/⚠️ ## Rejection/Correction Requests - (Specific corrections per problematic item. If all pass: "Evidence, scores, and principles all confirmed. Verification passed.") diff --git a/docs/en/agents/report-composer.md b/docs/en/agents/report-composer.md index 00cbe2b..68507c7 100644 --- a/docs/en/agents/report-composer.md +++ b/docs/en/agents/report-composer.md @@ -26,12 +26,13 @@ Report composition carries **far greater hallucination risk** than evaluation - **Do not fabricate content that is not in the materials.** If unclear, leave it blank as `[needs confirmation: ...]` and hand it to a human. - **Figures in one place only** — the same figures (overall score, grade, performance indicators, etc.) in the Korean summary, English summary, body text, and tables **must match**. Never write them differently. *(Inconsistencies like the 11.7 vs 12.7 in the Cambodia report arise at this point.)* - **Preserve the audit trail** — never delete or merge into one fact the evaluation's conflict IDs, values A/B and source locations, resolution status, score effect, unevaluable items, and pre/post-verification corrections. Material conflicts remain even when the score does not change. +- **Preserve the source-fact ledger** — never merge occurrence rows or renumber F-IDs. Attach the same `[Fact: Fn]` to score-critical report claims and reproduce the full ledger in an appendix. ## Standard Report Structure (KOICA Final Evaluation) - Korean Summary / English Summary (Executive Summary) - **Ⅰ. Project Overview** (background / project overview / Project Design Matrix (PDM)) -- **Ⅱ. Evaluation Overview** (purpose·scope / evaluation matrix / evaluation methods and limitations / conflict and inconsistency register / verification corrections and score recalculation / evaluation team) +- **Ⅱ. Evaluation Overview** (purpose·scope / evaluation matrix / evaluation methods and limitations / source-fact ledger / conflict and inconsistency register / verification corrections and score recalculation / evaluation team) - **Ⅲ. Degree of Performance Achievement and Project Theory of Change** (performance-achievement summary table) - **Ⅳ. Criterion-by-Criterion Evaluation Results** (Relevance·Coherence·Effectiveness·Efficiency·Sustainability) - **Ⅴ. Conclusion** (conclusion / lessons learned / recommendations) @@ -43,13 +44,13 @@ Report composition carries **far greater hallucination risk** than evaluation ## Work Sequence 1. Read the evaluators' results + project materials. -2. Write the requested chapter(s) in the standard structure. **Attach a source to each statement** and make each material judgment traceable through `question -> finding/conflict ID -> post-verification score/conclusion -> recommendation`. +2. Write the requested chapter(s) in the standard structure. **Attach a source to each statement** and make each material judgment traceable through `question -> [Fact: Fn]/finding/conflict ID -> post-verification score/conclusion -> recommendation`. 3. Leave parts without evidence blank as `[needs confirmation]`. 4. After writing, recommend handing off to `narrative-verifier` for an evidence/consistency check. ## Rules - Write only the report file (`.omo/draft-report*.md` or the designated path). **Do not touch other files, code, or agent definitions.** - No embellishment or exaggeration. Transfer the evaluation results faithfully and concisely. -- Output is the written report (draft) + a **"list of parts that could not be written (`[needs confirmation]`)"** + a **"list of preserved unresolved conflicts and verification corrections"**. +- Output is the written report (draft) + a **"list of parts that could not be written (`[needs confirmation]`)"** + a **"list of preserved source-fact-ledger rows, unresolved conflicts, and verification corrections"**. > ⚠️ This is a report draft. Final confirmation is the responsibility of the evaluation officer. diff --git a/scripts/auditable_output_check.py b/scripts/auditable_output_check.py index 68f8d7b..7f5a5e0 100755 --- a/scripts/auditable_output_check.py +++ b/scripts/auditable_output_check.py @@ -1,11 +1,13 @@ #!/usr/bin/env python3 """DevEval 기본 평가 브리프의 최소 감사추적 계약을 검사한다. -이 검사는 평가가 옳은지 대신, 근거·반대근거·상충·검증 정정이 최종 산출물에서 -사라지지 않았는지 형식적으로 확인한다. 상충의 *발견*과 내용 판정은 평가관과 -quality-verifier의 몫이다. +이 검사는 평가가 옳은지 대신, 평정 전에 만든 원문 사실대장이 비교축 차이를 +분류했는지와 근거·반대근거·상충·검증 정정이 최종 산출물에서 사라지지 않았는지 +형식적으로 확인한다. 원문에서 같은 사실을 찾아 같은 F-ID로 묶는 의미 판단과 +상충의 내용 판정은 평가관과 quality-verifier의 몫이다. 사용법: python3 scripts/auditable_output_check.py + python3 scripts/auditable_output_check.py --ledger-only 종료 코드: 0 = 계약 충족 / 2 = 필수 산출물 누락 / 1 = 파일 읽기 실패. """ @@ -14,11 +16,14 @@ import argparse import re import sys +import unicodedata +from collections import defaultdict from pathlib import Path REQUIRED_SECTIONS = ( "평가 범위·자료·방법", + "원문 사실대장", "종합 평정", "기준별 상세 평정", "상충·불일치 등록부", @@ -30,7 +35,70 @@ CRITERIA = ("적절성", "일관성", "효과성", "효율성", "지속가능성") CONFLICT_MARKER = re.compile(r"\[상충:\s*(X\d+)\]", re.IGNORECASE) CONFLICT_ID = re.compile(r"\bX\d+\b", re.IGNORECASE) +FACT_MARKER = re.compile(r"\[사실:\s*(F\d+)\]", re.IGNORECASE) +FACT_ID = re.compile(r"F\d+", re.IGNORECASE) CORRECTION_ID = re.compile(r"\bV\d+\b", re.IGNORECASE) +PLACEHOLDER_VALUES = {"-", "—", "–", "...", "…", "n/a", "na", "tbd"} +LEDGER_HEADERS = ( + "사실 ID", + "사실 키·정의", + "값·상태", + "단위", + "분모·대상", + "기준기간", + "집계규칙", + "상태기준일", + "문서·버전", + "원문 위치", + "대조 판정·상충 ID", +) +SIGNATURE_HEADERS = ( + "값·상태", + "단위", + "분모·대상", + "기준기간", + "집계규칙", + "상태기준일", + "문서·버전", +) +UNKNOWN_VALUES = {"미상", "불명", "확인 필요", "확인필요", "unknown"} +GENERIC_EXPLANATIONS = { + "근거", + "구체 근거", + "설명", + "구체 설명", + "사유", + "미작성", + "확인 필요", + "확인필요", + "tbd", +} +NO_MATERIAL_CONFLICT = re.compile( + r"(?m)^\s*(?:[-*]\s*)?(?:\*\*)?중대 상충 없음\s*[—-]\s*" + r".+(?:점검함|점검 완료|점검을 완료함)\.?\s*(?:\*\*)?\s*$" +) +CRITERION_UNAVAILABLE = re.compile( + r"(?m)^\s*(?:[-*]\s*)?(?:\*\*)?기준 전체 평가 불가(?:\*\*)?" + r"(?:\s*[—:-]\s*(?!아님|아니|불가하지)[^\n]+)?\s*$" +) +CONFLICT_HEADERS = ( + "ID", + "사실 ID", + "쟁점", + "값·진술 A(원문 위치)", + "값·진술 B(원문 위치)", + "대조 결과·가능한 설명", + "해결 상태", + "점수·결론 영향", + "후속 확인", +) + + +def criterion_heading_pattern(title: str) -> str: + return ( + rf"(?m)^###\s+(?:\d+[.)]\s*)?{re.escape(title)}" + rf"(?:\s*\([^\n)]*\))?(?:\s+(?:평정|평가))?\s*$" + ) def section(text: str, title: str) -> str | None: @@ -43,6 +111,261 @@ def section(text: str, title: str) -> str | None: return text[match.end():end] +def subsection(text: str, title: str) -> str | None: + """Return a level-3 Markdown subsection whose heading contains title.""" + match = re.search(criterion_heading_pattern(title), text) + if not match: + return None + next_heading = re.search(r"(?m)^###\s+", text[match.end():]) + end = match.end() + next_heading.start() if next_heading else len(text) + return text[match.end():end] + + +def markdown_cells(line: str) -> list[str] | None: + """Split one pipe table row, preserving escaped pipes inside cells.""" + stripped = line.strip() + if not (stripped.startswith("|") and stripped.endswith("|")): + return None + cells = re.split(r"(? str: + """Conservative comparison normalization; semantic equivalence stays explicit.""" + value = unicodedata.normalize("NFKC", value) + return re.sub(r"\s+", " ", value).strip().casefold() + + +def is_separator(cells: list[str]) -> bool: + return bool(cells) and all(re.fullmatch(r":?-{3,}:?", cell) for cell in cells) + + +def is_placeholder(cell: str) -> bool: + """Reject empty sentinels and fully bracketed template cells.""" + return normalized(cell) in PLACEHOLDER_VALUES or bool( + re.fullmatch(r"\[[^\]]*\]", cell) + ) + + +def is_single_source_verdict(value: str) -> bool: + return bool(re.fullmatch(r"단일 출처\s*[—-]\s*전체검색 완료", value.strip())) + + +def is_agreement_verdict(value: str) -> bool: + return normalized(value) == "일치" + + +def is_unknown_value(value: str) -> bool: + """Treat decorated unknown labels as unknown, not as confirmed agreement.""" + value = normalized(value) + if value in UNKNOWN_VALUES: + return True + return bool( + re.match( + r"^(?:미상|불명|확인\s*필요|unknown)(?:\s|\(|\[|/|:|-|—)", + value, + ) + ) + + +def has_concrete_explanation(value: str) -> bool: + """Require prose after ``설명됨:`` instead of a template placeholder.""" + match = re.search(r"설명됨\s*:\s*(.+)$", value.strip()) + if not match: + return False + explanation = match.group(1).strip() + if is_placeholder(explanation) or is_unknown_value(explanation): + return False + return normalized(explanation) not in GENERIC_EXPLANATIONS + + +def ledger_table(body: str) -> tuple[list[str] | None, list[list[str]]]: + """Find the source-fact-ledger table and return its header and data rows.""" + lines = body.splitlines() + for index, line in enumerate(lines): + header = markdown_cells(line) + if not header or header[0] != "사실 ID": + continue + if index + 1 >= len(lines): + return header, [] + separator = markdown_cells(lines[index + 1]) + if ( + separator is None + or len(separator) != len(header) + or not is_separator(separator) + ): + return header, [] + rows: list[list[str]] = [] + for candidate in lines[index + 2:]: + cells = markdown_cells(candidate) + if cells is None: + break + rows.append(cells) + return header, rows + return None, [] + + +def conflict_table(body: str) -> tuple[list[str] | None, list[list[str]], int]: + """Return the single conflict-register table and its header count.""" + + lines = body.splitlines() + header_indexes: list[int] = [] + for index, line in enumerate(lines): + cells = markdown_cells(line) + if cells and cells[0] == "ID": + header_indexes.append(index) + if not header_indexes: + return None, [], 0 + + index = header_indexes[0] + header = markdown_cells(lines[index]) + if index + 1 >= len(lines): + return header, [], len(header_indexes) + separator = markdown_cells(lines[index + 1]) + if ( + separator is None + or header is None + or len(separator) != len(header) + or not is_separator(separator) + ): + return header, [], len(header_indexes) + rows: list[list[str]] = [] + for candidate in lines[index + 2:]: + cells = markdown_cells(candidate) + if cells is None: + break + rows.append(cells) + return header, rows, len(header_indexes) + + +def inspect_ledger(text: str) -> tuple[list[str], set[str], dict[str, set[str]]]: + """Validate the early source-fact ledger and return F-IDs and X→F links.""" + violations: list[str] = [] + body = section(text, "원문 사실대장") + if body is None: + return ["필수 섹션 누락: 원문 사실대장"], set(), {} + + ledger_header_count = 0 + for line in body.splitlines(): + cells = markdown_cells(line) + if cells and cells[0] == "사실 ID": + ledger_header_count += 1 + if ledger_header_count > 1: + violations.append( + "원문 사실대장 표는 하나만 허용됨 — 모든 원문 표기를 같은 표에 합쳐야 함" + ) + + header, raw_rows = ledger_table(body) + if header is None: + return ["원문 사실대장 표 누락"], set(), {} + missing_headers = [field for field in LEDGER_HEADERS if field not in header] + for field in missing_headers: + violations.append(f"원문 사실대장 필드 누락: {field}") + if missing_headers: + return violations, set(), {} + + positions = {field: header.index(field) for field in LEDGER_HEADERS} + rows_by_fact: dict[str, list[dict[str, str]]] = defaultdict(list) + fact_ids_by_key: dict[str, set[str]] = defaultdict(set) + fact_key_labels: dict[str, str] = {} + ledger_conflicts: dict[str, set[str]] = defaultdict(set) + for row_number, cells in enumerate(raw_rows, 1): + if len(cells) != len(header): + violations.append( + f"원문 사실대장 행 {row_number} 열 수 불일치 — {len(header)}개 필드가 필요함" + ) + continue + fact_id = cells[positions["사실 ID"]].upper() + if not FACT_ID.fullmatch(fact_id): + violations.append(f"원문 사실대장 행 {row_number}의 사실 ID 형식 오류: {fact_id or '(빈칸)'}") + continue + record = {field: cells[index] for field, index in positions.items()} + for field, value in record.items(): + marker_value = field == "대조 판정·상충 ID" and bool( + CONFLICT_MARKER.fullmatch(value) + ) + if not value or (is_placeholder(value) and not marker_value): + violations.append(f"{fact_id}의 필수 필드가 비었거나 자리표시자임: {field}") + rows_by_fact[fact_id].append(record) + normalized_key = normalized(record["사실 키·정의"]) + fact_ids_by_key[normalized_key].add(fact_id) + fact_key_labels.setdefault(normalized_key, record["사실 키·정의"]) + for conflict_id in CONFLICT_MARKER.findall(record["대조 판정·상충 ID"]): + ledger_conflicts[conflict_id.upper()].add(fact_id) + + if not rows_by_fact: + violations.append("원문 사실대장에 F-ID 원문 표기 행이 하나 이상 필요함") + return violations, set(), dict(ledger_conflicts) + + for normalized_key, linked_ids in sorted(fact_ids_by_key.items()): + if len(linked_ids) > 1: + label = fact_key_labels[normalized_key] + violations.append( + f"동일 사실 키·정의 '{label}'가 여러 F-ID로 분할됨: " + + ", ".join(sorted(linked_ids)) + ) + + for fact_id, records in sorted(rows_by_fact.items()): + keys = {normalized(record["사실 키·정의"]) for record in records} + if len(keys) > 1: + violations.append(f"{fact_id}가 서로 다른 사실 키·정의에 재사용됨") + + occurrence_keys = [ + ( + normalized(record["문서·버전"]), + normalized(record["원문 위치"]), + ) + for record in records + ] + if len(set(occurrence_keys)) != len(occurrence_keys): + violations.append( + f"{fact_id}에 동일한 문서·버전과 원문 위치가 중복 등록됨" + ) + + comparisons = [record["대조 판정·상충 ID"] for record in records] + conflict_sets = [ + {item.upper() for item in CONFLICT_MARKER.findall(comparison)} + for comparison in comparisons + ] + signatures = { + tuple(normalized(record[field]) for field in SIGNATURE_HEADERS) + for record in records + } + has_unknown = any( + is_unknown_value(record[field]) + for record in records + for field in SIGNATURE_HEADERS + ) + + if len(records) == 1: + if conflict_sets[0]: + violations.append(f"{fact_id}는 상충 표기가 있으나 원문 표기 행이 한 개뿐임") + if not is_single_source_verdict(comparisons[0]): + violations.append(f"{fact_id}의 단일 행은 '단일 출처 — 전체검색 완료'로 대조 판정을 남겨야 함") + continue + + needs_explanation = len(signatures) > 1 or has_unknown + if needs_explanation: + all_conflicts = set().union(*conflict_sets) + explained = [ + has_concrete_explanation(comparison) + for comparison in comparisons + ] + if all_conflicts: + if len(all_conflicts) > 1 or any(ids != all_conflicts for ids in conflict_sets): + violations.append(f"{fact_id}의 모든 상충 행은 하나의 동일한 X-ID로 연결해야 함") + elif not all(explained): + axes = "/".join(SIGNATURE_HEADERS) + violations.append( + f"{fact_id}의 비교축({axes})이 다르거나 미상인데 " + "'설명됨: 구체 근거' 또는 '[상충: Xn]'이 없음" + ) + elif not all(is_agreement_verdict(comparison) for comparison in comparisons): + violations.append(f"{fact_id}의 동일한 반복 표기는 모든 행을 '일치'로 대조 판정해야 함") + + return violations, set(rows_by_fact), dict(ledger_conflicts) + + def inspect(text: str) -> list[str]: violations: list[str] = [] if re.search(r"\[\s*\]", text) or "[사업명]" in text: @@ -55,13 +378,44 @@ def inspect(text: str) -> list[str]: else: sections[title] = body + ledger_violations, fact_ids, ledger_conflicts = inspect_ledger(text) + # REQUIRED_SECTIONS already reports the same missing-section error. + if "원문 사실대장" not in sections: + ledger_violations = [item for item in ledger_violations if "필수 섹션 누락" not in item] + violations.extend(ledger_violations) + detail = sections.get("기준별 상세 평정", "") for criterion in CRITERIA: - if not re.search(rf"(?m)^###\s+[^\n]*{criterion}", detail): + if not re.search(criterion_heading_pattern(criterion), detail): violations.append(f"기준별 상세 평정 누락: {criterion}") for field in ("지지근거", "반대·제약근거", "근거 상태", "인접 점수"): if field not in detail: violations.append(f"상세 평정 필드 누락: {field}") + detail_facts = {item.upper() for item in FACT_MARKER.findall(detail)} + detail_conflict_links: dict[str, set[str]] = defaultdict(set) + for line_number, line in enumerate(detail.splitlines(), 1): + line_conflicts = {item.upper() for item in CONFLICT_MARKER.findall(line)} + if not line_conflicts: + continue + line_facts = {item.upper() for item in FACT_MARKER.findall(line)} + for conflict_id in line_conflicts: + if not line_facts: + violations.append( + f"상세 평정 행 {line_number}의 {conflict_id}에 [사실: Fn] 연결이 없음" + ) + detail_conflict_links[conflict_id].update(line_facts) + for fact_id in sorted(detail_facts - fact_ids): + violations.append(f"상세 평정의 {fact_id}가 원문 사실대장에 없음") + for fact_id in sorted(fact_ids - detail_facts): + violations.append(f"원문 사실대장의 {fact_id}가 어느 상세 평정에도 [사실: {fact_id}]로 연결되지 않음") + for criterion in CRITERIA: + criterion_body = subsection(detail, criterion) or "" + if not FACT_MARKER.search(criterion_body) and not CRITERION_UNAVAILABLE.search( + criterion_body + ): + violations.append( + f"{criterion} 상세 평정에 [사실: Fn] 근거 연결 또는 '기준 전체 평가 불가' 명시가 필요함" + ) overall = sections.get("종합 평정", "") for field in ("검증 후", "근거 상태", "검증 판정"): @@ -69,26 +423,92 @@ def inspect(text: str) -> list[str]: violations.append(f"종합 평정 필드 누락: {field}") register = sections.get("상충·불일치 등록부", "") - markers = {item.upper() for item in CONFLICT_MARKER.findall(text)} - register_ids = {item.upper() for item in CONFLICT_ID.findall(register)} - # Template instructions mention X1; only count IDs in actual table rows. - row_ids = { - match.group(1).upper() - for match in re.finditer(r"(?m)^\|\s*(X\d+)\s*\|", register, re.IGNORECASE) - } + detail_markers = {item.upper() for item in CONFLICT_MARKER.findall(detail)} + register_header, raw_register_rows, register_table_count = conflict_table(register) + if register_table_count > 1: + violations.append("상충 등록부 표는 하나만 허용됨") + register_records: dict[str, dict[str, str]] = {} + row_ids: set[str] = set() + if register_header is not None: + missing_headers = [field for field in CONFLICT_HEADERS if field not in register_header] + for field in missing_headers: + violations.append(f"상충 등록부 필드 누락: {field}") + positions = { + field: register_header.index(field) + for field in CONFLICT_HEADERS + if field in register_header + } + for row_number, cells in enumerate(raw_register_rows, 1): + if len(cells) != len(register_header): + violations.append( + f"상충 등록부 행 {row_number} 열 수 불일치 — {len(register_header)}개 필드가 필요함" + ) + continue + conflict_id = cells[0].upper() + if not CONFLICT_ID.fullmatch(conflict_id): + violations.append( + f"상충 등록부 행 {row_number}의 ID 형식 오류: {conflict_id or '(빈칸)'}" + ) + continue + if conflict_id in row_ids: + violations.append(f"상충 등록부 ID 중복: {conflict_id}") + continue + row_ids.add(conflict_id) + if missing_headers: + continue + record = {field: cells[index] for field, index in positions.items()} + for field, value in record.items(): + fact_marker_value = field == "사실 ID" and bool( + FACT_MARKER.fullmatch(value) + ) + if not value or (is_placeholder(value) and not fact_marker_value): + violations.append( + f"상충 등록부 {conflict_id}의 필수 필드가 비었거나 자리표시자임: {field}" + ) + register_records[conflict_id] = record + elif detail_markers or ledger_conflicts: + violations.append("상충·불일치 등록부 표 누락") + + register_ids = row_ids if row_ids: - register_ids = row_ids - for conflict_id in sorted(markers - register_ids): + for conflict_id in sorted(detail_markers - register_ids): violations.append(f"상세 평정의 {conflict_id}가 상충 등록부에 없음") - for conflict_id in sorted(register_ids - markers): + for conflict_id in sorted(register_ids - detail_markers): violations.append(f"상충 등록부의 {conflict_id}가 상세 평정에 [상충: {conflict_id}]로 연결되지 않음") - for field in ("해결 상태", "점수·결론 영향", "후속 확인"): - if field not in register: - violations.append(f"상충 등록부 필드 누락: {field}") - elif markers: - for conflict_id in sorted(markers): - violations.append(f"상세 평정의 {conflict_id}가 상충 등록부에 없음") - elif "중대 상충 없음" not in register or "점검" not in register: + for conflict_id in sorted(set(ledger_conflicts) - register_ids): + violations.append(f"원문 사실대장의 {conflict_id}가 상충 등록부에 없음") + for conflict_id in sorted(register_ids - set(ledger_conflicts)): + violations.append(f"상충 등록부의 {conflict_id}가 원문 사실대장에 [상충: {conflict_id}]로 연결되지 않음") + register_fact_links: dict[str, set[str]] = {} + for conflict_id, record in sorted(register_records.items()): + linked = {item.upper() for item in FACT_MARKER.findall(record["사실 ID"])} + register_fact_links[conflict_id] = linked + if not linked: + violations.append(f"상충 등록부의 {conflict_id}에 [사실: Fn] 연결이 없음") + for fact_id in sorted(linked - fact_ids): + violations.append(f"상충 등록부 {conflict_id}의 {fact_id}가 원문 사실대장에 없음") + for conflict_id, linked_facts in sorted(ledger_conflicts.items()): + register_linked = register_fact_links.get(conflict_id, set()) + missing = linked_facts - register_linked + for fact_id in sorted(missing): + violations.append(f"상충 등록부 {conflict_id}에 원문 사실대장 {fact_id} 연결이 없음") + for fact_id in sorted(register_linked - linked_facts): + violations.append( + f"상충 등록부 {conflict_id}의 {fact_id}가 원문 사실대장에서는 해당 상충과 연결되지 않음" + ) + detail_linked = detail_conflict_links.get(conflict_id, set()) + for fact_id in sorted(linked_facts - detail_linked): + violations.append( + f"상세 평정 {conflict_id}에 원문 사실대장 {fact_id} 연결이 없음" + ) + for fact_id in sorted(detail_linked - linked_facts): + violations.append( + f"상세 평정 {conflict_id}의 {fact_id}가 원문 사실대장에서는 해당 상충과 연결되지 않음" + ) + elif detail_markers or ledger_conflicts: + for conflict_id in sorted(detail_markers | set(ledger_conflicts)): + violations.append(f"상세 평정 또는 원문 사실대장의 {conflict_id}가 상충 등록부에 없음") + elif not NO_MATERIAL_CONFLICT.search(register): violations.append("상충 등록부에 상충 행 또는 '중대 상충 없음 — … 점검함' 확인문이 필요함") corrections = sections.get("검증 정정·점수 재산정", "") @@ -106,6 +526,11 @@ def inspect(text: str) -> list[str]: def main() -> int: parser = argparse.ArgumentParser(description="감사 가능한 평가 브리프 산출물 계약 검사") parser.add_argument("brief", help="검사할 Markdown 평가 브리프") + parser.add_argument( + "--ledger-only", + action="store_true", + help="평정 전 작업본에서 원문 사실대장 계약만 조기 검사", + ) args = parser.parse_args() try: text = Path(args.brief).read_text(encoding="utf-8") @@ -113,13 +538,20 @@ def main() -> int: print(f"읽기 실패: {exc}", file=sys.stderr) return 1 - violations = inspect(text) + if args.ledger_only: + violations, _fact_ids, _conflicts = inspect_ledger(text) + else: + violations = inspect(text) if violations: - print("감사 가능한 평가 브리프 계약 위반:", file=sys.stderr) + label = "원문 사실대장" if args.ledger_only else "감사 가능한 평가 브리프" + print(f"{label} 계약 위반:", file=sys.stderr) for index, item in enumerate(violations, 1): print(f" {index}. {item}", file=sys.stderr) return 2 - print("감사 가능한 평가 브리프 계약 통과.") + if args.ledger_only: + print("원문 사실대장 계약 통과.") + else: + print("감사 가능한 평가 브리프 계약 통과.") return 0 diff --git a/scripts/open_runner.py b/scripts/open_runner.py index 74f4f68..9ad0d14 100644 --- a/scripts/open_runner.py +++ b/scripts/open_runner.py @@ -80,7 +80,9 @@ def build_messages(target_path, reference_paths): target = read(target_path) user = ( "다음 사업 종료보고서를 KOICA 2024 평가지침의 DAC 기준으로 평가해줘.\n" - "기준별로 1~4점(근거 없으면 '평가 불가')과 지지·반대근거를 제시하고, " + "평정 전에 점수를 좌우하는 원문 표기를 한 행씩 F-ID 사실대장으로 만들고, " + "상세 평정은 [사실: Fn]으로 대장에 연결하라. 기준별로 1~4점(근거 없으면 " + "'평가 불가')과 지지·반대근거를 제시하고, " "같은 사실의 값·단위·횟수·기간·완료 상태가 다르면 상충 ID와 양쪽 위치를 " "상충·불일치 등록부에 남겨라. 검증 전후 점수와 정정 내역을 보인 뒤 종합점수와 " "A~F 등급(안)을 산정하고, 주입된 감사 가능한 평가 브리프 템플릿의 구조를 따르며, " diff --git a/skills/evaluate/SKILL.md b/skills/evaluate/SKILL.md index d838b5e..c58ea04 100644 --- a/skills/evaluate/SKILL.md +++ b/skills/evaluate/SKILL.md @@ -1,6 +1,6 @@ --- name: evaluate -description: ODA 사업을 OECD DAC/KOICA 기준으로 평가해 기준별 점수·근거, 원자료 상충 등록부, 검증 정정 내역과 종합점수·등급(안)을 산출한다. 사업 종료보고서나 사업 자료를 두고 "이 사업 평가해줘", "DAC 기준으로 평가", "종합점수/등급 내줘"라고 요청할 때 사용한다. 적절성·일관성·효과성·효율성·지속가능성 5기준을 독립 평정하고 근거를 검증해 감사 가능한 초안을 사람에게 넘긴다. 평가보고서 품질심사나 영향평가 방법론 검토에는 사용하지 않는다. +description: ODA 사업을 OECD DAC/KOICA 기준으로 평가해 구조화된 원문 사실대장, 기준별 점수·근거, 원자료 상충 등록부, 검증 정정 내역과 종합점수·등급(안)을 산출한다. 사업 종료보고서나 사업 자료를 두고 "이 사업 평가해줘", "DAC 기준으로 평가", "종합점수/등급 내줘"라고 요청할 때 사용한다. 적절성·일관성·효과성·효율성·지속가능성 5기준을 독립 평정하고 근거를 검증해 감사 가능한 초안을 사람에게 넘긴다. 평가보고서 품질심사나 영향평가 방법론 검토에는 사용하지 않는다. --- # KOICA 사업평가 (DAC 기준) @@ -16,7 +16,7 @@ description: ODA 사업을 OECD DAC/KOICA 기준으로 평가해 기준별 점 5. **등급 확정·공식 의견·환류 결정은 사람.** 6. **AI 평가의 한계 명시** — 정성적 영향력·수원국 맥락·정무적 판단은 "사람의 판단 필요"로 라벨링. 7. **평가윤리** — 조사대상자 익명(이니셜), 평가 독립성. -8. **감사추적 보존** — 같은 사실의 수치·단위·횟수·기간·완료 상태가 상충하면 유리한 값 하나를 골라 압축하지 않는다. 중대 상충, 반대근거, 평가 불가, 검증 정정은 점수가 바뀌지 않아도 최종 산출물에 남긴다. +8. **감사추적 보존** — 평정 전에 점수를 좌우하는 원문 표기를 전역 사실 ID `F1`, `F2`…로 구조화한다. 같은 사실의 수치·단위·횟수·기간·완료 상태가 상충하면 유리한 값 하나를 골라 압축하지 않는다. 중대 상충, 반대근거, 평가 불가, 검증 정정은 점수가 바뀌지 않아도 최종 산출물에 남긴다. ## 기준 체계 (2024) @@ -64,12 +64,13 @@ description: ODA 사업을 OECD DAC/KOICA 기준으로 평가해 기준별 점 사용자가 단순히 "평가해줘"라고 요청하면 `/templates/auditable-evaluation-brief-template.md`의 구조를 따른 **감사 가능한 평가 브리프**가 기본 산출물이다. 종합점수와 짧은 근거만 제시하는 한 단락 요약은 완료가 아니다. 최소한 다음을 모두 보여라. 1. 평가 범위·자료·방법과 자료 접근 한계 -2. 검증 전·**검증 후** 기준별 점수, 종합점수·등급(안), 검증 판정 -3. 기준별 핵심질문 판단, 지지근거의 원문 위치, 반대·제약근거, 근거 상태, 인접 점수가 아닌 이유 -4. **상충·불일치 등록부** — 값 A/B와 각각의 위치, 단위·분모·기간·버전·집계규칙 대조, 해결 상태, 점수·결론 영향, 후속 확인 -5. 평가 불가·미확인 등록부 -6. 검증자의 반려·정정과 이를 적용한 점수 재산정 -7. 결론과 ID로 연결된 제언, 한계, 사람 확정 게이트 +2. **원문 사실대장** — 한 원문 표기당 한 행, 같은 사실은 같은 `F` ID, 값·상태·단위·분모·기간·집계규칙·상태기준일·원문 위치와 대조 판정 +3. 검증 전·**검증 후** 기준별 점수, 종합점수·등급(안), 검증 판정 +4. 기준별 핵심질문 판단, `[사실: Fn]`으로 연결한 지지근거의 원문 위치, 반대·제약근거, 근거 상태, 인접 점수가 아닌 이유 +5. **상충·불일치 등록부** — 사실 ID, 값 A/B와 각각의 위치, 단위·분모·기간·버전·집계규칙 대조, 해결 상태, 점수·결론 영향, 후속 확인 +6. 평가 불가·미확인 등록부 +7. 검증자의 반려·정정과 이를 적용한 점수 재산정 +8. 결론과 ID로 연결된 제언, 한계, 사람 확정 게이트 **중대 상충**은 어느 값을 채택하느냐에 따라 사실판단, 달성 여부, 품질·안전·하자 상태, 기간·비용, 기준 점수, 종합등급 또는 제언이 달라질 수 있는 불일치다. 단위나 기준일 차이로 설명될 가능성이 있어도 실제로 대조·설명하기 전에는 상충 후보로 등록한다. 각 기준 평가관은 임시 ID를 쓰고, 검증자가 중복을 합쳐 전역 `X1`, `X2`…로 부여한다. 최종 상세 평정에서 영향을 받는 문장에는 `[상충: X1]`을 달아 등록부와 양방향 연결한다. @@ -85,8 +86,11 @@ description: ODA 사업을 OECD DAC/KOICA 기준으로 평가해 기준별 점 2. **자료 확인 + 사업유형 판별** — 평가 대상을 읽고 범위를 파악한다. - 평가시점과 결과 발생시점, 변화이론의 검증 가능성, 질문별 기초선·목표·자료 접근을 확인한다. - - 문서명·버전·작성일·대상기간을 자료목록에 기록한다. 핵심 지표·횟수·사업기간·예산·완료/하자 상태는 문서 전체에서 반복 표기를 찾아 **지표명 + 단위 + 분모 + 기준기간 + 집계규칙 + 상태기준일**로 정규화해 비교한다. - - 서로 다른 값·상태를 발견하면 이 단계에서 상충 후보로 남긴다. 표·본문·요약의 차이를 줄바꿈이나 오기로 추정해 임의 해결하지 않는다. + - 작업공간에 쓸 수 있으면 이 단계에서 `/templates/auditable-evaluation-brief-template.md`를 사용자 작업 폴더의 `.omo/auditable-evaluation-brief.md`로 복사해 **1. 평가 범위**와 **2. 원문 사실대장**부터 채운다. 기존 작업본이 있으면 덮어쓰거나 F-ID를 재번호화하지 말고 새 원문 표기만 추가한다. + - 문서명·버전·작성일·대상기간을 자료목록에 기록한다. 점수를 좌우하는 핵심 지표·횟수·사업기간·예산·수량·완료/하자 상태는 문서 전체에서 반복 표기를 찾아 **한 원문 표기당 한 행**으로 기록한다. 같은 판단대상·지표·목표/실적 역할은 같은 전역 `F1`, `F2`…를 반복 사용하고, 값이 다르다는 이유로 새 F-ID를 만들지 않는다. + - 각 행에 **값·상태 + 단위 + 분모·대상 + 기준기간 + 집계규칙 + 상태기준일 + 문서·버전 + 원문 위치**를 각각 명시한다. 해당하지 않거나 확인할 수 없는 축은 빈칸 대신 `해당 없음` 또는 `미상`으로 쓴다. + - 같은 F-ID의 비교축이 다르면 근거로 설명된 `설명됨: …` 또는 `[상충: Xn]`으로 분류한다. 표·본문·요약의 차이를 줄바꿈이나 오기로 추정해 임의 해결하지 않는다. + - 파일 작업본이 있으면 평정 위임 전에 `python3 /scripts/auditable_output_check.py --ledger-only <작업본>`을 실행한다. exit 2면 사실대장을 보완하고 다시 검사한다. 작업공간이 읽기 전용이면 같은 표를 내부 작업표로 유지하고 조기 자동검사를 못 했다는 한계를 최종 브리프에 명시한다. - 복잡한 평가면 위 설계 매트릭스를 채우고, 평가 불가 가능성이 있는 질문과 필요한 후속자료를 미리 분리한다. 3. **게이트웨이 증거 보강 (선택)** — `oda-intelligence` 커넥터의 도구(`oda_map_projects`, `country_report_context` 등)가 세션에 보이면, 위임 전에 외부 맥락 증거를 수집한다. 안 보이면 이 단계를 건너뛰고 7에서 한계로 명시한다(연동 안내: 저장소 `docs/oda-intelligence-integration.md`). @@ -96,15 +100,17 @@ description: ODA 사업을 OECD DAC/KOICA 기준으로 평가해 기준별 점 - ⚠️ **게이트웨이 증거는 보조 맥락이다.** 평가 대상 사업 문서가 1차 근거이며, 게이트웨이 증거로 사업 문서의 공백을 "달성"으로 메우지 마라. 4. **기준 평가관 병렬 위임** — 표준 5기준을 호스트의 서브에이전트 기능으로 가능한 범위에서 동시 위임한다. 영향력이 관련되면 `dac-impact-evaluator`도 병렬로 돌리되 **20점 종합에 합산하지 말고 별도 보고**한다. - - 위임 프롬프트에는 **담당 기준명 + 평가 대상 경로 + 규범 문서 절대경로 + 설계방법론 절대경로 + 감사 가능한 브리프 템플릿 절대경로**를 넣고, 작성한 설계 매트릭스가 있으면 그 절대경로도 넣는다. **효과성·효율성·영향력**에는 자료분석방법론도 항상 넣고, 나머지 기준도 설문·면담·표본·행정자료의 품질이 판단을 좌우하면 넣는다. + - 위임 프롬프트에는 **담당 기준명 + 평가 대상 경로 + 규범 문서 절대경로 + 설계방법론 절대경로 + 원문 사실대장을 채운 브리프 작업본(파일이 없으면 관련 행 전문) + 감사 가능한 브리프 템플릿 절대경로**를 넣고, 작성한 설계 매트릭스가 있으면 그 절대경로도 넣는다. **효과성·효율성·영향력**에는 자료분석방법론도 항상 넣고, 나머지 기준도 설문·면담·표본·행정자료의 품질이 판단을 좌우하면 넣는다. - 방법론 경로와 함께 다음 경계를 명시한다: *"채점은 현행 KOICA 규범 문서로만 한다. 방법론 문서는 근거의 적합성·강도·한계 점검용이며, 구형 DAC 기준이나 별도 점수규칙을 가져오지 마라."* - 3에서 만든 증거 블록이 있으면 **해당 기준과 관련된 부분만** 머리글 포함 그대로 덧붙인다. 그 외 기준의 점수·결론은 넣지 않는다. - - 모든 평가관에게 다음 공통 계약을 명시한다: *"핵심 주장마다 지지근거·반대/제약근거·원문 위치·근거 상태를 제시하라. 같은 사실의 값·단위·횟수·기간·완료 상태를 원문 전체에서 재검색하고, 차이가 나면 임시 상충 ID와 양쪽 위치를 기록하라. 상충을 임의 해결하거나 요약 과정에서 빼지 마라. 기준 종합점수는 질문 점수의 기계적 평균이 아니라 공식 루브릭에 대한 총체적 판단으로, 인접 점수가 아닌 이유를 써라."* + - 모든 평가관에게 다음 공통 계약을 명시한다: *"핵심 주장마다 `[사실: Fn]`과 지지근거·반대/제약근거·원문 위치·근거 상태를 제시하라. 사실대장의 관련 F-ID를 먼저 대조하고 같은 사실의 값·단위·횟수·기간·완료 상태를 원문 전체에서 재검색하라. 누락된 표기를 찾으면 기존 F-ID에 추가할 행으로 반환하고, 차이가 나면 임시 상충 ID와 양쪽 위치를 기록하라. F-ID를 임의 재번호화하거나 상충을 요약 과정에서 빼지 마라. 기준 종합점수는 질문 점수의 기계적 평균이 아니라 공식 루브릭에 대한 총체적 판단으로, 인접 점수가 아닌 이유를 써라."* - ⚠️ **다른 기준의 점수·결론·기대 등급을 언급하지 마라**(평가관 독립성 — 재위임 때도 동일). -5. **근거·점수·상충 검증** — `quality-verifier`에게 위임해 근거를 원문과 대조하고 점수-근거 정합성을 점검한다. 위임 프롬프트에 **모든 기준별 평가 초안 전문(파일로 저장했다면 각 절대경로) + 평가 대상 원자료 절대경로 + 규범층 2개 파일 + 방법론층 2개 파일 + 감사 가능한 브리프 템플릿의 절대경로**를 명시적으로 넣는다. 특히 일반 Codex 서브에이전트가 형제 평가관의 출력이나 대화 맥락을 상속한다고 가정하지 않는다. +5. **근거·점수·사실대장·상충 검증** — `quality-verifier`에게 위임해 근거를 원문과 대조하고 점수-근거 정합성을 점검한다. 위임 프롬프트에 **모든 기준별 평가 초안 전문(파일로 저장했다면 각 절대경로) + 원문 사실대장 작업본 + 평가 대상 원자료 절대경로 + 규범층 2개 파일 + 방법론층 2개 파일 + 감사 가능한 브리프 템플릿의 절대경로**를 명시적으로 넣는다. 특히 일반 Codex 서브에이전트가 형제 평가관의 출력이나 대화 맥락을 상속한다고 가정하지 않는다. - 검증자는 주장-원출처-방법-대상·표본-시점-비교-한계의 연결을 확인하고, 원문에 없는 근거나 허용 범위를 넘는 인과·일반화로 평정했으면 반려·정정한다. 공식 점수 판단은 여전히 KOICA 규범층으로만 한다. - - 평가관이 잡은 상충 후보를 확인하는 데 그치지 말고, **점수를 좌우하는 핵심 사실의 반복 표기를 원문에서 독립 재검색**한다. 임시 ID를 중복 제거해 전역 `X1`, `X2`…로 바꾸고, 값·위치·대조 결과·해결상태·점수영향을 갖춘 상충·불일치 등록부를 만든다. + - 평가관이 잡은 상충 후보를 확인하는 데 그치지 말고, **점수를 좌우하는 핵심 사실의 반복 표기를 원문에서 독립 재검색**한다. 사실대장의 F-ID가 같은 사실을 묶었는지, 각 원문 표기가 한 행인지, 필수 비교축과 위치가 채워졌는지, 다른 시그니처가 근거 있는 설명 또는 X-ID로 분류됐는지 검증한다. 누락 행은 기존 F-ID에 추가하고 새 사실만 새 F-ID를 붙이며 기존 ID를 재번호화하지 않는다. + - 임시 상충 ID를 중복 제거해 전역 `X1`, `X2`…로 바꾸고, 사실 ID·값·위치·대조 결과·해결상태·점수영향을 갖춘 상충·불일치 등록부를 만든다. 최종 상세 평정의 `[사실: Fn]`과 `[상충: Xn]`도 함께 대조한다. + - 검증 결과를 작업본에 적용한 뒤 `python3 /scripts/auditable_output_check.py --ledger-only <작업본>`을 다시 실행한다. 검증자는 읽기 전용이므로 실제 작업본 수정과 재검사는 평가총괄이 수행한다. - 각 기준에 **검증 전 점수 / 검증 후 점수 / 정정 이유**를 제시하고 전체 판정을 `통과 / 조건부 통과 / 반려`로 낸다. 미해결 중대 상충이 점수나 달성 여부를 바꿀 수 있으면 무조건 통과로 표시하지 않는다. 6. **종합점수 산정** — @@ -114,8 +120,8 @@ description: ODA 사업을 OECD DAC/KOICA 기준으로 평가해 기준별 점 - 미해결 중대 상충이 경계점수 또는 기준점수를 바꿀 수 있으면 가능한 점수 범위와 등급 민감도를 함께 제시한다. 유리한 값만 택해 단일 등급을 만들지 않는다. - **서술내용과 등급 배정 간 괴리 점검**(2024 p.7 의무). -7. **감사 가능한 브리프 작성·사람 인계** — `/templates/auditable-evaluation-brief-template.md`의 8개 섹션을 채워 제시한다. 모든 중대 상충은 점수 변화 여부와 무관하게 등록부에 남기고, 영향을 받는 상세 평정에 `[상충: Xn]`으로 연결한다. 검증자의 정정 내역과 재산정 전후도 숨기지 않는다. - - 긴/복잡한 평가 또는 파일 산출이 가능한 경우 사용자 작업 폴더의 `.omo/auditable-evaluation-brief.md`에 초안을 저장하고 `python3 /scripts/auditable_output_check.py <초안 경로>`를 실행한다. exit 2면 누락을 보완하고 다시 검사한다. 수치가 있는 초안은 `python3 /scripts/consistency_check.py <초안 경로> --mode project`도 실행한다. 산출물을 파일로 쓰지 못하면 같은 체크리스트를 대화 응답에 직접 적용하고 그 한계를 적는다. +7. **감사 가능한 브리프 작성·사람 인계** — `/templates/auditable-evaluation-brief-template.md`의 9개 섹션을 채워 제시한다. 원문 사실대장을 작업본에서 그대로 보존하고, 상세 평정은 `[사실: Fn]`, 모든 중대 상충은 사실대장·등록부·영향받는 상세 평정에서 같은 `[상충: Xn]`으로 연결한다. 검증자의 정정 내역과 재산정 전후도 숨기지 않는다. + - 파일 작업본이 있으면 `python3 /scripts/auditable_output_check.py <초안 경로>`를 실행한다. exit 2면 누락을 보완하고 다시 검사한다. 수치가 있는 초안은 `python3 /scripts/consistency_check.py <초안 경로> --mode project`도 실행한다. 산출물을 파일로 쓰지 못하면 같은 체크리스트를 대화 응답에 직접 적용하고 그 한계를 적는다. - 한계에는 외부 맥락 증거의 사용 여부도 적는다(3을 수행했으면 조회 소스·상태, 건너뛰었으면 "게이트웨이 미연결 — 외부 맥락 증거 미보강"). 끝에 **"최종 등급은 평가담당관이 확정해 주세요"**라고 명시한다. ## 다른 트랙과 혼동 금지 diff --git a/skills/write-report/SKILL.md b/skills/write-report/SKILL.md index 94bedf3..03fa131 100644 --- a/skills/write-report/SKILL.md +++ b/skills/write-report/SKILL.md @@ -15,12 +15,12 @@ description: 평가 결과로 KOICA 표준 종료평가보고서 초안을 장 ## 절차 -1. **전제 확인** — 평가 결과에는 기준별 점수·지지근거뿐 아니라 **반대·제약근거, 상충·불일치 등록부, 평가 불가 항목, 검증 정정과 검증 후 점수**가 있어야 한다. 없으면 없는 필드를 "상충 없음/정정 없음"으로 추정하지 말고 `deveval:evaluate`로 감사 가능한 브리프를 먼저 완성한다. +1. **전제 확인** — 평가 결과에는 기준별 점수·지지근거뿐 아니라 **원문 사실대장, 상세 평정의 `[사실: Fn]` 연결, 반대·제약근거, 상충·불일치 등록부, 평가 불가 항목, 검증 정정과 검증 후 점수**가 있어야 한다. 없으면 없는 필드를 "상충 없음/정정 없음"으로 추정하지 말고 `deveval:evaluate`로 감사 가능한 브리프를 먼저 완성한다. 2. **경로 확보** — 위 호스트 호환 절차로 ``를 구한다. **규범층은 기존에 검증된 평가 결과와 현행 KOICA 기준**이며, 보고서 구조 템플릿은 `/templates/evaluation-report-template.md`, 평가매트릭스 양식은 `/templates/evaluation-design-matrix-template.md`, 보고·윤리 방법론은 `/reference/개발평가-관리보고윤리-다이제스트.md`다. 변화이론·평가매트릭스를 재구성할 때만 `/reference/개발평가-설계방법론-다이제스트.md`, 조사방법·표본·분석·한계를 기술할 때만 `/reference/개발평가-자료분석방법론-다이제스트.md`를 추가한다. 방법론 파일은 구조·추적·표현을 돕는 보조자료이며 현행 KOICA 기준·점수를 바꾸지 않는다. 3. **`report-composer`(쓰기 권한)에게 위임** — 위임 프롬프트에 **평가 결과 전문(또는 각 절대경로) + 사업 원자료 절대경로 + 보고서 템플릿 절대경로 + 평가설계 매트릭스(작성본이 있으면 작성본, 없으면 양식) 절대경로 + 관리·보고·윤리 방법론 절대경로 + 필요시 설계/자료분석 방법론 절대경로 + 초안 출력 절대경로**를 명시하고, 템플릿 구조로 장별 초안을 작성시킨다. **모든 사실·평정 서술에 출처**, 미확인은 `[확인 필요]`, **국문/영문/표의 같은 수치는 반드시 일치**. 발견사항-결론-제언을 구분하고 핵심 결론은 평가질문과 원출처로 역추적 가능하게 한다. - - 평가 결과의 **상충·불일치 등록부와 ID, 해결 상태, 점수 영향, 검증 전후 점수, 평가 불가 항목**을 국문·영문 요약과 해당 기준 본문·첨부에 보존한다. 보고서 문체로 매끄럽게 만들면서 상충을 해소된 사실처럼 합치거나 누락하지 않는다. + - 평가 결과의 **원문 사실대장과 F-ID, 상충·불일치 등록부와 X-ID, 해결 상태, 점수 영향, 검증 전후 점수, 평가 불가 항목**을 국문·영문 요약과 해당 기준 본문·첨부에 보존한다. 사실대장 행을 요약 과정에서 합치거나 재번호화하지 않고, 보고서의 점수를 좌우하는 주장도 같은 `[사실: Fn]`으로 연결한다. 보고서 문체로 매끄럽게 만들면서 상충을 해소된 사실처럼 합치거나 누락하지 않는다. - 초안은 사용자 작업 폴더의 `.omo/draft-report*.md`에 저장한다(평가자의 로컬 산출물 — 플러그인 디렉토리에 쓰지 마라). 4. **수치 일관성 점검 (코드)** — 초안에 대해 실행: @@ -34,7 +34,7 @@ description: 평가 결과로 KOICA 표준 종료평가보고서 초안을 장 5. **규정 인용 검증 (선택 — 게이트웨이)** — `oda-intelligence` 커넥터가 세션에 보이면, 초안의 `{규정명} 제N조` 인용을 `verify_citation`으로 대조한다(`not_found` = 존재하지 않는 조문, `unknown_source` = 인덱스에 없는 규정명). 걸린 항목은 `report-composer`에게 **그 인용만** 정정·삭제시킨다. 조문을 원문 그대로 실어야 하면 `get_article`로 전문을 받아 쓴다. 이 검사는 KOICA 내부규정 인덱스만 대조하므로 외부 법령 인용은 이걸로 확정하지 마라. 커넥터가 없으면 건너뛰고 "규정 인용 미검증"을 한계에 남긴다(연동 안내: 저장소 `docs/oda-intelligence-integration.md`). -6. **`narrative-verifier`(읽기)에게 위임** — 위임 프롬프트에 **초안 절대경로 + 초안이 인용·사용한 평가 결과와 사업 원자료의 전문 또는 절대경로 + 3에서 사용한 방법론 파일의 절대경로**를 모두 넣는다. 특히 일반 Codex 서브에이전트가 `report-composer`의 대화 맥락을 상속한다고 가정하지 않는다. 서술-근거 *의미* 정합성(근거가 그 주장을 실제로 뒷받침하는가), 발견사항-결론-제언 추적, 인과·일반화 경계를 점검하고, 수치·등급의 기계적 일치는 4에서 코드가 봤으니 여기서는 **환각·해석 오류**에 집중한다. 평가 브리프의 **모든 중대 상충 ID·평가 불가·검증 정정이 요약과 본문에서 살아 있는지**도 대조한다. +6. **`narrative-verifier`(읽기)에게 위임** — 위임 프롬프트에 **초안 절대경로 + 초안이 인용·사용한 평가 결과와 사업 원자료의 전문 또는 절대경로 + 3에서 사용한 방법론 파일의 절대경로**를 모두 넣는다. 특히 일반 Codex 서브에이전트가 `report-composer`의 대화 맥락을 상속한다고 가정하지 않는다. 서술-근거 *의미* 정합성(근거가 그 주장을 실제로 뒷받침하는가), 발견사항-결론-제언 추적, 인과·일반화 경계를 점검하고, 수치·등급의 기계적 일치는 4에서 코드가 봤으니 여기서는 **환각·해석 오류**에 집중한다. 평가 브리프의 **모든 F-ID 사실대장 행·중대 상충 ID·평가 불가·검증 정정이 요약과 본문·첨부에서 살아 있는지**도 대조한다. 7. **(선택) 품질 자가심사** — `deveval:quality-review`로 24문항 심사를 돌려 미흡한 부분을 보완한다. @@ -44,15 +44,15 @@ description: 평가 결과로 KOICA 표준 종료평가보고서 초안을 장 - 국문 요약 / Executive Summary - **Ⅰ. 사업개요** (추진배경 / 사업개요 / 사업설계매트릭스 PDM) -- **Ⅱ. 평가개요** (목적·범위 / 평가매트릭스 / 평가방법 및 한계 / **상충·불일치 등록부 / 검증 정정·점수 재산정** / 평가팀) +- **Ⅱ. 평가개요** (목적·범위 / 평가매트릭스 / 평가방법 및 한계 / **원문 사실대장 / 상충·불일치 등록부 / 검증 정정·점수 재산정** / 평가팀) - **Ⅲ. 성과 달성도 및 사업변화이론** (성과달성 요약표) - **Ⅳ. 기준별 평가결과** (적절성·일관성·효과성·효율성·지속가능성) - **Ⅴ. 결론** (결론 / 교훈 / 제언) -- 첨부 +- 첨부 (**원문 사실대장 포함**) ## AI 작성 vs 사람 영역 -- **AI가 조립**: 사업개요(사실 정리), 성과달성 요약표, 기준별 평가결과(평가 결과를 보고서 문체로), 상충·불일치와 검증 정정의 감사추적 보존, 요약 생성, 구조·형식. +- **AI가 조립**: 사업개요(사실 정리), 성과달성 요약표, 기준별 평가결과(평가 결과를 보고서 문체로), 원문 사실대장·상충·불일치·검증 정정의 감사추적 보존, 요약 생성, 구조·형식. - **사람이 확정**: 최종 등급, 정무적·전략적 제언, 수원국 맥락 판단 → `[사람 판단 필요]`로 표시하거나 초안만 제시. > 원칙: **근거 없으면 서술 없음.** AI는 사실·구조·평가결과 정리를 돕고, 정무적 제언·최종 등급·맥락 판단은 사람이 확정한다. diff --git a/templates/auditable-evaluation-brief-template.md b/templates/auditable-evaluation-brief-template.md index ee6c032..722186f 100644 --- a/templates/auditable-evaluation-brief-template.md +++ b/templates/auditable-evaluation-brief-template.md @@ -1,6 +1,7 @@ # [사업명] 감사 가능한 평가 브리프 > `deveval:evaluate`의 기본 산출물이다. 단순 요약문이 아니라 **평정의 감사추적(audit trail)**을 사람에게 넘긴다. +> 점수를 좌우하는 원문 표기는 전역 사실 ID `F1`, `F2`…로 먼저 구조화하고, 상세 평정에서 `[사실: F1]`처럼 인용한다. > 모든 중대 상충은 전역 ID `X1`, `X2`…로 등록하고, 영향을 받는 상세 평정에 `[상충: X1]`처럼 연결한다. > `확인됨/단서 필요/불일치/근거 없음/평가 불가`는 DevEval 운영 라벨이며 KOICA 공식 점수가 아니다. @@ -14,7 +15,19 @@ | 검토 방법 | [문헌대조·지표 재계산·삼각측량 등 실제 수행한 방법] | | 자료 범위의 한계 | [원자료·회계·면담·외부자료 접근 여부] | -## 2. 종합 평정 +## 2. 원문 사실대장 + +점수를 좌우하는 지표·횟수·사업기간·예산·수량·완료/하자 상태를 **평정 전에** 기록한다. 한 원문 표기당 한 행을 쓰고, 같은 실제 사실의 반복 표기에는 같은 `F` ID를 반복한다. 값·단위·분모·기간 등이 다르다는 이유로 새 `F` ID를 만들지 않는다. + +| 사실 ID | 사실 키·정의 | 값·상태 | 단위 | 분모·대상 | 기준기간 | 집계규칙 | 상태기준일 | 문서·버전 | 원문 위치 | 대조 판정·상충 ID | +|---|---|---|---|---|---|---|---|---|---|---| +| F1 | [동일한 판단대상·지표·목표/실적 역할] | [원문 값 또는 상태] | [명/회/㎡/USD/해당 없음/미상] | [모집단·대상·분모/해당 없음/미상] | [측정기간/해당 없음/미상] | [누적·연간·고유인원·중복포함/미상] | [YYYY-MM-DD/해당 없음/미상] | [문서명·버전] | [쪽·표·절] | [단일 출처 — 전체검색 완료/일치/설명됨: 근거/[상충: Xn]] | + +- 한 행뿐이면 `단일 출처 — 전체검색 완료`, 같은 비교축으로 반복값이 같으면 `일치`라고 쓴다. +- 같은 `F` ID의 **값·상태, 단위, 분모·대상, 기준기간, 집계규칙, 상태기준일, 문서·버전** 중 하나라도 다르면 근거 있는 `설명됨: …` 또는 `[상충: Xn]`이 필요하다. 단위·기준일·문서버전 차이를 추정으로 해결하지 않는다. +- 값이 없는 축도 빈칸으로 두지 말고 `해당 없음` 또는 `미상`으로 명시한다. `미상`은 일치로 간주하지 않는다. + +## 3. 종합 평정 | 기준 | 검증 전 점수 | 검증 후 잠정 점수 | 근거 상태 | 중대 상충 ID | 핵심 판단 | |---|:---:|:---:|---|---|---| @@ -29,51 +42,51 @@ - 검증 판정: [통과 / 조건부 통과 / 반려] - 영향력(해당 시): [별도 잠정 평정 — 20점 종합 미포함] -## 3. 기준별 상세 평정 +## 4. 기준별 상세 평정 각 기준은 같은 구조를 유지한다. 질문별 점수를 기계적으로 평균하지 말고, 공식 루브릭에 따라 기준 전체의 잠정 점수를 판단한다. ### 적절성 -| 핵심질문 | 판단 | 지지근거·원문 위치 | 반대·제약근거 | 상충 ID | 근거 상태 | -|---|---|---|---|---|---| -| [질문] | [판단] | [문서·쪽·표·수치] | [반대근거·결측·대안설명] | [없음/Xn] | [확인됨/단서 필요/불일치/근거 없음] | +| 핵심질문 | 판단 | 사실 ID | 지지근거·원문 위치 | 반대·제약근거 | 상충 ID | 근거 상태 | +|---|---|---|---|---|---|---| +| [질문] | [판단] | [사실: F1] | [문서·쪽·표·수치] | [반대근거·결측·대안설명] | [없음/Xn] | [확인됨/단서 필요/불일치/근거 없음] | - 잠정 점수 및 이유: [왜 이 점수인지, 바로 위·아래 인접 점수가 아닌 이유] - 평가 불가 부분·한계: [ ] ### 일관성 -| 핵심질문 | 판단 | 지지근거·원문 위치 | 반대·제약근거 | 상충 ID | 근거 상태 | -|---|---|---|---|---|---| -| [질문] | [판단] | [문서·쪽·표·수치] | [반대근거·결측·대안설명] | [없음/Xn] | [상태] | +| 핵심질문 | 판단 | 사실 ID | 지지근거·원문 위치 | 반대·제약근거 | 상충 ID | 근거 상태 | +|---|---|---|---|---|---|---| +| [질문] | [판단] | [사실: F1] | [문서·쪽·표·수치] | [반대근거·결측·대안설명] | [없음/Xn] | [상태] | - 잠정 점수 및 이유: [왜 이 점수인지, 바로 위·아래 인접 점수가 아닌 이유] - 평가 불가 부분·한계: [ ] ### 효과성 -| 핵심질문 | 판단 | 지지근거·원문 위치 | 반대·제약근거 | 상충 ID | 근거 상태 | -|---|---|---|---|---|---| -| [질문] | [판단] | [문서·쪽·표·수치] | [품질·결과수준·형평성·상충] | [없음/Xn] | [상태] | +| 핵심질문 | 판단 | 사실 ID | 지지근거·원문 위치 | 반대·제약근거 | 상충 ID | 근거 상태 | +|---|---|---|---|---|---|---| +| [질문] | [판단] | [사실: F1] | [문서·쪽·표·수치] | [품질·결과수준·형평성·상충] | [없음/Xn] | [상태] | - 잠정 점수 및 이유: [왜 이 점수인지, 바로 위·아래 인접 점수가 아닌 이유] - 평가 불가 부분·한계: [ ] ### 효율성 -| 핵심질문 | 판단 | 지지근거·원문 위치 | 반대·제약근거 | 상충 ID | 근거 상태 | -|---|---|---|---|---|---| -| [질문] | [판단] | [문서·쪽·표·수치] | [비용·기간·품질·조달 제약] | [없음/Xn] | [상태] | +| 핵심질문 | 판단 | 사실 ID | 지지근거·원문 위치 | 반대·제약근거 | 상충 ID | 근거 상태 | +|---|---|---|---|---|---|---| +| [질문] | [판단] | [사실: F1] | [문서·쪽·표·수치] | [비용·기간·품질·조달 제약] | [없음/Xn] | [상태] | - 잠정 점수 및 이유: [왜 이 점수인지, 바로 위·아래 인접 점수가 아닌 이유] - 평가 불가 부분·한계: [ ] ### 지속가능성 -| 핵심질문 | 판단 | 지지근거·원문 위치 | 반대·제약근거 | 상충 ID | 근거 상태 | -|---|---|---|---|---|---| -| [질문] | [판단] | [문서·쪽·표·수치] | [예산·인력·제도·유지관리 제약] | [없음/Xn] | [상태] | +| 핵심질문 | 판단 | 사실 ID | 지지근거·원문 위치 | 반대·제약근거 | 상충 ID | 근거 상태 | +|---|---|---|---|---|---|---| +| [질문] | [판단] | [사실: F1] | [문서·쪽·표·수치] | [예산·인력·제도·유지관리 제약] | [없음/Xn] | [상태] | - 잠정 점수 및 이유: [왜 이 점수인지, 바로 위·아래 인접 점수가 아닌 이유] - 평가 불가 부분·한계: [ ] @@ -82,23 +95,23 @@ [영향력이 관련되면 위와 같은 표·점수 근거를 작성하고, 경쟁 설명·인과 한계와 20점 종합 미포함을 명시. 해당 없으면 N/A.] -## 4. 상충·불일치 등록부 +## 5. 상충·불일치 등록부 같은 지표·횟수·기간·예산·완료 상태가 다르게 기재되면 단위·분모·기준기간·문서버전·집계규칙·상태기준일을 대조한다. 설명되지 않은 차이는 유리한 값 하나를 택하지 않는다. -| ID | 쟁점 | 값·진술 A(원문 위치) | 값·진술 B(원문 위치) | 대조 결과·가능한 설명 | 해결 상태 | 점수·결론 영향 | 후속 확인 | -|---|---|---|---|---|---|---|---| -| X1 | [동일 사실의 상충] | [ ] | [ ] | [단위/시점/버전 차이 여부] | [해결/부분해결/미해결] | [기준·점수·결론 영향 또는 영향 없음의 이유] | [자료·담당] | +| ID | 사실 ID | 쟁점 | 값·진술 A(원문 위치) | 값·진술 B(원문 위치) | 대조 결과·가능한 설명 | 해결 상태 | 점수·결론 영향 | 후속 확인 | +|---|---|---|---|---|---|---|---|---| +| X1 | [사실: F1] | [동일 사실의 상충] | [ ] | [ ] | [단위/시점/버전 차이 여부] | [해결/부분해결/미해결] | [기준·점수·결론 영향 또는 영향 없음의 이유] | [자료·담당] | 상충이 없을 때만 위 예시행을 지우고 다음처럼 쓴다: **중대 상충 없음 — 핵심 지표·횟수·기간·예산·완료 상태의 반복 표기를 점검함.** -## 5. 평가 불가·미확인 등록부 +## 6. 평가 불가·미확인 등록부 | 질문·주장 ID | 현재 판단 | 부족하거나 미확인인 근거 | 점수·결론 영향 | 필요한 후속자료 | |---|---|---|---|---| | [ ] | [평가 불가/단서부] | [ ] | [ ] | [ ] | -## 6. 검증 정정·점수 재산정 +## 7. 검증 정정·점수 재산정 | 정정 ID | 검증 전 주장·점수 | 원문 대조 결과 | 적용한 정정 | 종합점수·등급 영향 | |---|---|---|---|---| @@ -106,13 +119,13 @@ 정정이 없을 때만 예시행을 지우고 **검증 정정 없음 — 검증 전·후 점수 동일**이라고 쓴다. 정정 요청을 반영하지 못했으면 숨기지 말고 사유와 함께 `[사람 판단 필요]`로 남긴다. -## 7. 결론 연계 제언 +## 8. 결론 연계 제언 | 우선순위 | 근거가 된 결론·상충·공백 ID | 제언 | 책임주체 | 시점·이행확인 | |---|---|---|---|---| | [상/중/하] | [C/X/Q/V ID] | [구체적 행동] | [ ] | [ ] | -## 8. 한계·사람 판단 +## 9. 한계·사람 판단 - 평가 한계: [자료·방법·대표성·외부 맥락·인과 해석 한계] - `[사람 판단 필요]`: [최종 등급·정무적 판단·미해결 상충] diff --git a/templates/eval-plan-template.md b/templates/eval-plan-template.md index 6e9f3e6..b62147e 100644 --- a/templates/eval-plan-template.md +++ b/templates/eval-plan-template.md @@ -12,14 +12,15 @@ ## 사업: <사업명> (사업유형: <유형>) - [ ] 자료 확인 및 사업유형 판별 -- [ ] 핵심 지표·횟수·기간·예산·완료 상태의 반복 표기 검색 및 상충 후보 등록 +- [ ] 원문 사실대장 작성 (한 표기당 한 행 / 같은 사실 F-ID 재사용 / 값·단위·분모·기간·집계규칙·기준일·위치) +- [ ] 사실대장 비교축 차이 설명 또는 상충 후보 등록 및 `auditable_output_check.py --ledger-only` 통과 - [ ] 적절성(Relevance) 평가 - [ ] 일관성(Coherence) 평가 - [ ] 효과성(Effectiveness) 평가 - [ ] 효율성(Efficiency) 평가 - [ ] 지속가능성(Sustainability) 평가 -- [ ] 근거·점수 검증 (quality-verifier) -- [ ] 전역 상충·불일치 등록부 확정 (X1… / 해결 상태 / 점수·결론 영향) +- [ ] 근거·점수·원문 사실대장 검증 (quality-verifier) +- [ ] 전역 상충·불일치 등록부 확정 (F-ID / X1… / 해결 상태 / 점수·결론 영향) - [ ] 검증 정정 적용 및 **검증 후 점수** 재산정 (5기준 20점 → A~F) -- [ ] 감사 가능한 브리프 작성 (기준별 상세근거·반대근거·상충·평가불가·정정·한계) +- [ ] 감사 가능한 브리프 작성 (원문 사실대장·[사실: Fn] 상세근거·반대근거·상충·평가불가·정정·한계) - [ ] 산출물 계약·수치 일관성 검사 후 사람 인계 diff --git a/templates/evaluation-report-template.md b/templates/evaluation-report-template.md index 20441bc..bb3c7ba 100644 --- a/templates/evaluation-report-template.md +++ b/templates/evaluation-report-template.md @@ -9,7 +9,7 @@ 1. **평가목적·범위·방법** (대상·시점·주요 자료·핵심 한계) 2. **기준별 평가 결과** (각 기준 핵심 결론 + 지지근거·반대근거·근거 상태) 3. **종합 평가등급** (검증 후 점수 → 등급 — *영문 요약과 동일 수치*) -4. **중대 상충·검증 정정** (상충 ID·해결 상태·점수영향, 검증 전후 점수) +4. **원문 사실대장·중대 상충·검증 정정** (F-ID·상충 ID·해결 상태·점수영향, 검증 전후 점수) 5. **평가 불가·사람 판단 필요 항목** 6. **교훈 및 제언** (우선순위·책임주체·시점) @@ -27,16 +27,17 @@ 1. 평가의 목적과 범위 2. 평가매트릭스 — `evaluation-design-matrix-template.md`의 질문·판단 및 자료·분석 표를 요약 3. 평가방법 및 한계 — 설계 / 모집단·표본 / 도구·시점 / 분석 / 자료품질 / 삼각측량 / 편향·결측·윤리·일반화 한계 -4. **상충·불일치 등록부** — 동일 사실의 값 A/B와 위치 / 단위·분모·기간·버전·집계규칙 대조 / 해결 상태 / 기준·점수·결론 영향 / 후속 확인 -5. 검증 정정·점수 재산정 — 검증 전 주장·점수 / 원문 대조 / 적용 정정 / 종합등급 영향 -6. 평가팀 구성 +4. **원문 사실대장** — 평가 브리프의 F-ID와 한 원문 표기당 한 행을 그대로 보존 / 값·상태·단위·분모·기간·집계규칙·기준일·원문 위치 / 대조 판정 +5. **상충·불일치 등록부** — F-ID / 동일 사실의 값 A/B와 위치 / 단위·분모·기간·버전·집계규칙 대조 / 해결 상태 / 기준·점수·결론 영향 / 후속 확인 +6. 검증 정정·점수 재산정 — 검증 전 주장·점수 / 원문 대조 / 적용 정정 / 종합등급 영향 +7. 평가팀 구성 ## Ⅲ. 성과 달성도 및 사업변화이론 -1. 성과달성 요약표 — | 지표 | 목표 | 실적 | 달성여부 | 입증자료 | +1. 성과달성 요약표 — | 사실 ID | 지표 | 목표 | 실적 | 달성여부 | 입증자료 | 2. 사업변화이론(ToC) — 단계별 가정·외부요인·검증 여부 포함 ## Ⅳ. 기준별 평가결과 -*(각 기준: 핵심질문별 평정 + 지지근거·반대/제약근거 + 상충 ID + 근거 상태 + 검증 후 잠정 점수와 인접 점수가 아닌 이유. 평가관 결과를 보고서 문체로 정리.)* +*(각 기준: 핵심질문별 평정 + `[사실: Fn]` + 지지근거·반대/제약근거 + 상충 ID + 근거 상태 + 검증 후 잠정 점수와 인접 점수가 아닌 이유. 평가관 결과를 보고서 문체로 정리.)* 1. 적절성 (Relevance) 2. 일관성 (Coherence) 3. 효과성 (Effectiveness) @@ -49,4 +50,4 @@ 3. 제언 — 각 결론과 연결하고 책임주체·우선순위·시점·이행확인 방법 명시. `[사람 판단 필요]` 정무적·전략적 부분은 초안만 ## 첨부 -- 국·영문 요약 / 평가설계 매트릭스 / 핵심 주장-근거 등록부 / **상충·불일치 등록부 / 검증 정정 이력** / 현지조사 개요 / 일별 활동내역 / 비식별 면담자 목록 및 질문 / 설문조사 결과 / 참고문헌 +- 국·영문 요약 / 평가설계 매트릭스 / 핵심 주장-근거 등록부 / **원문 사실대장 / 상충·불일치 등록부 / 검증 정정 이력** / 현지조사 개요 / 일별 활동내역 / 비식별 면담자 목록 및 질문 / 설문조사 결과 / 참고문헌 diff --git a/tests/fixtures/auditable-brief-clean.md b/tests/fixtures/auditable-brief-clean.md index fdd30ff..acccf1f 100644 --- a/tests/fixtures/auditable-brief-clean.md +++ b/tests/fixtures/auditable-brief-clean.md @@ -3,7 +3,17 @@ ## 1. 평가 범위·자료·방법 종료평가보고서 1건을 문헌대조했다. -## 2. 종합 평정 +## 2. 원문 사실대장 +| 사실 ID | 사실 키·정의 | 값·상태 | 단위 | 분모·대상 | 기준기간 | 집계규칙 | 상태기준일 | 문서·버전 | 원문 위치 | 대조 판정·상충 ID | +|---|---|---|---|---|---|---|---|---|---|---| +| F1 | 수요 부합 상태 | 부합 | 해당 없음 | 사업대상지역 | 설계기간 | 질적 판단 | 2023-12-01 | 종료보고서 v1 | p.10 | 단일 출처 — 전체검색 완료 | +| F2 | 공여기관 조율 상태 | 일부 조율 | 해당 없음 | 참여기관 | 사업기간 | 질적 판단 | 2023-12-01 | 종료보고서 v1 | p.12 | 단일 출처 — 전체검색 완료 | +| F3 | 현지연수 실시 횟수 실적 | 5 | 회 | 현지연수 | 사업누적 | 행사 횟수 | 2023-12-01 | 종료보고서 v1 | p.20 본문 | [상충: X1] | +| F3 | 현지연수 실시 횟수 실적 | 6 | 회 | 현지연수 | 사업누적 | 행사 횟수 | 2023-12-01 | 종료보고서 v1 | p.21 표 | [상충: X1] | +| F4 | 사업 일정 준수 상태 | 지연 | 해당 없음 | 전체 사업 | 사업기간 | 계획 대비 실제 | 2023-12-01 | 종료보고서 v1 | p.30 | 단일 출처 — 전체검색 완료 | +| F5 | 반복 운영재원 확보 상태 | 미확보 | 해당 없음 | 운영기관 | 종료 후 | 승인예산 기준 | 2023-12-01 | 종료보고서 v1 | p.40 | 단일 출처 — 전체검색 완료 | + +## 3. 종합 평정 | 기준 | 검증 전 점수 | 검증 후 잠정 점수 | 근거 상태 | 중대 상충 ID | 핵심 판단 | |---|---:|---:|---|---|---| | 적절성 | 3 | 3 | 확인됨 | 없음 | 수요 부합 | @@ -14,49 +24,49 @@ - 검증 후 종합점수·등급(안): 13/20, D(안) - 검증 판정: 조건부 통과 -## 3. 기준별 상세 평정 -| 핵심질문 | 판단 | 지지근거 | 반대·제약근거 | 상충 ID | 근거 상태 | -|---|---|---|---|---|---| +## 4. 기준별 상세 평정 +| 핵심질문 | 판단 | 사실 ID | 지지근거 | 반대·제약근거 | 상충 ID | 근거 상태 | +|---|---|---|---|---|---|---| ### 적절성 -| 수요 | 부합 | [근거: 보고서 p.10] | 주민자료 없음 | 없음 | 확인됨 | +| 수요 | 부합 | [사실: F1] | [근거: 보고서 p.10] | 주민자료 없음 | 없음 | 확인됨 | - 잠정 점수 및 이유: 3점. 전반적으로 부합하나 직접 수요조사가 없어 인접 점수 4는 아님. ### 일관성 -| 조율 | 일부 | [근거: 보고서 p.12] | 공동계획 없음 | 없음 | 단서 필요 | +| 조율 | 일부 | [사실: F2] | [근거: 보고서 p.12] | 공동계획 없음 | 없음 | 단서 필요 | - 잠정 점수 및 이유: 3점. 조율 흔적은 있으나 인접 점수 4의 시너지는 입증되지 않음. ### 효과성 -| 연수 | 대체로 달성 | [근거: 보고서 p.20] | 본문 5회, 표 6회 [상충: X1] | X1 | 단서 필요 | +| 연수 | 대체로 달성 | [사실: F3] | [근거: 보고서 p.20] | 본문 5회, 표 6회 [상충: X1] | X1 | 단서 필요 | - 잠정 점수 및 이유: 3점. 대체로 달성했으나 품질과 횟수가 상충해 인접 점수 4는 아님. ### 효율성 -| 일정 | 지연 | [근거: 보고서 p.30] | 비교비용 없음 | 없음 | 확인됨 | +| 일정 | 지연 | [사실: F4] | [근거: 보고서 p.30] | 비교비용 없음 | 없음 | 확인됨 | - 잠정 점수 및 이유: 2점. 지연 영향이 있어 인접 점수 3은 아님. ### 지속가능성 -| 재원 | 일부 준비 | [근거: 보고서 p.40] | 반복예산 미확인 | 없음 | 단서 필요 | +| 재원 | 일부 준비 | [사실: F5] | [근거: 보고서 p.40] | 반복예산 미확인 | 없음 | 단서 필요 | - 잠정 점수 및 이유: 2점. 운영계획은 있으나 인접 점수 3의 재원 확정은 없음. -## 4. 상충·불일치 등록부 -| ID | 쟁점 | 값 A | 값 B | 대조 결과 | 해결 상태 | 점수·결론 영향 | 후속 확인 | -|---|---|---|---|---|---|---|---| -| X1 | 연수 횟수 | 5회(p.20) | 6회(p.21) | 집계규칙 설명 없음 | 미해결 | 효과성 4→3 | 원자료 확인 | +## 5. 상충·불일치 등록부 +| ID | 사실 ID | 쟁점 | 값·진술 A(원문 위치) | 값·진술 B(원문 위치) | 대조 결과·가능한 설명 | 해결 상태 | 점수·결론 영향 | 후속 확인 | +|---|---|---|---|---|---|---|---|---| +| X1 | [사실: F3] | 연수 횟수 | 5회(p.20) | 6회(p.21) | 집계규칙 설명 없음 | 미해결 | 효과성 4→3 | 원자료 확인 | -## 5. 평가 불가·미확인 등록부 +## 6. 평가 불가·미확인 등록부 | 질문 | 현재 판단 | 부족한 근거 | 영향 | 후속자료 | |---|---|---|---|---| | 운영재원 | 단서부 | 예산서 없음 | 지속가능성 | 예산서 | -## 6. 검증 정정·점수 재산정 +## 7. 검증 정정·점수 재산정 | 정정 ID | 검증 전 주장·점수 | 원문 대조 결과 | 적용한 정정 | 종합점수·등급 영향 | |---|---|---|---|---| | V1 | 효과성 4 | 횟수 상충 미해결 | 효과성 3 | 14→13, C→D | -## 7. 결론 연계 제언 +## 8. 결론 연계 제언 | 우선순위 | 근거 ID | 제언 | 책임 | 시점 | |---|---|---|---|---| | 상 | X1 | 연수 원자료 대조 | 사업팀 | 즉시 | -## 8. 한계·사람 판단 +## 9. 한계·사람 판단 원자료가 없다. 최종 등급은 평가담당관이 확정해 주세요. diff --git a/tests/fixtures/auditable-brief-missing-register.md b/tests/fixtures/auditable-brief-missing-register.md index 1f44ab3..7158de2 100644 --- a/tests/fixtures/auditable-brief-missing-register.md +++ b/tests/fixtures/auditable-brief-missing-register.md @@ -2,27 +2,40 @@ ## 1. 평가 범위·자료·방법 자료 1건. -## 2. 종합 평정 + +## 2. 원문 사실대장 +| 사실 ID | 사실 키·정의 | 값·상태 | 단위 | 분모·대상 | 기준기간 | 집계규칙 | 상태기준일 | 문서·버전 | 원문 위치 | 대조 판정·상충 ID | +|---|---|---|---|---|---|---|---|---|---|---| +| F1 | 연수 실시 횟수 실적 | 5 | 회 | 현지연수 | 사업누적 | 행사 횟수 | 2023-12-01 | 보고서 v1 | p.20 | [상충: X1] | +| F1 | 연수 실시 횟수 실적 | 6 | 회 | 현지연수 | 사업누적 | 행사 횟수 | 2023-12-01 | 보고서 v1 | p.21 | [상충: X1] | + +## 3. 종합 평정 검증 후 근거 상태, 검증 판정. -## 3. 기준별 상세 평정 + +## 4. 기준별 상세 평정 지지근거 / 반대·제약근거 / 근거 상태 / 인접 점수 ### 적절성 -내용 +[사실: F1] 내용 ### 일관성 -내용 +[사실: F1] 내용 ### 효과성 -[상충: X1] +[사실: F1] [상충: X1] ### 효율성 -내용 +[사실: F1] 내용 ### 지속가능성 -내용 -## 4. 상충·불일치 등록부 +[사실: F1] 내용 + +## 5. 상충·불일치 등록부 중대 상충 없음 — 반복 표기를 점검함. -## 5. 평가 불가·미확인 등록부 + +## 6. 평가 불가·미확인 등록부 없음. -## 6. 검증 정정·점수 재산정 + +## 7. 검증 정정·점수 재산정 검증 전 / 원문 대조 / 종합점수·등급 영향. 검증 정정 없음. -## 7. 결론 연계 제언 + +## 8. 결론 연계 제언 없음. -## 8. 한계·사람 판단 + +## 9. 한계·사람 판단 최종 등급은 평가담당관이 확정해 주세요. diff --git a/tests/fixtures/source-fact-ledger-stage.md b/tests/fixtures/source-fact-ledger-stage.md new file mode 100644 index 0000000..678cb91 --- /dev/null +++ b/tests/fixtures/source-fact-ledger-stage.md @@ -0,0 +1,16 @@ +# 평정 전 작업본 + +## 1. 평가 범위·자료·방법 +원문 전체를 검색했다. + +## 2. 원문 사실대장 +| 사실 ID | 사실 키·정의 | 값·상태 | 단위 | 분모·대상 | 기준기간 | 집계규칙 | 상태기준일 | 문서·버전 | 원문 위치 | 대조 판정·상충 ID | +|---|---|---|---|---|---|---|---|---|---|---| +| F1 | 사업 승인예산 | 1000000 | USD | 전체 사업 | 사업기간 | 승인 총액 | 2021-12-31 | 보고서 v1 | p.10 | 단일 출처 — 전체검색 완료 | +| F2 | 초청연수 수료인원 실적 | 55 | 명 | 수료자 | 사업누적 | 고유인원 | 2021-12-31 | 보고서 v1 | p.20 본문 | 일치 | +| F2 | 초청연수 수료인원 실적 | 55 | 명 | 수료자 | 사업누적 | 고유인원 | 2021-12-31 | 보고서 v1 | p.21 표 | 일치 | +| F3 | 연도별 교육인원 실적 | 100 | 명 | 수료자 | 2020년 | 연간 고유인원 | 2020-12-31 | 보고서 v1 | p.30 | 설명됨: 서로 다른 연도별 실적 | +| F3 | 연도별 교육인원 실적 | 120 | 명 | 수료자 | 2021년 | 연간 고유인원 | 2021-12-31 | 보고서 v1 | p.31 | 설명됨: 서로 다른 연도별 실적 | + +## 3. 종합 평정 +평정 전. diff --git a/tests/fixtures/source-fact-ledger-unclassified.md b/tests/fixtures/source-fact-ledger-unclassified.md new file mode 100644 index 0000000..f45a5b5 --- /dev/null +++ b/tests/fixtures/source-fact-ledger-unclassified.md @@ -0,0 +1,13 @@ +# 미분류 비교축 차이 + +## 2. 원문 사실대장 +| 사실 ID | 사실 키·정의 | 값·상태 | 단위 | 분모·대상 | 기준기간 | 집계규칙 | 상태기준일 | 문서·버전 | 원문 위치 | 대조 판정·상충 ID | +|---|---|---|---|---|---|---|---|---|---|---| +| F1 | SICA 지표 목표 | 15 | 명 | 참가자 | 사업누적 | 고유인원 | 2021-12-31 | 보고서 v1 | p.10 | 일치 | +| F1 | SICA 지표 목표 | 15 | 회 | 컨퍼런스 | 사업누적 | 행사 횟수 | 2021-12-31 | 보고서 v1 | p.11 | 일치 | +| F2 | 현지연수 실시 횟수 실적 | 5 | 회 | 현지연수 | 사업누적 | 행사 횟수 | 2021-12-31 | 보고서 v1 | p.20 | 일치 | +| F2 | 현지연수 실시 횟수 실적 | 6 | 회 | 현지연수 | 사업누적 | 행사 횟수 | 2021-12-31 | 보고서 v1 | p.21 | 일치 | +| F3 | 사업 수행기간 | 2014~2020 | 년 | 전체 사업 | 승인 사업기간 | 시작·종료연도 | 2020-12-31 | 보고서 v1 | p.30 | 일치 | +| F3 | 사업 수행기간 | 2014~2021 | 년 | 전체 사업 | 승인 사업기간 | 시작·종료연도 | 2021-12-31 | 보고서 v1 | p.31 | 일치 | +| F4 | 시설 하자보수 상태 | 완료 | 해당 없음 | 전체 시설 | 종료점검 | 하자 전건 | 2022-11-30 | 보고서 v1 | p.40 | 일치 | +| F4 | 시설 하자보수 상태 | 진행 중 | 해당 없음 | 전체 시설 | 종료점검 | 하자 전건 | 2022-11-30 | 보고서 v1 | p.41 | 일치 | diff --git a/tests/test_auditable_output_contract.py b/tests/test_auditable_output_contract.py index 75e2f3f..95b534c 100644 --- a/tests/test_auditable_output_contract.py +++ b/tests/test_auditable_output_contract.py @@ -3,6 +3,7 @@ import os import subprocess import sys +import tempfile import unittest @@ -16,26 +17,258 @@ def read(*parts): return stream.read() -def run_fixture(name): +def run_fixture(name, *extra_args): return subprocess.run( - [sys.executable, CHECKER, os.path.join(FIXTURES, name)], + [sys.executable, CHECKER, *extra_args, os.path.join(FIXTURES, name)], check=False, text=True, capture_output=True, ) +def run_text(text, *extra_args): + """Write then close a temporary file before the checker reopens it (Windows-safe).""" + with tempfile.TemporaryDirectory() as directory: + path = os.path.join(directory, "brief.md") + with open(path, "w", encoding="utf-8") as stream: + stream.write(text) + return subprocess.run( + [sys.executable, CHECKER, *extra_args, path], + check=False, + text=True, + capture_output=True, + ) + + class CheckerBehavior(unittest.TestCase): def test_complete_brief_passes(self): result = run_fixture("auditable-brief-clean.md") self.assertEqual(result.returncode, 0, result.stdout + result.stderr) + def test_ledger_only_stage_passes_before_scoring(self): + result = run_fixture("source-fact-ledger-stage.md", "--ledger-only") + self.assertEqual(result.returncode, 0, result.stdout + result.stderr) + self.assertIn("원문 사실대장 계약 통과", result.stdout) + + def test_unit_count_period_and_status_differences_need_classification(self): + result = run_fixture("source-fact-ledger-unclassified.md", "--ledger-only") + self.assertEqual(result.returncode, 2, result.stdout + result.stderr) + for fact_id in ("F1", "F2", "F3", "F4"): + self.assertIn(fact_id, result.stderr) + self.assertIn("설명됨", result.stderr) + self.assertIn("상충", result.stderr) + + def test_disagreement_text_does_not_count_as_agreement(self): + text = read("tests", "fixtures", "source-fact-ledger-stage.md").replace( + "| 보고서 v1 | p.20 본문 | 일치 |", + "| 보고서 v1 | p.20 본문 | 불일치 |", + 1, + ) + result = run_text(text, "--ledger-only") + self.assertEqual(result.returncode, 2, result.stdout + result.stderr) + self.assertIn("F2", result.stderr) + self.assertIn("'일치'", result.stderr) + + def test_same_fact_key_cannot_be_split_across_fact_ids(self): + text = read("tests", "fixtures", "source-fact-ledger-stage.md") + text = text.replace( + "| F2 | 초청연수 수료인원 실적 | 55 | 명 | 수료자 | 사업누적 | 고유인원 | 2021-12-31 | 보고서 v1 | p.20 본문 | 일치 |", + "| F20 | 초청연수 수료인원 실적 | 55 | 명 | 수료자 | 사업누적 | 고유인원 | 2021-12-31 | 보고서 v1 | p.20 본문 | 단일 출처 — 전체검색 완료 |", + 1, + ).replace( + "| F2 | 초청연수 수료인원 실적 | 55 | 명 | 수료자 | 사업누적 | 고유인원 | 2021-12-31 | 보고서 v1 | p.21 표 | 일치 |", + "| F2 | 초청연수 수료인원 실적 | 55 | 명 | 수료자 | 사업누적 | 고유인원 | 2021-12-31 | 보고서 v1 | p.21 표 | 단일 출처 — 전체검색 완료 |", + 1, + ) + result = run_text(text, "--ledger-only") + self.assertEqual(result.returncode, 2, result.stdout + result.stderr) + self.assertIn("여러 F-ID로 분할", result.stderr) + self.assertIn("F20", result.stderr) + + def test_second_fact_ledger_table_cannot_hide_occurrences(self): + text = read("tests", "fixtures", "source-fact-ledger-stage.md").replace( + "\n## 3. 종합 평정", + "\n| 사실 ID | 사실 키·정의 | 값·상태 | 단위 | 분모·대상 | 기준기간 | 집계규칙 | 상태기준일 | 문서·버전 | 원문 위치 | 대조 판정·상충 ID |\n" + "|---|---|---|---|---|---|---|---|---|---|---|\n" + "| F9 | 숨긴 완료상태 | 진행 중 | 해당 없음 | 전체 시설 | 종료점검 | 하자 전건 | 2022-11-30 | 보고서 v1 | p.99 | 단일 출처 — 전체검색 완료 |\n" + "\n## 3. 종합 평정", + 1, + ) + result = run_text(text, "--ledger-only") + self.assertEqual(result.returncode, 2, result.stdout + result.stderr) + self.assertIn("표는 하나만 허용", result.stderr) + + def test_ledger_rows_must_be_contiguous_with_the_header(self): + text = read("tests", "fixtures", "source-fact-ledger-stage.md").replace( + "|---|---|---|---|---|---|---|---|---|---|---|\n| F1 |", + "|---|---|---|---|---|---|---|---|---|---|---|\n\n대장 표 종료\n| F1 |", + 1, + ) + result = run_text(text, "--ledger-only") + self.assertEqual(result.returncode, 2, result.stdout + result.stderr) + self.assertIn("원문 사실대장에 F-ID", result.stderr) + + def test_dash_is_not_a_completed_ledger_field(self): + text = read("tests", "fixtures", "source-fact-ledger-stage.md").replace( + "| F1 | 사업 승인예산 | 1000000 |", + "| F1 | 사업 승인예산 | - |", + 1, + ) + result = run_text(text, "--ledger-only") + self.assertEqual(result.returncode, 2, result.stdout + result.stderr) + self.assertIn("F1", result.stderr) + self.assertIn("값·상태", result.stderr) + + def test_decorated_unknown_cannot_be_confirmed_as_agreement(self): + text = read("tests", "fixtures", "source-fact-ledger-stage.md").replace( + "| 55 | 명 |", + "| 미상(원문 미기재) | 명 |", + ) + result = run_text(text, "--ledger-only") + self.assertEqual(result.returncode, 2, result.stdout + result.stderr) + self.assertIn("F2", result.stderr) + self.assertIn("설명됨", result.stderr) + + def test_explanation_placeholder_is_not_concrete_evidence(self): + text = read("tests", "fixtures", "source-fact-ledger-stage.md").replace( + "설명됨: 서로 다른 연도별 실적", + "설명됨: [근거]", + ) + result = run_text(text, "--ledger-only") + self.assertEqual(result.returncode, 2, result.stdout + result.stderr) + self.assertIn("F3", result.stderr) + self.assertIn("구체 근거", result.stderr) + + def test_fact_marker_cannot_fill_an_unrelated_ledger_field(self): + text = read("tests", "fixtures", "source-fact-ledger-stage.md").replace( + "| F1 | 사업 승인예산 |", + "| F1 | [사실: F1] |", + 1, + ) + result = run_text(text, "--ledger-only") + self.assertEqual(result.returncode, 2, result.stdout + result.stderr) + self.assertIn("사실 키·정의", result.stderr) + + def test_document_version_difference_needs_classification(self): + text = read("tests", "fixtures", "source-fact-ledger-stage.md").replace( + "| 보고서 v1 | p.21 표 | 일치 |", + "| 보고서 v2 | p.21 표 | 일치 |", + 1, + ) + result = run_text(text, "--ledger-only") + self.assertEqual(result.returncode, 2, result.stdout + result.stderr) + self.assertIn("F2", result.stderr) + self.assertIn("문서·버전", result.stderr) + + def test_same_source_occurrence_cannot_be_counted_twice(self): + text = read("tests", "fixtures", "source-fact-ledger-stage.md").replace( + "| 보고서 v1 | p.21 표 | 일치 |", + "| 보고서 v1 | p.20 본문 | 일치 |", + 1, + ) + result = run_text(text, "--ledger-only") + self.assertEqual(result.returncode, 2, result.stdout + result.stderr) + self.assertIn("F2", result.stderr) + self.assertIn("원문 위치가 중복", result.stderr) + + def test_detail_conflict_fact_pair_must_match_ledger(self): + text = read("tests", "fixtures", "auditable-brief-clean.md").replace( + "[사실: F3] | [근거: 보고서 p.20]", + "[사실: F4] | [근거: 보고서 p.20]", + 1, + ) + result = run_text(text) + self.assertEqual(result.returncode, 2, result.stdout + result.stderr) + self.assertIn("상세 평정 X1", result.stderr) + self.assertIn("F3", result.stderr) + self.assertIn("F4", result.stderr) + + def test_every_ledger_fact_must_reach_a_detailed_rating(self): + text = read("tests", "fixtures", "auditable-brief-clean.md").replace( + "[사실: F5] | [근거: 보고서 p.40]", + "[사실: F4] | [근거: 보고서 p.40]", + 1, + ) + result = run_text(text) + self.assertEqual(result.returncode, 2, result.stdout + result.stderr) + self.assertIn("F5", result.stderr) + self.assertIn("어느 상세 평정에도", result.stderr) + + def test_conflict_register_requires_the_full_row_schema(self): + text = read("tests", "fixtures", "auditable-brief-clean.md") + old = ( + "| ID | 사실 ID | 쟁점 | 값·진술 A(원문 위치) | 값·진술 B(원문 위치) | 대조 결과·가능한 설명 | 해결 상태 | 점수·결론 영향 | 후속 확인 |\n" + "|---|---|---|---|---|---|---|---|---|\n" + "| X1 | [사실: F3] | 연수 횟수 | 5회(p.20) | 6회(p.21) | 집계규칙 설명 없음 | 미해결 | 효과성 4→3 | 원자료 확인 |" + ) + collapsed = ( + "| ID | 사실 ID |\n" + "|---|---|\n" + "| X1 | [사실: F3] 해결 상태 점수·결론 영향 후속 확인 |" + ) + result = run_text(text.replace(old, collapsed, 1)) + self.assertEqual(result.returncode, 2, result.stdout + result.stderr) + self.assertIn("상충 등록부 필드 누락", result.stderr) + self.assertIn("값·진술 A", result.stderr) + + def test_combined_criterion_heading_cannot_satisfy_five_sections(self): + text = read("tests", "fixtures", "auditable-brief-clean.md").replace( + "### 적절성", + "### 적절성·일관성·효과성·효율성·지속가능성", + 1, + ) + for criterion in ("일관성", "효과성", "효율성", "지속가능성"): + text = text.replace(f"### {criterion}\n", "", 1) + result = run_text(text) + self.assertEqual(result.returncode, 2, result.stdout + result.stderr) + for criterion in ("적절성", "일관성", "효과성", "효율성", "지속가능성"): + self.assertIn(f"기준별 상세 평정 누락: {criterion}", result.stderr) + + def test_negated_unavailable_sentence_does_not_replace_fact_evidence(self): + text = read("tests", "fixtures", "auditable-brief-clean.md").replace( + "| 수요 | 부합 | [사실: F1] | [근거: 보고서 p.10] | 주민자료 없음 | 없음 | 확인됨 |", + "기준 전체 평가 불가 아님. 수요는 부합한다고 평가함.", + 1, + ).replace( + "| 조율 | 일부 | [사실: F2] |", + "| 조율 | 일부 | [사실: F2] [사실: F1] |", + 1, + ) + result = run_text(text) + self.assertEqual(result.returncode, 2, result.stdout + result.stderr) + self.assertIn("적절성 상세 평정", result.stderr) + def test_conflict_marker_missing_from_register_fails(self): result = run_fixture("auditable-brief-missing-register.md") self.assertEqual(result.returncode, 2, result.stdout + result.stderr) self.assertIn("X1", result.stderr) self.assertIn("상충 등록부", result.stderr) + def test_negated_no_conflict_sentence_does_not_pass(self): + text = read("tests", "fixtures", "auditable-brief-clean.md") + text = text.replace( + "| F3 | 현지연수 실시 횟수 실적 | 5 | 회 | 현지연수 | 사업누적 | 행사 횟수 | 2023-12-01 | 종료보고서 v1 | p.20 본문 | [상충: X1] |", + "| F3 | 현지연수 실시 횟수 실적 | 5 | 회 | 현지연수 | 사업누적 | 행사 횟수 | 2023-12-01 | 종료보고서 v1 | p.20 본문 | 일치 |", + 1, + ).replace( + "| F3 | 현지연수 실시 횟수 실적 | 6 | 회 | 현지연수 | 사업누적 | 행사 횟수 | 2023-12-01 | 종료보고서 v1 | p.21 표 | [상충: X1] |", + "| F3 | 현지연수 실시 횟수 실적 | 5 | 회 | 현지연수 | 사업누적 | 행사 횟수 | 2023-12-01 | 종료보고서 v1 | p.21 표 | 일치 |", + 1, + ).replace(" [상충: X1]", "", 1) + old_register = ( + "| ID | 사실 ID | 쟁점 | 값·진술 A(원문 위치) | 값·진술 B(원문 위치) | 대조 결과·가능한 설명 | 해결 상태 | 점수·결론 영향 | 후속 확인 |\n" + "|---|---|---|---|---|---|---|---|---|\n" + "| X1 | [사실: F3] | 연수 횟수 | 5회(p.20) | 6회(p.21) | 집계규칙 설명 없음 | 미해결 | 효과성 4→3 | 원자료 확인 |" + ) + text = text.replace( + old_register, + "중대 상충 없음이라고 확인할 수 없으며 반복 표기 점검도 수행하지 못했다.", + 1, + ) + result = run_text(text) + self.assertEqual(result.returncode, 2, result.stdout + result.stderr) + self.assertIn("중대 상충 없음", result.stderr) + def test_terse_score_only_answer_fails(self): result = run_fixture("auditable-brief-terse.md") self.assertEqual(result.returncode, 2, result.stdout + result.stderr) @@ -51,14 +284,45 @@ def test_unfilled_template_fails(self): self.assertEqual(result.returncode, 2, result.stdout + result.stderr) self.assertIn("자리표시자", result.stderr) + def test_unfilled_ledger_fails_early_gate(self): + result = subprocess.run( + [ + sys.executable, + CHECKER, + "--ledger-only", + os.path.join(ROOT, "templates", "auditable-evaluation-brief-template.md"), + ], + check=False, + text=True, + capture_output=True, + ) + self.assertEqual(result.returncode, 2, result.stdout + result.stderr) + self.assertIn("원문 사실대장", result.stderr) + self.assertIn("자리표시자", result.stderr) + + def test_detail_fact_reference_missing_from_ledger_fails(self): + text = read("tests", "fixtures", "auditable-brief-clean.md").replace( + "[사실: F5] | [근거: 보고서 p.40]", + "[사실: F99] | [근거: 보고서 p.40]", + 1, + ) + result = run_text(text) + self.assertEqual(result.returncode, 2, result.stdout + result.stderr) + self.assertIn("F99", result.stderr) + self.assertIn("원문 사실대장", result.stderr) + class RuntimeContract(unittest.TestCase): def test_evaluate_skill_routes_template_and_checker(self): skill = read("skills", "evaluate", "SKILL.md") self.assertIn("auditable-evaluation-brief-template.md", skill) self.assertIn("auditable_output_check.py", skill) + self.assertIn("원문 사실대장", skill) + self.assertIn("--ledger-only", skill) + self.assertIn("[사실: Fn]", skill) self.assertIn("상충·불일치 등록부", skill) self.assertIn("검증 후", skill) + self.assertLess(skill.index("--ledger-only"), skill.index("기준 평가관 병렬 위임")) def test_all_criterion_agents_capture_counterevidence_and_conflicts(self): agents = ( @@ -79,11 +343,37 @@ def test_all_criterion_agents_capture_counterevidence_and_conflicts(self): def test_verifier_aggregates_conflicts_and_corrected_scores(self): ko = read("agents", "quality-verifier.md") en = read("docs", "en", "agents", "quality-verifier.md") - for token in ("중대 상충", "상충·불일치 등록부", "검증 후 점수", "조건부 통과"): + for token in ( + "원문 사실대장", + "한 원문 표기당 한 행", + "[사실: Fn]", + "중대 상충", + "상충·불일치 등록부", + "검증 후 점수", + "조건부 통과", + ): self.assertIn(token, ko) - for token in ("material conflict", "Conflict and inconsistency register", "post-verification score"): + for token in ( + "source-fact ledger", + "one row per source occurrence", + "[Fact: Fn]", + "material conflict", + "Conflict and inconsistency register", + "post-verification score", + ): self.assertIn(token, en) + def test_fallback_and_plan_require_early_fact_ledger(self): + ko = read("AGENTS.md") + en = read("docs", "en", "AGENTS.md") + plan = read("templates", "eval-plan-template.md") + for token in ("원문 사실대장", "--ledger-only", "[사실: Fn]"): + self.assertIn(token, ko) + for token in ("source-fact ledger", "--ledger-only", "[Fact: Fn]"): + self.assertIn(token, en) + self.assertIn("원문 사실대장 작성", plan) + self.assertIn("--ledger-only", plan) + def test_known_conflict_classes_are_routed_to_the_right_roles(self): skill = read("skills", "evaluate", "SKILL.md") for token in ("단위", "분모", "집계규칙", "상태기준일"): @@ -99,6 +389,30 @@ def test_fallback_and_report_path_preserve_audit_trail(self): self.assertIn("상충·불일치 등록부", read("templates", "evaluation-report-template.md")) self.assertIn("상충·불일치 등록부", read("skills", "write-report", "SKILL.md")) + def test_report_path_preserves_source_fact_ledger(self): + active_paths = ( + ("skills", "write-report", "SKILL.md"), + ("agents", "report-composer.md"), + ("agents", "narrative-verifier.md"), + ("templates", "evaluation-report-template.md"), + ) + for path in active_paths: + body = read(*path) + self.assertIn("원문 사실대장", body, "/".join(path)) + self.assertIn("[사실: Fn]", body, "/".join(path)) + for path in ( + ("docs", "en", "agents", "report-composer.md"), + ("docs", "en", "agents", "narrative-verifier.md"), + ): + body = read(*path) + self.assertIn("source-fact ledger", body.lower(), "/".join(path)) + self.assertIn("[Fact: Fn]", body, "/".join(path)) + + def test_open_runner_requests_fact_ledger_traceability(self): + runner = read("scripts", "open_runner.py") + self.assertIn("F-ID 사실대장", runner) + self.assertIn("[사실: Fn]", runner) + def test_nonstandard_cts_criterion_is_not_routed(self): self.assertFalse(os.path.exists(os.path.join(ROOT, "agents", "cts-validity-evaluator.md"))) active_paths = ( From ec7b09a028aeea0c996738799c9ae90017bea10a Mon Sep 17 00:00:00 2001 From: amnotyoung Date: Wed, 5 Aug 2026 16:14:42 +0900 Subject: [PATCH 2/2] ci: reject duplicate files and canonical ID collisions --- .github/workflows/checks.yml | 8 +- .repository-hygiene-allow | 3 + CHANGELOG.md | 4 + CONTRIBUTING.md | 12 +- scripts/repository_hygiene_check.py | 479 ++++++++++++++++++++++++++++ tests/test_repository_hygiene.py | 271 ++++++++++++++++ 6 files changed, 774 insertions(+), 3 deletions(-) create mode 100644 .repository-hygiene-allow create mode 100755 scripts/repository_hygiene_check.py create mode 100644 tests/test_repository_hygiene.py diff --git a/.github/workflows/checks.yml b/.github/workflows/checks.yml index 10bbf51..8c6e742 100644 --- a/.github/workflows/checks.yml +++ b/.github/workflows/checks.yml @@ -1,7 +1,8 @@ name: checks # 결정적 컴포넌트의 회귀를 막는다 — 수치 일관성 검사기(consistency_check.py), -# 감사 가능한 산출물 계약, 공용 지식층 라우팅, 완료 엔진(hooks/boulder.sh), 이중 매니페스트 정체성. +# 감사 가능한 산출물 계약, 공용 지식층 라우팅, 저장소 중복·정본 ID, +# 완료 엔진(hooks/boulder.sh), 이중 매니페스트 정체성. # 픽스처는 실제 KOICA 종료평가 PDF 334건 스윕에서 관측된 사고·오탐 유형이다. on: @@ -18,7 +19,10 @@ jobs: steps: - uses: actions/checkout@v4 - - name: Python 회귀 테스트 (수치 검사기 + 감사 산출물 계약 + 지식층) + - name: 저장소 중복 사본·정본 ID 충돌 검사 + run: python3 scripts/repository_hygiene_check.py + + - name: Python 회귀 테스트 (수치 검사기 + 감사 산출물 계약 + 지식층 + 저장소 위생) run: python3 -m unittest discover -s tests -v - name: 완료 엔진(Stop hook) 동작 테스트 diff --git a/.repository-hygiene-allow b/.repository-hygiene-allow new file mode 100644 index 0000000..6375efe --- /dev/null +++ b/.repository-hygiene-allow @@ -0,0 +1,3 @@ +# 정확히 의도된 사본형 경로만 한 줄에 하나씩 적는다. +# 예: docs/phase 2.md +# 비워 두는 것이 기본이며, 예외 추가는 코드리뷰에서 정본과의 공존 이유를 설명한다. diff --git a/CHANGELOG.md b/CHANGELOG.md index ca9b1cd..2ee2ea8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -17,6 +17,10 @@ each slice below is recorded as a 0.x milestone. period, aggregation rule, or as-of date require a supported explanation or a linked conflict ID; final briefs cross-check `F` IDs across the ledger, detailed ratings, and conflict register. +- **Repository hygiene CI** — checks now reject copy-suffixed files and + directories when a canonical sibling exists, portable case/NFC path + collisions, and duplicate or path-mismatched agent/skill identities. An + exact-path allowlist documents the rare intentional suffix-shaped name. ## [0.12.0] — 2026-08-05 — Auditable briefs & normative criteria cleanup diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index d4176a0..3f5c58b 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -32,6 +32,13 @@ evaluation. Contributions of all sizes are welcome. anonymization principles intact (see [`docs/do-no-harm.md`](docs/do-no-harm.md)). 4. **Keep it model-agnostic.** Avoid hard dependencies on any single proprietary LLM or harness. If a change assumes one, note how it degrades/adapts on others. +5. **Keep canonical plugin identities unique.** An agent's frontmatter `name` + must exactly match its `agents/.md` filename; a skill's `name` must + exactly match its `skills//` directory. Do not commit operating-system + or sync-tool copies such as `file 2.md`, `file copy.md`, or `file (1).md` + alongside their canonical sibling. If a suffix-shaped path is genuinely + intentional, add its exact repository-relative path to + `.repository-hygiene-allow` and explain the exception in review. ## Process / 절차 @@ -43,7 +50,10 @@ evaluation. Contributions of all sizes are welcome. same PR — the `mirror-sync` check enforces this, and a deliberately one-sided PR needs the `mirror-sync-exempt` label. Run it yourself with `bash scripts/check-mirror-sync.sh` (or `--audit` for the whole repo). -4. Open a pull request with a clear description and, for methodology changes, the +4. Run the deterministic regression and repository-hygiene checks: + `python3 -m unittest discover -s tests -v` and + `python3 scripts/repository_hygiene_check.py`. +5. Open a pull request with a clear description and, for methodology changes, the supporting citation. ## Licensing of contributions / 기여물 라이선스 diff --git a/scripts/repository_hygiene_check.py b/scripts/repository_hygiene_check.py new file mode 100755 index 0000000..bafc694 --- /dev/null +++ b/scripts/repository_hygiene_check.py @@ -0,0 +1,479 @@ +#!/usr/bin/env python3 +"""추적 예정 파일의 중복 사본과 플러그인 정본 ID 충돌을 검사한다. + +파일시스템 전체를 훑지 않는다. Git이 추적하거나 추적할 예정이면서 ignore되지 +않은 파일만 검사해 ``.local/`` 자료, 로컬 worktree, ``.DS_Store`` 같은 개발자별 +파일이 CI 결과에 섞이지 않게 한다. + +종료 코드: + 0 = 통과 + 1 = Git/파일 읽기 오류 + 2 = 저장소 위생 계약 위반 +""" + +from __future__ import annotations + +import argparse +import hashlib +import os +from dataclasses import dataclass +from pathlib import Path, PurePosixPath +import re +import subprocess +import sys +import unicodedata +from collections import defaultdict +from collections.abc import Iterable, Sequence + + +ROOT = Path(__file__).resolve().parent.parent +COPY_ALLOWLIST = ".repository-hygiene-allow" + +# 운영체제·동기화 도구가 만드는 전형적인 사본 접미사. 숫자 제목의 오탐을 +# 피하려고 이 패턴만으로 실패시키지 않고, 접미사를 뺀 sibling이 실제로 있을 +# 때만 사본 충돌로 판정한다. +COPY_COMPONENT_RE = re.compile( + r"^(?P.+?)(?: - | )" + r"(?P(?:[2-9]|[1-9][0-9]+)|copy(?: [0-9]+)?|복사본(?: [0-9]+)?|\([1-9][0-9]*\))" + r"(?P(?:\.[^./]+)*)$", + re.IGNORECASE, +) +CANONICAL_ID_RE = re.compile(r"^[a-z0-9]+(?:-[a-z0-9]+)*$") +NAME_LINE_RE = re.compile(r"^name:\s*(.*?)\s*$") + + +class InventoryError(RuntimeError): + """Git 인벤토리를 만들지 못했을 때 발생한다.""" + + +@dataclass(frozen=True, order=True) +class Issue: + """결정적으로 정렬·출력할 수 있는 단일 위반.""" + + code: str + path: str + message: str + + +@dataclass(frozen=True) +class CanonicalEntry: + path: str + declared_name: str + expected_name: str + body: str + + +def _sort_key(path: str) -> tuple[str, str]: + normalized = unicodedata.normalize("NFC", path).casefold() + return normalized, path + + +def git_inventory(root: Path) -> list[str]: + """Git 추적 파일과 ignore되지 않은 미추적 파일을 NUL 안전하게 반환한다.""" + + try: + result = subprocess.run( + [ + "git", + "ls-files", + "--cached", + "--others", + "--exclude-standard", + "-z", + "--", + ], + cwd=root, + check=True, + stdout=subprocess.PIPE, + stderr=subprocess.PIPE, + ) + except (OSError, subprocess.CalledProcessError) as exc: + detail = getattr(exc, "stderr", b"") + if isinstance(detail, bytes): + detail = detail.decode("utf-8", errors="replace") + message = str(detail).strip() or str(exc) + raise InventoryError(f"Git 파일 인벤토리를 만들 수 없다: {message}") from exc + + paths: set[str] = set() + for raw_path in result.stdout.split(b"\0"): + if not raw_path: + continue + path = raw_path.decode("utf-8", errors="surrogateescape") + absolute = root / PurePosixPath(path) + # 삭제 중인 tracked 파일은 현재 작업트리 계약의 대상이 아니다. 심볼릭 + # 링크는 깨져 있어도 Git 엔트리로 남아 있으면 경로 충돌 검사에 포함한다. + if absolute.exists() or absolute.is_symlink(): + paths.add(PurePosixPath(path).as_posix()) + return sorted(paths, key=_sort_key) + + +def _all_nodes(paths: Iterable[str]) -> set[str]: + """파일과 그 상위 디렉터리의 POSIX 경로 집합을 만든다.""" + + nodes: set[str] = set() + for path in paths: + parts = PurePosixPath(path).parts + for end in range(1, len(parts) + 1): + nodes.add(PurePosixPath(*parts[:end]).as_posix()) + return nodes + + +def _portable_key(path: str) -> str: + return "/".join( + unicodedata.normalize("NFC", part).casefold() + for part in PurePosixPath(path).parts + ) + + +def _copy_collisions( + paths: Sequence[str], allowed_copy_paths: Iterable[str] = () +) -> list[Issue]: + nodes = _all_nodes(paths) + nodes_by_portable_key: dict[str, set[str]] = defaultdict(set) + for node in nodes: + nodes_by_portable_key[_portable_key(node)].add(node) + allowed_keys = {_portable_key(path) for path in allowed_copy_paths} + # 복제 디렉터리 안 파일이 여러 개여도 같은 디렉터리 사고는 한 번만 보고한다. + collisions: dict[tuple[str, str], str] = {} + + for path in paths: + parts = PurePosixPath(path).parts + for index, component in enumerate(parts): + match = COPY_COMPONENT_RE.fullmatch(component) + if not match: + continue + canonical_component = match.group("base") + match.group("extensions") + copied_prefix = PurePosixPath(*parts[: index + 1]).as_posix() + canonical_prefix = PurePosixPath( + *parts[:index], canonical_component + ).as_posix() + if _portable_key(copied_prefix) in allowed_keys: + continue + canonical_matches = nodes_by_portable_key.get( + _portable_key(canonical_prefix), set() + ) + if canonical_matches: + canonical = sorted(canonical_matches, key=_sort_key)[0] + collisions.setdefault((copied_prefix, canonical), path) + + return [ + Issue( + "copy-suffix-collision", + witness, + f"사본형 경로 '{copied}'가 정본형 sibling '{canonical}'와 충돌한다.", + ) + for (copied, canonical), witness in sorted(collisions.items()) + ] + + +def _portable_collisions(paths: Sequence[str]) -> list[Issue]: + groups: dict[str, set[str]] = defaultdict(set) + for node in _all_nodes(paths): + groups[_portable_key(node)].add(node) + + issues: list[Issue] = [] + for members in groups.values(): + if len(members) < 2: + continue + ordered = sorted(members, key=_sort_key) + issues.append( + Issue( + "portable-path-collision", + ordered[0], + "NFC·대소문자 정규화 후 같은 경로가 된다: " + ", ".join(ordered), + ) + ) + return issues + + +def _frontmatter( + text: str, *, allow_reference_preamble: bool = False +) -> tuple[str | None, str | None]: + """(name, body)를 반환한다. 형식이 잘못됐으면 (None, None)이다.""" + + normalized = text.replace("\r\n", "\n").replace("\r", "\n") + lines = normalized.split("\n") + if not lines: + return None, None + + # 실행 정본은 첫 줄부터 frontmatter여야 한다. docs/en/agents 번역 미러만 + # 정본을 가리키는 blockquote/빈줄 preamble을 첫 8줄 안에서 허용한다. + if allow_reference_preamble: + start = None + for index, line in enumerate(lines[:8]): + if line.lstrip("\ufeff") == "---": + start = index + break + if line and not line.lstrip("\ufeff").startswith(">"): + return None, None + if start is None: + return None, None + else: + if lines[0].lstrip("\ufeff") != "---": + return None, None + start = 0 + + try: + end = lines.index("---", start + 1) + except ValueError: + return None, None + + names: list[str] = [] + for line in lines[start + 1 : end]: + match = NAME_LINE_RE.fullmatch(line) + if match: + value = match.group(1).strip() + if len(value) >= 2 and value[0] == value[-1] and value[0] in "\"'": + value = value[1:-1] + names.append(value) + if len(names) != 1: + return None, None + + body = "\n".join(lines[end + 1 :]).rstrip("\n") + return names[0], body + + +def _canonical_specs(path: str) -> tuple[str, str] | None: + """(registry, expected_name)을 반환한다.""" + + parts = PurePosixPath(path).parts + if len(parts) == 2 and parts[0] == "agents" and parts[1].endswith(".md"): + return "agents", PurePosixPath(parts[1]).stem + if ( + len(parts) == 4 + and parts[:3] == ("docs", "en", "agents") + and parts[3].endswith(".md") + ): + # 영문 미러는 정본과 같은 ID를 갖지만 별도 registry로 검사한다. + return "docs/en/agents", PurePosixPath(parts[3]).stem + if len(parts) == 3 and parts[0] == "skills" and parts[2] == "SKILL.md": + return "skills", parts[1] + return None + + +def _semantic_id(value: str) -> str: + value = unicodedata.normalize("NFKC", value).casefold() + value = re.sub(r"[\s_]+", "-", value) + return re.sub(r"-+", "-", value).strip("-") + + +def _canonical_entries(root: Path, paths: Sequence[str]) -> tuple[list[CanonicalEntry], list[Issue]]: + entries: list[CanonicalEntry] = [] + issues: list[Issue] = [] + + skill_dirs = { + PurePosixPath(path).parts[1] + for path in paths + if len(PurePosixPath(path).parts) >= 3 + and PurePosixPath(path).parts[0] == "skills" + } + path_set = set(paths) + for skill_dir in sorted(skill_dirs, key=_sort_key): + entrypoint = f"skills/{skill_dir}/SKILL.md" + if entrypoint not in path_set: + issues.append( + Issue( + "missing-skill-entrypoint", + f"skills/{skill_dir}", + f"스킬 디렉터리에 정본 진입점이 없다: {entrypoint}", + ) + ) + + for path in paths: + spec = _canonical_specs(path) + if spec is None: + continue + registry, expected_name = spec + try: + text = (root / PurePosixPath(path)).read_text(encoding="utf-8") + except (OSError, UnicodeError) as exc: + issues.append( + Issue( + "unreadable-canonical-file", + path, + f"정본 UTF-8 파일을 읽을 수 없다: {exc}", + ) + ) + continue + + declared_name, body = _frontmatter( + text, allow_reference_preamble=registry == "docs/en/agents" + ) + if declared_name is None or body is None: + issues.append( + Issue( + "invalid-frontmatter-name", + path, + "YAML frontmatter에 단 하나의 평문 name 필드가 있어야 한다.", + ) + ) + continue + + if not CANONICAL_ID_RE.fullmatch(declared_name): + issues.append( + Issue( + "invalid-canonical-id", + path, + f"정본 ID '{declared_name}'는 소문자 kebab-case여야 한다.", + ) + ) + if declared_name != expected_name: + issues.append( + Issue( + "canonical-id-path-mismatch", + path, + f"frontmatter name '{declared_name}'가 경로 ID '{expected_name}'와 다르다.", + ) + ) + entries.append(CanonicalEntry(path, declared_name, expected_name, body)) + + return entries, issues + + +def _canonical_collisions(entries: Sequence[CanonicalEntry]) -> list[Issue]: + by_registry: dict[str, list[CanonicalEntry]] = defaultdict(list) + for entry in entries: + spec = _canonical_specs(entry.path) + if spec is not None: + by_registry[spec[0]].append(entry) + + issues: list[Issue] = [] + for registry, registry_entries in sorted(by_registry.items()): + id_groups: dict[str, list[CanonicalEntry]] = defaultdict(list) + body_groups: dict[str, list[CanonicalEntry]] = defaultdict(list) + for entry in registry_entries: + id_groups[_semantic_id(entry.declared_name)].append(entry) + digest = hashlib.sha256(entry.body.encode("utf-8")).hexdigest() + body_groups[digest].append(entry) + + for entries_with_id in id_groups.values(): + if len(entries_with_id) < 2: + continue + ordered = sorted(entries_with_id, key=lambda entry: _sort_key(entry.path)) + paths = [entry.path for entry in ordered] + issues.append( + Issue( + "duplicate-canonical-id", + paths[0], + f"{registry} registry에서 정규화 ID가 중복된다: " + ", ".join(paths), + ) + ) + + for entries_with_body in body_groups.values(): + if len(entries_with_body) < 2: + continue + ordered = sorted(entries_with_body, key=lambda entry: _sort_key(entry.path)) + paths = [entry.path for entry in ordered] + issues.append( + Issue( + "duplicate-canonical-body", + paths[0], + f"{registry} registry에서 frontmatter를 제외한 본문이 같다: " + + ", ".join(paths), + ) + ) + return issues + + +def inspect( + root: Path, + paths: Sequence[str], + allowed_copy_paths: Iterable[str] = (), +) -> list[Issue]: + """주어진 저장소 상대경로들을 검사한다. 테스트에서 순수하게 재사용한다.""" + + normalized_paths = sorted( + {PurePosixPath(path).as_posix() for path in paths}, key=_sort_key + ) + entries, entry_issues = _canonical_entries(root, normalized_paths) + issues = [ + *_copy_collisions(normalized_paths, allowed_copy_paths), + *_portable_collisions(normalized_paths), + *entry_issues, + *_canonical_collisions(entries), + ] + return sorted(set(issues)) + + +def load_copy_allowlist(root: Path) -> set[str]: + """검토 가능한 exact-path 사본 예외 목록을 읽는다.""" + + path = root / COPY_ALLOWLIST + try: + lines = path.read_text(encoding="utf-8").splitlines() + except FileNotFoundError: + return set() + except (OSError, UnicodeError) as exc: + raise InventoryError(f"{COPY_ALLOWLIST}를 읽을 수 없다: {exc}") from exc + + allowed: set[str] = set() + for line_number, raw_line in enumerate(lines, 1): + value = raw_line.strip() + if not value or value.startswith("#"): + continue + candidate = PurePosixPath(value) + if candidate.is_absolute() or ".." in candidate.parts or candidate.as_posix() in ("", "."): + raise InventoryError( + f"{COPY_ALLOWLIST}:{line_number}은 저장소 상대경로여야 한다: {value}" + ) + allowed.add(candidate.as_posix()) + return allowed + + +def _gha_property(value: str) -> str: + return ( + value.replace("%", "%25") + .replace("\r", "%0D") + .replace("\n", "%0A") + .replace(":", "%3A") + .replace(",", "%2C") + ) + + +def _gha_message(value: str) -> str: + return value.replace("%", "%25").replace("\r", "%0D").replace("\n", "%0A") + + +def report(issues: Sequence[Issue]) -> None: + github_actions = bool(os.environ.get("GITHUB_ACTIONS")) + for issue in issues: + message = f"[{issue.code}] {issue.message}" + if github_actions: + print( + f"::error file={_gha_property(issue.path)}::{_gha_message(message)}", + file=sys.stderr, + ) + else: + print(f"ERROR [{issue.code}] {issue.path}: {issue.message}", file=sys.stderr) + + +def main(argv: Sequence[str] | None = None) -> int: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument( + "--root", + type=Path, + default=ROOT, + help="검사할 Git 저장소 루트(기본: 이 스크립트의 상위 저장소)", + ) + args = parser.parse_args(argv) + root = args.root.resolve() + + try: + paths = git_inventory(root) + allowed_copy_paths = load_copy_allowlist(root) + except InventoryError as exc: + print(f"repository hygiene check 오류: {exc}", file=sys.stderr) + return 1 + + issues = inspect(root, paths, allowed_copy_paths) + if issues: + report(issues) + print(f"저장소 위생 계약 위반 {len(issues)}건.", file=sys.stderr) + return 2 + + print(f"저장소 위생 검사 통과 — Git 대상 파일 {len(paths)}개, 충돌 0건.") + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/tests/test_repository_hygiene.py b/tests/test_repository_hygiene.py new file mode 100644 index 0000000..6a3776e --- /dev/null +++ b/tests/test_repository_hygiene.py @@ -0,0 +1,271 @@ +"""중복 사본·플러그인 정본 ID 충돌 검사기의 회귀 테스트.""" + +from pathlib import Path +import subprocess +import sys +import tempfile +import unittest + + +ROOT = Path(__file__).resolve().parent.parent +sys.path.insert(0, str(ROOT / "scripts")) + +import repository_hygiene_check as hygiene # noqa: E402 + + +def frontmatter(name, body): + return f"---\nname: {name}\ndescription: test\n---\n\n{body}\n" + + +class Files: + """임시 파일 집합을 만들고 검사하는 작은 테스트 헬퍼.""" + + def __init__(self, files): + self.temp = tempfile.TemporaryDirectory() + self.root = Path(self.temp.name) + self.paths = sorted(files) + for relative, content in files.items(): + target = self.root / relative + target.parent.mkdir(parents=True, exist_ok=True) + target.write_text(content, encoding="utf-8") + + def inspect(self, allowed_copy_paths=()): + return hygiene.inspect(self.root, self.paths, allowed_copy_paths) + + def close(self): + self.temp.cleanup() + + +class RepositoryContract(unittest.TestCase): + def test_current_repository_passes(self): + paths = hygiene.git_inventory(ROOT) + allowed_copy_paths = hygiene.load_copy_allowlist(ROOT) + issues = hygiene.inspect(ROOT, paths, allowed_copy_paths) + self.assertEqual(issues, [], "\n".join(map(str, issues))) + + def test_git_inventory_excludes_ignored_copy_files(self): + with tempfile.TemporaryDirectory() as directory: + root = Path(directory) + subprocess.run(["git", "init", "-q"], cwd=root, check=True) + (root / ".gitignore").write_text(".local/\n", encoding="utf-8") + (root / "base.md").write_text("canonical\n", encoding="utf-8") + ignored = root / ".local" / "base 2.md" + ignored.parent.mkdir() + ignored.write_text("copy\n", encoding="utf-8") + + paths = hygiene.git_inventory(root) + self.assertIn("base.md", paths) + self.assertNotIn(".local/base 2.md", paths) + + +class CopySuffixContract(unittest.TestCase): + def assert_issue(self, issues, code): + self.assertIn(code, {issue.code for issue in issues}, issues) + + def test_numbered_file_with_canonical_sibling_fails(self): + files = Files( + { + "scripts/check-mirror-sync.sh": "canonical\n", + "scripts/check-mirror-sync 2.sh": "copy\n", + } + ) + try: + self.assert_issue(files.inspect(), "copy-suffix-collision") + finally: + files.close() + + def test_copied_skill_directory_fails(self): + files = Files( + { + "skills/evaluate/SKILL.md": frontmatter("evaluate", "# Evaluate"), + "skills/evaluate 2/SKILL.md": frontmatter("evaluate-2", "# Copy"), + } + ) + try: + self.assert_issue(files.inspect(), "copy-suffix-collision") + finally: + files.close() + + def test_two_digit_copy_suffix_fails(self): + files = Files( + { + "docs/guide.md": "canonical\n", + "docs/guide 10.md": "copy\n", + } + ) + try: + self.assert_issue(files.inspect(), "copy-suffix-collision") + finally: + files.close() + + def test_copy_suffix_finds_case_variant_canonical_sibling(self): + files = Files( + { + "docs/Guide.md": "canonical\n", + "docs/guide copy.md": "copy\n", + } + ) + try: + self.assert_issue(files.inspect(), "copy-suffix-collision") + finally: + files.close() + + def test_windows_style_copy_suffix_fails(self): + files = Files( + { + "docs/guide.md": "canonical\n", + "docs/guide - Copy.md": "copy\n", + } + ) + try: + self.assert_issue(files.inspect(), "copy-suffix-collision") + finally: + files.close() + + def test_explicit_exact_path_allowlist_permits_intentional_numeric_pair(self): + files = Files( + { + "docs/phase.md": "overview\n", + "docs/phase 2.md": "second phase\n", + } + ) + try: + issues = files.inspect({"docs/phase 2.md"}) + self.assertEqual(issues, []) + finally: + files.close() + + def test_standalone_numeric_title_passes(self): + with tempfile.TemporaryDirectory() as directory: + issues = hygiene.inspect(Path(directory), ["docs/phase 2.md"]) + self.assertEqual(issues, []) + + def test_intentional_pairs_and_mirrors_pass(self): + files = Files( + { + "assets/logo.svg": "svg\n", + "assets/logo.png": "png\n", + "README.md": "English\n", + "README.ko.md": "Korean\n", + "agents/example-agent.md": frontmatter("example-agent", "# 한국어"), + "docs/en/agents/example-agent.md": ( + "> English reference translation of the canonical agent.\n\n" + + frontmatter("example-agent", "# English") + ), + } + ) + try: + self.assertEqual(files.inspect(), []) + finally: + files.close() + + +class CanonicalIdentityContract(unittest.TestCase): + def assert_issue(self, issues, code): + self.assertIn(code, {issue.code for issue in issues}, issues) + + def test_duplicate_agent_id_fails(self): + files = Files( + { + "agents/alpha.md": frontmatter("alpha", "# Alpha"), + "agents/beta.md": frontmatter("alpha", "# Beta"), + } + ) + try: + issues = files.inspect() + self.assert_issue(issues, "duplicate-canonical-id") + self.assert_issue(issues, "canonical-id-path-mismatch") + finally: + files.close() + + def test_semantically_equivalent_agent_ids_fail(self): + files = Files( + { + "agents/example-agent.md": frontmatter("example-agent", "# Hyphen"), + "agents/example_agent.md": frontmatter("example_agent", "# Underscore"), + } + ) + try: + issues = files.inspect() + self.assert_issue(issues, "invalid-canonical-id") + self.assert_issue(issues, "duplicate-canonical-id") + finally: + files.close() + + def test_skill_name_must_match_directory(self): + files = Files( + { + "skills/evaluate/SKILL.md": frontmatter("quality-review", "# Evaluate") + } + ) + try: + self.assert_issue(files.inspect(), "canonical-id-path-mismatch") + finally: + files.close() + + def test_executable_agent_rejects_reference_preamble(self): + files = Files( + { + "agents/example.md": ( + "> reference-only preamble\n\n" + + frontmatter("example", "# Agent") + ) + } + ) + try: + self.assert_issue(files.inspect(), "invalid-frontmatter-name") + finally: + files.close() + + def test_skill_frontmatter_must_start_on_first_line(self): + files = Files( + {"skills/example/SKILL.md": "\n" + frontmatter("example", "# Skill")} + ) + try: + self.assert_issue(files.inspect(), "invalid-frontmatter-name") + finally: + files.close() + + def test_same_skill_body_after_frontmatter_fails(self): + files = Files( + { + "skills/alpha/SKILL.md": frontmatter("alpha", "# Same body"), + "skills/beta/SKILL.md": frontmatter("beta", "# Same body"), + } + ) + try: + self.assert_issue(files.inspect(), "duplicate-canonical-body") + finally: + files.close() + + def test_missing_skill_entrypoint_fails(self): + files = Files({"skills/example/references/guide.md": "support\n"}) + try: + self.assert_issue(files.inspect(), "missing-skill-entrypoint") + finally: + files.close() + + def test_direct_skills_document_is_not_a_skill_directory(self): + files = Files({"skills/README.md": "overview\n"}) + try: + self.assertEqual(files.inspect(), []) + finally: + files.close() + + def test_portable_case_collision_fails(self): + # 합성 경로를 직접 넘겨 대소문자 비구분 파일시스템에서도 재현 가능하게 한다. + with tempfile.TemporaryDirectory() as directory: + issues = hygiene.inspect(Path(directory), ["Docs/A.md", "docs/a.md"]) + self.assert_issue(issues, "portable-path-collision") + + def test_portable_unicode_normalization_collision_fails(self): + # 합성 경로를 사용해 호스트 파일시스템의 NFC/NFD 저장 방식과 무관하게 검사한다. + with tempfile.TemporaryDirectory() as directory: + issues = hygiene.inspect( + Path(directory), ["docs/caf\u00e9.md", "docs/cafe\u0301.md"] + ) + self.assert_issue(issues, "portable-path-collision") + + +if __name__ == "__main__": + unittest.main()