Skip to content

Add aisearch-v2 Korean analyzer/synonym comparison guide - #32

Closed
daeungo1 wants to merge 1 commit into
mainfrom
feature/aisearch-v2-analyzer-guide
Closed

daeungo1 wants to merge 1 commit into
mainfrom
feature/aisearch-v2-analyzer-guide

Conversation

@daeungo1

@daeungo1 daeungo1 commented Jul 6, 2026 •

Copy link
Copy Markdown
Collaborator

개요

Azure AI Search에서 필드 분석기(Analyzer) 선택이 한국어 검색 품질(recall/precision)에 미치는 영향을 실제 리소스로 재현·검증한 고객 가이드입니다. ko.microsoft가 복합명사·조사 처리에서 가장 우수함을, 정답/오답(distractor) 데이터로 정량 검증합니다.

포함 내용

  • analyzer_test.py — 4개 분석기(ko.microsoft, ko.lucene, standard.lucene, keyword) recall/precision 비교 (인덱스 생성 → 업로드 → Analyze 토큰 비교 → 검색 히트 → recall 리포트)
  • synonym_test.py — 동의어 맵 효과 검증 (동의어는 분석기를 대체하지 않고 보완함을 실측)
  • sample_data.json — 정답 6건 + distractor 4건 = 10건
  • README.md — 진단 → 조치 → 검증 playbook + FAQ
  • requirements.txt — 버전 고정: azure-search-documents==12.0.0, azure-core==1.41.0, python-dotenv==1.2.2
  • 포탈 스크린샷 2장, .env.example, .gitignore

검증 결과 (실제 실행, 두 스크립트 모두 EXIT=0)

쿼리 변형 완벽 recall ko.microsoft ko.lucene standard.lucene keyword
analyzer_test 5/5 4/5 0/5 0/5
  • synonym_test: msft+syn 5/5, standard 0/5, 문서에 없는 별칭 컬처랜드 0→3 회수
  • 결론: 한국어 필드는 ko.microsoft + 동의어 맵 조합 권장 (analyzer 변경 시 재색인 필수)

비고

  • .env(비밀키)와 .venv/는 .gitignore로 제외됨
  • SDK azure-search-documents 12.0.0(major)의 breaking change는 본 스크립트가 사용하는 API에 영향 없음을 확인

Summary by CodeRabbit

  • New Features

    • Added a Korean-language guide for comparing Azure AI Search analyzer behavior, including setup steps, validation results, and a customer-facing troubleshooting flow.
    • Added sample data plus two new scripts to compare search results and synonym handling across different analyzer configurations.
  • Documentation

    • Expanded setup guidance with environment variable examples and dependency notes.
  • Chores

    • Updated project ignore rules and pinned Python package versions for a more consistent local setup.

- analyzer_test.py: recall/precision comparison across ko.microsoft, ko.lucene, standard.lucene, keyword analyzers
- synonym_test.py: synonym map effect (complements, not replaces, analyzer)
- sample_data.json: 6 relevant + 4 distractor docs
- README.md: customer diagnosis/remediation/verification playbook
- deps pinned: azure-search-documents==12.0.0, azure-core==1.41.0, python-dotenv==1.2.2
@coderabbitai

coderabbitai Bot commented Jul 6, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Adds a new aisearch-v2 project comparing Azure AI Search Korean analyzers. Includes environment/config files, sample data, two standalone Python scripts benchmarking analyzer recall and synonym map effectiveness, pinned dependencies, and a Korean-language README documenting the methodology and results.

Changes

Azure AI Search Korean Analyzer Comparison

Layer / File(s) Summary
Project setup and sample data
aisearch-v2/.env.example, aisearch-v2/.gitignore, aisearch-v2/requirements.txt, aisearch-v2/sample_data.json
Adds Azure Search connection env vars, ignore rules, pinned dependencies, and a 10-document sample dataset with relevance labels.
Analyzer recall benchmark script
aisearch-v2/analyzer_test.py
Builds an index with per-analyzer fields, uploads documents, compares tokenization, runs searches, and reports recall/false positives across four analyzers.
Synonym map recall benchmark script
aisearch-v2/synonym_test.py
Creates a synonym map and index with synonym-enabled/disabled fields, uploads documents, runs searches, and reports recall comparing analyzers with synonyms on/off.
README documentation of experiments and results
aisearch-v2/README.md
Documents the problem, dataset, three reproductions (tokenization, recall, synonyms), portal verification, remediation guide, FAQ, run commands, and cleanup.

Estimated code review effort: 3 (Moderate) | ~25 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Script as Test Script (analyzer_test.py / synonym_test.py)
  participant IndexClient as SearchIndexClient
  participant SearchClient as SearchClient
  participant SampleData as sample_data.json

  Script->>SampleData: load_data()
  Script->>IndexClient: create_index() / create_synonym_map()
  Script->>SearchClient: upload_docs()
  Script->>SearchClient: search_hits() per query/analyzer/field
  SearchClient-->>Script: matched document ids
  Script->>Script: recall_report() computes recall and false positives
Loading

Compact metadata
Related issues: None found
Related PRs: None found
Suggested labels: documentation, enhancement
Suggested reviewers: hellices

Poem
A rabbit hopped through Hangul text so fine,
Testing analyzers, one query at a time.
Ko.microsoft won the recall race,
While standard.lucene lost its place.
With synonyms and scripts in tidy rows,
This bunny's README proudly shows! 🐰📊

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly matches the main change: a Korean analyzer and synonym comparison guide for aisearch-v2.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feature/aisearch-v2-analyzer-guide

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (2)
aisearch-v2/analyzer_test.py (1)

125-130: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Unused loop variable field_name.

Only analyzer is used in the loop body; static analysis correctly flags this.

♻️ Proposed fix
-    for field_name, analyzer in ANALYZERS.items():
+    for _field_name, analyzer in ANALYZERS.items():
         result = client.analyze_text(
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@aisearch-v2/analyzer_test.py` around lines 125 - 130, The loop in
analyzer_test.py iterates over ANALYZERS.items() but does not use the field_name
variable, so update the iteration in the test helper to avoid binding an unused
name. Keep the logic in the analyze_text loop centered on ANALYZERS and
analyzer, and either iterate only over the analyzer values or explicitly use
both names if the field name is needed later.

Source: Linters/SAST tools

aisearch-v2/synonym_test.py (1)

1-27: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Invalid escape sequences in docstring — use a raw string.

The Windows-path example (.\.venv\Scripts\python.exe) inside this non-raw docstring triggers deprecation warnings for \., \S, \p and is not forward-compatible with future Python releases that plan to make invalid escapes a hard error.

♻️ Proposed fix
-"""
+r"""
 Azure AI Search 동의어(Synonym Map) 효과 비교 테스트
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@aisearch-v2/synonym_test.py` around lines 1 - 27, The module docstring in
synonym_test.py contains Windows backslash sequences in the execution example,
which causes invalid escape warnings in a normal triple-quoted string. Update
the top-level docstring to be a raw string (or otherwise escape the backslashes)
so the path shown in the example is preserved without triggering invalid escape
handling. Keep the change localized to the module docstring near the execution
instructions.

Source: Linters/SAST tools

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@aisearch-v2/README.md`:
- Line 186: The README section heading contains a broken replacement character
before the heading text, indicating an encoding issue; update the affected
headings in the README content so the title renders cleanly without the stray
glyph. Check the duplicated heading entries referenced by the Synonym Map
section and remove any corrupted character sequences while preserving the emoji
and Korean text.
- Around line 226-236: The README example for SynonymMap creation is using the
wrong shape for the synonyms value. Update the SynonymMap example to pass a list
of rule strings via the SynonymMap.synonyms field instead of building a
newline-joined string, and keep the example aligned with the SDK usage shown by
the SynonymMap and index_client.create_or_update_synonym_map symbols.

---

Nitpick comments:
In `@aisearch-v2/analyzer_test.py`:
- Around line 125-130: The loop in analyzer_test.py iterates over
ANALYZERS.items() but does not use the field_name variable, so update the
iteration in the test helper to avoid binding an unused name. Keep the logic in
the analyze_text loop centered on ANALYZERS and analyzer, and either iterate
only over the analyzer values or explicitly use both names if the field name is
needed later.

In `@aisearch-v2/synonym_test.py`:
- Around line 1-27: The module docstring in synonym_test.py contains Windows
backslash sequences in the execution example, which causes invalid escape
warnings in a normal triple-quoted string. Update the top-level docstring to be
a raw string (or otherwise escape the backslashes) so the path shown in the
example is preserved without triggering invalid escape handling. Keep the change
localized to the module docstring near the execution instructions.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: a8a64f28-7679-47eb-a20f-52d14cd905eb

📥 Commits

Reviewing files that changed from the base of the PR and between 945f715 and a1ff28c.

⛔ Files ignored due to path filters (2)
  • aisearch-v2/images/portal_ko_microsoft.png is excluded by !**/*.png
  • aisearch-v2/images/portal_standard_lucene.png is excluded by !**/*.png
📒 Files selected for processing (8)
  • aisearch-v2/.env.example
  • aisearch-v2/.gitignore
  • aisearch-v2/README.md
  • aisearch-v2/analyzer_test.py
  • aisearch-v2/images/.gitkeep
  • aisearch-v2/requirements.txt
  • aisearch-v2/sample_data.json
  • aisearch-v2/synonym_test.py

Comment thread aisearch-v2/README.md

---

## � 재현 3 — 동의어(Synonym Map)로 정확도 더 높이기

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Broken/garbled characters in section headings.

Both headings contain a replacement character (�) before the emoji/text, indicating an encoding issue that will render as a broken glyph.

✏️ Proposed fix
-## � 재현 3 — 동의어(Synonym Map)로 정확도 더 높이기
+## 🧪 재현 3 — 동의어(Synonym Map)로 정확도 더 높이기
-## �🛠️ 고객 적용 가이드 (진단 → 조치 → 검증)
+## 🛠️ 고객 적용 가이드 (진단 → 조치 → 검증)

Also applies to: 250-250

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@aisearch-v2/README.md` at line 186, The README section heading contains a
broken replacement character before the heading text, indicating an encoding
issue; update the affected headings in the README content so the title renders
cleanly without the stray glyph. Check the duplicated heading entries referenced
by the Synonym Map section and remove any corrupted character sequences while
preserving the emoji and Korean text.

Comment thread aisearch-v2/README.md
Comment on lines +226 to +236
동의어 맵 등록/필드 연결(SDK) 예시:
```python
from azure.search.documents.indexes.models import SynonymMap, SearchableField, SearchFieldDataType

# 1) 동의어 맵 등록
sm = SynonymMap(name="promo-synonyms", synonyms="\n".join([
"온누리상품권, 온누리 상품권, 온누리",
"지역사랑상품권, 지역화폐",
]))
index_client.create_or_update_synonym_map(sm)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

# Locate the README and any related synonym-map examples or tests.
git ls-files 'aisearch-v2/README.md' 'aisearch-v2/**' | sed -n '1,200p'

printf '\n--- README excerpt ---\n'
sed -n '220,245p' aisearch-v2/README.md

printf '\n--- Search for SynonymMap usage ---\n'
rg -n "SynonymMap|synonyms=" aisearch-v2 -S

printf '\n--- Search for the specific warning / test mentioned in the review comment ---\n'
rg -n "character-by-character|join\\(self\\.synonyms\\)|synonym_test|broken synonym map" aisearch-v2 -S

Repository: hellices/devguidesample

Length of output: 2406


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '\n--- synonym_test.py excerpt ---\n'
sed -n '70,110p' aisearch-v2/synonym_test.py

printf '\n--- requirements.txt ---\n'
cat -n aisearch-v2/requirements.txt

printf '\n--- locate any SynonymMap docs or type hints in the repo ---\n'
rg -n "List\\[str\\]|SynonymMap\\(|synonym map|synonyms=" aisearch-v2 -S

Repository: hellices/devguidesample

Length of output: 2081


🏁 Script executed:

#!/bin/bash
set -euo pipefail

# Read the Azure Search SDK version and, if available, inspect installed metadata without running repo code.
python3 - <<'PY'
from pathlib import Path
p = Path("aisearch-v2/requirements.txt")
for line in p.read_text(encoding="utf-8").splitlines():
    if "azure-search-documents" in line:
        print(line)
PY

Repository: hellices/devguidesample

Length of output: 192


Use synonyms=[...] here

SynonymMap.synonyms should be a list of rule strings; "\n".join([...]) will serialize incorrectly and break the map. Use the list form instead.

Suggested fix
-sm = SynonymMap(name="promo-synonyms", synonyms="\n".join([
+sm = SynonymMap(name="promo-synonyms", synonyms=[
     "온누리상품권, 온누리 상품권, 온누리",
     "지역사랑상품권, 지역화폐",
-]))
+])
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
동의어 맵 등록/필드 연결(SDK) 예시:
```python
from azure.search.documents.indexes.models import SynonymMap, SearchableField, SearchFieldDataType
# 1) 동의어 맵 등록
sm = SynonymMap(name="promo-synonyms", synonyms="\n".join([
"온누리상품권, 온누리 상품권, 온누리",
"지역사랑상품권, 지역화폐",
]))
index_client.create_or_update_synonym_map(sm)
동의어 맵 등록/필드 연결(SDK) 예시:
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@aisearch-v2/README.md` around lines 226 - 236, The README example for
SynonymMap creation is using the wrong shape for the synonyms value. Update the
SynonymMap example to pass a list of rule strings via the SynonymMap.synonyms
field instead of building a newline-joined string, and keep the example aligned
with the SDK usage shown by the SynonymMap and
index_client.create_or_update_synonym_map symbols.

@daeungo1 daeungo1 closed this Jul 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant