diff --git a/docs/wireframes/out/benchmark-annotation.html b/docs/wireframes/out/benchmark-annotation.html
new file mode 100644
index 00000000..3929ed58
--- /dev/null
+++ b/docs/wireframes/out/benchmark-annotation.html
@@ -0,0 +1,897 @@
+
+
+
+
+
Benchmark · Annotation
+
Inspect evaluation answers and mark evidence quality.
+
+
+
+
+
+
+
+
+
+
Evaluation
+
+
ClimRetrieve v1 · Ranking
+
Question
+
+ - TCFD-01
+ - TCFD-02
+ - TCFD-03
+
+
12 / 16 evaluated · 3 annotations saved
+
+
+
Question
+
How does the board oversee climate-related risks?
+
Reference type: Governance · Expected chunks: 3
+
+
Retrieved evidence
+
+
+
+ | Rank |
+ Relevance |
+ Source passage |
+ Mark |
+
+
+
+
+ | 1 |
+ [Relevant v] |
+ “The Sustainability Committee meets quarterly to review climate risk…” · p.14 |
+ [x] |
+
+
+ | 2 |
+ [Partly relevant v] |
+ “The CRO reports climate metrics to the audit committee…” · p.18 |
+ [x] |
+
+
+ | 3 |
+ [Not relevant v] |
+ “Scenario analysis is presented annually to the full board…” · p.22 |
+ [ ] |
+
+
+
+
Relevance scale: 0 Not relevant · 1 Partly relevant · 2 Relevant
+
Annotation note
+
+
+
+
+
+
+
+
+
Evaluation details
+
Model answer: The board’s Sustainability Committee reviews climate risks quarterly, with reporting from the CRO.
+
Matched benchmark chunks: 2 / 3 · Evidence coverage: 67%
+
Annotation history: 3 saved · Last edited by you just now
+
+
+
+
+
+
+
+
+
+
+
\ No newline at end of file
diff --git a/docs/wireframes/out/benchmark-evaluation.html b/docs/wireframes/out/benchmark-evaluation.html
new file mode 100644
index 00000000..ac7b9f8e
--- /dev/null
+++ b/docs/wireframes/out/benchmark-evaluation.html
@@ -0,0 +1,919 @@
+
+
+
+
+
Benchmark · Evaluation
+
Run a ranking evaluation against a shared benchmark.
+
+
+
+
+
+
+
+
+
+
Evaluation mode
+
+
Ranking compares the order of retrieved chunks. Classification compares one predicted label per question.
+
+
+
Benchmark set
+
+
16 questions · 4 question types · shared reference set
+
+
+
Your dataset
+
+
Saved from Upload & map · ranking format validated
+
+
+
Question matching
+
12 / 16 questions will be evaluated
+
███████████████░░░ 75%
+
12 questions have matching IDs and compatible question types. 4 questions are not present in your dataset and will be excluded.
+
+
+
+ | Match status |
+ Questions |
+ Action |
+
+
+
+
+ | Ready to evaluate |
+ 12 |
+ Included |
+
+
+ | Missing from dataset |
+ 3 |
+ Excluded |
+
+
+ | Type mismatch |
+ 1 |
+ Excluded |
+
+
+
+
Higher question coverage generally gives a more reliable and accurate benchmark result.
+
+
+
Ranking configuration
+
+
+
+
+
K values:
+
Evaluation name:
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
\ No newline at end of file
diff --git a/docs/wireframes/out/benchmark-results-compare.html b/docs/wireframes/out/benchmark-results-compare.html
new file mode 100644
index 00000000..9f58714e
--- /dev/null
+++ b/docs/wireframes/out/benchmark-results-compare.html
@@ -0,0 +1,931 @@
+
+
+
+
+
Benchmark · Results
+
Review metrics and compare saved evaluations.
+
+
+
+
+
+
+
+
+
+
Microsoft retrieval · v2
+
ClimRetrieve v1 · Ranking · Completed 02 Sep 2026, 14:32
+
12 / 16 questions evaluated · 75% coverage · 4 excluded
+
Higher values indicate that relevant chunks were ranked earlier and the result is more accurate against the reference set.
+
+
Result
+
+
+
+ | MAP |
+ MRR |
+ Precision@5 |
+ Recall@10 |
+
+
+
+
+ | 0.742 |
+ 0.833 |
+ 0.780 |
+ 0.910 |
+
+
+
+
+
Compare evaluations
+
Select saved runs to compare side by side.
+
+
+
+ | Evaluation |
+ Dataset |
+ Questions |
+ MAP |
+ MRR |
+ Precision@5 |
+
+
+
+
+ | [x] Microsoft v2 |
+ Microsoft retrieval · v2 |
+ 12/16 |
+ 0.742 |
+ 0.833 |
+ 0.780 |
+
+
+ | [x] Northwind v1 |
+ Northwind retrieval · v1 |
+ 16/16 |
+ 0.681 |
+ 0.750 |
+ 0.702 |
+
+
+ | [ ] Chunks baseline |
+ chunks_data.csv |
+ 10/16 |
+ 0.554 |
+ 0.600 |
+ 0.571 |
+
+
+
+
Metric comparison
+
+
+
+ | Metric |
+ Microsoft v2 |
+ Northwind v1 |
+ Difference |
+
+
+
+
+ | MAP |
+ 0.742 |
+ 0.681 |
+ +0.061 |
+
+
+ | MRR |
+ 0.833 |
+ 0.750 |
+ +0.083 |
+
+
+ | Precision@5 |
+ 0.780 |
+ 0.702 |
+ +0.078 |
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
\ No newline at end of file
diff --git a/docs/wireframes/out/benchmark-upload-align-mapping.html b/docs/wireframes/out/benchmark-upload-align-mapping.html
new file mode 100644
index 00000000..3f28da15
--- /dev/null
+++ b/docs/wireframes/out/benchmark-upload-align-mapping.html
@@ -0,0 +1,937 @@
+
+
+
+
+
Benchmark · Upload and map
+
Prepare a dataset for evaluation.
+
+
+
+
+
+
+
+
+
+
1. Upload · 2. Align · 3. Review mapping
+
Upload your system output, align it with the benchmark format, then confirm the detected fields before evaluation.
+
+
+
Upload dataset
+
+
+
+
+
Selected file: chunks_Microsoft_2024.csv · 1,240 rows · Uploaded just now
+
Supported ranking fields: query/item ID, report ID, chunk ID, position/rank, optional score.
+
+
+
Benchmark format
+
Reference set:
+
+ - 16 benchmark questions
+ - 4 expected question types
+ - Ground-truth labels managed by the benchmark owner
+
+
+
+
+
Align columns
+
+
+
+ | Your column |
+ Benchmark field |
+ Status |
+
+
+
+
+ | [query_id________v] |
+ Query / question ID |
+ ✓ matched |
+
+
+ | [report_id_______v] |
+ Report ID |
+ ✓ matched |
+
+
+ | [chunk_id________v] |
+ Retrieved chunk ID |
+ ✓ matched |
+
+
+ | [position________v] |
+ Rank / position |
+ ✓ matched |
+
+
+ | [score___________v] |
+ Relevance score |
+ Optional |
+
+
+
+
+
+
+
+
+
+
Mapping preview
+
16 / 16 benchmark questions found · 1,240 / 1,240 rows valid · 0 errors
+
+
+
+ | Question ID |
+ Report |
+ Top retrieved chunks |
+ Question type |
+
+
+
+
+ | TCFD-01 |
+ Microsoft 2024 |
+ chunk_014, chunk_088, chunk_102 |
+ Governance |
+
+
+ | TCFD-02 |
+ Microsoft 2024 |
+ chunk_031, chunk_044, chunk_090 |
+ Strategy |
+
+
+ | TCFD-03 |
+ Microsoft 2024 |
+ chunk_052, chunk_067, chunk_071 |
+ Risk |
+
+
+
+
Dataset name
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
\ No newline at end of file
diff --git a/docs/wireframes/pages/benchmark-annotation.md b/docs/wireframes/pages/benchmark-annotation.md
new file mode 100644
index 00000000..e8f71464
--- /dev/null
+++ b/docs/wireframes/pages/benchmark-annotation.md
@@ -0,0 +1,71 @@
+# Benchmark · Annotation
+Inspect evaluation answers and mark evidence quality.
+
+::: card {#tabs}
+[Upload & map] [Evaluation] [Results] [Annotation]*
+:::
+
+::: card {#evaluation}
+### Evaluation
+[Microsoft retrieval · v2____________v]
+
+ClimRetrieve v1 · Ranking
+
+Question [TCFD-01________________v]
+- TCFD-01
+- TCFD-02
+- TCFD-03
+
+12 / 16 evaluated · 3 annotations saved
+:::
+
+::: card {#question}
+### Question
+**How does the board oversee climate-related risks?**
+
+Reference type: Governance · Expected chunks: 3
+
+
+::: callout {for:evaluation side:left}
+choose a saved run
+:::
+
+::: callout {for:question side:right}
+review one question
+:::
+
+::: card {#evidence}
+### Retrieved evidence
+
+| Rank | Relevance | Source passage | Mark |
+|---:|---|---|---|
+| 1 | [Relevant v] | “The Sustainability Committee meets quarterly to review climate risk…” · p.14 | [x] |
+| 2 | [Partly relevant v] | “The CRO reports climate metrics to the audit committee…” · p.18 | [x] |
+| 3 | [Not relevant v] | “Scenario analysis is presented annually to the full board…” · p.22 | [ ] |
+
+**Relevance scale:** 0 Not relevant · 1 Partly relevant · 2 Relevant
+
+### Annotation note
+[The first passage directly answers the governance question.______________________________]
+
+[Save annotation]* [Previous question] [Next question →]
+:::
+
+::: callout {for:evidence side:left}
+mark the chunks
+:::
+
+::: card {#answer}
+### Evaluation details
+
+Model answer: The board’s Sustainability Committee reviews climate risks quarterly, with reporting from the CRO.
+
+Matched benchmark chunks: 2 / 3 · Evidence coverage: 67%
+
+Annotation history: **3 saved** · Last edited by you just now
+:::
+
+::: callout {for:answer side:right}
+revisit saved
+annotations
+:::
diff --git a/docs/wireframes/pages/benchmark-evaluation.md b/docs/wireframes/pages/benchmark-evaluation.md
new file mode 100644
index 00000000..4865f445
--- /dev/null
+++ b/docs/wireframes/pages/benchmark-evaluation.md
@@ -0,0 +1,88 @@
+# Benchmark · Evaluation
+Run a ranking evaluation against a shared benchmark.
+
+::: card {#tabs}
+[Upload & map] [Evaluation]* [Results] [Annotation]
+:::
+
+::: card {#mode}
+### Evaluation mode
+
+- (x) Ranking
+- ( ) Classification
+
+Ranking compares the order of retrieved chunks. Classification compares one predicted label per question.
+:::
+
+::: callout {for:mode side:left}
+choose a track
+:::
+
+::: {#benchmark}
+### Benchmark set
+[ClimRetrieve v1________________v]
+- ClimRetrieve v1
+- ClimRetrieve report-level
+- ClimateFinanceBench
+
+16 questions · 4 question types · shared reference set
+:::
+
+::: {#dataset}
+### Your dataset
+[Microsoft retrieval · v2____________v]
+- Microsoft retrieval · v2
+- chunks_data.csv
+- Northwind retrieval · v1
+
+Saved from Upload & map · ranking format validated
+:::
+
+::: callout {for:benchmark side:left}
+same questions
+for every run
+:::
+
+::: callout {for:dataset side:right}
+your configured
+dataset
+:::
+
+::: card {#coverage}
+### Question matching
+
+**12 / 16 questions will be evaluated**
+
+███████████████░░░ 75%
+
+12 questions have matching IDs and compatible question types. 4 questions are not present in your dataset and will be excluded.
+
+| Match status | Questions | Action |
+|---|---:|---|
+| Ready to evaluate | 12 | Included |
+| Missing from dataset | 3 | Excluded |
+| Type mismatch | 1 | Excluded |
+
+Higher question coverage generally gives a more reliable and accurate benchmark result.
+:::
+
+::: callout {for:coverage side:left}
+coverage affects
+confidence
+:::
+
+::: card {#ranking}
+### Ranking configuration
+
+Ranking output: (x) Ordered chunks per question ( ) Relevance scores only
+
+K values: [1, 3, 5, 10________________v]
+
+Evaluation name: [Microsoft v2 · ClimRetrieve v1________________]
+
+[Review matched questions] [Start evaluation]*
+
+
+::: callout {for:ranking side:right}
+MAP · MRR · Precision@K · Recall@K
+:::
diff --git a/docs/wireframes/pages/benchmark-results-compare.md b/docs/wireframes/pages/benchmark-results-compare.md
new file mode 100644
index 00000000..ec4572b2
--- /dev/null
+++ b/docs/wireframes/pages/benchmark-results-compare.md
@@ -0,0 +1,52 @@
+# Benchmark · Results
+Review metrics and compare saved evaluations.
+
+::: card {#tabs}
+[Upload & map] [Evaluation] [Results]* [Annotation]
+:::
+
+::: card {#summary}
+### Microsoft retrieval · v2
+ClimRetrieve v1 · Ranking · Completed 02 Sep 2026, 14:32
+
+**12 / 16 questions evaluated** · 75% coverage · 4 excluded
+
+Higher values indicate that relevant chunks were ranked earlier and the result is more accurate against the reference set.
+:::
+
+### Result
+| MAP | MRR | Precision@5 | Recall@10 |
+|---|---|---|---|
+| 0.742 | 0.833 | 0.780 | 0.910 |
+
+
+
+::: callout {for:map side:left}
+primary ranking score
+:::
+
+::: card {#compare}
+### Compare evaluations
+
+Select saved runs to compare side by side.
+
+| Evaluation | Dataset | Questions | MAP | MRR | Precision@5 |
+|---|---|---:|---:|---:|---:|
+| [x] Microsoft v2 | Microsoft retrieval · v2 | 12/16 | **0.742** | **0.833** | **0.780** |
+| [x] Northwind v1 | Northwind retrieval · v1 | 16/16 | 0.681 | 0.750 | 0.702 |
+| [ ] Chunks baseline | chunks_data.csv | 10/16 | 0.554 | 0.600 | 0.571 |
+
+### Metric comparison
+
+| Metric | Microsoft v2 | Northwind v1 | Difference |
+|---|---:|---:|---:|
+| MAP | 0.742 | 0.681 | +0.061 |
+| MRR | 0.833 | 0.750 | +0.083 |
+| Precision@5 | 0.780 | 0.702 | +0.078 |
+
+[Open question-level comparison] [Save comparison] [Export CSV]
+
+::: callout {for:compare side:right}
+spot meaningful
+differences
+:::
diff --git a/docs/wireframes/pages/benchmark-upload-align-mapping.md b/docs/wireframes/pages/benchmark-upload-align-mapping.md
new file mode 100644
index 00000000..6be7ef13
--- /dev/null
+++ b/docs/wireframes/pages/benchmark-upload-align-mapping.md
@@ -0,0 +1,84 @@
+# Benchmark · Upload and map
+Prepare a dataset for evaluation.
+
+::: card {#tabs}
+[Upload & map]* [Evaluation] [Results] [Annotation]
+:::
+
+::: card {#steps}
+### 1. Upload · 2. Align · 3. Review mapping
+
+Upload your system output, align it with the benchmark format, then confirm the detected fields before evaluation.
+:::
+
+::: card {#upload}
+### Upload dataset
+
+[Drop CSV, Excel, YAML or JSON here____________________] [Browse files]
+
+Selected file: **chunks_Microsoft_2024.csv** · 1,240 rows · Uploaded just now
+
+Supported ranking fields: query/item ID, report ID, chunk ID, position/rank, optional score.
+:::
+
+::: card {#benchmark}
+### Benchmark format
+
+Reference set: [ClimRetrieve v1________________v]
+
+- 16 benchmark questions
+- 4 expected question types
+- Ground-truth labels managed by the benchmark owner
+
+[Download sample format]
+:::
+
+::: callout {for:upload side:left}
+your dataset
+:::
+
+::: callout {for:benchmark side:right}
+reference structure
+:::
+
+::: card {#align}
+### Align columns
+
+| Your column | Benchmark field | Status |
+|---|---|---|
+| [query_id________v] | Query / question ID | ✓ matched |
+| [report_id_______v] | Report ID | ✓ matched |
+| [chunk_id________v] | Retrieved chunk ID | ✓ matched |
+| [position________v] | Rank / position | ✓ matched |
+| [score___________v] | Relevance score | Optional |
+
+[Auto-detect fields] [Validate alignment]*
+:::
+
+::: callout {for:align side:left}
+map once,
+reuse later
+:::
+
+::: card {#mapping}
+### Mapping preview
+
+16 / 16 benchmark questions found · 1,240 / 1,240 rows valid · 0 errors
+
+| Question ID | Report | Top retrieved chunks | Question type |
+|---|---|---|---|
+| TCFD-01 | Microsoft 2024 | chunk_014, chunk_088, chunk_102 | Governance |
+| TCFD-02 | Microsoft 2024 | chunk_031, chunk_044, chunk_090 | Strategy |
+| TCFD-03 | Microsoft 2024 | chunk_052, chunk_067, chunk_071 | Risk |
+
+### Dataset name
+[Microsoft retrieval · v2________________________]
+
+[Save dataset]* [Continue to evaluation]*
+:::
+
+::: callout {for:mapping side:right}
+the user sees
+the mapped result
+:::
+
diff --git a/docs/wireframes/renders/benchmark-annotation.png b/docs/wireframes/renders/benchmark-annotation.png
new file mode 100644
index 00000000..32e81f03
Binary files /dev/null and b/docs/wireframes/renders/benchmark-annotation.png differ
diff --git a/docs/wireframes/renders/benchmark-evaluation.png b/docs/wireframes/renders/benchmark-evaluation.png
new file mode 100644
index 00000000..f8ffb7e6
Binary files /dev/null and b/docs/wireframes/renders/benchmark-evaluation.png differ
diff --git a/docs/wireframes/renders/benchmark-results-compare.png b/docs/wireframes/renders/benchmark-results-compare.png
new file mode 100644
index 00000000..beb05cf7
Binary files /dev/null and b/docs/wireframes/renders/benchmark-results-compare.png differ
diff --git a/docs/wireframes/renders/benchmark-upload-align-mapping.png b/docs/wireframes/renders/benchmark-upload-align-mapping.png
new file mode 100644
index 00000000..7772324f
Binary files /dev/null and b/docs/wireframes/renders/benchmark-upload-align-mapping.png differ