diff --git a/docs/wireframes/out/benchmark-annotation.html b/docs/wireframes/out/benchmark-annotation.html new file mode 100644 index 00000000..3929ed58 --- /dev/null +++ b/docs/wireframes/out/benchmark-annotation.html @@ -0,0 +1,897 @@ + + + + + + wiremd Mockup + + + +
+

Benchmark · Annotation

+

Inspect evaluation answers and mark evidence quality.

+
+
+ + + + +
+
+
+

Evaluation

+ +

ClimRetrieve v1 · Ranking

+

Question

+ +

12 / 16 evaluated · 3 annotations saved

+
+
+

Question

+

How does the board oversee climate-related risks?

+

Reference type: Governance · Expected chunks: 3

+
+

Retrieved evidence

+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
RankRelevanceSource passageMark
1[Relevant v]“The Sustainability Committee meets quarterly to review climate risk…” · p.14[x]
2[Partly relevant v]“The CRO reports climate metrics to the audit committee…” · p.18[x]
3[Not relevant v]“Scenario analysis is presented annually to the full board…” · p.22[ ]
+

Relevance scale: 0 Not relevant · 1 Partly relevant · 2 Relevant

+

Annotation note

+ +
+ + + +
+
+
+

Evaluation details

+

Model answer: The board’s Sustainability Committee reviews climate risks quarterly, with reporting from the CRO.

+

Matched benchmark chunks: 2 / 3 · Evidence coverage: 67%

+

Annotation history: 3 saved · Last edited by you just now

+
+
+ + + + + +
+ + + \ No newline at end of file diff --git a/docs/wireframes/out/benchmark-evaluation.html b/docs/wireframes/out/benchmark-evaluation.html new file mode 100644 index 00000000..ac7b9f8e --- /dev/null +++ b/docs/wireframes/out/benchmark-evaluation.html @@ -0,0 +1,919 @@ + + + + + + wiremd Mockup + + + +
+

Benchmark · Evaluation

+

Run a ranking evaluation against a shared benchmark.

+
+
+ + + + +
+
+
+

Evaluation mode

+ +

Ranking compares the order of retrieved chunks. Classification compares one predicted label per question.

+
+
+

Benchmark set

+ +

16 questions · 4 question types · shared reference set

+
+
+

Your dataset

+ +

Saved from Upload & map · ranking format validated

+
+
+

Question matching

+

12 / 16 questions will be evaluated

+

███████████████░░░ 75%

+

12 questions have matching IDs and compatible question types. 4 questions are not present in your dataset and will be excluded.

+ + + + + + + + + + + + + + + + + + + + + + + + + +
Match statusQuestionsAction
Ready to evaluate12Included
Missing from dataset3Excluded
Type mismatch1Excluded
+

Higher question coverage generally gives a more reliable and accurate benchmark result.

+
+
+

Ranking configuration

+
+ + +
+

K values:

+

Evaluation name:

+
+ + +
+
+ + + + + + +
+ + + \ No newline at end of file diff --git a/docs/wireframes/out/benchmark-results-compare.html b/docs/wireframes/out/benchmark-results-compare.html new file mode 100644 index 00000000..9f58714e --- /dev/null +++ b/docs/wireframes/out/benchmark-results-compare.html @@ -0,0 +1,931 @@ + + + + + + wiremd Mockup + + + +
+

Benchmark · Results

+

Review metrics and compare saved evaluations.

+
+
+ + + + +
+
+
+

Microsoft retrieval · v2

+

ClimRetrieve v1 · Ranking · Completed 02 Sep 2026, 14:32

+

12 / 16 questions evaluated · 75% coverage · 4 excluded

+

Higher values indicate that relevant chunks were ranked earlier and the result is more accurate against the reference set.

+
+

Result

+ + + + + + + + + + + + + + + + + +
MAPMRRPrecision@5Recall@10
0.7420.8330.7800.910
+
+

Compare evaluations

+

Select saved runs to compare side by side.

+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
EvaluationDatasetQuestionsMAPMRRPrecision@5
[x] Microsoft v2Microsoft retrieval · v212/160.7420.8330.780
[x] Northwind v1Northwind retrieval · v116/160.6810.7500.702
[ ] Chunks baselinechunks_data.csv10/160.5540.6000.571
+

Metric comparison

+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
MetricMicrosoft v2Northwind v1Difference
MAP0.7420.681+0.061
MRR0.8330.750+0.083
Precision@50.7800.702+0.078
+
+ + + +
+
+ + + +
+ + + \ No newline at end of file diff --git a/docs/wireframes/out/benchmark-upload-align-mapping.html b/docs/wireframes/out/benchmark-upload-align-mapping.html new file mode 100644 index 00000000..3f28da15 --- /dev/null +++ b/docs/wireframes/out/benchmark-upload-align-mapping.html @@ -0,0 +1,937 @@ + + + + + + wiremd Mockup + + + +
+

Benchmark · Upload and map

+

Prepare a dataset for evaluation.

+
+
+ + + + +
+
+
+

1. Upload · 2. Align · 3. Review mapping

+

Upload your system output, align it with the benchmark format, then confirm the detected fields before evaluation.

+
+
+

Upload dataset

+
+ + +
+

Selected file: chunks_Microsoft_2024.csv · 1,240 rows · Uploaded just now

+

Supported ranking fields: query/item ID, report ID, chunk ID, position/rank, optional score.

+
+
+

Benchmark format

+

Reference set:

+ + +
+
+

Align columns

+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
Your columnBenchmark fieldStatus
[query_id________v]Query / question ID✓ matched
[report_id_______v]Report ID✓ matched
[chunk_id________v]Retrieved chunk ID✓ matched
[position________v]Rank / position✓ matched
[score___________v]Relevance scoreOptional
+
+ + +
+
+
+

Mapping preview

+

16 / 16 benchmark questions found · 1,240 / 1,240 rows valid · 0 errors

+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
Question IDReportTop retrieved chunksQuestion type
TCFD-01Microsoft 2024chunk_014, chunk_088, chunk_102Governance
TCFD-02Microsoft 2024chunk_031, chunk_044, chunk_090Strategy
TCFD-03Microsoft 2024chunk_052, chunk_067, chunk_071Risk
+

Dataset name

+ +
+ + +
+
+ + + + + +
+ + + \ No newline at end of file diff --git a/docs/wireframes/pages/benchmark-annotation.md b/docs/wireframes/pages/benchmark-annotation.md new file mode 100644 index 00000000..e8f71464 --- /dev/null +++ b/docs/wireframes/pages/benchmark-annotation.md @@ -0,0 +1,71 @@ +# Benchmark · Annotation +Inspect evaluation answers and mark evidence quality. + +::: card {#tabs} +[Upload & map] [Evaluation] [Results] [Annotation]* +::: + +::: card {#evaluation} +### Evaluation +[Microsoft retrieval · v2____________v] + +ClimRetrieve v1 · Ranking + +Question [TCFD-01________________v] +- TCFD-01 +- TCFD-02 +- TCFD-03 + +12 / 16 evaluated · 3 annotations saved +::: + +::: card {#question} +### Question +**How does the board oversee climate-related risks?** + +Reference type: Governance · Expected chunks: 3 + + +::: callout {for:evaluation side:left} +choose a saved run +::: + +::: callout {for:question side:right} +review one question +::: + +::: card {#evidence} +### Retrieved evidence + +| Rank | Relevance | Source passage | Mark | +|---:|---|---|---| +| 1 | [Relevant v] | “The Sustainability Committee meets quarterly to review climate risk…” · p.14 | [x] | +| 2 | [Partly relevant v] | “The CRO reports climate metrics to the audit committee…” · p.18 | [x] | +| 3 | [Not relevant v] | “Scenario analysis is presented annually to the full board…” · p.22 | [ ] | + +**Relevance scale:** 0 Not relevant · 1 Partly relevant · 2 Relevant + +### Annotation note +[The first passage directly answers the governance question.______________________________] + +[Save annotation]* [Previous question] [Next question →] +::: + +::: callout {for:evidence side:left} +mark the chunks +::: + +::: card {#answer} +### Evaluation details + +Model answer: The board’s Sustainability Committee reviews climate risks quarterly, with reporting from the CRO. + +Matched benchmark chunks: 2 / 3 · Evidence coverage: 67% + +Annotation history: **3 saved** · Last edited by you just now +::: + +::: callout {for:answer side:right} +revisit saved +annotations +::: diff --git a/docs/wireframes/pages/benchmark-evaluation.md b/docs/wireframes/pages/benchmark-evaluation.md new file mode 100644 index 00000000..4865f445 --- /dev/null +++ b/docs/wireframes/pages/benchmark-evaluation.md @@ -0,0 +1,88 @@ +# Benchmark · Evaluation +Run a ranking evaluation against a shared benchmark. + +::: card {#tabs} +[Upload & map] [Evaluation]* [Results] [Annotation] +::: + +::: card {#mode} +### Evaluation mode + +- (x) Ranking +- ( ) Classification + +Ranking compares the order of retrieved chunks. Classification compares one predicted label per question. +::: + +::: callout {for:mode side:left} +choose a track +::: + +::: {#benchmark} +### Benchmark set +[ClimRetrieve v1________________v] +- ClimRetrieve v1 +- ClimRetrieve report-level +- ClimateFinanceBench + +16 questions · 4 question types · shared reference set +::: + +::: {#dataset} +### Your dataset +[Microsoft retrieval · v2____________v] +- Microsoft retrieval · v2 +- chunks_data.csv +- Northwind retrieval · v1 + +Saved from Upload & map · ranking format validated +::: + +::: callout {for:benchmark side:left} +same questions +for every run +::: + +::: callout {for:dataset side:right} +your configured +dataset +::: + +::: card {#coverage} +### Question matching + +**12 / 16 questions will be evaluated** + +███████████████░░░ 75% + +12 questions have matching IDs and compatible question types. 4 questions are not present in your dataset and will be excluded. + +| Match status | Questions | Action | +|---|---:|---| +| Ready to evaluate | 12 | Included | +| Missing from dataset | 3 | Excluded | +| Type mismatch | 1 | Excluded | + +Higher question coverage generally gives a more reliable and accurate benchmark result. +::: + +::: callout {for:coverage side:left} +coverage affects +confidence +::: + +::: card {#ranking} +### Ranking configuration + +Ranking output: (x) Ordered chunks per question ( ) Relevance scores only + +K values: [1, 3, 5, 10________________v] + +Evaluation name: [Microsoft v2 · ClimRetrieve v1________________] + +[Review matched questions] [Start evaluation]* + + +::: callout {for:ranking side:right} +MAP · MRR · Precision@K · Recall@K +::: diff --git a/docs/wireframes/pages/benchmark-results-compare.md b/docs/wireframes/pages/benchmark-results-compare.md new file mode 100644 index 00000000..ec4572b2 --- /dev/null +++ b/docs/wireframes/pages/benchmark-results-compare.md @@ -0,0 +1,52 @@ +# Benchmark · Results +Review metrics and compare saved evaluations. + +::: card {#tabs} +[Upload & map] [Evaluation] [Results]* [Annotation] +::: + +::: card {#summary} +### Microsoft retrieval · v2 +ClimRetrieve v1 · Ranking · Completed 02 Sep 2026, 14:32 + +**12 / 16 questions evaluated** · 75% coverage · 4 excluded + +Higher values indicate that relevant chunks were ranked earlier and the result is more accurate against the reference set. +::: + +### Result +| MAP | MRR | Precision@5 | Recall@10 | +|---|---|---|---| +| 0.742 | 0.833 | 0.780 | 0.910 | + + + +::: callout {for:map side:left} +primary ranking score +::: + +::: card {#compare} +### Compare evaluations + +Select saved runs to compare side by side. + +| Evaluation | Dataset | Questions | MAP | MRR | Precision@5 | +|---|---|---:|---:|---:|---:| +| [x] Microsoft v2 | Microsoft retrieval · v2 | 12/16 | **0.742** | **0.833** | **0.780** | +| [x] Northwind v1 | Northwind retrieval · v1 | 16/16 | 0.681 | 0.750 | 0.702 | +| [ ] Chunks baseline | chunks_data.csv | 10/16 | 0.554 | 0.600 | 0.571 | + +### Metric comparison + +| Metric | Microsoft v2 | Northwind v1 | Difference | +|---|---:|---:|---:| +| MAP | 0.742 | 0.681 | +0.061 | +| MRR | 0.833 | 0.750 | +0.083 | +| Precision@5 | 0.780 | 0.702 | +0.078 | + +[Open question-level comparison] [Save comparison] [Export CSV] + +::: callout {for:compare side:right} +spot meaningful +differences +::: diff --git a/docs/wireframes/pages/benchmark-upload-align-mapping.md b/docs/wireframes/pages/benchmark-upload-align-mapping.md new file mode 100644 index 00000000..6be7ef13 --- /dev/null +++ b/docs/wireframes/pages/benchmark-upload-align-mapping.md @@ -0,0 +1,84 @@ +# Benchmark · Upload and map +Prepare a dataset for evaluation. + +::: card {#tabs} +[Upload & map]* [Evaluation] [Results] [Annotation] +::: + +::: card {#steps} +### 1. Upload · 2. Align · 3. Review mapping + +Upload your system output, align it with the benchmark format, then confirm the detected fields before evaluation. +::: + +::: card {#upload} +### Upload dataset + +[Drop CSV, Excel, YAML or JSON here____________________] [Browse files] + +Selected file: **chunks_Microsoft_2024.csv** · 1,240 rows · Uploaded just now + +Supported ranking fields: query/item ID, report ID, chunk ID, position/rank, optional score. +::: + +::: card {#benchmark} +### Benchmark format + +Reference set: [ClimRetrieve v1________________v] + +- 16 benchmark questions +- 4 expected question types +- Ground-truth labels managed by the benchmark owner + +[Download sample format] +::: + +::: callout {for:upload side:left} +your dataset +::: + +::: callout {for:benchmark side:right} +reference structure +::: + +::: card {#align} +### Align columns + +| Your column | Benchmark field | Status | +|---|---|---| +| [query_id________v] | Query / question ID | ✓ matched | +| [report_id_______v] | Report ID | ✓ matched | +| [chunk_id________v] | Retrieved chunk ID | ✓ matched | +| [position________v] | Rank / position | ✓ matched | +| [score___________v] | Relevance score | Optional | + +[Auto-detect fields] [Validate alignment]* +::: + +::: callout {for:align side:left} +map once, +reuse later +::: + +::: card {#mapping} +### Mapping preview + +16 / 16 benchmark questions found · 1,240 / 1,240 rows valid · 0 errors + +| Question ID | Report | Top retrieved chunks | Question type | +|---|---|---|---| +| TCFD-01 | Microsoft 2024 | chunk_014, chunk_088, chunk_102 | Governance | +| TCFD-02 | Microsoft 2024 | chunk_031, chunk_044, chunk_090 | Strategy | +| TCFD-03 | Microsoft 2024 | chunk_052, chunk_067, chunk_071 | Risk | + +### Dataset name +[Microsoft retrieval · v2________________________] + +[Save dataset]* [Continue to evaluation]* +::: + +::: callout {for:mapping side:right} +the user sees +the mapped result +::: + diff --git a/docs/wireframes/renders/benchmark-annotation.png b/docs/wireframes/renders/benchmark-annotation.png new file mode 100644 index 00000000..32e81f03 Binary files /dev/null and b/docs/wireframes/renders/benchmark-annotation.png differ diff --git a/docs/wireframes/renders/benchmark-evaluation.png b/docs/wireframes/renders/benchmark-evaluation.png new file mode 100644 index 00000000..f8ffb7e6 Binary files /dev/null and b/docs/wireframes/renders/benchmark-evaluation.png differ diff --git a/docs/wireframes/renders/benchmark-results-compare.png b/docs/wireframes/renders/benchmark-results-compare.png new file mode 100644 index 00000000..beb05cf7 Binary files /dev/null and b/docs/wireframes/renders/benchmark-results-compare.png differ diff --git a/docs/wireframes/renders/benchmark-upload-align-mapping.png b/docs/wireframes/renders/benchmark-upload-align-mapping.png new file mode 100644 index 00000000..7772324f Binary files /dev/null and b/docs/wireframes/renders/benchmark-upload-align-mapping.png differ