Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
897 changes: 897 additions & 0 deletions docs/wireframes/out/benchmark-annotation.html

Large diffs are not rendered by default.

919 changes: 919 additions & 0 deletions docs/wireframes/out/benchmark-evaluation.html

Large diffs are not rendered by default.

931 changes: 931 additions & 0 deletions docs/wireframes/out/benchmark-results-compare.html

Large diffs are not rendered by default.

937 changes: 937 additions & 0 deletions docs/wireframes/out/benchmark-upload-align-mapping.html

Large diffs are not rendered by default.

71 changes: 71 additions & 0 deletions docs/wireframes/pages/benchmark-annotation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,71 @@
# Benchmark · Annotation
Inspect evaluation answers and mark evidence quality.

::: card {#tabs}
[Upload & map] [Evaluation] [Results] [Annotation]*
:::

::: card {#evaluation}
### Evaluation
[Microsoft retrieval · v2____________v]

ClimRetrieve v1 · Ranking

Question [TCFD-01________________v]
- TCFD-01
- TCFD-02
- TCFD-03

12 / 16 evaluated · 3 annotations saved
:::

::: card {#question}
### Question
**How does the board oversee climate-related risks?**

Reference type: Governance · Expected chunks: 3


::: callout {for:evaluation side:left}
choose a saved run
:::

::: callout {for:question side:right}
review one question
:::

::: card {#evidence}
### Retrieved evidence

| Rank | Relevance | Source passage | Mark |
|---:|---|---|---|
| 1 | [Relevant v] | “The Sustainability Committee meets quarterly to review climate risk…” · p.14 | [x] |
| 2 | [Partly relevant v] | “The CRO reports climate metrics to the audit committee…” · p.18 | [x] |
| 3 | [Not relevant v] | “Scenario analysis is presented annually to the full board…” · p.22 | [ ] |

**Relevance scale:** 0 Not relevant · 1 Partly relevant · 2 Relevant

### Annotation note
[The first passage directly answers the governance question.______________________________]

[Save annotation]* [Previous question] [Next question →]
:::

::: callout {for:evidence side:left}
mark the chunks
:::

::: card {#answer}
### Evaluation details

Model answer: The board’s Sustainability Committee reviews climate risks quarterly, with reporting from the CRO.

Matched benchmark chunks: 2 / 3 · Evidence coverage: 67%

Annotation history: **3 saved** · Last edited by you just now
:::

::: callout {for:answer side:right}
revisit saved
annotations
:::
88 changes: 88 additions & 0 deletions docs/wireframes/pages/benchmark-evaluation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,88 @@
# Benchmark · Evaluation
Run a ranking evaluation against a shared benchmark.

::: card {#tabs}
[Upload & map] [Evaluation]* [Results] [Annotation]
:::

::: card {#mode}
### Evaluation mode

- (x) Ranking
- ( ) Classification

Ranking compares the order of retrieved chunks. Classification compares one predicted label per question.
:::

::: callout {for:mode side:left}
choose a track
:::

::: {#benchmark}
### Benchmark set
[ClimRetrieve v1________________v]
- ClimRetrieve v1
- ClimRetrieve report-level
- ClimateFinanceBench

16 questions · 4 question types · shared reference set
:::

::: {#dataset}
### Your dataset
[Microsoft retrieval · v2____________v]
- Microsoft retrieval · v2
- chunks_data.csv
- Northwind retrieval · v1

Saved from Upload & map · ranking format validated
:::

::: callout {for:benchmark side:left}
same questions
for every run
:::

::: callout {for:dataset side:right}
your configured
dataset
:::

::: card {#coverage}
### Question matching

**12 / 16 questions will be evaluated**

███████████████░░░ 75%

12 questions have matching IDs and compatible question types. 4 questions are not present in your dataset and will be excluded.

| Match status | Questions | Action |
|---|---:|---|
| Ready to evaluate | 12 | Included |
| Missing from dataset | 3 | Excluded |
| Type mismatch | 1 | Excluded |

Higher question coverage generally gives a more reliable and accurate benchmark result.
:::

::: callout {for:coverage side:left}
coverage affects
confidence
:::

::: card {#ranking}
### Ranking configuration

Ranking output: (x) Ordered chunks per question ( ) Relevance scores only

K values: [1, 3, 5, 10________________v]

Evaluation name: [Microsoft v2 · ClimRetrieve v1________________]

[Review matched questions] [Start evaluation]*


::: callout {for:ranking side:right}
MAP · MRR · Precision@K · Recall@K
:::
52 changes: 52 additions & 0 deletions docs/wireframes/pages/benchmark-results-compare.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,52 @@
# Benchmark · Results
Review metrics and compare saved evaluations.

::: card {#tabs}
[Upload & map] [Evaluation] [Results]* [Annotation]
:::

::: card {#summary}
### Microsoft retrieval · v2
ClimRetrieve v1 · Ranking · Completed 02 Sep 2026, 14:32

**12 / 16 questions evaluated** · 75% coverage · 4 excluded

Higher values indicate that relevant chunks were ranked earlier and the result is more accurate against the reference set.
:::

### Result
| MAP | MRR | Precision@5 | Recall@10 |
|---|---|---|---|
| 0.742 | 0.833 | 0.780 | 0.910 |



::: callout {for:map side:left}
primary ranking score
:::

::: card {#compare}
### Compare evaluations

Select saved runs to compare side by side.

| Evaluation | Dataset | Questions | MAP | MRR | Precision@5 |
|---|---|---:|---:|---:|---:|
| [x] Microsoft v2 | Microsoft retrieval · v2 | 12/16 | **0.742** | **0.833** | **0.780** |
| [x] Northwind v1 | Northwind retrieval · v1 | 16/16 | 0.681 | 0.750 | 0.702 |
| [ ] Chunks baseline | chunks_data.csv | 10/16 | 0.554 | 0.600 | 0.571 |

### Metric comparison

| Metric | Microsoft v2 | Northwind v1 | Difference |
|---|---:|---:|---:|
| MAP | 0.742 | 0.681 | +0.061 |
| MRR | 0.833 | 0.750 | +0.083 |
| Precision@5 | 0.780 | 0.702 | +0.078 |

[Open question-level comparison] [Save comparison] [Export CSV]

::: callout {for:compare side:right}
spot meaningful
differences
:::
84 changes: 84 additions & 0 deletions docs/wireframes/pages/benchmark-upload-align-mapping.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,84 @@
# Benchmark · Upload and map
Prepare a dataset for evaluation.

::: card {#tabs}
[Upload & map]* [Evaluation] [Results] [Annotation]
:::

::: card {#steps}
### 1. Upload · 2. Align · 3. Review mapping

Upload your system output, align it with the benchmark format, then confirm the detected fields before evaluation.
:::

::: card {#upload}
### Upload dataset

[Drop CSV, Excel, YAML or JSON here____________________] [Browse files]

Selected file: **chunks_Microsoft_2024.csv** · 1,240 rows · Uploaded just now

Supported ranking fields: query/item ID, report ID, chunk ID, position/rank, optional score.
:::

::: card {#benchmark}
### Benchmark format

Reference set: [ClimRetrieve v1________________v]

- 16 benchmark questions
- 4 expected question types
- Ground-truth labels managed by the benchmark owner

[Download sample format]
:::

::: callout {for:upload side:left}
your dataset
:::

::: callout {for:benchmark side:right}
reference structure
:::

::: card {#align}
### Align columns

| Your column | Benchmark field | Status |
|---|---|---|
| [query_id________v] | Query / question ID | ✓ matched |
| [report_id_______v] | Report ID | ✓ matched |
| [chunk_id________v] | Retrieved chunk ID | ✓ matched |
| [position________v] | Rank / position | ✓ matched |
| [score___________v] | Relevance score | Optional |

[Auto-detect fields] [Validate alignment]*
:::

::: callout {for:align side:left}
map once,
reuse later
:::

::: card {#mapping}
### Mapping preview

16 / 16 benchmark questions found · 1,240 / 1,240 rows valid · 0 errors

| Question ID | Report | Top retrieved chunks | Question type |
|---|---|---|---|
| TCFD-01 | Microsoft 2024 | chunk_014, chunk_088, chunk_102 | Governance |
| TCFD-02 | Microsoft 2024 | chunk_031, chunk_044, chunk_090 | Strategy |
| TCFD-03 | Microsoft 2024 | chunk_052, chunk_067, chunk_071 | Risk |

### Dataset name
[Microsoft retrieval · v2________________________]

[Save dataset]* [Continue to evaluation]*
:::

::: callout {for:mapping side:right}
the user sees
the mapped result
:::

Binary file added docs/wireframes/renders/benchmark-annotation.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added docs/wireframes/renders/benchmark-evaluation.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading