Skip to content

Repository files navigation

turkish-rag-eval

validate licence: MIT + CC BY-SA 4.0 dataset on Hugging Face

Which parts of a RAG pipeline actually earn their cost on an agglutinative language? Three chunking strategies × four retrievers × six embedding models, measured on a hand-labelled Turkish gold set, with confidence intervals and a cost column — then pointed at your own documents.

pip install git+https://github.com/RizgarOzan/turkish-rag-eval   # not on PyPI yet
turkish-rag-eval run --corpus ./belgelerim --gold ./sorular.json
turkish-rag-eval report

Four findings

1. The embedding model matters more than any pipeline choice. Every model trained for retrieval beats stemmed BM25 (0.494) on its own. The best, Mursit-Large-TR-Retrieval, reaches 0.781 nDCG@10 on hierarchical chunks. A small E5 the same size as the default goes from 0.501 to 0.642 with nothing else changed, at about the same speed (22 ms against 23 ms). The popular default model is the weak link, not dense retrieval. See the Leaderboard.

2. Turkish stemming is the cheapest real win. Truncating tokens to a 5-character prefix before BM25 lifts nDCG@10 by 23–29% on every chunking strategy, and a paired bootstrap puts every one of those gains clear of zero (hierarchical +0.111, 95% CI [+0.026, +0.204]). "diyabet", "diyabetin", "diyabete", "diyabetli" are four surface forms of one concept; an unstemmed index almost never matches the query's form.

3. The top of the table is a tie, and the tie decides on cost. With 58 queries, a paired bootstrap cannot separate the best configuration from the other two hybrid ones; the remaining nine are measurably worse. So the real choice at the top is price: sentence + hybrid_rrf gives the same quality at 27 ms instead of 37 ms.

nDCG@10 per configuration with 95% bootstrap intervals; the best cannot be told apart from the other two hybrid configurations, the remaining nine are measurably worse

4. The intuitive confidence signal is the useless one. For deciding when not to answer, the obvious measure — how far ahead the top hit is — is worse than answering everything (0.25 selective accuracy against a 0.47 baseline). RRF fuses ranks as 1/(60+rank), so the top-two gap is ~2% on every query, confident or not. The dense retriever's raw cosine works; the margin does not.

Coverage against selective accuracy for two confidence signals; the top-1 margin performs worse than answering everything

Longer write-up of the first result: English · Türkçe.

Contents: Results · Leaderboard · Why not an existing benchmark? · Your own corpus · Running it · Gold set · Status · Limits · Contribute · Design notes

Results

58 queries, default embedding model paraphrase-multilingual-MiniLM-L12-v2, CPU only. Intervals are 95% bootstrap over queries; turkish-rag-eval report regenerates this table.

Chunking Retriever nDCG@10 95% CI R@5 MRR P95
hierarchical hybrid_rrf 0.613 [0.508, 0.716] 0.690 0.559 37 ms
sentence hybrid_rrf 0.558 [0.453, 0.662] 0.621 0.500 27 ms
fixed hybrid_rrf 0.517 [0.404, 0.629] 0.578 0.484 31 ms
sentence bm25_stem5 0.510 [0.403, 0.621] 0.638 0.462 4 ms
hierarchical dense 0.501 [0.396, 0.606] 0.569 0.441 30 ms
hierarchical bm25_stem5 0.494 [0.386, 0.605] 0.569 0.444 8 ms
fixed bm25_stem5 0.476 [0.371, 0.585] 0.526 0.421 4 ms
sentence dense 0.461 [0.356, 0.567] 0.552 0.410 21 ms
fixed dense 0.446 [0.341, 0.555] 0.491 0.402 25 ms
sentence bm25_nostem 0.411 [0.308, 0.515] 0.483 0.356 4 ms
fixed bm25_nostem 0.387 [0.290, 0.487] 0.414 0.319 3 ms
hierarchical bm25_nostem 0.383 [0.285, 0.482] 0.500 0.319 6 ms

turkish-rag-eval report does not stop at the table — it names the cheapest configuration the data cannot separate from the best:

Recommended: sentence + hybrid_rrf
  0.558 ndcg@10 against 0.613 for hierarchical + hybrid_rrf, a gap of 0.056
  that a paired bootstrap over 58 shared queries cannot distinguish from zero.
  It answers in 27 ms at P95 against 37 ms.

Comparisons are paired: both systems answered the same queries, so resampling per-query differences removes the variance of some queries being harder than others. It matters, and this data shows it in the sharpest way. hierarchical + hybrid_rrf beats fixed + hybrid_rrf by 0.097 and beats sentence + bm25_stem5 by 0.104 — yet only the larger gap clears zero, because its per-query differences are so much steadier (sd 0.34 against 0.40). Ranking by the gap alone gets this backwards.

Hierarchical chunking only helps the dense retriever, and fixed-size chunking — the most common default — lost on every retriever. More in the design notes.

Generation: groundedness

The R above feeds a G: the 58 questions answered from the top 5 hierarchical + bm25_stem5 passages by groq:openai/gpt-oss-120b, judged by a model from another family, nvidia:nvidia/nemotron-3-super-120b-a12b, at temperature 0 on 2026-09-30. The gold span was among the passages for 33 of 58 questions, and the two groups are scored apart:

Span retrieved (33)
answered when the span was retrieved 93.9%
answer rests on the passages 96.8%
answer conveys the gold span 96.8%
Span not retrieved (25)
declined when the span was not retrieved 80.0%
answered anyway 20.0%
of those, answer is right 60.0%
hallucinated (not in the passages) 4.0%

The generator mostly declines when the evidence is missing, and when it does answer, the answer is usually stated in some other retrieved passage — one of 25 is not. The first scoring rule called every answer without the span a hallucination and reported 20%; reading the five by hand showed four stated in a retrieved passage, so answers without the span now go to the judge too (why). 16 verdicts were checked by hand; 14 agreed. Of the other two, the judge was too strict once (an answer inverted "not recommended unless below 7 g/dl") and too lenient once (it accepted "hyperuricaemia" as the cause of gout).

This row depends on hosted APIs and cannot be reproduced offline; the models may change behind the same name. Only the rates are committed (results/groundedness.json) — the free API terms do not allow redistributing raw model output.

Leaderboard

Dense nDCG@10 per chunking strategy, the hybrid (dense + stemmed BM25, RRF) on hierarchical chunks, and the cost of each model on one 16-thread CPU, measured 2026-09-18. Stemmed BM25 alone scores 0.476 / 0.510 / 0.494. turkish-rag-eval leaderboard rebuilds this from results/models/.

Model Params Turkish-only Dense fixed Dense sentence Dense hierarchical Hybrid hierarchical Embed corpus Query P95
paraphrase-multilingual-MiniLM-L12-v2 (default) 118 M no 0.446 0.461 0.501 0.607 2 min 23 ms
emrecan/bert-base-turkish-cased-mean-nli-stsb-tr 111 M yes 0.408 0.431 0.497 0.654 4 min 43 ms
intfloat/multilingual-e5-small 118 M no 0.654 0.644 0.642 0.639 3 min 22 ms
intfloat/multilingual-e5-base 278 M no 0.631 0.677 0.668 0.648 10 min 49 ms
BAAI/bge-m3 568 M no 0.766 0.767 — — > 45 min —
newmindai/Mursit-Large-TR-Retrieval 404 M yes 0.746 0.740 0.781 0.673 35 min 214 ms

The two Turkish-only models are the most-downloaded Turkish entries in the Hugging Face sentence-similarity category. bge-m3 was stopped after 45 minutes, before the hierarchical chunks; its two numbers come from that partial run. google/embeddinggemma-300m is gated behind a licence click and was not run. Per-query results for every completed model are in results/models/.

These five rows were measured together on one machine before v0.1.0, which is why the timing columns are comparable with each other and not with the main table above. Their dense columns are unaffected by the stable tie-break, but each Hybrid hierarchical figure will move by roughly +0.006 when the model is re-run — the default row's went 0.607 → 0.613. turkish-rag-eval leaderboard --check reports them as missing provenance until then.

The top of this table is settled; the middle is not. Dense hierarchical, with 95% bootstrap intervals over the 58 queries: MiniLM 0.501 [0.396, 0.606], emrecan 0.497 [0.394, 0.601], e5-small 0.642 [0.542, 0.738], e5-base 0.668 [0.568, 0.765], Mursit 0.781 [0.701, 0.856]. The intervals overlap, but paired over the same queries Mursit beats the runner-up e5-base by +0.113 [+0.029, +0.202]. The two E5 models cannot be told apart (+0.026 [-0.039, +0.094]), and neither can MiniLM and emrecan (+0.004 [-0.123, +0.135]).

Hybrid fusion only pays for a weak dense model. RRF lifts the small default by +0.106 but pulls Mursit down from 0.781 to 0.673, a paired loss of -0.108 [-0.193, -0.030]; for the two E5 models it makes no measurable difference. "Turkish-only" is not enough either: the emrecan model was trained for sentence similarity and truncates input at 75 tokens.

To submit a model, open a pull request adding results/models/<org>__<name>/ (turkish-rag-eval run --model <org>/<name>, then leaderboard --check). CI re-derives every metric from the per-query relevance arrays committed beside it, and checks the harness version and corpus fingerprint.

Why not an existing benchmark?

MTEB-style retrieval benchmarks, TR-MTEB included, score an embedding model on passages that are already split. They answer "which model?", not "which chunker, is Turkish stemming worth it, does a hybrid help, and what does each cost on a CPU?". This harness keeps the articles whole, lets every chunker cut them its own way, and judges each chunk by the answer span, so pipeline choices are compared on the same labels. For a model-only comparison the same data exports to the BEIR layout MTEB reads (turkish-rag-eval export-hf), published as RizgarOzan/turkish-rag-eval; adding it to MTEB is proposed in embeddings-benchmark/mteb#5536.

Your own corpus

The harness is not tied to its own articles. Point it at a folder of .txt or .md files, a BEIR directory, or a JSON corpus:

turkish-rag-eval run --corpus ./belgelerim --gold ./sorular.json \
                     --retrievers bm25_stem5 bm25_nostem

A BM25-only run needs no model download and no torch — enough to answer "which chunker, and is stemming worth it on my documents" from the base install.

No labelled questions? That is the real wall, and the reason most benchmarks only ever measure themselves. bootstrap drafts a starting point (the name means drafting a gold set here, not the statistical resampling above):

export ANTHROPIC_API_KEY=...
turkish-rag-eval bootstrap ./belgelerim --out gold-draft.json --per-doc 3

It runs the same procedure the contributed questions in this repository used. One pass writes questions and marks the answer span; a second pass sees only the document and the questions and marks the span again. Where the passes disagree the item is written as "review": "needs-human" and never loads until a person settles it. Spans that are not verbatim, and questions copied out of their own answer, are dropped with a reason.

What you get is a draft to review, not a gold set.

Running it

Python 3.10+. CPU only — no GPU anywhere in this project.

pip install git+https://github.com/RizgarOzan/turkish-rag-eval                       # metrics, BM25, gold-set tooling
pip install 'turkish-rag-eval[all] @ git+https://github.com/RizgarOzan/turkish-rag-eval'  # + dense retrieval, charts, LLM commands
Command What it does
run every chunking × retriever combination; writes results/
report intervals, paired comparisons, and a recommendation
charts the two figures above, light and dark
leaderboard rebuild the model table; --check verifies every entry
abstain coverage / selective-accuracy curve for the best configuration (details)
groundedness score the generation half against the gold spans (details; results)
bootstrap draft a gold set for your own corpus
agreement inter-annotator agreement over the gold set
fetch-corpus download the Wikipedia snapshot; --verify checks the lock; --include-drafts for all 300 questions
export-hf the BEIR layout MTEB reads
validate every gold file's invariants

From a checkout:

git clone https://github.com/RizgarOzan/turkish-rag-eval
cd turkish-rag-eval
pip install -e '.[all]'
turkish-rag-eval fetch-corpus     # rebuilds data/raw/corpus.json
turkish-rag-eval run              # results/summary.json + per-query files
python -m pytest -q

run --model <name> swaps the embedding model; any sentence-transformers model works. Results for a non-default model go to results/models/<org>__<name>/, so the main table is never overwritten.

Gold set

Files Questions Labelled by In the results above
data/eval/gold.json (health) 58 one human yes
data/eval/contrib/llm-draft-*.json (history, geography, astronomy, biology, computing) 242 two independent LLM passes not yet

The set reached its planned 300 questions on 2026-09-26. Questions a model drafted are marked "source": "llm-draft". A second model then picked its own answer span for each one without seeing the first label (second_annotation). Two spans agree when one contains the other or their token F1 is at least 0.5; agreed items get "review": "agreed", the rest get "needs-human" and are never loaded. Batches 1 and 2 were 30 of 30 agreed, batch 3 was 29 of 30: for "what are several ribosomes working on one mRNA called?" the passes picked two different sentences that both name polysomes, so that question waits for a person. Batch 4 was 27 of 30: the second pass named the other claimant to the Hungarian throne, took the sentence beside the lysozyme result instead of the result itself, and answered "cross compilers" with the bare term where the first took its definition. Batch 5 was 30 of 30, and every pair is a containment pair: 9 identical, 15 differing only by a trailing full stop. One of its answers, Robert W. Holley's 1968 Nobel Prize, was cut to its last clause on 2026-09-27 and labelled again by the second pass: the sentence splitter breaks after "W.", so no hierarchical chunk held the whole sentence (#17). Batch 6 was 30 of 30 on reworded questions whose answers often run to two sentences; six first-pass spans were cut to one sentence before merging, because the two-sentence version fitted inside no chunk and so could never be retrieved. Batch 7 was 30 of 30 again, with 25 identical spans; one first-pass span was cut to its clause for the same reason. Like batch 5 its questions mostly point at a single sentence, so its high agreement says little about harder questions. Batch 8, the last 32, was written to be harder: 17 first-pass answers ran to two or more sentences. Seven of those crossed a sentence- or hierarchical-chunk boundary and were cut, leaving 12 multi-sentence answers. It was still 32 of 32, 29 identical: two LLMs reading the same article pick the same sentences even when the answer is long, which is one more reason these drafts need a human pass.

Drafts stay out of every number above until a re-run says otherwise: load_gold() skips them unless called with include_drafts=True. Synthetic test questions are common practice as long as they are declared and their agreement is measured. This section is that declaration.

The corpus is 54 Turkish Wikipedia articles, 1.09 M characters; 27 of them answer at least one question and the other 27 are distractors from the same domain.

Batch Questions Identical Mean IoU Mean token F1 Cohen's κ, fixed / sentence / hierarchical
1 — 2026-09-18 (Malazgirt, Kapadokya, Mars, Mitokondri, Linux) 30 16 0.816 0.878 1.00 / 1.00 / 1.00
2 — 2026-09-20 (İstanbul'un Fethi, Ağrı Dağı, Jüpiter, Fotosentez, İnternet)¹ 30 4 0.427 0.549 0.94 / 0.99 / 0.99
3 — 2026-09-21 (Çaldıran Muharebesi, Tuz Gölü, Satürn, Ribozom, Unix) 30 18 0.829 0.872 0.92 / 0.96 / 0.96
4 — 2026-09-24 (Mohaç Muharebesi, Kızılırmak, Venüs, Enzim, Derleyici) 30 15 0.747 0.792 0.96 / 0.96 / 0.96
5 — 2026-09-25 (Kösedağ Muharebesi, Uludağ, Neptün, RNA, İşletim sistemi) 30 24 0.922 0.943 1.00 / 1.00 / 1.00
6 — 2026-09-25 (Preveze Deniz Muharebesi, Erciyes, Uranüs, Hemoglobin, Veritabanı) 30 20 0.869 0.901 0.93 / 0.90 / 0.87
7 — 2026-09-26 (Ankara Muharebesi, Van Gölü, Merkür, DNA, World Wide Web) 30 25 0.944 0.960 0.93 / 0.93 / 0.97
8 — 2026-09-26 (Niğbolu Muharebesi, Fırat, Ay, Protein, Yapay zekâ) 32 29 0.964 0.975 0.94 / 0.97 / 1.00
All drafts 242 151 0.816 0.860 0.95 / 0.97 / 0.97

¹ Re-measured 2026-09-26: the Turkish Wikipedia article Jüpiter was rewritten that day to correct errors, and five of its spans no longer matched the live text. Three changed only in wording (a comma, "30,003" → "30" seconds, "dört uydu" → "uydular") and were edited in both labels; two changed in substance — the 40,000 km mantle thickness is gone and the Great Red Spot went from "at least 400 years" to "recorded since 1831" — so those two questions were rewritten and labelled again by both passes. The row was 0.419 / 0.540 / 0.96 / 1.00 / 1.00 before.

The passes disagree about how much of a sentence to take, rather than where the answer is. Two LLMs tend to pick the same sentence, so read this as a sanity check rather than human agreement. Full discussion in the design notes.

Scoring all 300

The drafted questions point at 40 articles outside the health snapshot, so they get a corpus of their own: the 54 health articles plus those 40, 1.89 M characters, pinned by data/corpus-full.lock.json. The published snapshot and its lock stay as they are.

turkish-rag-eval fetch-corpus --include-drafts   # data/raw/corpus-full.json
turkish-rag-eval run --include-drafts            # results/full/<model>/

Measured 2026-09-27 (re-run after the batch 5 fix; Mursit added 2026-10-01) on 296 questions (the 4 needs-human drafts never load), hierarchical chunks, nDCG@10:

Retriever 58 human questions 296 questions (238 drafted)
bm25_stem5 0.494 0.558
MiniLM (default), dense 0.501 0.405
MiniLM (default), hybrid_rrf 0.613 0.578
multilingual-e5-small, dense 0.642 0.568
multilingual-e5-small, hybrid_rrf 0.639 0.642
multilingual-e5-base, dense 0.668 0.646
multilingual-e5-base, hybrid_rrf 0.648 0.664
Mursit-Large-TR-Retrieval, dense 0.781 0.613
Mursit-Large-TR-Retrieval, hybrid_rrf 0.673 0.661

On the wider set stemmed BM25 gets stronger and every dense model weaker, so fusing the two now helps all four models, where on the 58 health questions it cost the E5 models and Mursit. Mursit drops the most, from 0.781 to 0.613, and falls below e5-base, so its lead on the human set does not carry over to the drafted questions yet. Those questions were written by an LLM from one sentence each; whether that style suits the smaller models better is open until people have checked them. These rows are not in the tables above: most of the questions have not been checked by a person yet. The run also names every question no chunk can answer; a span split by a chunk boundary is the usual reason. Here that is 17 questions with fixed-size chunks and none with sentence or hierarchical chunks (the one hierarchical miss was fixed on 2026-09-27, see batch 5 above). Results for all three chunkers are in results/full/.

Status

v0.1.0, tagged but not on PyPI yet — install from GitHub as above.

  • Works: the retrieval harness, groundedness on the 58 human questions, report, leaderboard --check, bootstrap, agreement, the Hugging Face export, scoring all 300 questions on a pinned corpus, and CI that validates every contributed question against live Wikipedia.
  • In progress: human review of the gold set. All 300 planned questions are in (58 human, 242 LLM-drafted); the drafts join the results once people have reviewed them.
  • Not yet: EmbeddingGemma, and groundedness on the 242 drafted questions.

Limits

  • 58 queries is a small set. The intervals above are the honest width of that; most of the table's ordering is not resolved by this much data. They replace the eyeballed rule of thumb this project used to carry — "treat differences under roughly 0.05 nDCG as noise" — which was both too strict for paired comparisons and too loose for unpaired ones.
  • One annotator for the human set. The 58 health questions are single-annotated, so they have no agreement figure. The 242 drafted questions are double-labelled, but by two LLM passes rather than by two people.
  • The corpus is Wikipedia, and so is much of the training data. Every embedding model ranked here was almost certainly trained on Turkish Wikipedia. Absolute scores are therefore optimistic; the comparison between pipeline choices on the same corpus is what this measures. A non-Wikipedia domain is the most valuable thing a contributor could add.
  • Six embedding models, one corpus. bge-m3 is a partial run and EmbeddingGemma is missing. abstain uses only the default model.
  • Encyclopaedic text, not clinical text. Nothing here transfers to a clinical setting without re-measurement. No patient data is used anywhere, and nothing here is a medical device.
  • The generation half is one generator, one judge, 58 questions. Groundedness was measured once, through hosted APIs, on the human set only; the judge was checked by hand on 16 verdicts.
  • The embedding model is pinned by name, not by revision.

Contribute

The first two limits shrink with every contributor. Adding questions needs no ML background — pick a Turkish Wikipedia article, write 5–10 paraphrased questions, and open a pull request with one JSON file. A validator checks each file against Wikipedia in CI. See CONTRIBUTING.md (Türkçe açıklama dahil) and the good first issues.

Submitting an embedding model is one command and a pull request — see Leaderboard.

Data

Turkish Wikipedia, CC BY-SA 4.0. See NOTICE.md.

Licence

Code MIT (LICENSE); data under data/ CC BY-SA 4.0 (NOTICE.md).

About

Which parts of a RAG pipeline earn their cost in Turkish? 12 retrieval configs measured on a hand-labelled Turkish gold set

Topics

Resources

Contributing

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages