Which parts of a RAG pipeline actually earn their cost on an agglutinative language? Three chunking strategies × four retrievers × six embedding models, measured on a hand-labelled Turkish gold set, with confidence intervals and a cost column — then pointed at your own documents.
pip install git+https://github.com/RizgarOzan/turkish-rag-eval # not on PyPI yet
turkish-rag-eval run --corpus ./belgelerim --gold ./sorular.json
turkish-rag-eval report1. The embedding model matters more than any pipeline choice. Every model trained for retrieval beats stemmed BM25 (0.494) on its own. The best, Mursit-Large-TR-Retrieval, reaches 0.781 nDCG@10 on hierarchical chunks. A small E5 the same size as the default goes from 0.501 to 0.642 with nothing else changed, at about the same speed (22 ms against 23 ms). The popular default model is the weak link, not dense retrieval. See the Leaderboard.
2. Turkish stemming is the cheapest real win. Truncating tokens to a
5-character prefix before BM25 lifts nDCG@10 by 23–29% on every chunking
strategy, and a paired bootstrap puts every one of those gains clear of zero
(hierarchical +0.111, 95% CI [+0.026, +0.204]). "diyabet", "diyabetin",
"diyabete", "diyabetli" are four surface forms of one concept; an unstemmed
index almost never matches the query's form.
3. The top of the table is a tie, and the tie decides on cost. With 58
queries, a paired bootstrap cannot separate the best configuration from the
other two hybrid ones; the remaining nine are measurably worse. So the
real choice at the top is price: sentence + hybrid_rrf gives the same
quality at 27 ms instead of 37 ms.
4. The intuitive confidence signal is the useless one. For deciding when
not to answer, the obvious measure — how far ahead the top hit is — is worse
than answering everything (0.25 selective accuracy against a 0.47 baseline).
RRF fuses ranks as 1/(60+rank), so the top-two gap is ~2% on every query,
confident or not. The dense retriever's raw cosine works; the margin does not.
Longer write-up of the first result: English · Türkçe.
Contents: Results · Leaderboard · Why not an existing benchmark? · Your own corpus · Running it · Gold set · Status · Limits · Contribute · Design notes
58 queries, default embedding model
paraphrase-multilingual-MiniLM-L12-v2, CPU only. Intervals are 95% bootstrap
over queries; turkish-rag-eval report regenerates this table.
| Chunking | Retriever | nDCG@10 | 95% CI | R@5 | MRR | P95 |
|---|---|---|---|---|---|---|
| hierarchical | hybrid_rrf | 0.613 | [0.508, 0.716] | 0.690 | 0.559 | 37 ms |
| sentence | hybrid_rrf | 0.558 | [0.453, 0.662] | 0.621 | 0.500 | 27 ms |
| fixed | hybrid_rrf | 0.517 | [0.404, 0.629] | 0.578 | 0.484 | 31 ms |
| sentence | bm25_stem5 | 0.510 | [0.403, 0.621] | 0.638 | 0.462 | 4 ms |
| hierarchical | dense | 0.501 | [0.396, 0.606] | 0.569 | 0.441 | 30 ms |
| hierarchical | bm25_stem5 | 0.494 | [0.386, 0.605] | 0.569 | 0.444 | 8 ms |
| fixed | bm25_stem5 | 0.476 | [0.371, 0.585] | 0.526 | 0.421 | 4 ms |
| sentence | dense | 0.461 | [0.356, 0.567] | 0.552 | 0.410 | 21 ms |
| fixed | dense | 0.446 | [0.341, 0.555] | 0.491 | 0.402 | 25 ms |
| sentence | bm25_nostem | 0.411 | [0.308, 0.515] | 0.483 | 0.356 | 4 ms |
| fixed | bm25_nostem | 0.387 | [0.290, 0.487] | 0.414 | 0.319 | 3 ms |
| hierarchical | bm25_nostem | 0.383 | [0.285, 0.482] | 0.500 | 0.319 | 6 ms |
turkish-rag-eval report does not stop at the table — it names the cheapest
configuration the data cannot separate from the best:
Recommended: sentence + hybrid_rrf
0.558 ndcg@10 against 0.613 for hierarchical + hybrid_rrf, a gap of 0.056
that a paired bootstrap over 58 shared queries cannot distinguish from zero.
It answers in 27 ms at P95 against 37 ms.
Comparisons are paired: both systems answered the same queries, so
resampling per-query differences removes the variance of some queries being
harder than others. It matters, and this data shows it in the sharpest way.
hierarchical + hybrid_rrf beats fixed + hybrid_rrf by 0.097 and beats
sentence + bm25_stem5 by 0.104 — yet only the larger gap clears zero,
because its per-query differences are so much steadier (sd 0.34 against 0.40).
Ranking by the gap alone gets this backwards.
Hierarchical chunking only helps the dense retriever, and fixed-size chunking — the most common default — lost on every retriever. More in the design notes.
The R above feeds a G: the 58 questions answered from the top 5
hierarchical + bm25_stem5 passages by groq:openai/gpt-oss-120b, judged by
a model from another family, nvidia:nvidia/nemotron-3-super-120b-a12b, at
temperature 0 on 2026-09-30. The gold span was among the passages for
33 of 58 questions, and the two groups are scored apart:
| Span retrieved (33) | |
|---|---|
| answered when the span was retrieved | 93.9% |
| answer rests on the passages | 96.8% |
| answer conveys the gold span | 96.8% |
| Span not retrieved (25) | |
|---|---|
| declined when the span was not retrieved | 80.0% |
| answered anyway | 20.0% |
| of those, answer is right | 60.0% |
| hallucinated (not in the passages) | 4.0% |
The generator mostly declines when the evidence is missing, and when it does answer, the answer is usually stated in some other retrieved passage — one of 25 is not. The first scoring rule called every answer without the span a hallucination and reported 20%; reading the five by hand showed four stated in a retrieved passage, so answers without the span now go to the judge too (why). 16 verdicts were checked by hand; 14 agreed. Of the other two, the judge was too strict once (an answer inverted "not recommended unless below 7 g/dl") and too lenient once (it accepted "hyperuricaemia" as the cause of gout).
This row depends on hosted APIs and cannot be reproduced offline; the models
may change behind the same name. Only the rates are committed
(results/groundedness.json) —
the free API terms do not allow redistributing raw model output.
Dense nDCG@10 per chunking strategy, the hybrid (dense + stemmed BM25, RRF)
on hierarchical chunks, and the cost of each model on one 16-thread CPU,
measured 2026-09-18. Stemmed BM25 alone scores 0.476 / 0.510 / 0.494.
turkish-rag-eval leaderboard rebuilds this from results/models/.
| Model | Params | Turkish-only | Dense fixed | Dense sentence | Dense hierarchical | Hybrid hierarchical | Embed corpus | Query P95 |
|---|---|---|---|---|---|---|---|---|
| paraphrase-multilingual-MiniLM-L12-v2 (default) | 118 M | no | 0.446 | 0.461 | 0.501 | 0.607 | 2 min | 23 ms |
| emrecan/bert-base-turkish-cased-mean-nli-stsb-tr | 111 M | yes | 0.408 | 0.431 | 0.497 | 0.654 | 4 min | 43 ms |
| intfloat/multilingual-e5-small | 118 M | no | 0.654 | 0.644 | 0.642 | 0.639 | 3 min | 22 ms |
| intfloat/multilingual-e5-base | 278 M | no | 0.631 | 0.677 | 0.668 | 0.648 | 10 min | 49 ms |
| BAAI/bge-m3 | 568 M | no | 0.766 | 0.767 | — | — | > 45 min | — |
| newmindai/Mursit-Large-TR-Retrieval | 404 M | yes | 0.746 | 0.740 | 0.781 | 0.673 | 35 min | 214 ms |
The two Turkish-only models are the most-downloaded Turkish entries in the
Hugging Face sentence-similarity category. bge-m3 was stopped after 45
minutes, before the hierarchical chunks; its two numbers come from that
partial run. google/embeddinggemma-300m is gated behind a licence click and
was not run. Per-query results for every completed model are in
results/models/.
These five rows were measured together on one machine before v0.1.0, which is why the timing columns are comparable with each other and not with the main table above. Their
densecolumns are unaffected by the stable tie-break, but eachHybrid hierarchicalfigure will move by roughly +0.006 when the model is re-run — the default row's went 0.607 → 0.613.turkish-rag-eval leaderboard --checkreports them as missing provenance until then.
The top of this table is settled; the middle is not. Dense hierarchical, with 95% bootstrap intervals over the 58 queries: MiniLM 0.501 [0.396, 0.606], emrecan 0.497 [0.394, 0.601], e5-small 0.642 [0.542, 0.738], e5-base 0.668 [0.568, 0.765], Mursit 0.781 [0.701, 0.856]. The intervals overlap, but paired over the same queries Mursit beats the runner-up e5-base by +0.113 [+0.029, +0.202]. The two E5 models cannot be told apart (+0.026 [-0.039, +0.094]), and neither can MiniLM and emrecan (+0.004 [-0.123, +0.135]).
Hybrid fusion only pays for a weak dense model. RRF lifts the small
default by +0.106 but pulls Mursit down from 0.781 to 0.673, a paired loss of
-0.108 [-0.193, -0.030]; for the two E5 models it makes no measurable
difference. "Turkish-only" is not enough either: the emrecan model was
trained for sentence similarity and truncates input at 75 tokens.
To submit a model, open a pull request adding results/models/<org>__<name>/
(turkish-rag-eval run --model <org>/<name>, then leaderboard --check).
CI re-derives every metric from the per-query relevance arrays committed
beside it, and checks the harness version and corpus fingerprint.
MTEB-style retrieval benchmarks, TR-MTEB included, score an embedding model on
passages that are already split. They answer "which model?", not "which
chunker, is Turkish stemming worth it, does a hybrid help, and what does each
cost on a CPU?". This harness keeps the articles whole, lets every chunker cut
them its own way, and judges each chunk by the answer span, so pipeline
choices are compared on the same labels. For a model-only comparison the same
data exports to the BEIR layout MTEB reads (turkish-rag-eval export-hf),
published as
RizgarOzan/turkish-rag-eval;
adding it to MTEB is proposed in
embeddings-benchmark/mteb#5536.
The harness is not tied to its own articles. Point it at a folder of .txt or
.md files, a BEIR directory, or a JSON corpus:
turkish-rag-eval run --corpus ./belgelerim --gold ./sorular.json \
--retrievers bm25_stem5 bm25_nostemA BM25-only run needs no model download and no torch — enough to answer "which chunker, and is stemming worth it on my documents" from the base install.
No labelled questions? That is the real wall, and the reason most
benchmarks only ever measure themselves. bootstrap drafts a starting point (the name means drafting a gold set here,
not the statistical resampling above):
export ANTHROPIC_API_KEY=...
turkish-rag-eval bootstrap ./belgelerim --out gold-draft.json --per-doc 3It runs the same procedure the contributed questions in this repository used.
One pass writes questions and marks the answer span; a second pass sees only
the document and the questions and marks the span again. Where the passes
disagree the item is written as "review": "needs-human" and never loads
until a person settles it. Spans that are not verbatim, and questions copied
out of their own answer, are dropped with a reason.
What you get is a draft to review, not a gold set.
Python 3.10+. CPU only — no GPU anywhere in this project.
pip install git+https://github.com/RizgarOzan/turkish-rag-eval # metrics, BM25, gold-set tooling
pip install 'turkish-rag-eval[all] @ git+https://github.com/RizgarOzan/turkish-rag-eval' # + dense retrieval, charts, LLM commands| Command | What it does |
|---|---|
run |
every chunking × retriever combination; writes results/ |
report |
intervals, paired comparisons, and a recommendation |
charts |
the two figures above, light and dark |
leaderboard |
rebuild the model table; --check verifies every entry |
abstain |
coverage / selective-accuracy curve for the best configuration (details) |
groundedness |
score the generation half against the gold spans (details; results) |
bootstrap |
draft a gold set for your own corpus |
agreement |
inter-annotator agreement over the gold set |
fetch-corpus |
download the Wikipedia snapshot; --verify checks the lock; --include-drafts for all 300 questions |
export-hf |
the BEIR layout MTEB reads |
validate |
every gold file's invariants |
From a checkout:
git clone https://github.com/RizgarOzan/turkish-rag-eval
cd turkish-rag-eval
pip install -e '.[all]'
turkish-rag-eval fetch-corpus # rebuilds data/raw/corpus.json
turkish-rag-eval run # results/summary.json + per-query files
python -m pytest -qrun --model <name> swaps the embedding model; any sentence-transformers
model works. Results for a non-default model go to
results/models/<org>__<name>/, so the main table is never overwritten.
| Files | Questions | Labelled by | In the results above |
|---|---|---|---|
data/eval/gold.json (health) |
58 | one human | yes |
data/eval/contrib/llm-draft-*.json (history, geography, astronomy, biology, computing) |
242 | two independent LLM passes | not yet |
The set reached its planned 300 questions on 2026-09-26. Questions a model
drafted are marked "source": "llm-draft". A second model then picked its own
answer span for each one without seeing the first label
(second_annotation). Two spans agree when one contains the other or their
token F1 is at least 0.5; agreed items get "review": "agreed", the rest get
"needs-human" and are never loaded. Batches 1 and 2 were 30 of 30 agreed,
batch 3 was 29 of 30: for "what are several ribosomes working on one mRNA
called?" the passes picked two different sentences that both name polysomes,
so that question waits for a person. Batch 4 was 27 of 30: the second pass
named the other claimant to the Hungarian throne, took the sentence beside the
lysozyme result instead of the result itself, and answered "cross compilers"
with the bare term where the first took its definition. Batch 5 was 30 of 30,
and every pair is a containment pair: 9 identical, 15 differing only by a
trailing full stop. One of its answers, Robert W. Holley's 1968 Nobel Prize,
was cut to its last clause on 2026-09-27 and labelled again by the second
pass: the sentence splitter breaks after "W.", so no hierarchical chunk held
the whole sentence (#17). Batch 6 was 30 of 30 on reworded questions whose answers
often run to two sentences; six first-pass spans were cut to one sentence
before merging, because the two-sentence version fitted inside no chunk and
so could never be retrieved. Batch 7 was 30 of 30 again, with 25 identical
spans; one first-pass span was cut to its clause for the same reason. Like
batch 5 its questions mostly point at a single sentence, so its high
agreement says little about harder questions. Batch 8, the last 32, was
written to be harder: 17 first-pass answers ran to two or more sentences.
Seven of those crossed a sentence- or hierarchical-chunk boundary and were
cut, leaving 12 multi-sentence answers. It was still 32 of 32, 29 identical:
two LLMs reading the same article pick the same sentences even when the
answer is long, which is one more reason these drafts need a human pass.
Drafts stay out of every number above until a re-run says otherwise:
load_gold() skips them unless called with include_drafts=True. Synthetic
test questions are common practice as long as they are declared and their
agreement is measured. This section is that declaration.
The corpus is 54 Turkish Wikipedia articles, 1.09 M characters; 27 of them answer at least one question and the other 27 are distractors from the same domain.
| Batch | Questions | Identical | Mean IoU | Mean token F1 | Cohen's κ, fixed / sentence / hierarchical |
|---|---|---|---|---|---|
| 1 — 2026-09-18 (Malazgirt, Kapadokya, Mars, Mitokondri, Linux) | 30 | 16 | 0.816 | 0.878 | 1.00 / 1.00 / 1.00 |
| 2 — 2026-09-20 (İstanbul'un Fethi, Ağrı Dağı, Jüpiter, Fotosentez, İnternet)¹ | 30 | 4 | 0.427 | 0.549 | 0.94 / 0.99 / 0.99 |
| 3 — 2026-09-21 (Çaldıran Muharebesi, Tuz Gölü, Satürn, Ribozom, Unix) | 30 | 18 | 0.829 | 0.872 | 0.92 / 0.96 / 0.96 |
| 4 — 2026-09-24 (Mohaç Muharebesi, Kızılırmak, Venüs, Enzim, Derleyici) | 30 | 15 | 0.747 | 0.792 | 0.96 / 0.96 / 0.96 |
| 5 — 2026-09-25 (Kösedağ Muharebesi, Uludağ, Neptün, RNA, İşletim sistemi) | 30 | 24 | 0.922 | 0.943 | 1.00 / 1.00 / 1.00 |
| 6 — 2026-09-25 (Preveze Deniz Muharebesi, Erciyes, Uranüs, Hemoglobin, Veritabanı) | 30 | 20 | 0.869 | 0.901 | 0.93 / 0.90 / 0.87 |
| 7 — 2026-09-26 (Ankara Muharebesi, Van Gölü, Merkür, DNA, World Wide Web) | 30 | 25 | 0.944 | 0.960 | 0.93 / 0.93 / 0.97 |
| 8 — 2026-09-26 (Niğbolu Muharebesi, Fırat, Ay, Protein, Yapay zekâ) | 32 | 29 | 0.964 | 0.975 | 0.94 / 0.97 / 1.00 |
| All drafts | 242 | 151 | 0.816 | 0.860 | 0.95 / 0.97 / 0.97 |
¹ Re-measured 2026-09-26: the Turkish Wikipedia article Jüpiter was rewritten that day to correct errors, and five of its spans no longer matched the live text. Three changed only in wording (a comma, "30,003" → "30" seconds, "dört uydu" → "uydular") and were edited in both labels; two changed in substance — the 40,000 km mantle thickness is gone and the Great Red Spot went from "at least 400 years" to "recorded since 1831" — so those two questions were rewritten and labelled again by both passes. The row was 0.419 / 0.540 / 0.96 / 1.00 / 1.00 before.
The passes disagree about how much of a sentence to take, rather than where the answer is. Two LLMs tend to pick the same sentence, so read this as a sanity check rather than human agreement. Full discussion in the design notes.
The drafted questions point at 40 articles outside the health snapshot, so
they get a corpus of their own: the 54 health articles plus those 40, 1.89 M
characters, pinned by data/corpus-full.lock.json. The published snapshot and
its lock stay as they are.
turkish-rag-eval fetch-corpus --include-drafts # data/raw/corpus-full.json
turkish-rag-eval run --include-drafts # results/full/<model>/Measured 2026-09-27 (re-run after the batch 5 fix; Mursit added 2026-10-01) on 296 questions (the 4 needs-human drafts never load),
hierarchical chunks, nDCG@10:
| Retriever | 58 human questions | 296 questions (238 drafted) |
|---|---|---|
| bm25_stem5 | 0.494 | 0.558 |
| MiniLM (default), dense | 0.501 | 0.405 |
| MiniLM (default), hybrid_rrf | 0.613 | 0.578 |
| multilingual-e5-small, dense | 0.642 | 0.568 |
| multilingual-e5-small, hybrid_rrf | 0.639 | 0.642 |
| multilingual-e5-base, dense | 0.668 | 0.646 |
| multilingual-e5-base, hybrid_rrf | 0.648 | 0.664 |
| Mursit-Large-TR-Retrieval, dense | 0.781 | 0.613 |
| Mursit-Large-TR-Retrieval, hybrid_rrf | 0.673 | 0.661 |
On the wider set stemmed BM25 gets stronger and every dense model weaker, so
fusing the two now helps all four models, where on the 58 health questions it
cost the E5 models and Mursit. Mursit drops the most, from 0.781 to 0.613, and
falls below e5-base, so its lead on the human set does not carry over to the
drafted questions yet. Those questions were written by an LLM from one sentence
each; whether that style suits the smaller models better is open until people
have checked them. These rows are not in the tables above: most of the
questions have not been checked by a person yet. The run also names every
question no chunk can answer; a span split by a chunk boundary is the usual
reason. Here that is 17 questions with fixed-size chunks and none with
sentence or hierarchical chunks (the one hierarchical miss was fixed on
2026-09-27, see batch 5 above). Results for all three chunkers are in
results/full/.
v0.1.0, tagged but not on PyPI yet — install from GitHub as above.
- Works: the retrieval harness,
groundednesson the 58 human questions,report,leaderboard --check,bootstrap,agreement, the Hugging Face export, scoring all 300 questions on a pinned corpus, and CI that validates every contributed question against live Wikipedia. - In progress: human review of the gold set. All 300 planned questions are in (58 human, 242 LLM-drafted); the drafts join the results once people have reviewed them.
- Not yet: EmbeddingGemma, and
groundednesson the 242 drafted questions.
- 58 queries is a small set. The intervals above are the honest width of that; most of the table's ordering is not resolved by this much data. They replace the eyeballed rule of thumb this project used to carry — "treat differences under roughly 0.05 nDCG as noise" — which was both too strict for paired comparisons and too loose for unpaired ones.
- One annotator for the human set. The 58 health questions are single-annotated, so they have no agreement figure. The 242 drafted questions are double-labelled, but by two LLM passes rather than by two people.
- The corpus is Wikipedia, and so is much of the training data. Every embedding model ranked here was almost certainly trained on Turkish Wikipedia. Absolute scores are therefore optimistic; the comparison between pipeline choices on the same corpus is what this measures. A non-Wikipedia domain is the most valuable thing a contributor could add.
- Six embedding models, one corpus.
bge-m3is a partial run and EmbeddingGemma is missing.abstainuses only the default model. - Encyclopaedic text, not clinical text. Nothing here transfers to a clinical setting without re-measurement. No patient data is used anywhere, and nothing here is a medical device.
- The generation half is one generator, one judge, 58 questions. Groundedness was measured once, through hosted APIs, on the human set only; the judge was checked by hand on 16 verdicts.
- The embedding model is pinned by name, not by revision.
The first two limits shrink with every contributor. Adding questions needs no ML background — pick a Turkish Wikipedia article, write 5–10 paraphrased questions, and open a pull request with one JSON file. A validator checks each file against Wikipedia in CI. See CONTRIBUTING.md (Türkçe açıklama dahil) and the good first issues.
Submitting an embedding model is one command and a pull request — see Leaderboard.
Turkish Wikipedia, CC BY-SA 4.0. See NOTICE.md.
Code MIT (LICENSE); data under data/ CC BY-SA 4.0
(NOTICE.md).