Skip to content

About

Domain-adapted small models for private-domain RAG: a 3B model vs an honestly measured frontier, plus a distilled groundedness guardrail. Open data, reproducible, negative results included.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Private-Domain RAG & Groundedness for Agents

Two reproducible, open-data studies on private-domain RAG methodology. All experiments run on public corpora used as a stand-in for a private domain, so every result here is reproducible from the code:

  • Study 1 — Beating the frontier on private-domain RAG (root of this repo) — a domain-adapted 3B model that matches or beats the honestly-measured frontier.
  • Study 2 — RAG-groundedness guardrail (groundedness-guardrail/) — a distilled verifier that flags unsupported answers (missed hallucinations).

Study 1 — Beating the frontier on private-domain RAG

A systematic method that lets a domain-adapted 3B model match or beat the honestly-measured frontier on private-domain retrieval QA — at ~1/5 the cost, ~94 ms/query, fully on-prem.

On a private corpus the frontier has no parametric advantage — it has never seen your data, so both a frontier model and a small model depend entirely on retrieval. This project shows that, in that regime, a small local model that is systematically adapted to the domain (retrieval + RAFT + distractor-augmentation) matches or beats the frontier on accuracy while winning decisively on cost, latency, and privacy. Every number is human-gold; the SLM numbers are held-out. Negative results and measurement artifacts are reported, not hidden.

PopQA is public data used here as a proxy for a private corpus: restricting the frontier model to the retrieved page approximates the information constraint of a domain it has never seen.


Headline result (PopQA-440, per-query latency on one 16 GB GPU)

config clean acc noisy acc latency
3B off-the-shelf + retrieval 77.1 67.3 65 ms
3B + RAFT 81.8 73.9 80 ms
3B + distractor-RAFT 82.3 75.2 94 ms
7B off-the-shelf + retrieval 82.3 71.8 139 ms
7B + RAFT 81.8 75.5 143 ms
frontier closed-book (no domain data) 73.4 — API
frontier-RAG, strict / context-only (honest private-domain) 77.0 76.4 API
frontier-RAG, hybrid (inflated by public-benchmark leak) 90.5 — API
  • Clean retrieval: every domain-adapted SLM config beats the honest frontier (82.3 vs 77.0, +5.3).
  • Noisy retrieval: a near-tie (75.2 vs 76.4) — and distractor-augmentation is what buys it (an off-the-shelf SLM collapses to 67).
  • Domain adaptation > model scaling: a fine-tuned 3B (82.3) equals a 2×-larger 7B off-the-shelf (82.3), at half the latency.
  • 3-seed robustness: 3B-RAFT = 82.5 ± 0.67 (seeds 42/43/44); every seed beats the honest frontier.

The key insight: measure the frontier honestly

Frontier-RAG is usually reported with a prompt that lets the model use its own world knowledge in addition to the retrieved context. On PopQA that scores 90.5% — but 77.0% once it is restricted to reading only the retrieved page (its true behaviour on data it has never memorized). So 13.5 of its 90.5 points were public-benchmark knowledge leakage, not reading. 77.0 is the honest private-domain baseline, and it ties an off-the-shelf 3B reading the same page. On a genuinely private corpus the 90.5 number simply does not exist.

The method

  1. Domain retrieval over the private corpus (here: the exact subject Wikipedia page; a real deployment uses a dense retriever — see below).
  2. RAFT-style SFT — fine-tune the SLM on (retrieved context, question) → gold answer, loss on the answer tokens only, so it learns to extract from the domain's documents, not to recite them.
  3. Distractor-augmented RAFT — train with the gold document plus distractor documents, shuffled, so it stays robust to the noisy retrieval real systems return.
  4. Deployment guards — a groundedness verifier to abstain on unsupported answers, and selective escalation of the reasoning-hard minority to the frontier.

Supplementary experiments (each is a runnable script)

  • Paraphrase robustness (eval_paraphrase.py): rewrite the question (natural + telegraphic "who direct X"), fixing entity/context/gold. The RAFT model is flat — 82.3 / 81.4 / 83.0 — so it learned to read, not to match PopQA's ~10 fixed templates.
  • Reasoning ceiling (eval_multihop.py): with context held fixed, the reader's limit is multi-hop composition, not size — 3B collapses on MuSiQue (12–17%), and scaling to 7B only partially helps (21–30%). Multi-hop needs query decomposition + iterative retrieval, not a bigger single-shot reader.
  • Retrieval recall@k (recall_eval.py, recall_dense.py): end-to-end accuracy = answer-recall@k × reader-acc. A real dense retriever (bge-small) reaches doc-recall@10 ≈ 99.5%.
  • Root-cause + fix of the recall ceiling (diagnose_recall.py, build_corpus_rich.py): answer-recall plateaued ~90% because the plaintext page extractor dropped infoboxes. Ingesting infobox key-values + full page text lifts answer-recall@20 90.9 → 95.4 and answer-in-gold-doc 88.6 → 93.6 — confirming the ceiling was a corpus artifact, not a retrieval limit. The residual ~6% is ~3% exact-match grading strictness + ~1.6% disambiguation + ~1.4% parse.

A prior negative result motivates the whole thing: an audited adaptive router (retrieve-or-not × model-tier) captures only 5.6–35.8% of the routing oracle's cost-quality advantage, and its reachable cost savings are unsafe (cutting cost 5× raises served confidently-wrong answers 0.1% → 33.5%). So the lever is domain adaptation, not clever routing. See RESULTS_M1.md–RESULTS_M3.md, RESULTS_VERIFY.md.

Reproduce

pip install -r requirements.txt          # torch, transformers, peft, bitsandbytes, datasets, scikit-learn, sentence-transformers
python build_ladder.py                    # build the difficulty ladder (PopQA/Trivia/Hotpot/MuSiQue/FreshQA)
python retrieve.py                        # retrieval contexts (MediaWiki API; resumable, throttle-aware)
python build_popqa_raft.py                # RAFT training set (held out from eval)
python build_raft_distractor.py           # distractor-augmented training set + noisy eval
python sft_popqa.py --seed 42             # RAFT LoRA SFT (add --qlora for 7B on 16 GB)
python ablation_eval.py                   # the headline table -> results_ablation.json
python recall_dense.py                    # retrieval recall@k (TF-IDF vs dense vs hybrid)

results_*.json in the repo are the exact outputs behind every number above. python = your interpreter with a CUDA PyTorch; set HF_HUB_OFFLINE=1 for cached models.

Honest limitations

  • PopQA is single-hop entity extraction — the SLM's sweet spot. Multi-hop / synthesis is its real weakness (escalate).
  • Under noisy retrieval the frontier is marginally ahead on accuracy (76.4 vs 75.2); the SLM's win there is cost/latency/privacy.
  • Costs are nominal relative units; latency is local-GPU-batched vs an API round-trip.
  • PopQA is a public-data proxy for a private domain; its 90.5 hybrid number would not exist on a real internal corpus.

Paper

paper/PAPER.md (full write-up), paper/main.pdf (compiled), paper/verification_report.md (every headline number traced to its source file).

Study 2 — RAG-groundedness guardrail

groundedness-guardrail/ — the groundedness-verification component referenced in Study 1's method (step 4). Given a retrieved context and a candidate answer, predict GROUNDED vs UNSUPPORTED (UNSUPPORTED = positive class; the headline is recall of missed hallucinations), evaluated on human gold with contamination controls (FaithBench 749 + TofuEval-MeetingBank 777).

  • A distilled ModernBERT cross-encoder student (trained on RAGTruth) reaches 58.89 min-AUROC across the panel — above the open-detector reference LettuceDetect (57.3) — i.e. it matches/beats open-SOTA discrimination.
  • Standalone it still trails a frontier zero-shot judge at the frozen operating point; a confidence-gated cascade recovers balanced accuracy as uncertain cases defer to the frontier — but its efficiency is capped by student miscalibration (reported as an honest negative, not hidden).
  • Multi-seed (student), human-gold, frozen-threshold, contamination-controlled. Details in groundedness-guardrail/paper/ and groundedness-guardrail/RESULTS_*.md; shares the top-level requirements.txt.

License

MIT — see LICENSE.

About

Domain-adapted small models for private-domain RAG: a 3B model vs an honestly measured frontier, plus a distilled groundedness guardrail. Open data, reproducible, negative results included.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages